Building AI Agents with Multimodal Models
- Código del Curso GK847003
- Duración 1 Día
Otros Métodos de Impartición
Salta a:
Método de Impartición
Este curso está disponible en los siguientes formatos:
-
Cerrado
Cerrado
Solicitar este curso en un formato de entrega diferente.
Temario
Parte superiorCompany Events
These events can be delivered exclusively for your company at our locations or yours, specifically for your delegates and your needs. The Company Events can be tailored or standard course deliveries.
Calendario
Parte superiorObjetivos del Curso
Parte superior- In this course, you will learn about:
- Different data types and how to make them neural network ready
- Model fusion, and the differences between early, late, and intermediate fusion
- PDF extraction using OCR
- The difference between modality and agent orchestration
- Customization of NVIDIA AI Blueprints with Video Search and Summarization (VSS)
Contenido
Parte superiorModule 1: Early and Late Fusion
- Use camera and LiDAR data to predict object positions.
- Convert various datatypes to make them neural network ready.
Module 2: Intermediate Fusion
- Explore the theory behind effective multimodal model architecture.
- Train a Contrastive Pretraining model.
- Create a vector database.
Module 3: Cross-modal Projection
- Converting a Language model into a Vision Language Model (VLM).
- Process PDFs with Optical Character Recognition (OCR) tools.
Module 4: Model Orchestration
- Analyze video using Cosmos Nemotron.
- Use VSS to answer user queries about video content.
- Orchestrate with NVIDIA AI Blueprints.
Module 5: Assessment
- Convert a pre-trained model to input a different datatype using projection.