Building AI Agents with Multimodal Models
- Course Code GK847003
- Duration 1 day
Course Delivery
Course Delivery
This course is available in the following formats:
-
Company Event
Event at company
Request this course in a different delivery format.
Course Overview
Top
Learn how to build neural network agents that reason across multiple data types using advanced fusion techniques, OCR, and NVIDIA AI Blueprints for real-world applications like robotics and healthcare.
Company Events
These events can be delivered exclusively for your company at our locations or yours, specifically for your delegates and your needs. The Company Events can be tailored or standard course deliveries.
Course Schedule
TopCourse Objectives
Top- In this course, you will learn about:
- Different data types and how to make them neural network ready
- Model fusion, and the differences between early, late, and intermediate fusion
- PDF extraction using OCR
- The difference between modality and agent orchestration
- Customization of NVIDIA AI Blueprints with Video Search and Summarization (VSS)
Course Content
TopModule 1: Early and Late Fusion
- Use camera and LiDAR data to predict object positions.
- Convert various datatypes to make them neural network ready.
Module 2: Intermediate Fusion
- Explore the theory behind effective multimodal model architecture.
- Train a Contrastive Pretraining model.
- Create a vector database.
Module 3: Cross-modal Projection
- Converting a Language model into a Vision Language Model (VLM).
- Process PDFs with Optical Character Recognition (OCR) tools.
Module 4: Model Orchestration
- Analyze video using Cosmos Nemotron.
- Use VSS to answer user queries about video content.
- Orchestrate with NVIDIA AI Blueprints.
Module 5: Assessment
- Convert a pre-trained model to input a different datatype using projection.