Skip to main Content

Building AI Agents with Multimodal Models

  • Course Code GK847003
  • Duration 1 day

Course Delivery

Company Event Price

Please call

Request Group Training Add to Cart

Course Delivery

This course is available in the following formats:

  • Company Event

    Event at company

Request this course in a different delivery format.

Course Overview

Top
Learn how to build neural network agents that reason across multiple data types using advanced fusion techniques, OCR, and NVIDIA AI Blueprints for real-world applications like robotics and healthcare.

Company Events

These events can be delivered exclusively for your company at our locations or yours, specifically for your delegates and your needs. The Company Events can be tailored or standard course deliveries.

Course Schedule

Top

Course Objectives

Top
  • In this course, you will learn about:
  • Different data types and how to make them neural network ready
  • Model fusion, and the differences between early, late, and intermediate fusion
  • PDF extraction using OCR
  • The difference between modality and agent orchestration
  • Customization of NVIDIA AI Blueprints with Video Search and Summarization (VSS)

Course Content

Top

Module 1:  Early and Late Fusion

  • Use camera and LiDAR data to predict object positions.
  • Convert various datatypes to make them neural network ready.

Module 2: Intermediate Fusion

  • Explore the theory behind effective multimodal model architecture.
  • Train a Contrastive Pretraining model.
  • Create a vector database.

Module 3: Cross-modal Projection

  • Converting a Language model into a Vision Language Model (VLM).
  • Process PDFs with Optical Character Recognition (OCR) tools.

Module 4: Model Orchestration

  • Analyze video using Cosmos Nemotron.
  • Use VSS to answer user queries about video content.
  • Orchestrate with NVIDIA AI Blueprints.

Module 5:  Assessment

  • Convert a pre-trained model to input a different datatype using projection.
Cookie Control toggle icon