Skip to main Content

Implementing and Optimizing LLM Inferencing Systems with Cisco AI Pods and NVIDIA Data Center Technologies (DCLLM)

  • Course Code N1_DCLLM
  • Duration 5 days

Course Delivery

Virtual Learning Price

USD1,675.00

excl. VAT

Request Group Training Add to Cart

Course Delivery

This course is available in the following formats:

  • Company Event

    Event at company

  • Virtual Learning

    Learning that is virtual

Request this course in a different delivery format.

Course Overview

Top

This comprehensive training equips participants with the knowledge and skills required to design, deploy, and optimize Large Language Model (LLMs) operations using NVIDIA GPUs and Cisco AI Pods infrastructure. Through in-depth modules, hands-on labs, and real-world case studies, participants will learn how to manage data preparation, build scalable pipelines, optimize performance, ensure security, and migrate from cloud to on-premises deployments. The course provides a holistic approach for mastering the technical complexities of LLM systems and their infrastructure management while leveraging cutting-edge NVIDIA and Cisco technologies for scalability, efficiency, and security.

Updated 05/05/2026

Virtual Learning

This interactive training can be taken from any location, your office or home and is delivered by a trainer. This training does not have any delegates in the class with the instructor, since all delegates are virtually connected. Virtual delegates do not travel to this course, Global Knowledge will send you all the information needed before the start of the course and you can test the logins.

Course Schedule

Top

Target Audience

Top

This course is tailored for professionals involved in designing, implementing, optimizing, and managing AI and data infrastructure, including:

- Systems Architects: To understand the integration of LLM systems into broader IT environments

- Network Architects: To optimize network configurations for high-speed LLM training and inferencing

- Storage Architects: To manage the storage and retrieval of large-scale datasets used in LLM systems

- AI Infrastructure Architects: To build robust and scalable AI platforms optimized for LLM workloads

- Data Scientists: To prepare high-quality datasets and fine-tune LLMs for specific use cases

- Machine Learning Engineers: To deploy and optimize LLMs for real-world applications with low latency and high throughput

Course Objectives

Top

By the end of this workshop, participants will:

1.Examine the Foundations of LLMs: Gain an in-depth understanding of LLM architecture, scaling principles, and design trade-offs

2.Prepare and Manage Large Datasets: Learn techniques for sourcing, preprocessing, and managing large-scale, high-quality datasets for LLM training

3.Deploy LLMs for Production: Use NVIDIA TensorRT and Cisco Nexus Dashboard to build efficient, low-latency inferencing pipelines

4.Optimize LLM Performance: Apply advanced optimization techniques like quantization, pruning, and dynamic batching to improve throughput and reduce latency

5.Design Scalable Pipelines: Build fault-tolerant, high-performance pipelines for real-time and batch inferencing

6.Monitor and Maintain AI Systems: Use NVIDIA and Cisco tools to monitor GPU and network performance, ensuring reliability and uptime

7.Ensure Security and Privacy: Implement robust security measures using Cisco Nexus Dashboard, Cisco XDR, and NVIDIA encryption tools

8.Build On-Premises Data Center Skills: Design and implement LLM inferencing systems using NVIDIA GPUs and Cisco AI Pods & Secure AI Factories for maximum scalability and efficiency

9.Migrate Cloud Models to On-Premise: Transition cloud-trained LLMs to on-premise infrastructure while optimizing performance and costs

Course Content

Top

Module 1: On-Premises Data Center Design for LLM Inferencing Systems

Objective of Module 1: Design an on-premises data center with Cisco and NVIDIA technologies

Topics of Module 1:

  • Cisco UCS and NVIDIA GPUs for high-performance compute
  • Network design and automation with Cisco Nexus Dashboard
  • Storage solutions for large-scale data management

Module 2: On-Premises Data Center Implementation for LLM Inferencing Systems

Objective Module 2: Implement and configure an LLM inferencing data center using NVIDIA and Cisco technologies

Topics of Module 2:

  • Physical setup: NVIDIA GPUs on Cisco AI Pods and Secure AI Factories and Nexus networking configuration
  • Performance testing and validation of inferencing pipelines

Module 3: Large Language Model (LLM) Foundations

Objectives of Module 3:

  • Understand the architecture and mathematical principles of LLMs
  • Learn design trade-offs for scalability and performance
  • Explore emerging innovations in LLM development

 Topics of Module 3:

  • Transformer architecture, self-attention mechanism, and positional encoding
  • Types of LLMs: Encoder-only, decoder-only, and encoder-decoder
  • Training objectives: Masked language modeling (MLM), causal language modeling (CLM), and sequence-to-sequence modeling
  • Scaling laws and challenges: Parameter size, dataset size, and compute
  • Emerging architectures: Reformer, Longformer, and multi-modal LLMs

Module 4: Deployment of LLMs for Inferencing

Objectives of Module 4:

  • Deploy LLMs for production inferencing with high performance and scalability
  • Use NVIDIA TensorRT and Cisco Nexus Dashboard for optimized deployment

Topics of Module 4:

  • Deployment architectures: On-premises, cloud, and hybrid
  • Optimizing inferencing with NVIDIA TensorRT: Precision calibration, layer fusion, and batching
  • Traffic management and load balancing with Cisco Nexus Dashboard
  • Exposing LLM APIs: RESTful and gRPC endpoints with security mechanisms

Module 5: Monitoring, Logging, and Maintenance for LLM Systems

Objectives of Module 5: Monitor and maintain LLM deployments using NVIDIA and Cisco tools

Topics of Module 5:

  • Key metrics: Latency, throughput, GPU utilization, and memory usage
  • Monitoring tools: NVIDIA DCGM and Cisco Nexus Dashboard Insights
  • Maintenance workflows for hardware and software reliability

Module 6: Optimizing LLM Models and their performance for Inferencing

Objectives of Module 6:

  • Optimize LLM inferencing pipelines for low latency and high throughput
  • Learn techniques like quantization, pruning, and model compression

Topics of Module 6:

  • Quantization: FP16, INT8, and mixed precision
  • Pruning and knowledge distillation for lightweight models
  • TensorRT optimization: Dynamic batching and asynchronous execution
  • Benchmarking tools: NVIDIA Triton Inference Server, TensorRT Profiler

Module 7: Data Collection, Retrieval and Preparation for LLM Applications

Objectives of Module 7:

  • Understand data requirements for LLMs and their impact on performance
  • Learn techniques for sourcing, cleaning, and managing large-scale datasets
  • Explore NVIDIA and Cisco tools for efficient data handling

Topics of Module 7:

  • Data sourcing: Open-source, proprietary, and domain-specific datasets
  • Preprocessing: Cleaning, deduplication, tokenization, and filtering
  • Data management: Sharding, scalable storage, and high-speed data transfer
  • Ethical considerations: Bias detection, privacy compliance, and fairness

Module 8: Scalable Pipeline Design for LLM Inferencing

Objectives Module 8:

  • Build robust, scalable, and fault-tolerant pipelines for inferencing
  • Use batching, caching, and dynamic scaling for efficient pipelines

Topics of Module 8:

  • Pipeline components: Batching, caching, and queuing
  • Load balancing with Cisco Nexus Dashboard for traffic optimization
  • Fault tolerance: Automatic failover and disaster recovery plans
  • Monitoring pipeline performance with NVIDIA DCGM and Cisco Nexus Dashboard

Module 9: Hybrid Operations for LLM Inference

Objective of Module 9: Implement a hybrid LLM inference approach that integrates on-premises and cloud-based models for flexible, secure, and scalable AI service delivery

Topics of Module 9:

  • Deploying LLM inference endpoints on-premises and in the cloud
  • Configuring hybrid model access and request-routing policies
  • Managing security, visibility, and operational consistency across hybrid AI services

Module 10: Security and Privacy Considerations in LLM Training and Inferencing

Objectives of Module 10:

  • Secure LLM pipelines using Cisco Nexus Dashboard, Cisco XDR, and NVIDIA tools

Topics of Module 10:

  • NVIDIA runtime encryption and secure boot
  • Cisco Robust Intelligence for adversarial defense and vulnerability detection
  • Cisco XDR for unified threat detection and automated response
  • Traffic segmentation and endpoint authentication

Lab Outline:

  • Design a complete data center architecture for LLM inferencing
  • Deploying Transformer Model Inference Services on Red Hat OpenShift AI
  • Tokenization
  • Build LLM-powered applications using inference services
  • Configuring GPU Monitoring for LLM Inference Service
  • Implementing Observability for LLM Systems with Splunk
  • Benchmarking LLM Inference on Cisco AI Pods
  • Implementing a Retrieval-Augmented Generation Pipeline
  • Deploying Multi-Model Routing
  • Configuring API Gateway and Rate Limiting for LLM Inference Services
  • Implementing Multi-Model Hybrid Access for LLM Inference
  • Enabling AI Runtime Protection
  • Discovering AI Access
  • Enforcing AI Guardrails
  • Operating AI Security

Course Prerequisites

Top
  • No specific requirement for this course

Test Certification

Top
  • None

Follow on Courses

Top
  • None recommended

Further Information

Top
"Course book in English provided to participants"
Cookie Control toggle icon