Implementing and Optimizing LLM Inferencing Systems with Cisco AI Pods and NVIDIA Data Center Technologies (DCLLM)
- Course Code N1_DCLLM
- Duration 5 days
Course Delivery
Jump to:
Course Delivery
This course is available in the following formats:
-
Company Event
Event at company
-
Virtual Learning
Learning that is virtual
Request this course in a different delivery format.
Course Overview
TopThis comprehensive training equips participants with the knowledge and skills required to design, deploy, and optimize Large Language Model (LLMs) operations using NVIDIA GPUs and Cisco AI Pods infrastructure. Through in-depth modules, hands-on labs, and real-world case studies, participants will learn how to manage data preparation, build scalable pipelines, optimize performance, ensure security, and migrate from cloud to on-premises deployments. The course provides a holistic approach for mastering the technical complexities of LLM systems and their infrastructure management while leveraging cutting-edge NVIDIA and Cisco technologies for scalability, efficiency, and security.
Updated 05/05/2026
Virtual Learning
This interactive training can be taken from any location, your office or home and is delivered by a trainer. This training does not have any delegates in the class with the instructor, since all delegates are virtually connected. Virtual delegates do not travel to this course, Global Knowledge will send you all the information needed before the start of the course and you can test the logins.
Course Schedule
TopTarget Audience
TopThis course is tailored for professionals involved in designing, implementing, optimizing, and managing AI and data infrastructure, including:
- Systems Architects: To understand the integration of LLM systems into broader IT environments
- Network Architects: To optimize network configurations for high-speed LLM training and inferencing
- Storage Architects: To manage the storage and retrieval of large-scale datasets used in LLM systems
- AI Infrastructure Architects: To build robust and scalable AI platforms optimized for LLM workloads
- Data Scientists: To prepare high-quality datasets and fine-tune LLMs for specific use cases
- Machine Learning Engineers: To deploy and optimize LLMs for real-world applications with low latency and high throughput
Course Objectives
TopBy the end of this workshop, participants will:
1.Examine the Foundations of LLMs: Gain an in-depth understanding of LLM architecture, scaling principles, and design trade-offs
2.Prepare and Manage Large Datasets: Learn techniques for sourcing, preprocessing, and managing large-scale, high-quality datasets for LLM training
3.Deploy LLMs for Production: Use NVIDIA TensorRT and Cisco Nexus Dashboard to build efficient, low-latency inferencing pipelines
4.Optimize LLM Performance: Apply advanced optimization techniques like quantization, pruning, and dynamic batching to improve throughput and reduce latency
5.Design Scalable Pipelines: Build fault-tolerant, high-performance pipelines for real-time and batch inferencing
6.Monitor and Maintain AI Systems: Use NVIDIA and Cisco tools to monitor GPU and network performance, ensuring reliability and uptime
7.Ensure Security and Privacy: Implement robust security measures using Cisco Nexus Dashboard, Cisco XDR, and NVIDIA encryption tools
8.Build On-Premises Data Center Skills: Design and implement LLM inferencing systems using NVIDIA GPUs and Cisco AI Pods & Secure AI Factories for maximum scalability and efficiency
9.Migrate Cloud Models to On-Premise: Transition cloud-trained LLMs to on-premise infrastructure while optimizing performance and costs
Course Content
TopModule 1: On-Premises Data Center Design for LLM Inferencing Systems
Objective of Module 1: Design an on-premises data center with Cisco and NVIDIA technologies
Topics of Module 1:
- Cisco UCS and NVIDIA GPUs for high-performance compute
- Network design and automation with Cisco Nexus Dashboard
- Storage solutions for large-scale data management
Module 2: On-Premises Data Center Implementation for LLM Inferencing Systems
Objective Module 2: Implement and configure an LLM inferencing data center using NVIDIA and Cisco technologies
Topics of Module 2:
- Physical setup: NVIDIA GPUs on Cisco AI Pods and Secure AI Factories and Nexus networking configuration
- Performance testing and validation of inferencing pipelines
Module 3: Large Language Model (LLM) Foundations
Objectives of Module 3:
- Understand the architecture and mathematical principles of LLMs
- Learn design trade-offs for scalability and performance
- Explore emerging innovations in LLM development
Topics of Module 3:
- Transformer architecture, self-attention mechanism, and positional encoding
- Types of LLMs: Encoder-only, decoder-only, and encoder-decoder
- Training objectives: Masked language modeling (MLM), causal language modeling (CLM), and sequence-to-sequence modeling
- Scaling laws and challenges: Parameter size, dataset size, and compute
- Emerging architectures: Reformer, Longformer, and multi-modal LLMs
Module 4: Deployment of LLMs for Inferencing
Objectives of Module 4:
- Deploy LLMs for production inferencing with high performance and scalability
- Use NVIDIA TensorRT and Cisco Nexus Dashboard for optimized deployment
Topics of Module 4:
- Deployment architectures: On-premises, cloud, and hybrid
- Optimizing inferencing with NVIDIA TensorRT: Precision calibration, layer fusion, and batching
- Traffic management and load balancing with Cisco Nexus Dashboard
- Exposing LLM APIs: RESTful and gRPC endpoints with security mechanisms
Module 5: Monitoring, Logging, and Maintenance for LLM Systems
Objectives of Module 5: Monitor and maintain LLM deployments using NVIDIA and Cisco tools
Topics of Module 5:
- Key metrics: Latency, throughput, GPU utilization, and memory usage
- Monitoring tools: NVIDIA DCGM and Cisco Nexus Dashboard Insights
- Maintenance workflows for hardware and software reliability
Module 6: Optimizing LLM Models and their performance for Inferencing
Objectives of Module 6:
- Optimize LLM inferencing pipelines for low latency and high throughput
- Learn techniques like quantization, pruning, and model compression
Topics of Module 6:
- Quantization: FP16, INT8, and mixed precision
- Pruning and knowledge distillation for lightweight models
- TensorRT optimization: Dynamic batching and asynchronous execution
- Benchmarking tools: NVIDIA Triton Inference Server, TensorRT Profiler
Module 7: Data Collection, Retrieval and Preparation for LLM Applications
Objectives of Module 7:
- Understand data requirements for LLMs and their impact on performance
- Learn techniques for sourcing, cleaning, and managing large-scale datasets
- Explore NVIDIA and Cisco tools for efficient data handling
Topics of Module 7:
- Data sourcing: Open-source, proprietary, and domain-specific datasets
- Preprocessing: Cleaning, deduplication, tokenization, and filtering
- Data management: Sharding, scalable storage, and high-speed data transfer
- Ethical considerations: Bias detection, privacy compliance, and fairness
Module 8: Scalable Pipeline Design for LLM Inferencing
Objectives Module 8:
- Build robust, scalable, and fault-tolerant pipelines for inferencing
- Use batching, caching, and dynamic scaling for efficient pipelines
Topics of Module 8:
- Pipeline components: Batching, caching, and queuing
- Load balancing with Cisco Nexus Dashboard for traffic optimization
- Fault tolerance: Automatic failover and disaster recovery plans
- Monitoring pipeline performance with NVIDIA DCGM and Cisco Nexus Dashboard
Module 9: Hybrid Operations for LLM Inference
Objective of Module 9: Implement a hybrid LLM inference approach that integrates on-premises and cloud-based models for flexible, secure, and scalable AI service delivery
Topics of Module 9:
- Deploying LLM inference endpoints on-premises and in the cloud
- Configuring hybrid model access and request-routing policies
- Managing security, visibility, and operational consistency across hybrid AI services
Module 10: Security and Privacy Considerations in LLM Training and Inferencing
Objectives of Module 10:
- Secure LLM pipelines using Cisco Nexus Dashboard, Cisco XDR, and NVIDIA tools
Topics of Module 10:
- NVIDIA runtime encryption and secure boot
- Cisco Robust Intelligence for adversarial defense and vulnerability detection
- Cisco XDR for unified threat detection and automated response
- Traffic segmentation and endpoint authentication
Lab Outline:
- Design a complete data center architecture for LLM inferencing
- Deploying Transformer Model Inference Services on Red Hat OpenShift AI
- Tokenization
- Build LLM-powered applications using inference services
- Configuring GPU Monitoring for LLM Inference Service
- Implementing Observability for LLM Systems with Splunk
- Benchmarking LLM Inference on Cisco AI Pods
- Implementing a Retrieval-Augmented Generation Pipeline
- Deploying Multi-Model Routing
- Configuring API Gateway and Rate Limiting for LLM Inference Services
- Implementing Multi-Model Hybrid Access for LLM Inference
- Enabling AI Runtime Protection
- Discovering AI Access
- Enforcing AI Guardrails
- Operating AI Security
Course Prerequisites
Top- No specific requirement for this course
Test Certification
Top- None
Follow on Courses
Top- None recommended