Flash Sale

Special Discount Available

Summer discount is ending soon!

00 Days:00:00:01
HomeBootcampsGeneralAI Infrastructure/Networking Bootcamp by Orhan Ergun
Intermediate

AI Infrastructure/Networking Bootcamp by Orhan Ergun

5-Day Live Training — GPU Clusters, AI Fabrics, RDMA, RoCEv2, InfiniBand & AI Data Center Design

Bootcamps Certification
Language
English
Level
Intermediate
Full Lifetime Access
Certificate of Completion

Instructor-Led AI Infrastructure & Networking Training

5-Day Live Training — GPU Clusters, AI Fabrics, RDMA, RoCEv2, InfiniBand & AI Data Center Design

AI is changing data center networking.

Traditional enterprise and data center networks were designed around applications, servers, and relatively predictable traffic patterns. Large-scale AI training and inference introduce a very different environment: thousands of GPUs communicating simultaneously, massive east-west traffic, collective operations, microbursts, incast, extremely high bandwidth requirements, and sensitivity to congestion and latency.

The AI Infrastructure & Networking Bootcamp for Network Engineers is an intensive 5-day live training designed to help experienced network engineers and architects understand how modern AI infrastructure, GPU clusters, AI fabrics, and AI data center networks are designed.

The bootcamp connects the entire architecture from GPUs, parallelism, and collective communication to RDMA, RoCEv2, InfiniBand, lossless Ethernet, PFC, ECN, DCQCN, load balancing, AI fabric topologies, Ultra Ethernet, MRC, Scale-up, Scale-out, scale-across networks, power, cooling, and storage.

The goal is not simply to learn a list of new protocols and technologies. You will understand why AI networks are designed differently, how the technologies interact, what design trade-offs exist, and how architectural decisions affect GPU cluster performance and scalability.


Bootcamp Type

Number of Days

Location

Status

Timing (Based on Istanbul)

Nov 9-13, 2026

5 days in a row

Webex

🟡Nearly Confirmed

6 pm to 9 pm UTC+3

(Istanbul time) 

Jan 18-22, 2027

5 days in a row

Webex

🔵 Tentative

6 pm to 9 pm UTC +3

(Istanbul time)


🟢 Guaranteed:
Confirmed and scheduled to run on the announced dates.

🟡 Nearly Confirmed:
Approaching confirmation. Secure today’s pricing and join any guaranteed session within 24 months if rescheduled.

🔵 Tentative:
Planned based on demand and scheduling. Register now to lock in current pricing and receive priority placement in a guaranteed session within 24 months.

🟣Rescheduled:
This session has been moved to a future date. Your registration remains fully valid for the next guaranteed session within 24 months.

When you register, your seat and discounted price are secured. If your selected bootcamp is rescheduled, you can join any future guaranteed session of the same bootcamp within 24 months at no additional cost.


What You Will Learn

1. AI Fundamentals for Network Engineers

  • AI, Machine Learning, and Deep Learning fundamentals
  • Training vs. inference
  • Why GPUs are used for AI workloads
  • Model and GPU scaling
  • Why the network becomes critical as GPU clusters scale

2. AI Workloads & Network Traffic

  • North-South vs. East-West traffic
  • Elephant flows and microbursts
  • Incast behavior
  • AI traffic lifecycle
  • How AI workloads change traditional data center networking assumptions

3. AI Parallelism

  • Data Parallelism
  • Tensor Parallelism
  • Pipeline Parallelism
  • When and why different parallelism strategies are used
  • How parallelism choices affect network traffic

4. Collective Communications

  • Broadcast, Gather, and Scatter
  • Reduce and All-Reduce
  • All-Gather and Reduce-Scatter
  • Ring vs. Tree algorithms
  • How collective communication patterns influence AI fabric design

5. GPU & Compute System Architecture

  • CPU, GPU, TPU, DPU, and xPU
  • GPU memory hierarchy
  • High Bandwidth Memory (HBM)
  • PCIe fundamentals
  • Scale-Up vs. Scale-Out architectures

6. Scale-Up Networking

  • NVLink
  • NVSwitch
  • UALink
  • SUE-T
  • Scale-up network design principles
  • Connecting GPUs inside high-performance compute systems

7. Scale-Out AI Networks

  • Ethernet AI fabrics
  • InfiniBand fabrics
  • Clos architectures
  • Scale-out network design principles
  • Building networks for large GPU clusters

8. AI Network Topologies

  • Clos and high-radix switching
  • Dragonfly and Dragonfly+
  • Slimfly
  • 1D, 2D, and 3D Torus
  • Topology scalability and design trade-offs

9. RDMA Fundamentals

  • Remote Direct Memory Access (RDMA)
  • Kernel bypass
  • Zero-copy networking
  • Queue Pairs and Work Queues
  • Completion Queues
  • Memory Registration
  • RDMA Read, Write, and Atomic operations

10. RDMA Transport Technologies

  • InfiniBand
  • RoCEv1
  • RoCEv2
  • iWARP
  • Technology comparisons
  • Choosing the appropriate transport for AI infrastructure

11. Lossless Ethernet for AI

  • Why AI fabrics require lossless behavior
  • Priority Flow Control (PFC)
  • Head-of-Line Blocking
  • Congestion spreading
  • Data Center Bridging (DCB)
  • DCBX
  • Enhanced Transmission Selection (ETS)

12. Congestion Control

  • Explicit Congestion Notification (ECN)
  • Data Center Quantized Congestion Notification (DCQCN)
  • TIMELY
  • HPCC
  • In-Band Network Telemetry (INT)
  • Congestion-control deployment considerations

13. Load Balancing in AI Fabrics

  • Flow hashing
  • Limitations of traditional ECMP
  • Flowlet switching
  • Packet spraying
  • Adaptive routing
  • RDMA packet reordering challenges
  • Why load balancing becomes critical in large GPU clusters

14. Training vs. Inference Networks

  • Training traffic characteristics
  • Inference traffic characteristics
  • Frontend networks
  • Backend networks
  • Different networking requirements for training and inference workloads

15. AI Data Center & GPU Cluster Design

  • Rail-optimized architectures
  • Rail-unified designs
  • Single-rail vs. multi-rail networks
  • GPU pods
  • Large-scale AI clusters
  • Frontend vs. backend fabric design
  • Architectural trade-offs and their consequences

16. Ultra Ethernet

  • Ultra Ethernet Consortium (UEC) vision
  • Limitations of traditional Ethernet for AI workloads
  • Ultra Ethernet enhancements
  • UEC vs. RoCEv2 vs. InfiniBand
  • The evolution of Ethernet for AI and HPC environments

17. Power for AI Data Centers

  • Traditional vs. AI data center power requirements
  • High-density GPU racks
  • Power delivery
  • Redundancy
  • Power Usage Effectiveness (PUE)

18. Cooling

  • Why traditional air cooling becomes challenging for AI systems
  • Air vs. liquid cooling
  • Direct-to-Chip cooling
  • Immersion cooling
  • Cooling Distribution Units (CDUs)
  • Facility water loops

19. Storage for AI Infrastructure

  • Where storage is used in the AI lifecycle
  • HBM and its role in AI systems
  • Storage tiers
  • Checkpointing
  • GPUDirect Storage
  • Relationship between storage, compute, and network performance

Prerequisites

This bootcamp is designed primarily for network engineers, network architects, data center engineers, infrastructure architects, and experienced IT professionals moving into AI infrastructure and AI networking.

You should already have a working understanding of:

  • Ethernet and IP networking
  • Routing and switching fundamentals
  • Data center networking concepts
  • Basic leaf-spine / Clos architectures

You do not need to be an AI/ML engineer, data scientist, or GPU programmer.

The AI, GPU, parallelism, and collective communication concepts required to understand the networking architecture are introduced during the bootcamp from a network engineer's perspective.

Bootcamp Outcome

By the end of this 5-day bootcamp, you should be able to look at an AI data center or GPU cluster architecture and reason about it end-to-end.

Instead of seeing GPUs, NVLink, InfiniBand, RoCEv2, PFC, ECN, DCQCN, Clos fabrics, rail-optimized designs, Ultra Ethernet, storage, cooling, and power as independent technologies, you will understand how they fit together and how a decision in one part of the architecture affects the others.

You will be able to:

  • Understand how AI training and inference workloads generate network traffic
  • Explain how GPU-to-GPU communication affects network architecture
  • Understand scale-up and scale-out networking
  • Compare Ethernet, RoCEv2, and InfiniBand approaches
  • Understand RDMA and lossless Ethernet architectures
  • Reason about PFC, ECN, DCQCN, congestion, and load-balancing behavior
  • Evaluate AI fabric and GPU cluster topologies
  • Understand rail-optimized and multi-rail AI network designs
  • Compare traditional ECMP, flowlets, packet spraying, and adaptive routing
  • Understand the role of Ultra Ethernet in next-generation AI infrastructure
  • Discuss the architectural trade-offs behind modern AI data center designs
  • Understand how networking interacts with compute, storage, power, and cooling

Most importantly, you will develop the architectural foundation needed to design, evaluate, and discuss modern AI infrastructure and AI networking environments as a network engineer or network architect.

Technologies Covered

Bootcamps Certification

This Bootcamp Includes

  • Lifetime Access
  • Certificate of Completion

Subscribe for Exclusive Deals & Promotions

Stay informed about special discounts, limited-time offers, and promotional campaigns. Be the first to know when we launch new deals!