Private team bootcamp

Private instructor-led training

AI Infrastructure & AI Networking

Build practical AI networking expertise with GPU clusters, RDMA, lossless Ethernet, congestion control, scale-out fabrics, and AI data center design.

Available Languages
English, Turkish

Program details

Curriculum and delivery details

Scroll tables horizontally to view all columns.

AI Infrastructure & AI Networking

AI Infrastructure & AI Networking is a technical bootcamp for network engineers, data center architects, infrastructure teams, and technical leaders who need to understand how modern AI workloads change compute, network, storage, power, and cooling design. The course connects AI fundamentals with the practical realities of building GPU clusters, lossless fabrics, RDMA-based transport, and scale-out architectures for training and inference environments.

Companies invest in AI networking and infrastructure training because AI projects often fail to scale when the underlying fabric, compute architecture, or data center design is treated like a traditional enterprise workload. Large GPU clusters create intense east-west traffic, collective communication patterns, congestion sensitivity, power density challenges, and cooling requirements that demand specialized design decisions. Training helps engineering teams make better architectural trade-offs, reduce avoidable design mistakes, and communicate more effectively across networking, platform, storage, and facilities teams.

Learning from Orhan Ergun gives teams a structured, network-engineering-focused path through a rapidly evolving domain. The bootcamp emphasizes the “why” behind AI fabric choices, not only the terminology, so participants can evaluate Ethernet, InfiniBand, RoCEv2, Ultra Ethernet, topology options, congestion-control methods, and GPU cluster designs with clearer technical reasoning.

Who This Bootcamp Is For

  • Network engineers moving from traditional data center networking into AI and GPU-cluster environments
  • Data center architects evaluating scale-up and scale-out designs for AI workloads
  • Infrastructure, cloud, and platform teams supporting training or inference systems
  • Technical decision makers who need to understand AI fabric, power, cooling, and storage trade-offs
  • Engineers comparing InfiniBand, RoCE, iWARP, Ultra Ethernet, and Ethernet-based AI fabrics

Technical Focus

The program begins with AI concepts for network engineers, including training versus inference, GPU scaling, and why the network becomes critical as GPU clusters grow. It then moves into AI workload behavior, parallelism models, collective communications, GPU and compute system architecture, scale-up interconnects, scale-out fabrics, RDMA, lossless Ethernet, congestion control, load balancing, and complete AI data center design considerations.

Curriculum at a Glance

Module Area

Core Topics

Practical Design Focus

AI and Workload Foundations

AI, Machine Learning, Deep Learning, training vs. inference, GPU scaling, north-south and east-west traffic, elephant flows, microbursts, incast

Understand why AI workloads break traditional data center networking assumptions

Parallelism and Collective Communication

Data, tensor, and pipeline parallelism; broadcast, gather, scatter, reduce, All-Reduce, All-Gather, Reduce-Scatter; ring and tree algorithms

Map AI communication patterns to network bandwidth, latency, and topology requirements

GPU Systems and Scale-Up Networking

CPU, GPU, TPU, DPU, xPU, HBM, PCIe, NVLink, NVSwitch, UALink, SUE-T, scale-up architecture

Evaluate how GPUs connect inside high-performance compute systems

Scale-Out Fabrics and Topologies

Ethernet AI fabrics, InfiniBand, Clos, high-radix switching, Dragonfly, Dragonfly+, Slimfly, 1D/2D/3D Torus

Compare topology scalability, design trade-offs, and GPU-cluster fabric options

RDMA, Lossless Ethernet, and Congestion

RDMA, kernel bypass, zero-copy, Queue Pairs, Work Queues, Completion Queues, memory registration, RoCEv1, RoCEv2, iWARP, PFC, DCB, DCBX, ETS, ECN, DCQCN, TIMELY, HPCC, INT

Design and assess transport behavior, lossless operation, and congestion control in AI fabrics

AI Data Center Infrastructure

Training and inference networks, frontend and backend fabrics, rail-optimized and rail-unified designs, single-rail and multi-rail networks, GPU pods, Ultra Ethernet, power, PUE, air and liquid cooling, Direct-to-Chip, immersion cooling, CDUs, storage tiers, checkpointing, GPUDirect Storage

Connect network architecture with compute, facility, power, cooling, and storage requirements

Why Companies Take This Training

  • To prepare network and infrastructure teams for the traffic patterns produced by large AI training and inference workloads
  • To evaluate AI fabric choices such as InfiniBand, RoCEv2, iWARP, Ethernet AI fabrics, and Ultra Ethernet with more technical confidence
  • To understand the impact of collective communication, RDMA behavior, congestion, and load balancing on GPU utilization
  • To align networking decisions with GPU architecture, storage access, power density, and cooling constraints
  • To build a common vocabulary between network engineering, AI platform, data center, and facilities teams

What Makes the Training Practical

Rather than treating AI infrastructure as a single technology choice, the bootcamp frames it as a system design problem. Participants examine how workload behavior influences fabric design, how scale-up and scale-out networking differ, why lossless behavior matters, where traditional ECMP can fall short, and how architectural choices such as rail-optimized designs, multi-rail networking, or frontend/backend fabric separation affect large-scale AI clusters.

What your team will gain

  • Analyze how AI training and inference workloads affect data center network design
  • Compare scale-up and scale-out architectures for GPU-based AI infrastructure
  • Evaluate RDMA transports, lossless Ethernet mechanisms, and congestion-control options for AI fabrics
  • Design AI network topologies using Clos, Dragonfly, Torus, and related scalability trade-offs
  • Apply AI data center design considerations across networking, storage, power, and cooling

Recommended preparation

  • Working knowledge of data center networking fundamentals
  • Familiarity with Ethernet, IP networking, and basic traffic-flow concepts
  • Willingness to learn AI infrastructure terminology, GPU-cluster concepts, and fabric design trade-offs

Technology scope

Network Design

Get Network Training Updates

Hear about new bootcamps, course launches, and limited offers.