Anton Nazaruk

CTO

Cloud Combinator

Aastha Paul

Solutions Architect

Cloud Combinator

From Zero to HyperPod: Choosing, Operating, and Cost-Controlling Distributed Model Training Infrastructure on AWS

Session Details

Event Date: 26/08/2026

Learning Areas


Services & Technologies


Content Type


Audience Level

Abstract

Training large models on AWS is no longer just an ML problem - it is a capacity, infrastructure, reliability, and cost-control problem.

This talk gives engineers a practical framework for choosing between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, SageMaker Training Plans, and SageMaker HyperPod. We'll look at when each option makes sense, what tradeoffs they introduce, and how to avoid common mistakes around GPU availability, quota planning, interruptions, and runaway cost.

We'll then walk through a repeatable distributed training blueprint: compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, checkpointing, and failure recovery. The session includes a demo-style walkthrough of launching a distributed training job and showing how node failure and recovery should be handled in a production-ready setup.

The goal is for attendees to leave with a clear mental model of how to run distributed model training on AWS reliably, how to choose the right service or capacity model, and how to make cost and failure recovery part of the architecture from day one.

What you'll learn

  • Choose between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, Training Plans, and HyperPod with a clear framework for when each makes sense.
  • Build a repeatable distributed training blueprint - compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, and checkpointing.
  • Design cost control and node failure recovery into your training architecture from day one.

About the Speaker

Anton Nazaruk is CTO at Cloud Combinator, where he works on cloud architecture, AI infrastructure, and distributed systems. He helps teams design production-ready platforms for data and AI workloads on AWS, with a focus on reliability, cost control, and repeatable infrastructure patterns.

Thank you to our Sponsors

Huge thanks to our sponsors for powering this community - so let’s return the favour - please take a moment to check out their websites, learn what they do, and say hello at our next event.

Secure foundations for bold innovation. Cloudscaler delivers the trusted cloud platforms that power data and AI transformation.

Rayo leads and supports organisations through the complex journey of technology transformation, whilst helping to build strategic resilience throughout the technology stack.

Get ahead with The Scale Factory, the award-winning AWS partner dedicated to helping ambitious businesses to grow, fast.