Below are our latest videos, you can find more great content on our YouTube Channel
If you have a story you would like to share with our community, check out our Call for Papers.
If you would like to get email notifications for our new videos, subscribe below.

Martin Thwaites
Principal Developer AdvocateHoneycomb.ioEvery Player Gets a VM: Prepping Battleships for re:Invent Scale
My job is not to build games, but it is to make our booths engaging, and Battleships is my latest creation. The first version ran on EKS, which meant a mostly idle cluster for something that only runs a few days at a time.
And then I needed to think about re:Invent and its 65,000 attendees. When Lambda MicroVMs launched in June I rebuilt it, giving every player their own VM with its own dedicated memory and no shared state to fight over, scaling out as far as the queue at the booth demands.
You will learn the nuances I hit along the way, how the move changed the architecture once a game session had a whole machine to itself rather than a slice of a pod, and the constraints I am designing around for the moment a keynote ends and the expo hall fills up. I’ll show the cost modelling against EKS and the instrumentation going in so that I will know the game is working before an attendee tells me it isn’t.
What you’ll learn
- Compare the cost of a persistently provisioned EKS cluster against per-session Lambda MicroVMs for spiky, short-lived workloads, using real modelling rather than intuition.
- Rework an architecture once a session owns a whole machine rather than a slice of a pod – which shared state disappears, and which new constraints take its place.
- Instrument a system so you know it is degrading before a user tells you, and design around the failure modes of a hard, unmovable traffic spike.
About the Speaker
Martin Thwaites is a Principal Developer Advocate at Honeycomb and a contributor to OpenTelemetry’s .NET libraries. He has spent the last seven years working in observability, following a twenty-year engineering career building large-scale systems. His focus is distributed tracing and telemetry pipelines in cloud-native environments. He advocates for observability as a core part of software engineering, something owned by the engineers building software, not just the teams operating it.
Learning Area
Services & Technologies
Content Type
Audience Level

Yan Cui
AWS Serverless HeroIndependent ConsultantEvent-driven architecture: the hard parts
Event-driven architectures are great at helping teams build loosely coupled, scalable systems – but the neat architecture diagrams rarely show the difficult bits. In practice, they hand you a fresh set of problems that nobody warns you about at the whiteboard stage.
Once events start flowing across multiple services and bounded contexts, new challenges emerge. How do you trace a request end-to-end when there’s no longer a simple call stack? How do you change an event without unexpectedly breaking consumers you don’t control? And how do you test and operate a system where failures can happen asynchronously, perhaps minutes after the original request?
In this talk, Yan will explore the practical lessons and patterns that make event-driven architectures work in production, including:
- End-to-end observability as events flow across services and asynchronous boundaries
- Testing strategies within bounded contexts, and deciding what should, and shouldn’t, be tested end-to-end
- Schema evolution and how to change events safely without breaking existing consumers
- Catching integration problems early, before they reach production
- Error handling, retries and idempotency, including dealing with duplicate and out-of-order events
- Documentation and governance so teams can understand which events exist, who owns them and how they should be used
Drawing on real-world experience building event-driven systems on AWS, Yan will look beyond the happy path and focus on the trade-offs, failure modes and engineering practices that help these architectures remain manageable as systems and teams grow.
What you’ll learn
- How to design an event envelope that gives you versioning, idempotency and tracing for free
- Why “no breaking changes” often beats every event versioning scheme, and how to enforce it
- A practical way to scope end-to-end tests so they don’t spill into other teams’ systems
- Where governance is worth the operational overhead, and where it just slows you down
About the Speaker
Yan Cui is an AWS Serverless Hero and independent consultant who has run production workloads on AWS since 2009. He helps teams go faster for less with serverless, and has trained thousands of developers through his Production-Ready Serverless workshops. A prolific writer and speaker, his articles on AWS and serverless have been read millions of times.
Learning Area
Services & Technologies
Content Type
Audience Level

Anton Nazaruk
CTOCloud Combinator
Aastha Paul
Solutions ArchitectAWSFrom Zero to HyperPod: Choosing, Operating, and Cost-Controlling Distributed Model Training Infrastructure on AWS
Training large models on AWS is no longer just an ML problem – it is a capacity, infrastructure, reliability, and cost-control problem.
This talk gives engineers a practical framework for choosing between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, SageMaker Training Plans, and SageMaker HyperPod. We’ll look at when each option makes sense, what tradeoffs they introduce, and how to avoid common mistakes around GPU availability, quota planning, interruptions, and runaway cost.
We’ll then walk through a repeatable distributed training blueprint: compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, checkpointing, and failure recovery. The session includes a demo-style walkthrough of launching a distributed training job and showing how node failure and recovery should be handled in a production-ready setup.
The goal is for attendees to leave with a clear mental model of how to run distributed model training on AWS reliably, how to choose the right service or capacity model, and how to make cost and failure recovery part of the architecture from day one.
What you’ll learn
- Choose between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, Training Plans, and HyperPod with a clear framework for when each makes sense.
- Build a repeatable distributed training blueprint – compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, and checkpointing.
- Design cost control and node failure recovery into your training architecture from day one.
About the Speaker
Anton Nazaruk is CTO at Cloud Combinator, where he works on cloud architecture, AI infrastructure, and distributed systems. He helps teams design production-ready platforms for data and AI workloads on AWS, with a focus on reliability, cost control, and repeatable infrastructure patterns.
Learning Area
Services & Technologies
Content Type
Audience Level

Silvia Lehnis
Chief AI OfficerUBDS DigitalBeyond Prompt Optimisation: Improving Accuracy in LLM-Based Document Extraction
Variable document formats make structured data extraction difficult to solve with prompting alone. Especially when the documents include everything from images, handwriting, complex financial tables repeated multiple times and without a particular form or naming conventions.
This talk explores how we optimised an intelligent document processing solution on Amazon Bedrock by testing 17 different variants across prompting, model selection, document handling and validation. We will share the methods to optimise an LLM based solution that apply to many use-cases beyond documents, as well as the key trade-offs, failure modes and design decisions that had the greatest impact on extraction accuracy and reliability.
What you’ll learn
- How to test and compare AI solution options, including model selection, experiment design and evaluation methods.
- How the findings shaped key design decisions and the final Amazon Bedrock solution released into production.
- Which methods can improve the performance of an LLM-based solution, and the trade-offs between value and implementation effort.
About the Speaker
Silvia Lehnis helps high-impact organisations use data and AI to reach their goals faster and safer. She has led global and national data and AI transformations in sensitive environments across the public sector, finance, energy and academia, taking strategy through to implementation across people, process and technology – with solutions reaching up to 110,000 users and delivering £27m in savings over three years. She’s also a board member of the charity Care in Action.
Learning Area
Services & Technologies
Content Type
Audience Level

Guilherme Dalla Rosa
CTOMerCloudBuilding Secure and Efficient SaaS Platforms on AWS Serverless
Let’s go on a journey through the world of multi-tenant architectures on AWS using serverless technologies. In this talk, we will uncover the key aspects of multi-tenancy, including security, tenant isolation, and performance. We will learn how to utilise Cognito for authentication, DynamoDB to store millions of tenant-partitioned records and lambda for compute. We will also explore different deployment models and their tradeoffs, and, finally, we will learn how to implement policy-based isolation with IAM to keep our execution context tied to one specific tenant and avoid data leakage. By the end of this talk, you will feel more confident building SaaS applications on AWS with serverless technologies and you will have learned some of the many insights that come from the AWS Well-Architected SaaS Lens.
What you’ll learn
- Implement IAM policy-based isolation to scope each Lambda execution context to a single tenant and prevent data leakage
- Evaluate pool, silo, and bridge deployment models – the cost, complexity, and isolation trade-offs of each
- Use Amazon Cognito and DynamoDB together for tenant-partitioned authentication and millions of tenant-scoped records at scale
- Apply AWS Well-Architected SaaS Lens patterns to make defensible, production-ready architectural decisions
- Automate tenant provisioning from the start – the patterns that work at five tenants break at fifty
About the Speaker
Guilherme Dalla Rosa is a seasoned software engineer with extensive industry experience, having contributed to projects across Brazil, Ireland, and the UK. He currently serves as the CTO of MerCloud, a B2B e-commerce platform that streamlines the sales process for companies. In addition to his leadership role at MerCloud, Guilherme is also an AWS Community Builder, where he actively shares his expertise and passion for leveraging technology to help businesses enhance their processes and achieve their goals.
Learning Area
Services & Technologies
Content Type
Audience Level
Thank you to our Sponsors
Huge thanks to our sponsors for powering this community – so let’s return the favour – please take a moment to check out their websites, learn what they do, and say hello at our next event.
Secure foundations for bold innovation. Cloudscaler delivers the trusted cloud platforms that power data and AI transformation.
Rayo leads and supports organisations through the complex journey of technology transformation, whilst helping to build strategic resilience throughout the technology stack.


