Waitlist · Enrollment not open
Applied AI Engineering Cohort
Build AI systems beyond the demo.
A live engineering cohort on context, RAG, tools, agents, evaluation, durable execution, observability and the system-design decisions behind production AI.
8 live sessions · 3 hours each · Online
Planned format. Dates and price have not been announced.
Waitlist capture is temporarily unavailable. See the waitlist section below for its current status.
AI engineering gets difficult after the first working demo
Calling a model is usually the easy part.
The difficult part starts when the system needs to retrieve the right context, call tools safely, recover after failures, evaluate its own behavior, stay within latency and cost limits, and keep working when a task runs for minutes or hours.
That is what this cohort is about. We will build systems, break them, inspect what happened, and change the design.
- User
- Context
- Model / Planner
- Tools
- Execution State
- Evaluation
- Observability
We will look at how these parts interact in a real system: what the model knows, what it can change, what survives a crash, and what evidence tells us whether it worked. This is a reasoning path through the system, not a claim that evaluation and observability happen only at the end.
We will not only talk about these systems
The planned learning experience is hands-on engineering:
- Build a RAG system and inspect why retrieval fails.
- Change context construction and measure the result.
- Add tools to an agent and reason about permissions and failure boundaries.
- Introduce a failure after an external action succeeds.
- Recover the agent without executing the action twice.
- Add evaluation and inspect where the system is still wrong.
- Trace a multi-step execution.
- Reason about latency, cost and rate limits.
- Compare single-agent and multi-agent designs.
- Review architecture decisions as you would in a production design review.
8 live engineering sessions
These are planned session themes, not a confirmed calendar schedule.
Session 1
From LLM Call to AI System
Model calls vs AI systems; structured outputs; context; tools; execution; evaluation; production boundaries.
Build a prototype that exposes the difference between a model response and a system decision.
Session 2
Context Engineering
Context windows and assembly; retrieval; conversation state; memory; what should and should not enter the prompt; context failure modes.
Run an experiment where adding more context makes the system worse.
Session 3
Production RAG
Ingestion; chunking; lexical, vector and hybrid retrieval; reranking; metadata; permission-aware retrieval; answerability; abstention; evaluation.
Inspect retrieval evidence rather than only the final answer.
Session 4
Tools and Agent Execution
Tool schemas and selection; permissions; external actions; idempotency; retry boundaries; failure semantics.
Reason through a refund workflow with an external side effect and an explicit permission boundary.
Session 5
Durable Agents
Long-running tasks; execution state; checkpoints; retries; leases and fencing where relevant; recovery; reconciliation; exactly-once illusions.
The external action succeeds, but the worker crashes before recording success. Decide what happens next.
Session 6
Evaluation and Observability
Traces; model outputs; tool calls; retrieval evidence; task success; evaluation sets; online vs offline evaluation; failure classification; debugging AI systems.
Investigate how an apparently good final answer can hide a broken system.
Session 7
Multi-Agent and Production Architecture
When multiple agents make sense; planner vs runtime; specialization; shared context; coordination; failure propagation; stopping; cost; concurrency.
Compare a single-agent design with a multi-agent design. More agents are not automatically better.
Session 8
Designing the Production System
Routing; rate limiting; cost control; tail latency; permissions; observability; evaluation; durable execution; failure recovery.
Bring the pieces together in an architecture review of a complete AI system.
How the cohort will work
Planned format: 8 live sessions, 3 hours per session, 2 sessions per week, planned across four weekends (4 weekends), online.
Each session will move approximately through this sequence:
- Problem
- How the system works
- Architecture
- Build
- Break it
- Inspect the evidence
- Improve the design
Exact dates will be announced before enrollment opens.
Illustrative lab · Planned exercise
Example: the refund succeeded, but your agent thinks it failed
- Agent decides a refund should be issued.
- Tool calls the payment provider.
- Provider successfully issues the refund.
- Worker crashes before recording the result.
- Another worker resumes the task.
- It sees no recorded success.
Should it call the refund API again?
We will inspect durable state, idempotency, external side effects, retries, reconciliation and recovery. Then we will design recovery without repeating the external action.
This is the kind of problem we will work through. Not just how to make an agent call a tool.
This describes planned teaching. No starter repository, downloadable lab pack, recording or hosted environment is promised. No real payment or refund is performed on this website.
This will probably be useful if...
- You already know software engineering and want to move deeper into AI engineering.
- You can build an LLM prototype but are less confident designing it for production.
- You work on backend, platform or distributed systems and want to understand AI-native architecture.
- You want to reason about RAG, agents, context, evaluation and execution as systems problems.
- You are preparing for senior AI engineering or system-design discussions.
Software engineers, senior engineers and Staff/Principal engineers moving into AI engineering are the intended audience.
Probably not the right cohort if...
- You are looking mainly for prompt-writing techniques.
- You want a no-code AI-tools course.
- You are completely new to programming.
- You want a course centered on a single framework.
This is not a certification program or a job guarantee.
Every session needs evidence
We will start with a recognizable task, a visible output and an engineering question. You will make a decision, inspect the evidence and explain what you would change.
We will not stop at “Here is an agent architecture.” We will inspect retrieved passages, traces, state changes, tool calls, evaluation results and failure behavior. Where available, a reproducible teaching artifact will let us revisit the result.
The goal is to understand why the system behaved the way it did.

Taught by Amit Kumar
Principal Software Architect
Works on backend and distributed systems and writes ArchCrux about how production AI systems fail, scale and recover.
Join the waitlist
Dates have not been announced and enrollment is not open. Waitlist signup is temporarily unavailable. Individual or team interest does not reserve a place.
Waitlist signup is temporarily unavailable. No request is collected here; dates and fees remain unannounced.
Only the interests you choose. A global unsubscribe or provider suppression is never reversed by this form. Privacy