Reading paths
Choose where to start
ArchCrux essays are written as connected parts of a larger production engineering system. Start with the path closest to what you are currently building.
I am building AI systems
RAG, agents, evaluations, permissions, generation control, and engineering knowledge.
RAG in Production Is Mostly Not About Retrieval
Most RAG discussions obsess over retrieval. In production, the harder problems are ingestion, data lifecycle, generation control and evaluation.
Your RAG Evals Are Measuring the Wrong Thing
Most RAG evals measure answer quality at the end. Production failures usually begin much earlier: ingestion, freshness, permissions, retrieval boundaries and generation control.
The Model Is Not the Product
The model is only one unreliable dependency inside a production AI system. The product is the full path around it: retrieval, tools, permissions, latency, evaluation, fallback, and operational control.
Designing the Context Layer for Production AI Agents
Production AI agents need a context layer that assembles evidence, preserves authority and freshness, and separates model reasoning from durable workflow state.
Designing Long-Running AI Agents That Survive Failures
Long-running AI agents need durable execution, explicit state, idempotent actions, bounded retries, recovery workflows, and operational control to survive real production failures.
I am designing systems at scale
Multi-tenancy, rate limiting, queues, scheduling, tail latency, reliability, and capacity.
LLM Inference Is a Scheduler Problem
LLM inference performance is shaped by prefill, decode, KV cache, continuous batching, batching policy, queueing and scheduler tradeoffs, not just model speed.
Your LLM Rate Limiter Is Counting the Wrong Thing
The system design of LLM Rate Limiter.
One Tenant Can Break Your LLM Platform Without Sending Many Requests
The system design of multi-tenant LLM serving.
I want stronger architecture judgment
Failure modes, correctness boundaries, operational tradeoffs, and how systems behave under pressure.
Retries Make LLM Systems Less Reliable
Retries help normal distributed systems recover from transient failure. In LLM systems, they can also amplify cost, latency, duplicate work, fallback drift and wrong product behavior.
Your LLM System Isnt Slow. Your Tail Is Slow.
LLM systems rarely fail because the average request is slow. They fail because the tail gets worse when retrieval, model calls, tools, retries, and streaming delays stack together.
The Truth Is Not in the Docs
Production knowledge is not only in documentation. It is spread across code, tickets, incidents, ownership, permissions, dashboards, and the decisions engineers make under pressure.
Interview Prep
Preparing for Staff or Principal AI interviews?
Study system decisions through original questions on agents, retrieval, inference, reliability, and evaluation.
Explore Interview Prep →Newsletter
New essays by email
Email signup is temporarily unavailable. You can still read every essay in the archive.
Explore essays