Question Library
Search 495 Staff and Principal-level questions, or browse by engineering area.
Search and browse questions
495 questions
Staff · Work through production mechanics, failure modes and trade-offs.
Principal · Reason through ambiguity, platform boundaries, invariants and second-order effects.
Data and Knowledge Systems
48 questions- How would you design retrieval when permissions change faster than the index?Permissions, identity and deletion ·Principal
- The document was deleted yesterday. Why is the assistant still quoting it?Permissions, identity and deletion ·Staff
- How would you change the parser and embedding model without breaking search?Versioning and migration ·Principal
- Which policy source wins when the newer indexed copy is wrong?Knowledge representation and evidence ·Staff
- One tenant owns forty percent of ingestion. How do you keep everyone fresh?Ingestion and freshness ·Principal
- A connector event is replayed after the file changes. Which version gets indexed?Ingestion and freshness ·Staff
- A connector says success, but eighteen percent of new grants are missing. Where is the gap?Permissions, identity and deletion ·Principal
- Sentence chunks improve recall. Would you embed every sentence?Knowledge representation and evidence ·Staff
- A connector uses file paths as document IDs. What breaks on a move?Versioning and migration ·Staff
- Can a search index safely expand nested groups into user IDs?Permissions, identity and deletion ·Principal
- How do you stop an answer from mixing two document versions?Versioning and migration ·Principal
- A retrieved table gives the right row and the wrong numberKnowledge representation and evidence ·Staff
- Which code actually consumes a changed event?Knowledge representation and evidence ·Principal
- Can the assistant say no consumers exist?Knowledge representation and evidence ·Principal
- A bulk snapshot races the change streamIngestion and freshness ·Principal
- A user updates a document and search answers from the old revisionIngestion and freshness ·Principal
- The source corrected yesterday's data. Should the assistant rewrite yesterday's answer?Knowledge representation and evidence ·Principal
- The incident agent read today's code. What was actually running during Tuesday's outage?Versioning and migration ·Staff
- Three data sources are current and one is a day behind. Is the AI report current?Ingestion and freshness ·Principal
- The Delta table deleted a customer row. Why does vector search still return it?Permissions, identity and deletion ·Principal
- The training feature knew about a payment that had not arrived yet. How?Ingestion and freshness ·Staff
- The source changed while the report was being written. Which facts can it publish?Versioning and migration ·Staff
- The CDC consumer is healthy again. Did it miss three days of changes?Ingestion and freshness ·Staff
- The PDF parser captured every word. Why did the answer reverse the chart's trend?Knowledge representation and evidence ·Staff
- The CDC field is still called `amount`. When did its units change?Versioning and migration ·Principal
- A deleted document still appears in the next training run. Which copy did the loader read?Permissions, identity and deletion ·Principal
- Train and eval have different document IDs. Why did the model see the test answers?Ingestion and freshness ·Principal
- A deleted customer document returned after an index rebuild. Which deletion did replay miss?Permissions, identity and deletion ·Principal
- A source row changed its primary key. Why does search show both identities?Ingestion and freshness ·Staff
- One database transaction updated two tables. Why did the AI report see only half of it?Ingestion and freshness ·Principal
- The incident happened last Sunday. Why did the AI report count it this week?Ingestion and freshness ·Principal
- The event schema is backward compatible. Why did old requests become approved?Versioning and migration ·Principal
- The newer customer update was processed first. Why did the old value win?Ingestion and freshness ·Principal
- The incident happened at 1:30 AM twice. Which event did the AI report use?Ingestion and freshness ·Principal
- The update changed a name. Why did search erase the customer's address?Ingestion and freshness ·Principal
- The indexing job says success. Why are new policy documents missing?Ingestion and freshness ·Principal
- The document sync cursor expired. Can we just request a fresh one and continue?Ingestion and freshness ·Principal
- The table was emptied. Why does the AI index still contain every old row?Ingestion and freshness ·Principal
- Search found the newest policy. Why did the assistant apply next month's rule today?Ingestion and freshness ·Principal
- The AI indexer is paused. Why did the source Postgres disk fill up?Ingestion and freshness ·Principal
- The hourly AI report was final. Why did a late event change its count?Ingestion and freshness ·Principal
- We deleted a customer record from Cassandra. Why did repair bring it back?Permissions, identity and deletion ·Principal
- We saved the Delta table version. Why can't we reproduce last month's training data?Versioning and migration ·Principal
- We retained the Kafka topic. Why can't the agent reconstruct last year's policy changes?Versioning and migration ·Principal
- We reuploaded a deleted file at the same path. Why did search delete the new one?Permissions, identity and deletion ·Principal
- OCR read every word on the page. Why did RAG combine two different clauses?Knowledge representation and evidence ·Staff
- The PDF parser extracted every sentence. Why did the summary join two unrelated paragraphs?Knowledge representation and evidence ·Staff
- The form extractor read every option. Why did RAG say an unchecked box was selected?Knowledge representation and evidence ·Staff
Retrieval and RAG
37 questions- Search code and docs when exact names and fuzzy questions mixQuery understanding and retrieval design ·Principal
- The query rewrite removed a version number. Why did RAG still look healthy?Query understanding and retrieval design ·Staff
- Hybrid search p99 climbs at 10,000 QPS. One shard is hotScale and federated operations ·Principal
- Recall at twenty rises, but citations get worseRanking, evidence and grounding ·Staff
- How do you retrieve a contract clause and the exception that changes it?Ranking, evidence and grounding ·Principal
- Would you move retrieval from a managed vector service to pgvector?Query understanding and retrieval design ·Staff
- Only newly ingested PDFs got worse. How do you find the broken stage?Scale and federated operations ·Principal
- The best search hit is restricted. What reaches the model?Access, completeness and answerability ·Staff
- Why did hybrid search fill every slot with one old document?Ranking, evidence and grounding ·Principal
- Clicks improve the reranker while the exception disappearsRanking, evidence and grounding ·Staff
- The search returned results, but the policy source timed outAccess, completeness and answerability ·Principal
- Filtered ANN returns no permitted hits from a populated corpusAccess, completeness and answerability ·Principal
- The reranker never sees the policy exceptionRanking, evidence and grounding ·Staff
- A federated search merges incomparable scoresScale and federated operations ·Principal
- Would token-level retrieval beat one vector per passage?Query understanding and retrieval design ·Principal
- RAG found five rows. Can it answer how many accounts breached a limit?Access, completeness and answerability ·Staff
- Half the index uses the new embedding model. Are the search scores still meaningful?Scale and federated operations ·Staff
- Two retrieved facts are true. Why is the answer built from them false?Ranking, evidence and grounding ·Staff
- A Hindi-English query names a product in Latin letters. Why does search miss its policy?Query understanding and retrieval design ·Staff
- The answer cites the right document, but the quoted paragraph moved. Which span is authoritative?Ranking, evidence and grounding ·Staff
- The SQL replica says zero incidents while search has incident notes. What can the assistant claim?Access, completeness and answerability ·Staff
- The embedding model is unchanged. Why did cosine to dot product change the top result?Ranking, evidence and grounding ·Staff
- New documents are indexed. Why did ANN recall fall after a week of updates?Scale and federated operations ·Principal
- The new retriever finds fresh documents. Why does offline NDCG say it got worse?Ranking, evidence and grounding ·Principal
- The retriever found the exception. Why did the reranker put it last?Ranking, evidence and grounding ·Staff
- The embedding model is the same. Why does using `encode()` for queries hurt search?Query understanding and retrieval design ·Staff
- Search filters for scores at least 80. Why does it return 9 and miss 100?Ranking, evidence and grounding ·Staff
- The assistant counted 12 invoices. Why did joining payment rows make it say 30?Query understanding and retrieval design ·Principal
- The exact error code ranks first in keyword search. Why did hybrid search bury it?Ranking, evidence and grounding ·Staff
- The user asked who is not eligible. Why did search return the eligibility policy?Query understanding and retrieval design ·Staff
- The document did not change. Why did its BM25 score fall after another tenant indexed files?Ranking, evidence and grounding ·Principal
- The customer typed the exact incident ID. Why did keyword search miss it?Query understanding and retrieval design ·Staff
- The retrieved child chunk is permitted. Why did its parent expose private text?Access, completeness and answerability ·Principal
- The RAG answer found the right table. Why did it assign the limit to the wrong plan?Query understanding and retrieval design ·Staff
- Search returned HTTP 200. Why can't the assistant say no incident exists?Scale and federated operations ·Principal
- The agent paged through every search result. Why did it skip and repeat documents?Scale and federated operations ·Principal
- Reranking improved answers. Why did search p99 collapse at peak traffic?Ranking, evidence and grounding ·Principal
Context Engineering
21 questions- What should enter the next 64K context of a coding agent?Context assembly and efficiency ·Principal
- A long conversation summary dropped 'never call delete'. Where should that rule live?Persistent state and memory ·Staff
- Why can a larger context window produce worse answers?Context assembly and efficiency ·Staff
- What should a support agent keep as state, context, or memory?Persistent state and memory ·Principal
- Cache the common prompt while tenant instructions changeContext assembly and efficiency ·Staff
- Correct tool result, wrong citation from a parallel branchTool and branch provenance ·Principal
- You have eighty passages and room for six. Which six belong in context?Context assembly and efficiency ·Staff
- The test log was truncated before the failureContext assembly and efficiency ·Staff
- The answer survives compaction. Does its evidence?Persistent state and memory ·Principal
- Two investigation agents disagree. What goes into the next context?Tool and branch provenance ·Principal
- The user corrected the agent's memory. Why does the old fact return next week?Persistent state and memory ·Principal
- A human takes over an agent case. What exactly should the handoff contain?Persistent state and memory ·Staff
- A long-lived agent resumes under a new system instruction. Which version governs its old plan?Persistent state and memory ·Principal
- A retrieved page's instruction was compressed into the agent's summary. Is it now trusted?Context assembly and efficiency ·Principal
- A user switches accounts. Why does the agent remember the previous customer's preference?Persistent state and memory ·Principal
- The coding agent read the new tests and the old API. Which repository did it actually see?Tool and branch provenance ·Principal
- A tool guessed the customer's address. Why did the agent remember it as a fact?Persistent state and memory ·Principal
- The memory expired yesterday. Why does the agent still act on it today?Persistent state and memory ·Principal
- The user said 'if I move to London.' Why does the agent think they live there?Persistent state and memory ·Staff
- Two customers share a name. Why did the agent merge their memories?Persistent state and memory ·Principal
- The agent replayed Monday's case. Why did it use a preference learned on Tuesday?Persistent state and memory ·Principal
Model and Inference Engineering
203 questions- KV cache is nearly full and p99 doubled. What would you investigate?KV cache and capacity ·Principal
- When can continuous batching improve throughput and make a short request slower?Scheduling and workload mix ·Staff
- Interactive and batch inference share a GPU fleet. Who gets capacity?Scheduling and workload mix ·Principal
- Quantization frees memory but hurts rare enterprise queries. Ship it?Quality and decoding trade-offs ·Staff
- The stream starts promptly, then stalls halfway. Where is time going?End-to-end serving and modalities ·Principal
- What is actually in the KV cache, and why does it fill up?KV cache and capacity ·Staff
- Route 50,000 requests a second across three unequal modelsRouting, tenancy and rollout ·Principal
- Speculative decoding wins a benchmark but raises cost per successful taskQuality and decoding trade-offs ·Staff
- Should prefill and decode run on different GPUs?Parallelism and placement ·Principal
- Eight GPUs are available. How should a model be split?Parallelism and placement ·Principal
- Prefix cache hits rise while first token p99 gets worseKV cache and capacity ·Staff
- A model advertises a million token window. Should we serve it?KV cache and capacity ·Principal
- One MoE expert group is hot while GPU averages look fineParallelism and placement ·Staff
- A tokens per second benchmark looks great. What did it measure?End-to-end serving and modalities ·Staff
- Swap KV to host memory or recompute it?KV cache and capacity ·Principal
- JSON always parses, but the agent still makes bad callsQuality and decoding trade-offs ·Staff
- Hundreds of tenants share one base model and many LoRA adaptersRouting, tenancy and rollout ·Principal
- Video requests hide behind a normal text QPS chartEnd-to-end serving and modalities ·Principal
- Scale to zero makes the first model request miss its deadlineScheduling and workload mix ·Staff
- The output length is unknown when the scheduler admits a requestScheduling and workload mix ·Principal
- One slow tensor-parallel rank holds every tokenParallelism and placement ·Principal
- A rolling model deploy runs out of GPU capacityRouting, tenancy and rollout ·Principal
- FP8 KV saves memory but changes long-context answersKV cache and capacity ·Principal
- A speculative decoder gets faster by sampling from the wrong modelQuality and decoding trade-offs ·Staff
- Prefill and decode disagree about the KV cacheKV cache and capacity ·Principal
- The GPUs are idle while first-token latency climbsEnd-to-end serving and modalities ·Staff
- Prompt cache hits rose. Why did cost per completed task double?KV cache and capacity ·Staff
- The tenant deleted its adapter. Why can the serving fleet still use it?Routing, tenancy and rollout ·Principal
- First token got faster. Why did the agent take longer to finish?Scheduling and workload mix ·Staff
- The router saves money until hard requests exhaust the premium-model quota. What now?Routing, tenancy and rollout ·Principal
- The model weights passed evaluation. Why did production tool calls break?End-to-end serving and modalities ·Staff
- The new model is ready. Why does the first unusual prompt take 30 seconds?End-to-end serving and modalities ·Staff
- An image request fits the text token budget. Why does the GPU run out of room?End-to-end serving and modalities ·Staff
- The benchmark uses common sequence shapes. Why does the production engine fall back on real traffic?End-to-end serving and modalities ·Staff
- GPU memory is free on average. Why does swapping base models crush p99?Parallelism and placement ·Staff
- CUDA graph replay is faster. Why does it return the wrong tokens for some batches?End-to-end serving and modalities ·Staff
- FlashAttention avoids the giant attention matrix. Why is 64K prefill still expensive?Transformer and model mechanics ·Staff
- Requests time out, but GPU memory stays high. Who still owns the KV blocks?KV cache and capacity ·Principal
- A stop string appears inside the answer. Why did the server return half a JSON object?Quality and decoding trade-offs ·Staff
- The 4-bit weights fit on one GPU. Why did tokens per second get worse?Quality and decoding trade-offs ·Staff
- The sliding-window KV cache fits. Why did evicting every layer break long prompts?KV cache and capacity ·Staff
- The same weights and greedy decoding give different tokens on eight GPUs. Is that a bug?End-to-end serving and modalities ·Staff
- The chat template already adds BOS. Why did the serving tokenizer add another?Quality and decoding trade-offs ·Staff
- The request has the same seed. Why did a new inference engine choose different tokens?End-to-end serving and modalities ·Staff
- Decode latency is great. Why do long prompts wait forever for their first token?Scheduling and workload mix ·Principal
- GPU utilization improved. Why is one long answer repeatedly starting over?KV cache and capacity ·Staff
- Only two MoE experts run per token. Why did moving experts across GPUs increase latency?Parallelism and placement ·Principal
- Two requests share a cached prefix. Why did cancelling one corrupt the other?KV cache and capacity ·Principal
- A rolling model upgrade kept the same conversation. Why did its next token go wrong?KV cache and capacity ·Principal
- Prefill and decode are on separate GPUs. Why is the first token slower?KV cache and capacity ·Principal
- The verifier rejected a draft token. Why does the next decode still remember it?KV cache and capacity ·Principal
- The prompt passed quota. Why did one long answer consume the decode budget?Scheduling and workload mix ·Principal
- Each GPU sampled from its own vocabulary shard. Why did top-p change the model?End-to-end serving and modalities ·Principal
- One sequence emitted EOS. Why did every request in the batch stop?Scheduling and workload mix ·Staff
- Tensor parallelism doubled, but KV cache per GPU did not halve. Why?KV cache and capacity ·Principal
- The LoRA adapter changed. Why did a cached prefix still act like the old one?KV cache and capacity ·Principal
- KV memory is allocated in blocks. Why can short requests waste so much of it?KV cache and capacity ·Staff
- Speculative decoding accepts most draft tokens. Why is latency unchanged?Quality and decoding trade-offs ·Staff
- The KV cache budget fits. Why does a long prefill still run out of GPU memory?KV cache and capacity ·Staff
- Two beam-search branches share a prefix. Why did one branch change the other's answer?KV cache and capacity ·Staff
- The LoRA adapter works in serving. Why did merging it into 4-bit weights change answers?End-to-end serving and modalities ·Principal
- The GPU has 24 GiB for KV cache. Can it hold six 32K conversations?KV cache and capacity ·Staff
- Decode uses little GPU compute. Why doesn't that mean the GPU is underloaded?KV cache and capacity ·Staff
- Eight GPUs make decode slower than four. Did tensor parallelism cross the node boundary?Parallelism and placement ·Principal
- FP8 serving passes the average eval. Why do a few prompts change the answer?Quality and decoding trade-offs ·Principal
- CUDA graphs sped up decode. Why is GPU memory still occupied with no requests?KV cache and capacity ·Staff
- The logits and seed stayed the same. Why did moving temperature before top-p change the answer?Quality and decoding trade-offs ·Staff
- QPS is flat. Why did the inference autoscaler run out of GPUs?Scheduling and workload mix ·Principal
- The draft and target tokenize the same sentence differently. Can speculative decoding stay exact?Quality and decoding trade-offs ·Principal
- Sequence packing doubled training throughput. Why can one document attend to another?Transformer and model mechanics ·Staff
- Grouped query attention cuts KV memory. Why might prefill still be expensive?Transformer and model mechanics ·Staff
- RoPE scaling accepts 128K tokens. Why does the model lose track near the end?Transformer and model mechanics ·Staff
- The MoE has plenty of total capacity. Why are tokens still dropped?Transformer and model mechanics ·Staff
- A 12-layer Transformer trains. Why does the 48-layer version fail after moving LayerNorm?Transformer and model mechanics ·Staff
- The attention map points to the right token. Why does the model still answer wrongly?Transformer and model mechanics ·Staff
- The attention scores use model width instead of head width. What changes?Transformer and model mechanics ·Staff
- Only two MoE experts run per token. Why does serving still need room for all the experts?Transformer and model mechanics ·Staff
- The same prompt answers differently alone and in a padded batch. Which position did its first token get?Transformer and model mechanics ·Staff
- The checkpoint loads after a SwiGLU kernel rewrite. Why did model quality fall?Transformer and model mechanics ·Staff
- A fused RMSNorm kernel is faster. Why do small activations change the answer?Transformer and model mechanics ·Staff
- Training loss is excellent. Why does generation fail when the attention mask sees the future?Transformer and model mechanics ·Staff
- The GQA cache has the right shape. Why do answers change after tensor parallel reshaping?Transformer and model mechanics ·Staff
- The tiled attention kernel returns NaNs. What did it forget when the row maximum changed?Transformer and model mechanics ·Staff
- A 96-layer model diverges after one residual-path rewrite. Why can a tiny scale matter?Transformer and model mechanics ·Staff
- The converted checkpoint loads. Why does RoPE change the model's answers?Transformer and model mechanics ·Principal
- One packed attention row has no valid keys. Why did the training loss become NaN?Transformer and model mechanics ·Staff
- The MoE picks the same two experts. Why did changing router normalization change the answer?Transformer and model mechanics ·Staff
- Every attention head is correct. Why is the tensor-parallel output wrong?Transformer and model mechanics ·Staff
- Each tensor-parallel GPU runs RMSNorm. Why does the full model disagree?Transformer and model mechanics ·Staff
- The prompt is identical. Why does chunked prefill change the model's answer?Transformer and model mechanics ·Staff
- The MHA checkpoint loads after converting to GQA. Why did answer quality fall?Transformer and model mechanics ·Principal
- The sliding-window attention kernel passes prefill. Why does one-token decode read the wrong keys?Transformer and model mechanics ·Staff
- The MoE router chose the right experts. Why did a token receive another token's output?Transformer and model mechanics ·Staff
- The model predicts four future tokens. Why can't the server just stream all four?Transformer and model mechanics ·Principal
- MLA compresses the KV cache. Why does it still store a positional key?Transformer and model mechanics ·Principal
- All MoE experts are equally busy. Why did answer quality get worse?Transformer and model mechanics ·Principal
- We replaced GELU with SwiGLU at the same FFN width. Why did the model get larger?Transformer and model mechanics ·Staff
- Prefill matches the reference. Why does decode rotate the cached keys twice?Transformer and model mechanics ·Staff
- We added tokens to the vocabulary. Why did old prompts change without using them?Transformer and model mechanics ·Staff
- We disabled half the attention heads. Why didn't inference become faster?Transformer and model mechanics ·Staff
- The JSON grammar allowed a token that ended generation mid-string. How?Transformer and model mechanics ·Principal
- The reward model score rises while people prefer the old assistant. What did optimization learn?Post-training and alignment ·Principal
- The preference pairs stayed the same. Why did a new DPO reference model change the policy?Post-training and alignment ·Staff
- The tool fine tune is fluent, but the model invents tool results. Which tokens were training targets?Post-training and alignment ·Staff
- A coding policy earns full reward by changing the tests. What was the verifier allowed to trust?Post-training and alignment ·Principal
- Annotators split 50/50 on refusal. What should the preference label teach the model?Post-training and alignment ·Principal
- RL workers are three checkpoints behind. Are their rollouts still usable?Post-training and alignment ·Principal
- The reward model prefers longer answers. Did humans prefer length or correctness?Post-training and alignment ·Staff
- The DPO pair has a chosen and rejected answer. Did both answer the same prompt?Post-training and alignment ·Staff
- The RL policy stays within its KL budget. Why does it break tool calls on rare prompts?Post-training and alignment ·Principal
- The DPO reference and policy score different prompt tokens. What does the margin mean?Post-training and alignment ·Principal
- The reward model ranks answers the same. Why did multiplying scores change the RL policy?Post-training and alignment ·Principal
- The code solution passes, but its verifier timed out. What reward should it get?Post-training and alignment ·Principal
- Every sampled solution gets zero reward. Why is the GRPO training run doing almost no learning?Post-training and alignment ·Principal
- The DPO pairs did not change. Why did length-normalizing log probabilities change the model?Post-training and alignment ·Principal
- The coding policy passes every training test. Why does it fail hidden cases?Post-training and alignment ·Principal
- The tool returned JSON. Why is the RL policy getting credit for those tokens?Post-training and alignment ·Principal
- The preference data was valid. Why did training reward the rejected answer?Post-training and alignment ·Principal
- One answer scores 9 and another scores 6. Can you compare them across prompts?Post-training and alignment ·Principal
- The SFT loss looks good. Why is it training on the user's words?Post-training and alignment ·Staff
- The RL run rewards short proofs. Did the model finish, or did generation stop it?Post-training and alignment ·Principal
- The reward model praises a safe answer. Did it read the last 2,000 tokens?Post-training and alignment ·Principal
- The reward differences are tiny. Why can GRPO give them a large update?Post-training and alignment ·Principal
- The DPO pair is valid in the dataset. Did truncation remove the preference?Post-training and alignment ·Staff
- The DPO pairs are unchanged. Why did raising beta stall learning?Post-training and alignment ·Staff
- The reward says correct. Why did the model's final answer give another number?Post-training and alignment ·Principal
- The DPO reference model is frozen. Why do its log probabilities change on the same pair?Post-training and alignment ·Staff
- Reviewers prefer A to B, B to C, and C to A. What should the reward model learn?Post-training and alignment ·Principal
- The agent got a good final reward. Did RL learn which earlier tool choice helped?Post-training and alignment ·Principal
- PPO clipped almost no tokens. Did we recompute the old log probabilities after updating?Post-training and alignment ·Staff
- The agent's old trajectories scored well. Can we replay them after the tool changed?Post-training and alignment ·Principal
- GRPO rewards the easy answer and punishes the best hard answer. Were the groups mixed?Post-training and alignment ·Staff
- The reward model wins on held-out pairs. Did it see the same prompts in training?Post-training and alignment ·Principal
- One GRPO candidate timed out. Why did the other candidates' advantages change?Post-training and alignment ·Principal
- DPO loss fell. Why did the chosen answer become less likely?Post-training and alignment ·Principal
- The vision model reads the number correctly. Why is its citation on the wrong part of the image?Multimodal evidence and voice ·Staff
- The transcript changes ‘do’ to ‘don't’ after the assistant starts speaking. What was committed?Multimodal evidence and voice ·Staff
- The video summary reverses two events. Which timestamps reached the model?Multimodal evidence and voice ·Staff
- The speaker says fifteen, but the slide shows fifty. What should the assistant report?Multimodal evidence and voice ·Staff
- The transcript says the chair approved it. Did the right speaker say yes?Multimodal evidence and voice ·Principal
- The video model says nothing happened. What if the event lasted one frame?Multimodal evidence and voice ·Principal
- The spoken words are right. Why does the video citation point ten seconds early?Multimodal evidence and voice ·Staff
- The call recording has two channels. Why did transcription lose the customer's reply?Multimodal evidence and voice ·Staff
- The call opens in English. Why did the transcript miss the Hindi complaint later?Multimodal evidence and voice ·Staff
- The screenshot contains the error code. Why did the vision assistant read a different one?Multimodal evidence and voice ·Staff
- The voice agent heard 'approve the transfer.' Did it miss the next two words?Multimodal evidence and voice ·Principal
- The voice agent interrupted itself. Who did its microphone actually hear?Multimodal evidence and voice ·Staff
- The model found the right moment in a video. Why is its timestamp wrong?Multimodal evidence and voice ·Staff
- The caller said fifteen. Why did the voice agent send fifty?Multimodal evidence and voice ·Principal
- The call lasted ten seconds. Why did the speech pipeline hear only three?Multimodal evidence and voice ·Staff
- Both audio channels contain the caller. Why does the mono transcript lose their voice?Multimodal evidence and voice ·Staff
- The meeting transcript captured the louder speaker. Where did the interruption go?Multimodal evidence and voice ·Staff
- The speech model recognizes names in recordings. Why does the live agent lose their first syllable?Multimodal evidence and voice ·Staff
- The video found the right badge. Why did it say the wrong person wore it?Multimodal evidence and voice ·Staff
- The screen shows $1.50. Why did the voice agent say $150?Multimodal evidence and voice ·Staff
- The voice agent logged 'confirmation spoken.' Did the caller hear it?Multimodal evidence and voice ·Staff
- Audio and video align at the start. Why is the live agent seconds off after twenty minutes?Multimodal evidence and voice ·Staff
- The PDF page is high resolution. Why did the vision model miss its small footnote?Multimodal evidence and voice ·Staff
- The caller interrupted the voice agent. Why did ASR hear silence?Multimodal evidence and voice ·Staff
- Rank 7 dies while the training checkpoint is being written. Which checkpoint can you resume?Training and model lifecycle ·Staff
- Training loss drops after adding a corpus. Why do rare customers get worse?Training and model lifecycle ·Principal
- A training job restarts with half as many GPUs. Can its optimizer state be loaded safely?Training and model lifecycle ·Staff
- The tokenizer changes midway through training, but the weight files still load. What has been mixed?Training and model lifecycle ·Staff
- A teacher model generated a million training examples. Which ones are safe to keep?Training and model lifecycle ·Principal
- One training rank has nonfinite gradients while the global loss looks normal. Can the step commit?Training execution and recovery ·Principal
- The data mixture says 20% code. Why did code dominate the training tokens?Training data and objectives ·Principal
- More pipeline stages fit the model. Why did training tokens per second fall?Training execution and recovery ·Staff
- We doubled gradient accumulation after losing GPUs. Why did the training recipe change?Training execution and recovery ·Staff
- FP8 training is faster until one outlier batch breaks the run. Which scale was stale?Training execution and recovery ·Staff
- The checkpoint resumes at the right step. Why does training repeat old samples?Training execution and recovery ·Principal
- The gradient norm is clipped on every microbatch. Why is the accumulated norm still too large?Training execution and recovery ·Staff
- Every training rank has a different shard. Why are its random augmentations identical?Training data and objectives ·Staff
- The optimizer skipped a step. Why did the learning rate schedule advance?Training execution and recovery ·Staff
- The new LoRA adapter has gradients. Why do its weights never change?Training execution and recovery ·Staff
- Ranks process different token counts. Why does averaging their losses change the objective?Training data and objectives ·Principal
- The checkpoint loads and loss starts normally. Why do tied embeddings drift apart after training resumes?Training execution and recovery ·Staff
- The optimizer checkpoint loads. Did its moments attach to the right parameters?Training execution and recovery ·Principal
- The optimizer update is valid. Why are normalization scales shrinking throughout training?Training execution and recovery ·Staff
- Activation checkpointing fits the model. Why do its dropout gradients no longer match?Training execution and recovery ·Staff
- The all-reduce was launched. Why did one rank step before it finished?Training execution and recovery ·Staff
- The fine-tune learned the answers. Why won't it stop generating?Training data and objectives ·Staff
- Every FSDP rank clipped its gradients. Why was the global update still too large?Training execution and recovery ·Staff
- The crawl has unique URLs. Why does pretraining keep seeing the same article?Training data and objectives ·Principal
- The vocabulary is sharded across GPUs. Why is each rank's softmax the wrong loss?Training execution and recovery ·Staff
- The effective batch stayed at 256. Why did the retriever lose most of its negatives?Training data and objectives ·Principal
- The pipeline updated a stage before backward reached it. Which weights made the gradient?Training execution and recovery ·Principal
- The Adam moments restored. Why is the next update wrong?Training execution and recovery ·Staff
- You quadrupled LoRA rank and kept alpha fixed. Why did training change?Training execution and recovery ·Staff
- The retriever got more in-batch negatives. Why did recall for valid answers fall?Training data and objectives ·Principal
- Activation checkpointing saved GPU memory. Why did FSDP training slow down?Training execution and recovery ·Staff
- The base model is frozen. Why do its weights drift during LoRA training?Training execution and recovery ·Staff
- Both training runs saw the same number of tokens. Why did the 32K run cost much more?Training execution and recovery ·Staff
- The data loader got faster. Why did the model see an easier training set?Training data and objectives ·Principal
- A 20-token answer and a 2,000-token answer get equal weight. What did the loss train?Training data and objectives ·Staff
- The learning rate decays at the planned step. Why has the model seen half the intended tokens?Training execution and recovery ·Staff
- The gradient norm is capped at one. Why does the mixed-precision run barely learn?Training execution and recovery ·Staff
- FSDP trains with bf16. Can we export its current compute weights as a resumable checkpoint?Training execution and recovery ·Staff
- Validation loss rose after switching tokenizers. Did the new model get worse?Training data and objectives ·Staff
- Every training rank got equal work. Why did the epoch repeat some examples?Training data and objectives ·Staff
- Sequence parallelism reduced memory. Why does 64K training still OOM in attention?Training execution and recovery ·Principal
- We enabled shuffle. Why did training still see one data source for hours?Training data and objectives ·Principal
- The 7B weights take 14 GB. Why won't Adam training fit on a 40 GB GPU?Training execution and recovery ·Staff
- The fine-tune loss falls. Why is the model learning the token after next?Training data and objectives ·Staff
Agent Architecture
38 questions- The refund succeeded but the agent crashed. What happens next?Runtime and effect recovery ·Principal
- What does cancelling a two-day agent run actually mean?Durable runs and workflow change ·Staff
- Several coding agents edit one repository. What can you safely merge?Multi-agent coordination ·Principal
- The planner says a timed-out tool succeededRuntime and effect recovery ·Staff
- An investigation subagent wants to issue a creditAuthority and approval ·Principal
- Retry or compensate a read, an email, and a payment?Runtime and effect recovery ·Staff
- How do you deploy new agent workflow code while 200,000 runs are waiting?Durable runs and workflow change ·Principal
- The multi-agent demo is better, but production pays for itMulti-agent coordination ·Staff
- The agent keeps searching but makes no progressRuntime and effect recovery ·Staff
- A tool returns 202. Has the agent completed the action?Runtime and effect recovery ·Principal
- The human approved a preview. The world changed before commit.Authority and approval ·Principal
- The coding agent made CI green by changing the test. Do you merge it?Coding agent change validation ·Principal
- The browser agent clicked Approve after the page changed. What did it approve?Authority and approval ·Staff
- An agent changes an API used by twelve repositories. In what order can it ship?Coding agent change validation ·Principal
- The provider fails after a tool result. What can the next model actually continue?Runtime and effect recovery ·Staff
- The coding agent fixed a CVE. What else changed in the dependency tree?Coding agent change validation ·Staff
- The agent can run for months. What happens when its workflow history keeps growing?Durable runs and workflow change ·Principal
- The coding agent passes CI but changes what a 409 means. Do you merge it?Coding agent change validation ·Staff
- The user changes their mind while the agent's tool call is in flight. What stops?Runtime and effect recovery ·Staff
- The coding agent's patch passes tests, but the production feature flag takes a different branch. What did it test?Coding agent change validation ·Principal
- Two subagents update the same customer case note. Which changes survive?Multi-agent coordination ·Staff
- The agent checked 100 invoices and found no duplicates. Were there only 100 invoices?Runtime and effect recovery ·Principal
- Each subagent stayed under its budget. Why did the parent run spend ten times more?Multi-agent coordination ·Principal
- Two investigation agents agree. What if both read the same stale document?Multi-agent coordination ·Principal
- The coding subagent says tests passed. Which patch did it test?Multi-agent coordination ·Principal
- Two agents both found the last available unit. Can both promise it to customers?Multi-agent coordination ·Principal
- The parent agent can read one case. Why can its subagent search the whole tenant?Authority and approval ·Principal
- The batch tool returned 200. Why did the agent say all 20 invites succeeded?Runtime and effect recovery ·Staff
- The agent rolled back a failed branch. Why does the next model call remember it?Multi-agent coordination ·Principal
- Two invoice tool calls finished out of order. Why were their results swapped?Runtime and effect recovery ·Staff
- An agent resumes after approval. Why does its saved tool call now fail?Authority and approval ·Principal
- The agent sent the right account ID. Why did the tool edit a different account?Runtime and effect recovery ·Principal
- The agent approved an action on account 42. Why did it hit a new account 42?Runtime and effect recovery ·Principal
- The agent reused the idempotency key. Why did the payment API reject its retry?Runtime and effect recovery ·Principal
- An old tool callback arrived after the agent restarted. Which run gets the result?Runtime and effect recovery ·Principal
- The workflow replayed after a crash. Why did the agent send the email again?Durable runs and workflow change ·Principal
- The planner discarded a branch. Why did that branch's tool still charge the card?Multi-agent coordination ·Principal
- A reviewer approved a $50 refund. Why did the resumed agent issue $500?Authority and approval ·Principal
Distributed Reliability
45 questions- A model provider returns 429s at peak. What happens over the first hour?Admission, fairness and overload ·Principal
- How can three retries turn a small model outage into a platform incident?Retry, idempotency and ownership ·Staff
- Two regions resume the same agent step. How do you stop a duplicate write?Retry, idempotency and ownership ·Principal
- RAG median is flat, but p99 doubles for three percent of trafficLatency and load evidence ·Staff
- One tenant sends ordinary QPS and huge prompts. How is that fair?Admission, fairness and overload ·Principal
- Would you hedge a model call after 800 milliseconds?Latency and load evidence ·Staff
- Workers are healthy, but new jobs wait forty minutes. What now?Admission, fairness and overload ·Principal
- The circuit opens, then every instance probes at onceQueue recovery and health restoration ·Staff
- One malformed agent job keeps coming back from the queueQueue recovery and health restoration ·Staff
- The load test p99 looks fine during a five-second stallLatency and load evidence ·Principal
- The browser reconnects to an agent stream. Does the action restart?Retry, idempotency and ownership ·Principal
- A model stream fails after the user has seen half an answerLatency and load evidence ·Principal
- One agent run opens a hundred tool calls at onceAdmission, fairness and overload ·Principal
- Restoring yesterday's search index resurrects deleted documentsQueue recovery and health restoration ·Principal
- A global rate limit cannot be guessed independently in each regionAdmission, fairness and overload ·Principal
- The queue lease expires while an agent tool call is runningRetry, idempotency and ownership ·Principal
- One tenant's quota errors open a breaker for everyoneAdmission, fairness and overload ·Staff
- The caller says stop while the voice agent is submitting a refundRetry, idempotency and ownership ·Staff
- The binary rolled back. Can the old version read the data the new version already wrote?Queue recovery and health restoration ·Principal
- A million-row model batch fails at 70 percent. What exactly can you publish?Queue recovery and health restoration ·Staff
- A zone fails during a model rollout. Why do both fleets miss their deadlines?Admission, fairness and overload ·Staff
- Database and workflow store are restored to different minutes. Which agent actions still exist?Queue recovery and health restoration ·Principal
- The user closed the chat. Why are GPUs still finishing their answer?Queue recovery and health restoration ·Staff
- The model provider recovered. Why did 50,000 agents take it down again?Retry, idempotency and ownership ·Principal
- The index event arrived, but the document transaction rolled back. What should search believe?Retry, idempotency and ownership ·Principal
- The second region is healthy. Why does failover still take the AI service down?Admission, fairness and overload ·Principal
- The agent had a 20-second deadline. Why is its tool still running a minute later?Queue recovery and health restoration ·Principal
- Failover restored the agent. Why did it miss the approval revocation?Queue recovery and health restoration ·Principal
- A failed agent job was replayed next week. Does its old approval still count?Queue recovery and health restoration ·Principal
- The queue has a deduplication ID. Why did a retry six minutes later create two jobs?Retry, idempotency and ownership ·Principal
- Kafka lag is zero. Why is the search index missing yesterday's updates?Retry, idempotency and ownership ·Principal
- Hedged model calls improved p95. Why did peak-hour p99 collapse?Latency and load evidence ·Principal
- The client reads tokens slowly. Where does the generated text pile up?Latency and load evidence ·Principal
- We gave each agent its own queue. Why did one slow tool still stall everyone?Latency and load evidence ·Principal
- The model emits a token in 200 ms. Why does the browser see nothing for five seconds?Latency and load evidence ·Staff
- The gRPC channel is healthy. Why do short tool calls wait behind long streams?Latency and load evidence ·Staff
- Kafka lag grows while consumers are alive. Why do partitions keep moving?Retry, idempotency and ownership ·Staff
- Kafka says exactly once. Why did the search index apply the document twice?Retry, idempotency and ownership ·Principal
- Kafka acknowledged the event. Why did it vanish after the leader failed?Retry, idempotency and ownership ·Principal
- The outbox row says sent. Why did no index event reach Kafka?Retry, idempotency and ownership ·Principal
- The policy update succeeded. Why did the next read return the old rule?Queue recovery and health restoration ·Principal
- The Kafka transaction aborted. Why did the agent still act on its event?Retry, idempotency and ownership ·Principal
- R plus W is greater than N. Why did a read still miss the latest write?Retry, idempotency and ownership ·Principal
- The Postgres standby acknowledged the write. Why did failover lose it?Retry, idempotency and ownership ·Principal
- Both databases prepared the transaction. Why are writes now blocked?Retry, idempotency and ownership ·Principal
Evaluation and Quality
51 questions- Retrieval recall improved, so why did the answers get worse?Metric validity, judges and human signals ·Principal
- How would you evaluate a docs assistant that should sometimes refuse to answer?Answerability, labels and datasets ·Staff
- The new model wins preferences but breaks a critical tool workflow. Who ships it?Experiments and causal evidence ·Principal
- The citation is relevant, but it does not support the number. Is the answer correct?Metric validity, judges and human signals ·Staff
- How do you evaluate an agent when several trajectories are correct?Agent task and trajectory outcomes ·Principal
- Your synthetic eval grew a hundredfold. Why might you trust it less?Statistical reliability ·Staff
- Online win rate fell, but offline eval is flat. What do you believe?Experiments and causal evidence ·Principal
- Should an LLM judge score be the SLO for an assistant?Metric validity, judges and human signals ·Staff
- The judge score rises while human reviewers find less supportMetric validity, judges and human signals ·Principal
- The coding agent wins at pass@8. Does that help one user?Agent task and trajectory outcomes ·Staff
- How do you replay an agent eval when tools change the world?Agent task and trajectory outcomes ·Principal
- What is the unit of an online experiment for a multi-turn agent?Agent task and trajectory outcomes ·Principal
- Can old router logs prove the new router saves money?Experiments and causal evidence ·Principal
- What happens to gold labels when the documents change?Answerability, labels and datasets ·Principal
- Thumbs-up rate improves while audited answer quality fallsMetric validity, judges and human signals ·Staff
- Zero failures in 200 tests does not establish a rare-event guaranteeStatistical reliability ·Principal
- Choose an answer-or-abstain threshold without gaming accuracyAnswerability, labels and datasets ·Principal
- The judge changes its winner when answer order swapsMetric validity, judges and human signals ·Staff
- A coding benchmark score jumps, but new private tasks do notExperiments and causal evidence ·Principal
- Two reviewers disagree whether the question is answerableAnswerability, labels and datasets ·Principal
- The support agent closes more tickets. Did it solve more problems?Agent task and trajectory outcomes ·Principal
- The chart says growth, but the axes were cropped. What can the assistant claim?Metric validity, judges and human signals ·Staff
- The model improves globally but harms one small customer cohort. Do you stop the rollout?Experiments and causal evidence ·Principal
- The answer streamed a claim before the evidence check finished. Can a late citation fix it?Answerability, labels and datasets ·Principal
- The grader says the agent improved, but customers still call back. Which result wins?Agent task and trajectory outcomes ·Principal
- You tuned against the holdout for six months. Is it still a test set?Answerability, labels and datasets ·Principal
- The model A/B test shares one GPU pool. Can you trust its latency result?Experiments and causal evidence ·Principal
- The safety update refuses more users but blocks more harmful requests. Do you roll it back?Metric validity, judges and human signals ·Principal
- The model looks accurate on labeled cases. What happened to cases still waiting for an outcome?Answerability, labels and datasets ·Staff
- Reviewers agree more after calibration. Why do they all miss the same failure?Metric validity, judges and human signals ·Principal
- The safety set passes, then attackers adapt after launch. How should the evaluation change?Metric validity, judges and human signals ·Principal
- An A/B test assigns users separately, but their generated artifacts are shared. Is the effect identifiable?Experiments and causal evidence ·Principal
- The agent succeeds after humans rescue its hardest cases. Whose success are we measuring?Agent task and trajectory outcomes ·Principal
- The safety filter catches 99% of harmful requests. Why is its review queue mostly harmless?Statistical reliability ·Staff
- Your eval has 10,000 cases from 50 customers. How many independent wins did you measure?Metric validity, judges and human signals ·Principal
- The assistant won the A/B test on Tuesday. Why did the win disappear on Friday?Experiments and causal evidence ·Principal
- The agent's final answer is correct. Why should this run still fail evaluation?Agent task and trajectory outcomes ·Principal
- Human-reviewed accuracy rose. Did the assistant improve, or did review routing change?Answerability, labels and datasets ·Principal
- The judge prefers the new assistant. Does it prefer its writing style?Metric validity, judges and human signals ·Principal
- The new model wins on completed eval cases. Where did its timed-out cases go?Metric validity, judges and human signals ·Principal
- The new RAG stack wins. Was it the model, the retriever, or their combination?Metric validity, judges and human signals ·Principal
- The agent solves 90% with ten tries. What will one production run solve?Agent task and trajectory outcomes ·Principal
- The same agent patch passes once and fails twice. Did the model change?Agent task and trajectory outcomes ·Principal
- The new assistant follows the updated policy. Why does the eval call it wrong?Answerability, labels and datasets ·Principal
- Overall answer quality rose. Why did the highest-risk users get worse?Metric validity, judges and human signals ·Principal
- The assistant is 90% accurate when it answers. Why are its 95% confidence answers risky?Statistical reliability ·Principal
- The RAG model scores 95% with evidence. Why does the live assistant miss those answers?Answerability, labels and datasets ·Principal
- Safety reviewers found 11% violations. Why isn't the production violation rate 11%?Statistical reliability ·Principal
- The defense blocks 99% of attacks. What happens when one user can try a hundred times?Metric validity, judges and human signals ·Principal
- Model B has lower overall p99. Why is it slower in every prompt-length bucket?Statistical reliability ·Principal
- The coding agent passes each task alone. Why does the benchmark fail when tasks run together?Agent task and trajectory outcomes ·Principal
Security, Governance and Platform
52 questions- The PDF looks redacted. Why did RAG reveal the hidden text?Tenant and data boundaries ·Staff
- What belongs in a common AI platform for 200 teams?Platform contracts and execution isolation ·Principal
- A relevant incident note tells the agent to upload private logs. Where do you stop it?Untrusted input and tools ·Staff
- Who holds the credentials when an agent can issue refunds?Action authorization and policy ·Principal
- The tool result validates, but its notes try to command the agentUntrusted input and tools ·Staff
- A semantic cache served tenant A's private answer to tenant B. What now?Tenant and data boundaries ·Principal
- What must a model gateway expose about provider differences?Platform contracts and execution isolation ·Staff
- A new model changes tool-call format. How do hundreds of apps migrate?Platform contracts and execution isolation ·Principal
- Should one logging pipeline keep every retrieved passage?Audit and observability ·Staff
- A coding agent runs repository tests that can read its secretsPlatform contracts and execution isolation ·Principal
- Can an incident operator let an agent bypass a failed policy service?Action authorization and policy ·Principal
- A new tool field can bypass the old authorization checkAction authorization and policy ·Staff
- A tool catalog changes while an agent is runningUntrusted input and tools ·Principal
- Model failover crosses the customer's data boundaryTenant and data boundaries ·Principal
- Tenant traces become a shared evaluation datasetTenant and data boundaries ·Principal
- The assistant's answer loads an external URLUntrusted input and tools ·Staff
- Audit an agent's actions without dumping private conversationsAudit and observability ·Principal
- An agent splits a capped action into smaller callsAction authorization and policy ·Principal
- Can an MCP server pass the user's token to its downstream API?Platform contracts and execution isolation ·Principal
- The support agent knows the customer's last invoice. Can it reset MFA?Action authorization and policy ·Principal
- The tool URL was approved. May the agent follow its redirect to a new host?Untrusted input and tools ·Staff
- The database is failing. Should the incident agent run its proposed migration?Action authorization and policy ·Principal
- The batch passed policy at 9 AM. Can it publish results after the rule changes at noon?Action authorization and policy ·Principal
- The audit log says approved. Can you prove what the reviewer saw?Audit and observability ·Principal
- The training data license expires tomorrow. Can the model stay online?Tenant and data boundaries ·Principal
- The agent can read private files and call public search. Which bytes may leave?Tenant and data boundaries ·Principal
- The agent's task was approved yesterday. The employee loses access today. Does the scheduled action run?Action authorization and policy ·Principal
- Half the fleet enforces the new policy and half the old one. When can rollout be called complete?Action authorization and policy ·Principal
- A tool gives the agent a signed file URL. Should that URL enter long-term memory?Platform contracts and execution isolation ·Principal
- The agent wrote a CSV report. Why did opening it run a formula from a customer name?Untrusted input and tools ·Principal
- The SQL tool uses parameters. Why can the agent still choose the wrong table?Untrusted input and tools ·Principal
- The agent used a valid tool. Why did a retrieved tenant ID expose another customer's data?Tenant and data boundaries ·Principal
- No private document was returned. Why did search still reveal that it exists?Tenant and data boundaries ·Principal
- The agent's tool is read-only. Why did its GET request send a reminder?Untrusted input and tools ·Principal
- The coding agent can read only the repo. How did a symlink expose a private key?Platform contracts and execution isolation ·Principal
- The coding agent only opened a PR. Why did CI run its code with secrets?Platform contracts and execution isolation ·Principal
- The prefix cache never returns another tenant's text. What can latency still reveal?Tenant and data boundaries ·Principal
- Two MCP servers expose a tool called `search`. Which one did the agent approve?Untrusted input and tools ·Principal
- The agent's URL passed the allowlist. Why did its fetch reach an internal service?Action authorization and policy ·Principal
- A retrieved transcript says `system`. Why did the assistant treat it as an instruction?Platform contracts and execution isolation ·Principal
- The agent has a valid user token. Why can the second tool reject it?Action authorization and policy ·Principal
- The identity provider rotated its signing key. Why did every agent tool call fail?Platform contracts and execution isolation ·Principal
- The webhook JSON is valid. Why does signature verification fail after a refactor?Platform contracts and execution isolation ·Staff
- Two agent runs refresh the same OAuth credential. Why does one lose access?Action authorization and policy ·Principal
- User A was allowed to read a case. Why did the cached decision let user B in?Action authorization and policy ·Principal
- The API authenticated the user. Why can a queued agent action run as someone else?Platform contracts and execution isolation ·Principal
- The document was revoked. Why is its generated preview still public?Tenant and data boundaries ·Principal
- The webhook signature is valid. Why did the agent grant access twice?Platform contracts and execution isolation ·Principal
- The agent has a valid refund token. Why can it refund another customer's order?Action authorization and policy ·Principal
- The user's access was revoked. Why is their open answer stream still revealing the document?Action authorization and policy ·Principal
- The encrypted dataset snapshot is intact. Why can't we decrypt it after key rotation?Tenant and data boundaries ·Principal
- The model artifact has a valid signature. Why did staging weights reach production?Platform contracts and execution isolation ·Principal