Why “chat with your codebase” disappoints, and what a real engineering knowledge layer needs.

Every team eventually builds some version of the same thing: dump the design docs, README files, API specs, and code chunks into a vector store, put a chat box on top, and call it an engineering knowledge assistant.

For the first few days, it feels useful. It can explain what a service does, find a config file, summarize an old design doc, and answer onboarding questions well enough that people start believing the tool has understood the system.

Then someone asks a real engineering question.

Is warehouseId mapped from the inventory response to the public order response?

That question has a precise answer. It is not a vibes question and it is not a documentation question. It lives in a mapper class, a response model, an OpenAPI schema, maybe a feature flag, probably a test, and maybe a column somewhere downstream. The design doc, if it mentions the field at all, probably describes what someone intended two years ago.

That is the whole problem in one line.

Docs explain intent. Code explains truth.

Even that sentence is incomplete, because runtime explains what actually happened, config explains what is enabled, the live database explains what exists now, tests explain what someone cared enough to protect, and ownership metadata explains who should be asked when the answer is not obvious.

The truth is not in one place, and that is why most “chat with your codebase” systems disappoint. They are built like search boxes, but engineers ask system questions.

Chat with docs answers a question nobody is asking#

Document RAG works when the answer was written down. That is the boundary, not a criticism, just the boundary. Feed it onboarding docs and it can explain why a service exists. Feed it architecture notes and it can summarize old decisions. Feed it runbooks and it can help with known operational steps.

But engineers usually ask a different class of questions.

Where is this field validated? Is the rule checked at the API edge, or three services deep? Which services break if I remove this column? What columns does this table have right now, not what a migration file suggests it should have? Is this behavior controlled by code, config, or a flag that is only on in one region? Which consumer reads this Kafka event? Which tests actually cover this path?

These are not document questions. They are system questions.

The answer is implemented, not written. It lives across code, schema, config, migrations, tests, traces, ownership metadata, and sometimes stale documentation that now disagrees with reality.

A better model does not fix that by itself. A stronger model may explore better, phrase the answer better, and even guess correctly more often, but it still cannot prove a fact it cannot retrieve. So the first correction is not technical. It is about what you are actually building.

You are not building a better chatbot over documents. You are building a system that finds the truth wherever it actually lives.

Existing tools solve pieces, not the joined answer#

It is too easy to pretend nothing exists today, but that is not true. Teams already have useful pieces of this stack. Exact search can find text. Code intelligence can follow symbols and references. Coding assistants can work through local context. Service catalogs can tell you ownership. Enterprise search can find docs. Observability can show runtime behavior. Schema systems can show contracts.

The problem is what happens after that. The engineer still has to open five tools, compare half-trusted facts, check freshness, check permissions, and reconstruct the path by hand.

All of that is useful, but it still leaves the engineer stitching things together manually.

The missing piece is not another search box. The missing piece is the joined answer.

An engineer does not want ten tabs open across code search, wiki, OpenAPI, migration history, feature flag UI, trace explorer, service catalog, and test files. They want the path: this field came from here, it was mapped here, it appears in this public schema, it is gated by this flag, it is persisted to this column, it is covered by this test, this doc says something different, and the whole answer is based on main indexed eight minutes ago with production traces from the last seven days.

That is not search. That is an engineering knowledge layer.

The model is not the knowledge#

The rule is simple, and most systems violate it: the language model is never the source of truth.

It can plan retrieval, ask follow up questions to tools, assemble evidence, and explain the answer in human language, but it does not know whether warehouseId is mapped. Your codebase knows that. Your schema knows that. Your tests may know that. Your traces may confirm that the path is used in production. The model is only the thing reading it back.

Engineering knowledge architecture split into a deterministic truth plane, a fuzzy recall plane, and a reasoning plane where the LLM plans retrieval and writes cited answers.

The model is not the knowledge. It plans and explains, while truth comes from deterministic indexes, schemas, code, config, tests, and runtime evidence.

A useful architecture splits the system into three planes.

The truth plane is deterministic. It contains the knowledge graph, symbol index, lexical index, schema metadata, migrations, ownership data, and runtime evidence. This plane answers “what is true right now” with facts you can point at: these are the three call sites, this endpoint returns this response model, this column is NOT NULL, this migration added this field, this test covers this mapper, and this trace observed service A calling service B.

The recall plane is fuzzy. It is a vector store over docs, code chunks, comments, tickets, and long form engineering prose. It answers “where should I look?” That is useful, but it should not be allowed to assert truth. Embeddings can tell you that a chunk looks related to your question, they cannot prove that service A calls service B, that a field is persisted, or that a flag gates a behavior. Similar text is not evidence.

The reasoning plane is the model. It classifies the question, chooses tools, decomposes the query, follows evidence, and writes the final answer. When evidence is missing, it should abstain.

The hard rule is that a claim the truth plane cannot cite is dropped. Not weakened, not hidden behind “likely,” just dropped. That one rule changes the system. Without it, you have a fluent assistant that sometimes invents. With it, you have something engineers can start trusting.

Truth is scoped#

One thing the word “truth” hides is scope. There is no single universal truth in a large engineering system.

There is branch truth, deployed truth, production truth, region-specific truth, flag-on truth, flag-off truth, migration-file truth, live-catalog truth, static-analysis truth, and runtime-observed truth. A sampled trace proves that a path happened, but it does not prove that no other path exists. A migration proves that someone intended to change the database, but it does not prove the column exists in the live environment. A config file proves the default, but it does not prove the flag is enabled for one region, one tenant, or one experiment cohort.

So the system should not answer “true” without saying true where, from which source, and as of when.

A weak answer says:

warehouseId is returned.

A better answer says:

warehouseId is returned by the public order response on main, based on the OpenAPI schema indexed 12 minutes ago. The mapper exists in code. The path is gated by enableWarehouseInOrderResponse. I found production traces for the order endpoint in the last seven days, but traces do not prove this specific field was populated.

That is the level of honesty engineers need. Not perfect truth, but scoped truth.

The graph is the product#

Vector search finds plausibly relevant chunks. The graph encodes relationships. That is the difference between a system that can guess and a system that can trace.

Knowledge graph tracing warehouseId from an inventory response through a mapper, public order response, feature flag, database column, test, and missing design documentation.

A real engineering answer is a cited path across code, schema, config, tests, docs, and runtime evidence, not a paragraph generated from similar chunks.

Model the entities engineers actually care about: service, endpoint, request model, response model, field, mapper, validator, table, column, migration, config, feature flag, event topic, test, owner, design doc, and runtime trace.

Then connect them with edges that mean something. A service exposes an endpoint. An endpoint returns a model. A model has a field. A mapper maps one field to another. A field is stored as a column. A flag gates a mapper. A test covers a validator. A doc documents an endpoint. A trace confirms a runtime call.

Now the warehouseId question is no longer a search problem. It is a graph traversal.

Start from the upstream field, follow the mapper, check the downstream model, check the schema, check whether a feature flag gates the path, check whether a test covers it, and check whether the table has a matching column if the field is persisted. Then return the path with a citation at every hop.

A real answer should look more like this:

Question:
Is warehouseId returned in the public order response?

Answer:
Yes, but only when enableWarehouseInOrderResponse is enabled.

Evidence path:
1. InventoryResponse.warehouseId exists in inventory-api/.../InventoryResponse.java
2. InventoryToOrderMapper maps InventoryResponse.warehouseId to OrderResponse.warehouseId
3. OrderResponse.warehouseId exists in the public OpenAPI schema for GET /orders/{id}
4. The mapper branch is gated by enableWarehouseInOrderResponse
5. OrderResponseMapperTest covers the flag-on case
6. No test found for the flag-off case
7. No design doc mentioning this field was found

Confidence:
Medium-high. Symbol, mapper, and schema were resolved. Flag was found. Missing negative-path test.

Freshness:
Code indexed from main 8 minutes ago.
OpenAPI schema indexed 12 minutes ago.
Flag config indexed 1 hour ago.

That is very different from “based on the code, it seems warehouseId is probably returned.” The first answer gives you a path. The second gives you prose.

Engineering tools need paths.

The stale doc case matters too. If the doc says the field is not returned but the code maps it, the system should say that. That is not confusion. That is drift.

One implementation detail matters a lot here: keep static and runtime evidence separate. “Static analysis says A calls B” and “production traces confirm A calls B” are not the same fact, and both are useful. If static analysis says A calls B and production never observes it, that is interesting. Maybe the path is dead, maybe traffic is low, maybe tracing missed it, or maybe the call only happens in one region. The system should preserve that distinction instead of pretending there is one clean answer.

The entity linker is where the pain lives#

This is where most demos quietly stop. Finding a chunk is easy enough. The messy part is connecting the same business concept across code, API schema, database column, config, generated mapper, and test without pretending that a name match is proof.

It has to understand that warehouseId in Java, warehouseId in OpenAPI, warehouse_id in a database column, and maybe wh_id in an event schema are related but not automatically the same thing. A name match is not proof.

The linker should start with exact identity where it can: class fields, generated schema fields, method references, symbol definitions, imports, annotations, and generated mapper code. Then it can fall back to constrained heuristics like compatible type, same owning flow, nearby mapper method, matching serialization annotation, matching OpenAPI path, matching migration history, and matching test fixture.

But the edge should carry how it was derived. mapsTo from a mapper method is stronger than “same name after snake_case conversion.” A schema edge is stronger than a doc mention. A runtime observed call is different from a static call. An inferred link should not become a proven link just because the model liked the story.

This is where a real system separates itself from a demo. Everything above the graph is a chatbot. The graph, the linker, and the evidence model are the product.

Trust is the product, not a feature#

Staff engineers do not abandon these tools because one feature is missing. They abandon them the first time the tool says something confidently wrong.

After that, every answer gets double checked manually. Once that happens, the tool is dead. Maybe people still use it for summaries and maybe juniors still ask it onboarding questions, but it is no longer part of serious engineering work.

Trust cannot be a layer added later. It has to be designed into the answer path. Every claim needs a source: file at commit and line, schema path, migration id, config key, trace id, test name, owner source, indexed timestamp. No citation, no claim.

Confidence should also come from evidence, not from the model. Asking a model “how confident are you?” is not a trust system, it is another generated answer.

Confidence should be derived from retrieval signals: was the symbol resolved exactly or guessed by name, was the schema read from the live contract or from old docs, was a test found, do code and docs agree, was the static call confirmed by runtime traces, is the source fresh, and did the query touch a repo the caller can actually access?

Same evidence should produce the same confidence every time. Otherwise confidence is decoration.

Abstention also has to be a normal answer. “I could not find evidence that this field is mapped” is a good answer when the system really could not prove it. A system that never abstains is lying somewhere.

Freshness belongs in the answer too. “Based on main indexed six minutes ago, schema introspected one hour ago, traces from the last seven days” is not boring metadata, it is what lets an engineer decide whether to trust the result.

Permissions are part of grounding as well. If the caller cannot access a repo, doc, schema, or trace, that evidence should not exist for the query. Filtering after synthesis is too late because the model should never see unauthorized evidence in the first place. An index that answers from private repos the user cannot access is not an assistant. It is a data leak with a chat interface.

The eval is not whether the answer sounds right#

A system like this cannot be evaluated with “the answer looked useful.” That is how demos pass and production tools rot.

The eval needs to check whether the system found the right evidence. For a field mapping question, the expected answer is not just prose. It is a path: upstream field, mapper, downstream field, schema, flag, and test if one exists.

For 100 field mapping questions, measure whether the system returned the exact field path, whether every hop had a citation, whether inferred edges were marked as inferred, and whether the system abstained when the mapper could not be proven. The score is not answer similarity. The score is path correctness.

For an impact question, the expected answer is a dependency slice. For an ownership question, the expected answer is the owner source, not the model’s guess from naming conventions. For a schema question, the expected answer is the current schema source, not a migration file that may or may not have shipped.

Track the boring numbers: citation precision, path correctness, stale answer rate, unsafe answer rate, permission filtering failures, abstain-when-should-answer, answer-when-should-abstain, and cost per answered question.

The dangerous failure is not abstaining too often. The dangerous failure is answering when the system did not have evidence.

Corrections should become eval cases. If an engineer says “this answer is wrong because the mapper is behind a flag,” that should not disappear into a thumbs down metric. It should become a regression test, otherwise the system keeps making the same mistake with better wording.

Adoption is a cost problem, not only an accuracy problem#

At some point in the review, someone will ask the obvious question: “Why do we need another assistant when engineers already use coding agents?”

You will not win that argument by saying your answers are better. You win by changing the economics.

The knowledge layer should make engineering questions cheap to answer and should also make existing coding assistants cheaper to use. That is not a pricing decision, it is an architecture decision.

Cost aware query flow where structured questions use zero-LLM graph templates, cached answers avoid repeated work, local models handle common synthesis, and frontier models are reserved for hard cases.

Adoption depends on cost as much as accuracy. Most questions should be answered by deterministic tools, caches, or cheaper models before reaching a frontier model.

First, route structured questions to zero-LLM answers. “What columns are in this table?”, “Who owns this service?”, “What endpoints does this service expose?”, and “Which services consume this topic?” are graph or schema queries rendered through templates. No model needed. The result is deterministic, cheap, and auditable.

Not magically always correct. The index can still be stale and the extractor can still miss something, but the answer is correctable because every fact points back to evidence.

Second, cache in layers, keyed to the index version. Precompute cards for services, endpoints, tables, and common flows. Cache graph traversals because they are deterministic for a given commit and permission scope. Cache final answers only when the normalized question, caller permission scope, and freshness vector match. Never cache across ACL scopes, and never let an answer outlive the evidence it cited.

Provider prompt caching helps too, but only for the stable prefix: system prompt, tool definitions, output schema, maybe fixed instructions. It does not remove the cost of fresh per query evidence. Self hosted inference is not free either, it is only low marginal cost when the GPU capacity is already paid for and available.

Third, route models by difficulty. A frontier model should not be required for every question. If the evidence set is tight and cited, the model’s job is smaller because it is not discovering the whole codebase from scratch, it is summarizing known evidence and refusing to go beyond it. That is a job a cheaper or local model can often do well.

Use the stronger model for low confidence, multi hop, ambiguous, or high impact queries. Use cheaper models for the bulk. Use no model for structured lookups. Retrieval quality substitutes for model size more often than people admit.

But the saving has to be measured inside your repo. Count how many files a coding agent opens before a task, how many input tokens it spends exploring, how often a context pack avoids that exploration, how many cache hits you get, and how many questions never reach a model. Without those numbers, “this reduces coding agent cost” is only a guess. With those numbers, it becomes something you can defend in an adoption review.

Context packs are the bridge to coding agents#

This is where the knowledge layer becomes more than a chat interface. It starts acting as context infrastructure for coding agents.

Coding agents spend a lot of input tokens exploring the repo. They search, open files, follow imports, inspect tests, read docs, and slowly build the context they need. Your knowledge layer already has much of that context, so expose it as a context pack, not just as a chatbot answer.

For a task, return the relevant files, why each file matters, the endpoints involved, the models and fields touched, the flags, the tests, the owning service, and the dependency slice.

A coding agent should be able to ask:

get_context_for_task("remove warehouseId from public order response")

and receive a small pack:

Likely files:
- OrderResponse.java: contains public field
- InventoryToOrderMapper.java: maps upstream warehouseId
- order-openapi.yaml: public contract
- order_response_mapper_test.java: covers flag-on case
- feature-flags.yaml: enableWarehouseInOrderResponse
- OrderServiceController.java: endpoint returning response

Dependency slice:
- inventory-service provides upstream field
- order-service exposes public response
- reporting-service consumes OrderCreated event with warehouse_id

Risks:
- Public API contract change
- Flag-off behavior not tested
- Design doc does not mention this field

Now the coding agent starts from a smaller, better context. You are not competing with coding agents, you are reducing how much repo exploration they have to do. That is easier to sell than “we built another assistant.”

Compile once, keep current#

Plain retrieval has another problem: it rederives the same cross service understanding on every query.

If ten engineers ask about the order fulfillment flow, the system should not rediscover that flow ten times. A better shape is to compile expensive synthesis once into durable, cited knowledge pages, then keep those pages current.

Think of them as generated engineering pages, not replacement docs. A page for an end to end flow. A page for a cross service field map. A page for an impact map. A page for service ownership and dependencies.

The important part is that these pages are built from evidence and carry citations. They are not free form wiki prose. They are compiled views over the graph, symbols, schemas, docs, and traces.

Three operations matter. Ingest updates the affected pages when code changes. Query reads the compiled page first, then drills into the graph only when needed. Lint runs on a schedule and looks for stale claims, broken citations, contradictions, orphaned pages, and doccode drift.

The caution is simple: do not compile everything. A generated page that mirrors one small greppable file is negative value because it is bigger than the source and now the agent may read both. Compile pages only where synthesis actually compresses scattered knowledge: cross service flows, field mappings, ownership maps, impact maps, API behavior across services. Answer file level questions from live retrieval and deterministic tools.

The compiled layer is only useful if it compounds understanding. Otherwise it becomes another stale documentation system.

Start where the system can be trusted#

Do not start with the magic demo.

Everyone wants to ask, “If I remove this field, what breaks?” That is a valuable question, but it is also where the system is most likely to lie early. The graph is incomplete, the linker is immature, runtime evidence is partial, and the model will happily fill the missing parts if you let it.

Start with questions that are boring but easy to verify.

Who owns this service? What endpoints does it expose? What request and response models are in the contract? What columns exist in the current schema? Which topics does the service produce and consume? Which docs mention this endpoint?

These questions prove the basics: retrieval works, citations resolve, permissions are enforced, freshness is visible, and abstention is allowed. Once that is working, add traversal: validation location, service calls, consumer discovery, doc code drift, and static call plus runtime confirmation.

Only then add the hard stuff: field level dataflow, impact analysis, reflection heavy mappings, dynamic dispatch, and config dependent behavior. These answers should be evidence bounded from day one. If the system inferred a link, say it inferred the link. If it could not prove the mapping, say that. If traces confirm one path but not another, say that too.

This architecture also has a cost. You are maintaining connectors, extractors, indexes, graph edges, freshness jobs, permission filters, evals, caches, and drift checks. For a small codebase, good docs, grep, code search, and ownership discipline may be enough.

The graph starts earning its keep when the same questions repeatedly cross service boundaries, schemas, flags, teams, migrations, contracts, and runtime behavior. If the answer is in one file, open the file. If the answer needs five systems and three people to reconstruct every time, build the knowledge layer.

Before adding a bigger model, ask what evidence is missing. If the question needs a call edge, add symbol or runtime evidence. If it needs schema truth, add live catalog introspection. If it needs behavior truth, add config and traces. If it needs ownership, add catalog metadata. If it needs a cross service field path, improve the entity linker.

Model quality is the last mile. The missing evidence layer is usually the real problem.

Where this breaks in real codebases#

Field level dataflow is where the system will struggle first. Reflection based mappers, builders, renamed fields, generated code, MapStruct, ModelMapper, and framework magic all make static analysis messy. The system should trace what it can prove statically, use schema links where they exist, use generated mapper metadata if the framework exposes it, and treat tests or runtime traces as supporting evidence. If the mapping cannot be proven, the answer should say that instead of turning a guess into a fact.

Call graphs have the same problem. Dependency injection, dynamic dispatch, reflection, async queues, and framework annotations break simple static call graphs. A Java service does not always call another service through a clean method reference you can follow, so the graph needs separate edge types: static call, inferred call, runtime observed call, event production, and event consumption. These should not be collapsed into one generic “calls” edge.

Schema truth is also not as clean as it looks. Migration files tell you what should have happened. Live catalog introspection tells you what exists now. Both are useful, but they are not the same source. If they disagree, the answer should surface the disagreement instead of silently picking one.

Freshness is another place where this system can quietly become dangerous. A stale index gives confident wrong answers, which is worse than no answer. Incremental indexing helps latency, but a full reconcile is still needed for correctness because webhooks get missed, parsers have bugs, files move, and old edges need to be deleted.

Permissions are not a side detail either. If the system indexes everything and filters loosely at the end, it will leak. Query time ACL enforcement is not enterprise polish. It is the difference between a useful internal tool and a data exposure incident.

And the model will try to fill gaps. That is what models do. The architecture has to make unsupported claims difficult to ship through strict evidence sets, citation checks, abstention, eval cases, and audit logs.

None of this is the demo. But this is the part that decides whether engineers still trust the tool after the first few wrong looking answers.

The information already exists#

Your codebase already knows whether the field is mapped. It knows which service owns the endpoint, which table has the column, which migration added it, which flag gates the behavior, which test covers the path, and which service called which service last week.

The knowledge is already there. It is just scattered across systems that were never designed to answer questions together.

That is the system worth building: not a smarter chatbot, but a truth layer for engineering work, fresh enough to trust, cheap enough to ask often, permission aware enough to use inside a real company, and honest enough to say when it does not know.

The model was never the knowledge. It was only ever the thing that reads it back to you.