Skip to content
sDEN

Manufacturing knowledge

RAG for engineering documents

Naive RAG splits documents into text chunks and loses drawings, tables and revisions. What structured extraction and a governed knowledge graph change.

A stack of documents flowing through a processing unit into one structured record
On this page

Retrieval-augmented generation works well on prose. Split a policy manual into passages, embed them, retrieve the closest ones for a question and let a model summarize them with citations. Point the same pipeline at an engineering archive and the answers degrade in ways that are hard to see: plausible, well written, and wrong about the revision, the part or the value that matters.

The reason is structural. Engineering knowledge is not mostly in sentences. It sits in title blocks, revision tables, bills of materials, dimension callouts, notes columns and scanned tables, and in the relationships between documents: this drawing belongs to that assembly, this revision supersedes that one, this part appears in those inspection reports. A text chunk captures almost none of that.

This piece explains where naive RAG breaks on engineering documents, what structured extraction recovers, and why a governed knowledge graph is the layer that lets private AI answer engineering questions with sources a person can check.

Where it breaks

Four ways chunk-and-embed fails on engineering content

Each failure looks like a model problem. Each is a data problem that happens before the model sees anything.

Drawings come first. A drawing is a spatial document: the title block, the revision table, the bill of materials and the notes are separate regions, each with its own meaning. Text extraction flattens them into a stream where a part number, a material callout and a revision letter end up side by side without the labels that gave them meaning. Embedding that stream produces a vector that is close to many questions and correct for few of them. Scanned drawings are worse, because the text layer may be missing or noisy, and the retrieval step has nothing reliable to match.

Tables are the second failure. A bill of materials or an inspection table is a grid where each value depends on its row and its column header. Fixed-size chunking cuts tables across rows, separates values from headers and splits a single item across two chunks. The model then receives a fragment, such as a quantity and a material without the part they belong to, and fills the gap with something reasonable. Reasonable is not the same as correct when the question is which material was specified for a given part.

Revisions are the third failure, and the most dangerous. An engineering archive usually holds several revisions of the same drawing or specification, often with nearly identical text. Vector similarity has no notion of which one is current, which was superseded or which was released for production. It retrieves whichever passage is closest to the question, and an answer built on a superseded revision can look exactly like an answer built on the released one. The fourth failure follows from the other three: the answer cites a chunk, but a chunk is not a source location an engineer can verify against a drawing sheet, a table row or a revision entry.

What changes

Structured extraction: entities, attributes, revisions and source location

The alternative is to read engineering documents the way engineers do, region by region, and to keep what was read as records rather than as passages.

Structured extraction starts from a schema: the entities and attributes the questions depend on. For a drawing register, that means drawing number, title, revision, release status, parts, materials and notes. For a bill of materials, it means parent assembly, item, part number, quantity and material. The extraction pipeline reads each region with the method suited to it, combining OCR for scans, deterministic parsing where layouts are regular and private open-weight models where content needs interpretation, and writes the result as typed records.

Every record keeps two things a chunk never has. The first is its source location: the document, the page and the region or table row it came from, so any value can be checked against the original. The second is its review status: whether the value was verified automatically or confirmed by a person. Human review is configured where accuracy matters, for example on values that feed release or quality decisions, and is not spent on fields where an error would be harmless.

Revisions become explicit data rather than a property buried in text. Each drawing revision is its own record, linked to the one it supersedes, with its release status attached. A question about the current material for a part can then be answered from the released revision by design, and a question about what changed between revisions can be answered by comparing records instead of asking a model to spot differences between two similar passages.

Why a graph

The governed knowledge graph becomes the retrieval layer

Extracted records become far more useful once they are linked to each other and to the rest of the engineering estate.

A knowledge graph links entities across documents and systems: the part on this drawing is the same part in that bill of materials, the item on that certificate and the subject of those inspection reports. Questions that span documents, such as which assemblies use a part that changed in its latest revision, become traversals of explicit relationships rather than hopes that the right chunks land in the same context window.

Governance is what makes the graph safe to put behind an assistant. Lineage records where each fact came from and how it was produced. Classifications and existing permissions travel with the data, so a user or an agent querying the graph never receives more than that user may see. The graph is monitored and refreshed as documents change, so a new revision updates the records instead of sitting next to the old one in an index.

Text retrieval does not disappear. Notes, procedures and reports still contain prose worth retrieving by meaning, and passage search remains the right tool for it. The difference is that passages now hang off entities in the graph, so retrieval can start from the right part and revision and then fetch the text that explains it. Private AI then answers from a context that is both structured and sourced.

How sDEN approaches it

Three principles for engineering retrieval

sDEN Foundation and sDEN Solutions apply the same discipline to engineering archives: read the document properly, keep the source, and govern what AI can reach.

Three principles for engineering retrievalExplore the perimeter

Within this scope

Read the structure, not just the text

Title blocks, revision tables, bills of materials and notes are extracted as regions with their own meaning, using OCR, private open-weight models and deterministic methods chosen per content and task.

Within this scope

Every value keeps its source

Extracted records carry their document, page and region, plus a review status, so an engineer can check any field against the drawing or table it came from.

Within this scope

Revisions and permissions are data

Revisions are linked records with a release status, and existing permissions follow the data into the graph, so answers draw on the right revision and never on content a user may not see.

Illustrative map. Scope, responsibilities and controls are agreed for each engagement.

What good looks like

Answers an engineer can check in one step

The test is simple: for any answer, can a person open the cited drawing, table row or revision entry and confirm it?

When engineering documents are extracted into structured, source-linked records, a question about a part, a material or a revision returns specific values with their origin attached. An engineer does not need to trust the assistant; they can verify it. That changes how AI is used in engineering, from an occasional search shortcut to something people rely on for routine questions, because the cost of checking is low.

Document control benefits independently of any assistant. Drawing registers can be built from title blocks and revision tables, superseded revisions and missing documents can be flagged, and part references can be linked across projects and sites. The same records feed PLM, analytics and search, so the work done for retrieval is not trapped inside one AI application.

Most importantly, the knowledge stays yours. Extraction runs inside an environment you control, the graph belongs to your organization, and models remain replaceable. When a better model arrives, it reads the same governed records.

Questions

Manufacturing knowledge, answered.

Why does standard RAG struggle with engineering drawings?

Standard RAG converts documents to text and splits that text into chunks. Drawings are spatial documents whose meaning depends on regions such as the title block, the revision table and the bill of materials. Flattening them to text separates values from their labels, and scanned drawings may have no reliable text at all, so retrieval has little accurate material to work from.

Can a larger context window solve the problem?

A larger context window lets a model read more text, but it does not tell the model which revision is current, which table row a value belongs to or which documents describe the same part. Those are properties of the data, not of the model. Structured extraction and explicit links address them directly and keep the context small and relevant.

How are multiple revisions of the same document handled?

Each revision is extracted as its own record, linked to the revision it supersedes and carrying its release status. Questions about the current state are answered from released revisions, and questions about change history compare records rather than similar passages of text.

Does this approach replace vector search?

No. Passage retrieval stays useful for notes, procedures and reports written in prose. The change is that passages are linked to entities in a governed knowledge graph, so retrieval can start from the right part and revision and then bring in the text that explains it.

Do we have to migrate our documents?

No. sDEN connects to existing repositories, keeps existing permissions and controls in place, and delivers structured data to the systems your teams already use. Documents stay where they are, and the graph enriches them where they live.

Ready to build AI you can govern, audit and own?

Discuss your data estate, institutional priorities and infrastructure requirements. Together, we can define the next decisions and the scope of a suitable engagement.