
The first demo of document extraction with a language model is almost always impressive. You put in a contract and get parties, term, amounts and notice periods back as clean data. After an afternoon the problem looks solved.
It is not. Extraction is the easy part. The hard part is modelling what was read in context: what belongs together, what replaces what, what applies from when – and what contradicts what?
The real world contradicts itself
In real document collections, documents say different things. The contract states one term, the amendment another. The invoice shows an amount that does not match the order. A valuation from 2021 and one from 2023 arrive at different figures. Sometimes a single document even contradicts itself – the text says something different from the table.
A language model resolves such contradictions silently. It picks an answer, usually a plausible one, and presents it with the same confidence as everything else. That is the real risk: not the obvious error, but the credible wrong answer.
Embeddings help you find, not decide
Retrieval with embeddings finds relevant passages. It does not answer whether two passages mean the same thing, complement each other or contradict each other. That requires a model of the domain: what is an amendment, what does it replace, which document takes precedence, from when does something apply?
How a reliable system handles it
- Every value with provenance: page, section, document. No source, no value.
- Rules before trust: deterministic checks for everything that can be checked – totals, deadlines, dependencies, precedence.
- Conflicts as objects: a contradiction is not resolved but stored – with both sources.
- People decide, the system remembers: a domain expert resolves the conflict, and the decision becomes part of the data.
That sounds like more work than “ask the model, store the answer”. It is the difference between a demo and a system whose data you can base decisions on.