[{"data":1,"prerenderedAt":24},["ShallowReactive",2],{"page:blog-id:en:311":3},{"id":4,"order":5,"postTopCategoryName":6,"componentName":7,"dateInserted":8,"publishDate":9,"dateChanged":9,"ratings":10,"ratingAmount":10,"ratingAverage":10,"ratingOwn":10,"files":11,"videos":12,"documents":13,"postTexts":14,"distance":10,"effectiveStatus":21,"effectiveAudience":10,"listed":23},311,10003,"blog","insight-contradictions-not-extraction","0001-01-01T00:00:00","2026-08-25T14:55:00",0,[],[],[],[15],{"id":10,"locale":16,"header":17,"contentShort":18,"content":17,"contentHTML":19,"slug":20,"postId":4,"languageId":21,"automaticTranslation":22},"en","LLM extraction isn’t the hard part – modelling the knowledge in context is","A language model reads values from documents remarkably well. The hard part is modelling what was read in context – and when two documents say different things, that is where it is decided whether the system can be trusted.","\u003Cp>The first demo of document extraction with a language model is almost always impressive. You put in a contract and get parties, term, amounts and notice periods back as clean data. After an afternoon the problem looks solved.\u003C\u002Fp>\u003Cp>It is not. Extraction is the easy part. The hard part is modelling what was read in context: what belongs together, what replaces what, what applies from when – and what contradicts what?\u003C\u002Fp>\u003Ch2>The real world contradicts itself\u003C\u002Fh2>\u003Cp>In real document collections, documents say different things. The contract states one term, the amendment another. The invoice shows an amount that does not match the order. A valuation from 2021 and one from 2023 arrive at different figures. Sometimes a single document even contradicts itself – the text says something different from the table.\u003C\u002Fp>\u003Cp>A language model resolves such contradictions silently. It picks an answer, usually a plausible one, and presents it with the same confidence as everything else. That is the real risk: not the obvious error, but the credible wrong answer.\u003C\u002Fp>\u003Ch2>Embeddings help you find, not decide\u003C\u002Fh2>\u003Cp>Retrieval with embeddings finds relevant passages. It does not answer whether two passages mean the same thing, complement each other or contradict each other. That requires a model of the domain: what is an amendment, what does it replace, which document takes precedence, from when does something apply?\u003C\u002Fp>\u003Ch2>How a reliable system handles it\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Every value with provenance:\u003C\u002Fstrong> page, section, document. No source, no value.\u003C\u002Fli>\u003Cli>\u003Cstrong>Rules before trust:\u003C\u002Fstrong> deterministic checks for everything that can be checked – totals, deadlines, dependencies, precedence.\u003C\u002Fli>\u003Cli>\u003Cstrong>Conflicts as objects:\u003C\u002Fstrong> a contradiction is not resolved but stored – with both sources.\u003C\u002Fli>\u003Cli>\u003Cstrong>People decide, the system remembers:\u003C\u002Fstrong> a domain expert resolves the conflict, and the decision becomes part of the data.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That sounds like more work than “ask the model, store the answer”. It is the difference between a demo and a system whose data you can base decisions on.\u003C\u002Fp>","insight-modelling-knowledge-in-context",1,false,true,1791147997938]