kb-bench

Which platform preserves figures in contracts and specifications?

Short answer

fastgpt — FastGPT. It retained every planted fact in the corpus with exact chunk visibility, and unlike the platform that scored higher on retrieval, you can audit what it stored.

What this use case actually needs

How the candidates measured up

Platform Documents ingested Parsing Retrieval Peak memory Ready
fastgpt 36/36 100% 74% 3,478.8 MB 239 s
dify 36/36 100% 68% 2,619.7 MB 39 s
ragflow 36/36 99% 73% 9,959.6 MB 511 s
anythingllm 36/36 98% 83% 542.2 MB 146 s

Every figure comes from the same run against the same documents, so these are directly comparable. The full per-document and per-query rows are in the raw data.

Why FastGPT wins this one

Two platforms retained 100% of the facts planted across the corpus — FastGPT and Dify. FastGPT is the recommendation because it pairs that with the higher retrieval accuracy of the two, 74% against 68%.

FastGPTDifyAnythingLLM
Parsing fidelity100%100%98% (approximate)
Chunk visibilityexactexactnone
Retrieval accuracy74%68%83%
Peak memory3,479 MB2,620 MB542 MB

Why not AnythingLLM, which scored higher

It retrieved better — 83% against 74% — and on most workloads that would settle it. It does not settle this one, because of what “98% parsing” means for that platform.

AnythingLLM exposes no API that lists a document’s stored chunks. Its parsing figure had to be approximated by sweeping retrieval and collecting whatever came back, which can only under-report and, more importantly, cannot be audited. For a contract or a specification, “we think the figure survived” is a different proposition from “here is the stored text containing it”.

FastGPT and Dify both let you read back exactly what was stored. If your documents are the kind where a wrong digit is worse than a missing answer, that visibility is the feature you are buying.

What “100% parsing” actually measured

Facts were planted in real public documents — version strings, specification identifiers, configuration keys, precise quantities, and distinctive phrases from running prose — each verified to occur exactly once in its document. The score is the share of those found in the chunks the platform itself stored. No model judged it; it is a string search you can repeat against the published data.

Both FastGPT and Dify found all 180 of them.

The caveat that used to be here

An earlier version of this page recommended FastGPT with a warning that it dropped three of four large filings on a time limit. On this corpus it ingested all thirty-six documents, so that warning no longer applies to the workload described here — but it does still apply if your contracts run to hundreds of pages. Both runs are published; the ranking page shows the two side by side.

Where this recommendation stops applying