kb-bench

Self-hosted AI knowledge base platforms, ranked

Six platforms installed on a clean host from their own published images, pointed at the same embedding model, fed the same 36 public documents and asked the same 93 questions.

The ranking in one line

AnythingLLM scores highest — highest retrieval accuracy of anything measured, on a tenth of the memory. Every platform ingested the whole corpus this time, so the ranking turns on retrieval accuracy and what each one costs to run. Which is right still depends on the job — and more than that, it depends on your documents: we ran this twice on different corpora and the order inverted.

Ranking

# Platform Score Verdict Best for
1 AnythingLLM 1.16.0 9.4 Highest retrieval accuracy of anything measured, on a tenth of the memory Most workloads on a small host
2 FastGPT 4.15.8 8.1 Perfect parsing fidelity and strong retrieval, at a moderate memory cost Documents where exactness matters
3 Dify 1.16.1 8.1 Perfect parsing, but the lowest retrieval accuracy of the five Agents and workflows, not search
4 RAGFlow 0.26.4 7.5 Complete and accurate, and the only permissive licence — at ten times the memory Resale and redistribution
5 Open WebUI 0.11.0 7.4 Fast and light, but the weakest parsing observed by a wide margin Chat first, retrieval second
— Flowise 3.0.6 n/a Current release does not start, and the API cannot be automated Not recommended

Scored out of 10 on the four dimensions this run measured: retrieval accuracy 35%, ingestion completeness 25%, parsing fidelity 25%, memory footprint 15%. Parsing is multiplied by the share of the corpus a platform actually ingested, so a high rate earned by skipping the hard documents does not bank. Memory is scored on a log scale because it behaves like a threshold rather than a linear cost. Flowise is unranked because its current release does not start.

What this score does not include: workflow orchestration, integration breadth and licence terms are part of the methodology but were assessed by reading rather than measured in this run, so folding them into a composite would put a number on something that has none behind it. They are covered on each platform's page instead — and they matter: the platform ranked third here is the only one whose licence permits reselling a hosted deployment, and the one ranked highest treats retrieval as a feature of a chat product rather than as the product. The weighting is a judgement and the raw rows are published so you can apply your own.

The measurements behind the ranking

Platform Ingested Parsing Retrieval Chunks Peak memory Ready Corpus time
anythingllm 36/36 98% 83% 1,900 542.2 MB 146 s 17 s/doc
fastgpt 36/36 100% 74% 1,198 3,478.8 MB 239 s 30 s/doc
dify 36/36 100% 68% 4,263 2,619.7 MB 39 s 20 s/doc
ragflow 36/36 99% 73% 477 9,959.6 MB 511 s 20 s/doc
openwebui 36/36 46% 73% 5,060 980.7 MB 18 s 25 s/doc

A tilde marks a figure the platform's API cannot report exactly — two expose no way to list stored chunks, so their parsing is approximated by exhaustive retrieval and can only under-report. An asterisk marks a rate covering only part of the corpus: the highest parsing figure here belongs to the platform that ingested the fewest documents.

Which one for which job

No platform won every dimension, and the one at the top is the wrong choice for at least two of these:

Internal documentation

Mixed formats, recall matters most — Dify

Contracts and statements

A figure must survive exactly — FastGPT

A small VPS

Under a gigabyte, small documents — AnythingLLM

Provisioned by CI

Everything scriptable, no browser — FastGPT

The finding that matters most: your documents decide the answer

We ran this benchmark twice on the same platforms, the same host and the same embedding model, changing only the corpus. The first run used large documents — SEC filings of one to eleven megabytes. The second used the shape a real internal knowledge base has: thirty-six documents with a median size of fifteen kilobytes.

The ranking did not shift. It inverted.

Platform Large documents: ingested Large: retrieval Small documents: ingested Small: retrieval
AnythingLLM6/1141%36/3683%
FastGPT8/1141%36/3674%
RAGFlow11/1145%36/3673%
Open WebUI11/1162%36/3673%
Dify11/1166%36/3668%

Dify led the first run and comes last on retrieval in this one. AnythingLLM could not finish half the first corpus and tops this one. Nothing about the software changed between the two runs — only the documents.

So treat any single benchmark number, including ours, as conditional. If your documents look like large filings, the first column is the relevant one and this page's ranking is misleading for you. We publish the corpus, both manifests and the harness precisely so you can check which case you are in — or re-run it on your own documents, which is the only comparison that fully answers the question.

A claim from our first run has not survived this one, and is worth retracting explicitly: we reported that finer chunking predicted better retrieval, because the two lined up almost perfectly across five platforms. On this corpus they do not line up at all — the platform storing the most chunks and the one storing the fewest both scored 73%. That relationship held for documents large enough to bury an answer inside one chunk, and does not generalise beyond them.

How this was measured

Deployed, not read about

Every figure came from a running install, not a requirements page.

No model judges anything

Facts planted in real documents, then searched for as exact strings.

Failures published

Documents that would not ingest are in the data, not filtered out.

Reproducible

The harness is open source and rebuilds the same corpus.