kb-bench

Which platform is best for searching internal documentation?

Short answer

anythingllm — AnythingLLM. It answered 83% of the questions correctly — the highest of anything measured — while holding a tenth of the memory of the heaviest platform and a fifth of the next best performer.

What this use case actually needs

How the candidates measured up

Platform Documents ingested Parsing Retrieval Peak memory Ready
anythingllm 36/36 98% 83% 542.2 MB 146 s
fastgpt 36/36 100% 74% 3,478.8 MB 239 s
openwebui 36/36 46% 73% 980.7 MB 18 s
ragflow 36/36 99% 73% 9,959.6 MB 511 s
dify 36/36 100% 68% 2,619.7 MB 39 s

Every figure comes from the same run against the same documents, so these are directly comparable. The full per-document and per-query rows are in the raw data.

Why AnythingLLM wins this one

On a corpus shaped like a real internal knowledge base — thirty-six documents, median size fifteen kilobytes, a mix of technical docs, policy pages, specifications and forms — it came first on the dimension that matters most and cost the least to run:

AnythingLLMBest of the rest
Retrieval accuracy83%74% (FastGPT)
Documents ingested36/3636/36 (all five)
Peak memory542 MB981 MB (Open WebUI)
Ready to serve146 s18 s (Open WebUI)
Containers11 (Open WebUI)

Nine points of retrieval accuracy over the next best, on a fifth of the memory of the platform in second place overall.

The reversal you should know about

This is the same platform that, on our first corpus, ingested six documents out of eleven and scored 41% — the worst result on that board. Nothing about AnythingLLM changed. The documents did.

Its weakness is large files: it embeds a whole document synchronously inside the upload call, and on megabyte-scale filings that call does not finish. Give it the small documents an internal wiki actually contains and the weakness never appears, while its retrieval quality does.

That is why this page is scoped to internal documentation rather than to knowledge bases in general, and why the ranking page carries the same warning.

What it costs you

No visibility into chunking. There is no route that lists what it stored, so its 98% parsing figure is approximated through exhaustive retrieval and can only under-report. If you need to audit how a document was split, FastGPT and Dify both expose that and AnythingLLM does not.

An open default. A fresh deployment runs with authentication disabled and will hand a full-access API key to an unauthenticated caller. That is fine on a laptop and unacceptable on anything reachable. Put authentication in front of it before you expose the port.

If those matter more than the nine points

FastGPT is the safer institutional choice: 74% retrieval, 100% parsing fidelity with exact chunk visibility, and a fully scriptable API. It costs 3.5 GB against AnythingLLM’s 542 MB and takes longer to start.

Dify parses just as accurately but retrieved the least of the five here at 68%, which makes it hard to recommend for search specifically — its strengths are elsewhere.

Where this recommendation stops applying