kb-bench

Which platform runs on a small VPS?

Short answer

anythingllm — AnythingLLM, without a trade this time. It holds 542 MB in one container and also had the highest retrieval accuracy of anything measured on this corpus.

What this use case actually needs

How the candidates measured up

Platform Documents ingested Parsing Retrieval Peak memory Ready
anythingllm 36/36 98% 83% 542.2 MB 146 s
openwebui 36/36 46% 73% 980.7 MB 18 s
dify 36/36 100% 68% 2,619.7 MB 39 s
fastgpt 36/36 100% 74% 3,478.8 MB 239 s

Every figure comes from the same run against the same documents, so these are directly comparable. The full per-document and per-query rows are in the raw data.

The trade that used to exist here has gone

On our first corpus this page had to hedge: AnythingLLM won on footprint but ingested six documents out of eleven, so choosing it meant accepting an incomplete knowledge base. On a corpus of ordinary internal documents it ingested all thirty-six and scored highest on retrieval as well.

So on this workload it is simply the recommendation, on both dimensions at once.

PlatformContainersPeak memoryIngestedRetrieval
AnythingLLM1542 MB36/3683%
Open WebUI1981 MB36/3673%
Dify152,620 MB36/3668%
FastGPT133,479 MB36/3674%
RAGFlow59,960 MB36/3673%

The spread is eighteen-fold, and container count does not predict it

RAGFlow runs five containers and needs 9,960 MB. FastGPT runs thirteen and needs 3,479 MB. Dify runs fifteen and needs 2,620 MB — the most containers and the third-lowest memory.

The reason is what is inside them: RAGFlow bundles Elasticsearch, which holds several gigabytes before a document arrives. A requirements page quoting either number alone tells you very little about whether a platform fits your host.

Open WebUI is the alternative worth weighing

981 MB in one container, ready to serve in eighteen seconds — the fastest start of anything measured — and 73% retrieval. It costs roughly twice AnythingLLM’s memory for ten points less accuracy.

Its weakness is parsing: 46%, by far the lowest observed. That figure is approximate, since it exposes no chunk listing and can only under-report, but the gap to the others is too wide to be entirely an artefact. If your documents carry precise values you need back intact, this is the wrong end of the trade.

What a small host cannot have

None of the platforms that expose exact chunk visibility — FastGPT, Dify, RAGFlow — fit comfortably in under two gigabytes. If auditability of stored text is a requirement, a small VPS is not the right host, and that is a real constraint rather than a preference.

Where this recommendation stops applying