kb-bench

Open WebUI 0.11.0

A chat front-end that grew retrieval features, deployed and measured against the same corpus as every other platform here.

In one paragraph

Deployed from its published image, Open WebUI runs 1 container and held 980.7 MB at peak during ingestion. It retained 46% of the facts planted in the corpus and answered 73% of the retrieval questions correctly.

Measured

MeasurementResultHow
Time to first served request 18 s Includes image pull, from a clean host
Containers 1 Running after the stack settles
Idle memory 664.3 MB Sum across containers, 60s after ready
Peak memory during ingestion 980.7 MB Sampled every 5s across the whole corpus
Documents ingested 36/36 Failures counted, not excluded
Chunks stored 5,060 Across the whole corpus
Parsing fidelity 46% Approximate: no chunk-listing API, measured by exhaustive retrieval
Retrieval accuracy 73% Answer present in top-5 context, 93 questions
Retrieval latency 271.7 / 305.8 ms p50 / p95, CPU embedding
Licence Open WebUI License (BSD-3-Clause with branding conditions) Permits offering it as a service

What it does well

Where it falls short

Choose it when

Choose something else when

Deployment notes

The upstream compose builds from a Dockerfile and bundles its own Ollama. For a comparison where every platform must use the same embedding model, that is not usable as published, so the compose here pins the released image and points at the shared embedding service instead. The file says so at the top.

Why its parsing figure carries a tilde

Like AnythingLLM, there is no route listing a document’s stored chunks. The parsing measurement is taken by sweeping retrieval, which can only under-report — a missing fact means not observed, not dropped.