kb-bench

Methodology

Everything on this site comes from deploying the software and measuring it. This page states exactly how, including the parts that weaken the result.

The short version

One host, one platform at a time, one shared embedding model. Facts are planted in real public documents; a platform scores by whether those facts survive its parser and whether its own retrieval can find them again. No language model judges any result.

Why no model scores anything

A benchmark where a model writes the questions, answers them and marks them is grading its own homework: a confident wrong answer gets full marks. So the answer to every question here is a string lifted verbatim from the source document, chosen by a deterministic rule and verified to occur exactly once in that document.

A model is used in exactly one place — phrasing a natural question around a fact that was already extracted. It never decides what the answer is and never sees a platform's response. A badly phrased question is simply harder for every platform equally; it cannot turn a wrong answer into a right one.

Controls

What each dimension measures

DimensionWeightHow it is measuredModel involved
Setup and time to first knowledge base20%Seconds from compose up to a served request, plus the manual edits required before first startnone
Parsing fidelity20%Share of planted facts found in the chunks the platform storednone
Retrieval accuracy15%Share of questions whose answer appears in the platform's own top-k contextnone
Workflow orchestration15%Feature audit against published documentation, stated per platformnone
Self-host footprint10%Container count and peak memory sampled every 5s during ingestionnone
Ecosystem and integrations10%Count of first-party integrations from public documentationnone
Licence and commercial terms10%Reading the licence textnone

The weighting is a judgement and reasonable people would weight it differently. That is why the raw per-item results are published: anyone who disagrees can recompute the table with their own weights.

The corpus

36 public documents with a median size of about fifteen kilobytes: product and configuration documentation, security and process pages, protocol specifications and tax forms. Every source URL is listed in the manifest.

That shape is a deliberate correction. An earlier corpus built from SEC filings was dominated by four documents that produced more chunks than forty others combined, so it measured how each platform copes with outliers rather than with the workload anyone has. A second attempt using documentation pages fetched as HTML failed the same way for a different reason: a rendered page is two megabytes of navigation and scripting around thirty kilobytes of content. Documentation is therefore taken as the markdown it is written in.

Both runs are published. The corpus change reversed the ranking, which is itself the most useful thing this benchmark has produced, and the front page shows the two side by side rather than quietly replacing one with the other.

Nothing is written for the test: synthetic documents are too uniform to separate one parser from another, and a reader cannot verify them. Each document carries probe markers — strings verified to occur exactly once in it, ranging from version numbers and configuration keys to distinctive phrases from running prose. A document yielding no unambiguous marker is dropped rather than scored on guesswork.

What this cannot tell you

Reproducing it

The harness is open source. On a Linux host with Docker, roughly 32 GB of RAM and 100 GB of disk, bash scripts/reproduce.sh rebuilds the corpus from its public sources and re-runs every platform. Expect several hours; CPU embedding is the bottleneck.