Evaluation and benchmarks
What the numbers measure, which ones are reproducible, and the rules that keep them honest.
Numbers on this page come from the engine's benchmark register; the public mirror lives on the product benchmarks page. They are retrieval-evidence numbers unless labeled otherwise — treating them as answer accuracy is a category error (see below).
Two kinds of number
Retrieval recall (R@K) — did the right evidence land in the top-K? Measured against gold evidence, no LLM in the loop.
Response accuracy — did an LLM judge accept the answer? What commercial leaderboards usually publish.
Comparing ValorBrain's R@10 with someone else's judged accuracy is the most common wrong read of this page. When the engine publishes accuracy (BEAM via AMB), it says so.
Reproducible register
| Benchmark | Metric | Result | Notes |
|---|---|---|---|
| LoCoMo | evidence R@10 | 96.58% | 1986 QAs, two identical clean runs (2026-07-29). The honest base. |
| LoCoMo | R@1 / R@5 | 69.4% / 91.3% | Same runs. |
| LongMemEval-S (full 500) | R@5 / R@10 | 95.2% / 97.6% | Pending re-verification with the corrected index gate. |
| LoCoMo, dense-only | R@10 | 63.5% | The embedding leg alone — a component measurement, not the pipeline. |
| BEAM-100K (AMB harness) | accuracy | see the public scoreboard | Fixed reader/judge pair; comparable with the AMB leaderboard. |
The older LoCoMo register of 97.4% R@10 is explicitly not reproducible — it was measured on a half-indexed corpus (documents counted before their vectors reached the hybrid index) and the engine's docs mark it do not cite. The delta between 96.58% and 97.4% is systematic, not noise; it is documented rather than quietly forgotten.
Reference points that are not ValorBrain's: Engram's R@5 93.9% (LoCoMo) and 98.4% (LongMemEval) belong to that system. Comparing them against the table above is recall-vs-recall and is fair; adopting them as ValorBrain's is not.
The pipeline behind the numbers
The R@10 96.58% is the full stack — BM25 (pgturbohybrid) + dense (LFM2.5-Embedding-350M-finetuned-v3, 1024-d) + RRF + PPR graph rerank + BGE-Reranker-v2-m3 cross-encoder. Each leg is measurable alone (dense-only: 63.5%); the fusion is where the rest comes from. The embedding swap that matters historically: Jina → LFM2.5 took dense-only R@10 from 32.7% to 63.5% (+30.8pp) on the same corpus.
Rules the engine holds itself to
These came from real, expensive mistakes; a benchmark that breaks them measures the harness, not the system.
- Wait for the index, not the embedding. A document is searchable by the dense leg only after it lands in the hybrid index. Gates that waited on embedding state measured a half-indexed corpus and produced the non-reproducible 97.4%.
- Never benchmark against a moving system. No restarts or deploys mid-run; no second benchmark sharing the BM25 statistics or the dedup budget.
- Fingerprint the corpus. The corpus is alive; runs are comparable only under the same fingerprint.
- Know the noise floor before comparing. On LoCoMo it is ±2 documents in 10 conversations; below that, a "regression" is weather.
- The judge is never the responder. Answer-quality benchmarks use a judge model from a different family than the reader — self-preference bias is real and measurable.
- An empty reader response is a harness defect — counted, not averaged away.
Where the numbers come from
- Canonical public runs: the AMB harness (agent-memory-benchmark,
vectorize-io), same arrears the public scoreboard uses. - Research iterations inside the engine (
beam-qa.ts) use a different prompt/judge pair and are labeled internal — they are for choosing what to fix, never for publishing.
On an on-premise (Enterprise) installation, to measure your own instance: bun run gate:locomo (nightly recall floor, R@10 ≥ 0.90 default) and bun run gate:beam (accuracy via AMB) ship with the engine. They cost LLM tokens and time; they are gates, not unit tests.