Concepts

Evaluation and benchmarks

What the numbers measure, which ones are reproducible, and the rules that keep them honest.

Numbers on this page come from the engine's benchmark register; the public mirror lives on the product benchmarks page. They are retrieval-evidence numbers unless labeled otherwise — treating them as answer accuracy is a category error (see below).

Two kinds of number

Retrieval recall (R@K) — did the right evidence land in the top-K? Measured against gold evidence, no LLM in the loop.

Response accuracy — did an LLM judge accept the answer? What commercial leaderboards usually publish.

Comparing ValorBrain's R@10 with someone else's judged accuracy is the most common wrong read of this page. When the engine publishes accuracy (BEAM via AMB), it says so.

Reproducible register

BenchmarkMetricResultNotes
LoCoMoevidence R@1096.58%1986 QAs, two identical clean runs (2026-07-29). The honest base.
LoCoMoR@1 / R@569.4% / 91.3%Same runs.
LongMemEval-S (full 500)R@5 / R@1095.2% / 97.6%Pending re-verification with the corrected index gate.
LoCoMo, dense-onlyR@1063.5%The embedding leg alone — a component measurement, not the pipeline.
BEAM-100K (AMB harness)accuracysee the public scoreboardFixed reader/judge pair; comparable with the AMB leaderboard.

The older LoCoMo register of 97.4% R@10 is explicitly not reproducible — it was measured on a half-indexed corpus (documents counted before their vectors reached the hybrid index) and the engine's docs mark it do not cite. The delta between 96.58% and 97.4% is systematic, not noise; it is documented rather than quietly forgotten.

Reference points that are not ValorBrain's: Engram's R@5 93.9% (LoCoMo) and 98.4% (LongMemEval) belong to that system. Comparing them against the table above is recall-vs-recall and is fair; adopting them as ValorBrain's is not.

The pipeline behind the numbers

The R@10 96.58% is the full stack — BM25 (pgturbohybrid) + dense (LFM2.5-Embedding-350M-finetuned-v3, 1024-d) + RRF + PPR graph rerank + BGE-Reranker-v2-m3 cross-encoder. Each leg is measurable alone (dense-only: 63.5%); the fusion is where the rest comes from. The embedding swap that matters historically: Jina → LFM2.5 took dense-only R@10 from 32.7% to 63.5% (+30.8pp) on the same corpus.

Rules the engine holds itself to

These came from real, expensive mistakes; a benchmark that breaks them measures the harness, not the system.

  1. Wait for the index, not the embedding. A document is searchable by the dense leg only after it lands in the hybrid index. Gates that waited on embedding state measured a half-indexed corpus and produced the non-reproducible 97.4%.
  2. Never benchmark against a moving system. No restarts or deploys mid-run; no second benchmark sharing the BM25 statistics or the dedup budget.
  3. Fingerprint the corpus. The corpus is alive; runs are comparable only under the same fingerprint.
  4. Know the noise floor before comparing. On LoCoMo it is ±2 documents in 10 conversations; below that, a "regression" is weather.
  5. The judge is never the responder. Answer-quality benchmarks use a judge model from a different family than the reader — self-preference bias is real and measurable.
  6. An empty reader response is a harness defect — counted, not averaged away.

Where the numbers come from

  • Canonical public runs: the AMB harness (agent-memory-benchmark, vectorize-io), same arrears the public scoreboard uses.
  • Research iterations inside the engine (beam-qa.ts) use a different prompt/judge pair and are labeled internal — they are for choosing what to fix, never for publishing.

On an on-premise (Enterprise) installation, to measure your own instance: bun run gate:locomo (nightly recall floor, R@10 ≥ 0.90 default) and bun run gate:beam (accuracy via AMB) ship with the engine. They cost LLM tokens and time; they are gates, not unit tests.

On this page