flmnt → benchmark
One question: does retrieving your recorded decisions change what a model gets right?
Each pack seeds a corpus of recorded decisions, then asks the model probe questions. The same model answers the same question five ways — with no context, with three different retrieval strategies, and with the answer handed to it directly. Every probe, every seeded decision, the judge rubric and every score are public.
Experimental design
Five arms. Same model, same question — only the context differs.
The two controls bound the result. Without the ceiling, a high retrieval score could mean the answer was never in the stream. Without the floor, it could mean the model already knew it.
Arm 01 · the floor
cold
Control
No tools, no memory. The model answers from parametric knowledge only — what it knows with no project context.
Arm 02
warm
Snapshot tier
A recency snapshot read of the stream.
Arm 03
flat
Vector-only baseline
Vector-similarity retrieval over the same stream, with no graph traversal. The comparator for graph-based retrieval.
Arm 04
rlm
The production path
Query-driven deep retrieval: embedding-anchored causal traversal, escalating to a graph-neighbourhood walk when the direct hit isn't enough.
Arm 05 · the ceiling
oracle
Control
The correct decisions injected directly into the prompt — the answer is in the stream and the model can use it when handed over.
Why the vector-only arm mattersThe three middle arms read the same stream and differ only in retrieval strategy — snapshot, vector similarity, and graph-based query-driven retrieval. That makes the gap between them a measurement of strategy rather than of corpus.
Results
Median answer accuracy, eight packs.
Scores are identical across both models except where noted. Answer accuracy is the primary metric, graded by the judge against each probe's expected answer.
| Pack | cold | warm | flat | rlm | oracle |
|---|---|---|---|---|---|
| Core context recovery | 0.08–0.09 | 0.31–0.47 | 0.95–1.00 | 0.95–1.00 | 0.97–1.00 |
| Decision supersession | 0.00 | 0.10–0.20 | 1.00 | 1.00 | 1.00 |
| Causal traversal | 0.14 | 0.13 | 0.99–1.00 | 0.99–1.00 | 0.99–1.00 |
| Multi-session handoff | 0.04–0.05 | 0.97–1.00 | 0.97–1.00 | 0.93–1.00 | 0.98–1.00 |
| Project isolation | 0.13–0.14 | 1.00 | 1.00 | 1.00 | 0.97–1.00 |
| Abstention | 0.33 | 1.00 | 1.00 | 1.00 | 1.00 |
| Adversarial supersession | 0.00 | 0.00 | 0.94–0.98 | 0.94–0.97 | 0.95–0.97 |
| Supersession edge | 0.00 | 0.00–0.27 | 0.00 | 1.00 | 1.00 |
Pack 8, warm arm: 0.27 on claude-opus-4-8, 0.00 on claude-haiku-4-5-20251001.
Every per-probe score is in the repository's results file.
The row that matters
Pack 8: a decision replaced with nothing in the words to say so.
In pack 2 the current decision is identifiable from the stream's content and recency — and vector similarity finds it. In pack 8 the stale entry carries no textual marker at all. The only signal that it has been replaced is the recorded supersession relationship.
Served the superseded decision. Not broken — the stale entry was still the closest match, and nothing in its words said otherwise.
Identified the current decision on every probe, by following the recorded supersession relationship rather than the text.
The oracle arm confirms the answer was in the stream. The gap is retrieval strategy, not corpus.
Why this is the one to readEvery other pack is answerable by similarity, so every tool in this category can pass them. Pack 8 removes the textual signal and the vector-only arm goes to zero while the graph path stays at one. That difference is the product.
Method
What each probe asks, and how it's graded.
- recallprobe type
- Retrieve a recorded decision and its rationale. Scored on the fraction of the expected answer's key claims present — no credit for hedging or style.
- contradictionprobe type
- Reject a false premise stated in the question and correct it. Scored 1.00 only on an explicit rejection; "it depends" scores zero.
- abstentionprobe type
- Admit the information isn't recorded rather than fabricate it. Fabricating a specific fact scores zero; hedging without acknowledging absence scores 0.50.
- famaprobe type · supersession currency
- Recognise a decision was superseded and not cite the stale one as current. Scored on coverage of the current decision, with a flag when the model treats a superseded decision as current.
- AA · answer accuracyprimary metric · llm judge
- Does the answer match the expected answer? This is the figure in the table above.
- RF · rationale fidelityllm judge
- Does the stated reasoning match the recorded rationale? Composite is the mean of the two.
- sub_emdeterministic
- Substring exact-match, recorded alongside as a judge-independent secondary signal.
- n = 5, medianaggregation
- Every probe runs five times per arm per model; the published value is the median, which is robust to single-run variance. The judge is never told which arm produced a response, and the same per-type rubric is applied to every arm.
- corpus sweepthe consumption axis
- Pack 1 additionally runs across corpus sizes of 10, 100, 500 and 1,500 seeded decisions, so tokens processed per query can be measured against corpus size. Token usage — input, output, cache reads and writes, and retrieved context — is recorded per probe.
Transparency boundary
The test is published. The runtime isn't.
Published in full
- Every probe question and every seeded decision, per pack
- The verbatim LLM-as-judge rubric, reproduced from the scoring function
- Every per-probe score — one row per probe, arm and model
- Token counts per probe, including retrieved context
- Per-run history, with the fixes applied before each run. No run is deleted
Not published
- The retrieval algorithm and its ranking weights
- The embedding and traversal engineering
- The retrieval system prompts
- Any runtime source code
Reproducibility, and its limitThe inputs and the results are generated from the same source, so the corpus and the scores describe the same suite. You can audit every number and re-grade every answer against the published rubric. You cannot rebuild the retrieval path from this repository — that stays on flmnt's infrastructure.
Coverage
Models are added pass by pass, and each one's full results publish when it completes.
If a number here looks wrong, the raw rows are in the repository.