flmnt → benchmark

One question: does retrieving your recorded decisions change what a model gets right?

Each pack seeds a corpus of recorded decisions, then asks the model probe questions. The same model answers the same question five ways — with no context, with three different retrieval strategies, and with the answer handed to it directly. Every probe, every seeded decision, the judge rubric and every score are public.

run 2026-06-25models claude-haiku-4-5-20251001 · claude-opus-4-8judge claude-haiku-4-5-202510018 packs5 armsn = 5 median

Experimental design

Five arms. Same model, same question — only the context differs.

The two controls bound the result. Without the ceiling, a high retrieval score could mean the answer was never in the stream. Without the floor, it could mean the model already knew it.

Arm 01 · the floor

cold

Control

No tools, no memory. The model answers from parametric knowledge only — what it knows with no project context.

Arm 02

warm

Snapshot tier

A recency snapshot read of the stream.

Arm 03

flat

Vector-only baseline

Vector-similarity retrieval over the same stream, with no graph traversal. The comparator for graph-based retrieval.

Arm 04

rlm

The production path

Query-driven deep retrieval: embedding-anchored causal traversal, escalating to a graph-neighbourhood walk when the direct hit isn't enough.

Arm 05 · the ceiling

oracle

Control

The correct decisions injected directly into the prompt — the answer is in the stream and the model can use it when handed over.

Why the vector-only arm mattersThe three middle arms read the same stream and differ only in retrieval strategy — snapshot, vector similarity, and graph-based query-driven retrieval. That makes the gap between them a measurement of strategy rather than of corpus.

Results

Median answer accuracy, eight packs.

Scores are identical across both models except where noted. Answer accuracy is the primary metric, graded by the judge against each probe's expected answer.

Packcoldwarmflatrlmoracle
Core context recoverypack 1 · 15 probes0.08–0.090.31–0.470.95–1.000.95–1.000.97–1.00
Decision supersessionpack 2 · 5 probes0.000.10–0.201.001.001.00
Causal traversalpack 3 · 8 probes0.140.130.99–1.000.99–1.000.99–1.00
Multi-session handoffpack 4 · 6 probes0.04–0.050.97–1.000.97–1.000.93–1.000.98–1.00
Project isolationpack 5 · 8 probes0.13–0.141.001.001.000.97–1.00
Abstentionpack 6 · 6 probes0.331.001.001.001.00
Adversarial supersessionpack 7 · 5 probes0.000.000.94–0.980.94–0.970.95–0.97
Supersession edgepack 8 · 3 probes · no textual tell0.000.00–0.270.001.001.00

Pack 8, warm arm: 0.27 on claude-opus-4-8, 0.00 on claude-haiku-4-5-20251001.

Every per-probe score is in the repository's results file.

The row that matters

Pack 8: a decision replaced with nothing in the words to say so.

In pack 2 the current decision is identifiable from the stream's content and recency — and vector similarity finds it. In pack 8 the stale entry carries no textual marker at all. The only signal that it has been replaced is the recorded supersession relationship.

0.00vector-only retrieval

Served the superseded decision. Not broken — the stale entry was still the closest match, and nothing in its words said otherwise.

1.00the production path

Identified the current decision on every probe, by following the recorded supersession relationship rather than the text.

1.00ceiling

The oracle arm confirms the answer was in the stream. The gap is retrieval strategy, not corpus.

Why this is the one to readEvery other pack is answerable by similarity, so every tool in this category can pass them. Pack 8 removes the textual signal and the vector-only arm goes to zero while the graph path stays at one. That difference is the product.

Method

What each probe asks, and how it's graded.

recallprobe type
Retrieve a recorded decision and its rationale. Scored on the fraction of the expected answer's key claims present — no credit for hedging or style.
contradictionprobe type
Reject a false premise stated in the question and correct it. Scored 1.00 only on an explicit rejection; "it depends" scores zero.
abstentionprobe type
Admit the information isn't recorded rather than fabricate it. Fabricating a specific fact scores zero; hedging without acknowledging absence scores 0.50.
famaprobe type · supersession currency
Recognise a decision was superseded and not cite the stale one as current. Scored on coverage of the current decision, with a flag when the model treats a superseded decision as current.
AA · answer accuracyprimary metric · llm judge
Does the answer match the expected answer? This is the figure in the table above.
RF · rationale fidelityllm judge
Does the stated reasoning match the recorded rationale? Composite is the mean of the two.
sub_emdeterministic
Substring exact-match, recorded alongside as a judge-independent secondary signal.
n = 5, medianaggregation
Every probe runs five times per arm per model; the published value is the median, which is robust to single-run variance. The judge is never told which arm produced a response, and the same per-type rubric is applied to every arm.
corpus sweepthe consumption axis
Pack 1 additionally runs across corpus sizes of 10, 100, 500 and 1,500 seeded decisions, so tokens processed per query can be measured against corpus size. Token usage — input, output, cache reads and writes, and retrieved context — is recorded per probe.

Transparency boundary

The test is published. The runtime isn't.

Published in full

  • Every probe question and every seeded decision, per pack
  • The verbatim LLM-as-judge rubric, reproduced from the scoring function
  • Every per-probe score — one row per probe, arm and model
  • Token counts per probe, including retrieved context
  • Per-run history, with the fixes applied before each run. No run is deleted

Not published

  • The retrieval algorithm and its ranking weights
  • The embedding and traversal engineering
  • The retrieval system prompts
  • Any runtime source code

Reproducibility, and its limitThe inputs and the results are generated from the same source, so the corpus and the scores describe the same suite. You can audit every number and re-grade every answer against the published rubric. You cannot rebuild the retrieval path from this repository — that stays on flmnt's infrastructure.

Coverage

Models are added pass by pass, and each one's full results publish when it completes.

If a number here looks wrong, the raw rows are in the repository.