Model-prior leakage in query answers¶
The query operation answers in prose: it retrieves the top-k particles for a question and asks a model to compose an answer from them. That composing step is the one place the engine runs inference over what it retrieved, and the model brings its own background with it. Anything it adds that the retrieved particles did not supply reaches the reader with no record of where it came from. When the store holds nothing relevant, the relevance floor and the answer step's own refusal rule say so. When the store holds something and the model fills in the rest, nothing used to say so.
particles benchmark leakage measures that blended case. It reports the
share of answer sentences that the particles the answer was composed from do
not support.
What it does¶
- Answer. Each question runs through the query operation exactly as
configured: the same retrieval depth, the same relevance floor, the same
composing model (
llm.query_response). A refused answer is never judged. Refusals are reported as their own rate. - Split. The answer is split into sentences by a deterministic, markdown-aware splitter. List items, headings, table rows, and code blocks are their own units, and abbreviations, initials, decimals, and version strings never end a sentence. The same answer always yields the same sentences.
- Judge. Each sentence goes to the entailment judge that abstraction promotion already uses, with a rubric for answer sentences. The judge sees the sentence, the question and the two preceding sentences as context (to resolve an "it"), and the claims the composer was given: the rendered top-k plus any narrative's constituent claims. It first decides whether the sentence asserts anything. A lead-in, a remark about the knowledge base, or a caveat is counted apart and is outside every rate. For a sentence that does assert something, it decides whether the claims support every part of it. A paraphrase, a summary, a less specific restatement, a combination of claims, or an inference that follows directly is supported. Anything else is unsupported, including things that are true: the measure is attribution, not accuracy.
- Report. The headline is the pooled unsupported rate over claim-bearing sentences. Beside it are the rate per question source, the mean per-answer fraction (each answer weighted equally), and the share of answers with at least one unsupported sentence.
The judge runs on llm.benchmark and must resolve to a different model
from the composer. A model judging its own answers shares the background being
measured and tends to find its own additions supported, so a run where the two
resolve to the same model is refused unless
benchmark_leakage.require_distinct_judge is turned off. A self-judged report
says its number is a lower bound.
The questions are the private held-out set the
relevance-floor benchmark harvests from an operator's
own agent transcripts. --question measures ad hoc questions instead and
prints one row per question. The rendered table never prints a question, an
answer, a sentence, or a judge's reason.
The first run (2026-10-03)¶
Judged under rubric protocol 1, which asks for the verdict before the reason.
The owner's store, 39,571 ACTIVE particles. All 338 held-out questions, at the
query surfaces' defaults: top-k 40, the general audience, a relevance floor of
0.25. The composer was claude-sonnet-5 and the judge was claude-opus-5-5,
under rubric protocol 1 with two context sentences.
| Measure | Value |
|---|---|
| Unsupported sentences, pooled | 23.6% (379 of 1,609 claim-bearing) |
| Mean per-answer fraction | 23.2% |
| Answers with at least one unsupported sentence | 66.8% (157 of 235) |
| Refused answers, never judged | 30.5% (103 of 338) |
| Sentences that assert nothing, counted apart | 519 |
| Sentences the judge could not score | 0 |
| Question source | Unsupported sentences |
|---|---|
user_prompt |
23.6% (376 of 1,596) |
cli_query |
60.0% (3 of 5) |
mcp_query |
0.0% (0 of 8) |
Almost every judged sentence comes from typed prompts, which are a proxy for memory queries. The explicit-query rows are too small to read.
What the unsupported sentences are. This store is about a private software project, so the composing model cannot have memorised its facts. A sample of unsupported sentences shows the leakage is the model's own interpretation, written in the same voice as the recorded claims. Typically the first half of a sentence restates a claim and the second half adds a gloss the store does not hold:
- a motive ("presumably a dev-only default", a refactor "could suggest the current structure was seen as a problem");
- a status ("this issue seems to remain a live concern");
- a conclusion ("not yet reliable enough … pending further tuning");
- occasionally an inference the store contradicts (that the project had moved past a milestone the store records as re-ranked to come later).
On a store about public knowledge, the same measurement would also catch memorised facts. On this one, it is catching editorial inference passed off as recalled fact.
How to read it. The judge samples: the Opus-class judge rejects the
temperature parameter, so verdicts are not greedy, and a repeat run over the
same answers would differ by a few sentences. The rubric is strict by design.
"Follows directly" is supported, a plausible reading is not. The number is one
operator's store and one question set, not a benchmark.
Grounded answers and the reason-first judge (2026-10-03)¶
Both tables in this section were judged under rubric protocol 2, which asks for the reason first and the verdicts last. Protocol 2 is now the default.
Two changes were measured together, so that the comparison between them is made once, on the corrected judge.
- Protocol 2. The first run's rubric asked for the verdict before the reason, the order that lets a judge commit before it reasons. Protocol 2 is the same rubric with the reason first. Protocol 1 is kept unchanged so the first number stays reproducible.
- Grounded answers.
particles query --groundedhas the composer cite the retrieved particle IDs behind every sentence and label each uncited sentence as its owninferenceorbackground(see Grounded answers). The benchmark's--groundedmode judges a cited sentence against the particles it cites, specifically, and every other sentence against everything the composer saw. Its headline counts only silent unsupported sentences, cited or unattributed, over every claim-bearing sentence. A labelled sentence is disclosed, so it is reported beside the headline rather than in it.
Both runs used the first run's frozen copy of the store, its questions, and
its settings: top-k 40, the general audience, a relevance floor of 0.25,
composer claude-sonnet-5, judge claude-opus-5-5.
Both runs are partial. The account's API credit ran out with about four fifths of the questions measured, and the remainder was not re-run because it could not change the decision below. The tables report the 286 questions both runs completed, compared question for question.
| Measure | Ungrounded | Grounded |
|---|---|---|
| Silent unsupported sentences | 25.3% (367 of 1,452) | 21.7% (344 of 1,585) |
| Mean per-answer fraction | 26.0% | 19.8% |
| Answers with at least one silent unsupported sentence | 72.0% (144 of 200) | 70.5% (141 of 200) |
| Refused answers, never judged | 30.1% (86 of 286) | 30.1% (86 of 286) |
| Sentences that assert nothing, counted apart | 397 | 66 |
The grounded mode lowers the silent unsupported rate by 3.6 points. A paired bootstrap over the questions puts the 95% interval at 0.1 to 7.2 points.
| Grounded attribution | Count |
|---|---|
| Sentences citing a retrieved particle | 1,370, of which 344 are unsupported by the particles they cite |
Sentences labelled inference |
262 (16.5% of claim-bearing), of which the judge found 139 supported anyway |
Sentences labelled background |
0 |
| Sentences with no valid citation and no label | 19, none unsupported |
| Citations naming a particle that was not retrieved | 2 |
What it shows. The reason-first judge reads about the same as the first
run's: 25.3% here against 23.6% under protocol 1, over a slightly different
question set. The grounded mode does not materially reduce silent leakage.
About one sentence in five is still unsupported by what it cites, every one of
them a sentence that cites a particle, and the model's own labels are loose:
it marked a sixth of its sentences as inference, and the judge found half of
those were supported after all. It never used background, which fits this
store: the leakage here is interpretation, not memorised fact. The default
therefore stays off (query.grounded_answers: false).
What it cannot show. A grounded sentence is checked only against the particles it cites, which is a stricter test than an ungrounded sentence gets. Part of the 344 may be miscitation, a sentence supported somewhere in the retrieved set but not by the particle it names, rather than content from nowhere. Separating the two needs a second judge call per such sentence and was not run.