Skip to content

Subject resolution: the judge sees a deeper search (2026-10-02)

A dated snapshot, as of 2026-10-02. This page measures one change to the subject-link judge: it is now shown seven Wikidata search hits instead of five, and never a hit that stands for a name or a disambiguation page. It follows the 2026-10-01 judge page, uses its two gold sets and its measurement script, and re-measures the unchanged judge on the same day so before and after share one state of live Wikidata.

The failure

On the ordinary-prose gold set the judge left four real names bare local: Jest, Lodge, Chelsea and Ottolenghi. The resolver's one search asked wbsearchentities for five hits, and none of the four right items was among them. Shown five hits, the judge answered "none" each time, which was correct from what it saw.

Ten hits from the same search show where each item sits:

Name Right item Rank in the search
Jest Q65121527, the JavaScript test framework 6
Chelsea Q9616, the football club 6
Lodge Q17034962, the cookware maker not in the first 10
Ottolenghi Q8056413, the chef not in the first 10

The change

A deeper search for the judge only. A new setting, subjects.wikidata_judge_search_limit (default 7), sets how many hits the one search request asks for under llm_judge. The ambiguity gate still reads only the hits within the first five of the response, and top_hit and best_description still search five. The set of names that reach the model is therefore unchanged, and so is the resolution of every name that never reaches it. The extra hits come back in the same request and pass through the prefix-expansion filter and its alias read exactly as the first five do.

No name or disambiguation items in what the judge is shown. A search hit whose description is a given name, a family name, "name", or "Wikimedia disambiguation page" is dropped from the judge's list. Such an item never denotes the person, place or thing a claim names. The rule reads the description the search response already carries, so it costs no request. When every candidate is such an item, the name stays bare local with no model call, which is the answer the judge could only have given.

The second rule exists because the first one alone broke the no-correct-link-lost rule, as the next section shows.

What the deeper search did alone

The first round raised the depth and changed nothing else. Ten runs per name, each in a fresh store, on the names whose outcome moved:

Name, correct runs of 10 5 hits 7 hits 10 hits
Jest 0 10 10
Chelsea 0 10 10
Jordan, an invented colleague 10 3 3
Brown Palace 10 7 8
Braga 1 1 0

Jest and Chelsea linked every time. Jordan did not hold: "Jordan, who joined from Salesforce last year" was linked to Q14021944, "unisex given name", in 7 of 10 runs at either depth, against none at five. The deeper list adds two painters named Jordaens and Giordano, and with people in view the judge reads the given-name item as a fit. That lost a correct outcome, so the plain depth change could not ship.

Brown Palace's misses were a pick of Q4976226, a disambiguation page. Its search returns the same five hits at any depth, so the prompts at 5, 7 and 10 hits were identical apart from the per-call nonce. Its misses are the judge's own variation, which the depth did not cause.

A clause in the prompt saying that a given-name, family-name or disambiguation item is never what a claim names fixed Jordan in 10 of 10 runs. It did not hold for Priya, "my neighbour Priya", whose "female given name" item was picked in 2 of 10 runs at seven hits and in 1 of 2 full passes. A rule the model must follow is a rate, not a guarantee, so the clause was dropped and the filter above was built instead. The shipped prompt is the 2026-10-01 one, version 3614d60fbfb2fb1d, unchanged.

Result

Parameter Value
Gold sets ordinary-prose-001 v0.1.0 (49 names), prose-article-seed-001 v0.3.0 (35 names)
Resolver particles 1.169.1 with this change (released in 1.169.6), llm_judge
Judge anthropic/claude-sonnet-4-6, temperature 0, prompt 3614d60fbfb2fb1d
Wikidata live, 2026-10-02
Repeats 2 full passes per arm, plus 10 runs per name on 10 names
Command uv run python scripts/measure_subject_resolution.py <gold> --selection llm_judge --judge-search-limit <n> --out <file>

Ordinary prose. "Before" is the shipped resolver: five hits, no filter.

Arm Accuracy Wrong ref Bare local Passes match
Before 0.898 / 0.898 1 / 2 4 / 3 no
5 hits, filter 0.918 / 0.918 0 / 0 4 / 4 yes
7 hits, filter (shipped) 0.959 / 0.959 0 / 0 2 / 2 yes
10 hits, filter 0.959 / 0.939 0 / 1 2 / 2 no

The two values per cell are the two passes. Against the before arm, the shipped arm links Jest and Chelsea and holds Braga on the municipality the gold names, and it loses no correct outcome in either pass. The two names left bare local are Lodge and Ottolenghi.

  • Chelsea under the before arm was bare local in one pass and linked to the London district in the other: shown the district and four places, the judge sometimes took the district for "Chelsea's ground". Ten runs gave 7 bare and 3 wrong. With the club in view it linked the club 10 times of 10.
  • Braga has two items for the one city, the municipality (the gold) and its city seat. The filter drops a family-name hit from its list, and with that gone the judge chose the municipality in 10 of 10 runs at five and seven hits, against 1 of 10 before. At ten hits, which add three civil parishes, it chose the city seat in 2 of 10.

The first suite. All four arms scored 0.971 (34 of 35) in both passes, and no correct link was lost. Its one miss is Lantern, an invented payroll vendor, which the judge links to a wrong item in every arm; the 2026-10-01 page explains why the claim gives it nothing to go on.

Ten runs per name at seven hits, with the filter: all ten names held their outcome 10 times of 10. That covers the six invented names (Mark Ellison, Vitest, Lena Marsh, Rui Matos, Priya and Jordan), all left without a link, and Jest, Chelsea, Braga and Brown Palace, all linked to the gold item.

Reach and cost

The gate is unchanged, so the same names reached the model in every arm: 45 of 49 on ordinary prose, 14 of 35 on the first suite.

Measure, per judged name Before 7 hits, filter 10 hits, filter
Input tokens, ordinary prose 636 659 711
Input tokens, first suite 615 637 679
Cost per 100 names, ordinary prose US$0.199 US$0.201 US$0.217
Cost per 100 names, first suite US$0.081 US$0.084 US$0.089

Seven hits add about 20 input tokens per judged name, because the filter takes back some of what the two extra hits add. The search itself is the same one request; a deeper response is a few hundred bytes more.

Seven, not ten

Ten hits link the same two names as seven and cost 40 to 50 more input tokens per judged name. In the measured runs they also did worse on the one name with two valid items: Braga went to the city seat in one full pass and in 2 of 10 repeats, against none at seven. Seven reaches the rank-6 items with one hit to spare.

The two names a deeper search cannot reach

Lodge and Ottolenghi are not in the first ten hits for the bare name. The register entry proposed a second query built from the claim: the name plus a word that says what kind of thing it is. Its premise was measured before anything was built, by searching for the name plus each content word of its sentence and checking whether the right item came back in the top five.

Name Claim words tried Entity search finds it Full-text search finds it
Lodge 8 with no word only with "cast-iron"
Ottolenghi 1 ("recipe") with no word with no word
Jest 7 with no word with no word
Chelsea 4 with no word only with "Fulham"

The word that finds Ottolenghi ("chef") and the word that finds Lodge through the entity search ("cookware", or the label "Lodge Manufacturing") are not in the sentence. A query drawn from the claim therefore buys Lodge at most, only through the full-text endpoint, which returns no labels or descriptions and so needs a third request, and only if the right word of eight is chosen. It was not built. The two names stay open in the register with this measurement.

Upgrading

The candidate set is part of the key the judge's verdicts are recorded under. A name whose search returns more than five hits, or whose list included a name or disambiguation item, has a new key, so the first resolution of it after the upgrade asks the model once more. On ordinary prose that is nearly every judged name. The old verdict rows stay in the ledger and are never read again; the ledger only grows by one row per judged name each time its candidate set changes, as it already did whenever Wikidata's search response changed.

Method

Each arm ran the column on its own through scripts/measure_subject_resolution.py, with subjects.wikidata_judge_search_limit set per arm, and every gold subject was resolved in a throwaway store with an empty verdict ledger, so every judged name was asked afresh in every pass. The ten-run repeats used the same function, each repeat in its own store. The before arm was measured on the shipped code, the rounds with the clause and the filter on the code they describe, all on 2026-10-02.