Subject resolution: baseline and the embedding-only selector (2026-09-30)¶
A dated snapshot, as of 2026-09-30. This page records how often the subject resolver links a name to the right external entity, measured for the first time. The baseline was taken on the resolver as it shipped in 1.168.1, before any change. The same page then measures the first candidate fix, choosing among Wikidata's search candidates by how well each description matches the claim text, and records that it made the resolver worse.
Later
The Postgres failure below was fixed on 2026-10-01; see the prefix-expansion fix. Harbor and Rust were fixed by an optional model judge the same day; see the judged selector.
It measures the subject resolver (the cascade that turns an extracted name such as "Harbor" into a Subject, consulting the local store, then in-name identifiers, then Wikidata, then falling back to a bare local Subject). It does not measure the extractor and it is not the agent-memory benchmark on the Benchmarks page.
What the column can and cannot see
Twenty-two of the thirty-five gold subjects are invented entities with no Wikidata item, and most of them return no Wikidata search hit at all, so they score correct for any resolver that does not invent a link. The column's ability to separate one resolver from another rests on the other thirteen and on the handful of invented names that are also ordinary words. Read a change of one subject as a change of 2.9 points, and read the per-subject table rather than the headline alone.
What was measured¶
| Parameter | Value |
|---|---|
| Suite | prose-article-seed-001 v0.3.0, its gold_subjects block |
| Gold subjects | 35: 13 with a Wikidata QID, 22 with no item (ref: null) |
| Resolver | the production cascade, particles 1.168.1 |
| Context per subject | its gold mention (the first gold claim naming it) |
| Store | a throwaway SQLite store per subject |
| Encoder | all-MiniLM-L6-v2 (the default) |
| Thresholds | subjects.external_link_abstain_threshold 0.15, subjects.wikidata_link_suppress_threshold 0.25 (defaults) |
| Wikidata | live wbsearchentities, 2026-09-30 |
| Repeats | 2, byte-identical per-subject results |
The two cases v0.3.0 adds exist for this column. web-article-003 is a trade
brief about two invented companies named with ordinary English words, Harbor
and Lantern: the recorded "POET" failure in a form a fixture can hold.
web-article-004 is an engineering post naming real software (Go,
Python, Postgres, Kafka, Rust, Signal, Notion) whose names collide with other
Wikidata items.
Which names are gold follows one rule: every proper-named entity a gold claim names, whether or not today's extractor emits it as a subject. Each subject is resolved from its mention rather than from the extractor's output, so extraction sampling never moves the column.
A link counts at any confidence. Exporters hide a link below 0.25 and lint
flags it, but the store's join (find_by_external_ref) reads no confidence, so
the next mention that resolves to the same QID lands on that Subject either
way. Counting only the displayed links would score this baseline 0.771 and
hide one wrong link and five right ones.
Baseline¶
| Metric | Value | Count |
|---|---|---|
resolution_accuracy |
0.886 | 31 / 35 |
resolution_wrong_ref |
0.086 | 3 |
resolution_bare_local |
0.029 | 1 |
The four failures, and why each happens:
- Harbor →
Q283202, "sheltered body of water" (0.30). This is the recorded "POET" failure exactly. The company has no item, the common noun is the top hit, and its description scores 0.30 against "Harbor has acquired Lantern.", above both thresholds. The Subject is renamed "harbor". - Lantern →
Q862454, "fixed or portable enclosed lighting device" (0.20). The same failure in the band between the two thresholds: the link is hidden from exports and flagged by lint, but it is stored, and the Subject is renamed "lantern". - Postgres →
Q28975208, "discontinued database software, predecessor to PostgreSQL" (0.33). Wikidata's top hit is PostgreSQL (Q192490), the right answer. The resolver's prefix-expansion filter discards it, because "PostgreSQL" continues "Postgres" with a letter, the pattern the filter uses to reject "micrograd" → "Microgradients …". The next hit is taken instead. - Rust → no link. The top hit is right (
Q575650, the programming language), but its description scores below 0.15 against "Ostrander Freight considered Rust and chose Go for the billing service rewrite because …", so the resolver abstains.
Five correct links (Python 0.20, Go 0.23, Kafka 0.20, Signal 0.16, Notion 0.24) sit in the band where exporters hide them. They are correct and they join, so they count.
Every gold subject¶
| Subject | Gold | Stored ref (confidence) | Outcome |
|---|---|---|---|
| Fernwood Systems | none | none | correct |
| SHA-256 | Q110651361 |
Q110651361 (0.32) |
correct |
| Halifax, Nova Scotia | Q2141 |
Q2141 (0.58) |
correct |
| Dana Okonkwo | none | none | correct |
| Halcyon Grid Cooperative | none | none | correct |
| Duluth, Minnesota | Q485708 |
Q485708 (0.43) |
correct |
| Pike Lake array | none | none | correct |
| Tomas Ilves | none | none | correct |
| Priya Raghunathan | none | none | correct |
| Bracken & Doyle | none | none | correct |
| Salt Marsh Notes | none | none | correct |
| Maren Kaldestad | none | none | correct |
| Kvitholm Institute | none | none | correct |
| Meridian Rose | none | none | correct |
| Stray Light Foundation | none | none | correct |
| Bergen | Q26793 |
Q26793 (0.27) |
correct |
| Jonas Brekke | none | none | correct |
| Tidewrack Quarterly | none | none | correct |
| Harbor | none | Q283202 (0.30) |
wrong ref |
| Lantern | none | Q862454 (0.20) |
wrong ref |
| Manchester | Q18125 |
Q18125 (0.56) |
correct |
| Leeds | Q39121 |
Q39121 (0.27) |
correct |
| Odile Fanshawe | none | none | correct |
| Wendell Asquith | none | none | correct |
| Counting House Review | none | none | correct |
| Ostrander Freight | none | none | correct |
| Python | Q28865 |
Q28865 (0.20) |
correct |
| Go | Q37227 |
Q37227 (0.23) |
correct |
| Postgres | Q192490 |
Q28975208 (0.33) |
wrong ref |
| Kafka | Q16235208 |
Q16235208 (0.20) |
correct |
| Rust | Q575650 |
none | bare local |
| Signal | Q19718090 |
Q19718090 (0.16) |
correct |
| Notion | Q60747998 |
Q60747998 (0.24) |
correct |
| Ines Marchetti | none | none | correct |
| Rewriting billing in Go | none | none | correct |
After: choosing the candidate by its description¶
The selector, subjects.wikidata_candidate_selection: best_description,
changes one step. Where the resolver took Wikidata's top search hit, it now
scores every candidate the same search returns (up to five; each carries its
English description, so this costs no further network call) against the claim
text, with the local encoder and no LLM call. The best score is adopted only at
or above 0.25, the line above which a link is shown. Below it the Subject keeps
the extracted name and records the best candidate as a low-confidence link,
and below 0.15 the resolver still records nothing. The choice is a
function of the search response, the claim text and the encoder alone, so it
is as reproducible as the baseline: two passes gave identical rows here too.
| Metric | top_hit (baseline) |
best_description |
Change |
|---|---|---|---|
resolution_accuracy |
0.886 (31) | 0.800 (28) | −3 subjects |
resolution_wrong_ref |
0.086 (3) | 0.171 (6) | +3 subjects |
resolution_bare_local |
0.029 (1) | 0.029 (1) | none |
It fixed none of the four baseline failures and broke three correct links:
| Subject | Gold | top_hit |
best_description |
|---|---|---|---|
| SHA-256 | Q110651361 |
Q110651361 "cryptographic hash function" (0.32), correct |
Q124624061 "double SHA-256" (0.46), wrong |
| Leeds | Q39121 |
Q39121 "city in West Yorkshire" (0.27), correct |
Q1128631 "Leeds United F.C." (0.38), wrong |
| Signal | Q19718090 |
Q19718090 "privacy-focused encrypted messaging app" (0.16), correct |
Q828130 "signal transduction" (0.18), wrong |
| Lantern | none | Q862454 lighting device (0.20), wrong |
Q2618398 lighthouse lantern room (0.32), wrong and now shown |
| Harbor | none | Q283202 (0.30), wrong |
unchanged |
| Postgres | Q192490 |
Q28975208 (0.33), wrong |
unchanged |
| Rust | Q575650 |
no link | unchanged |
Every other row is unchanged. The one difference the column does not count is naming: below the 0.25 line the Subject now keeps the extracted name, so "Kafka" stays "Kafka" where the baseline renamed it "Apache Kafka".
Why it cannot work with this encoder¶
The losses share one cause. A longer, more specific description that shares words with the claim outscores a shorter, correct one: "double SHA-256" beats "cryptographic hash function" on a sentence about SHA-256 hashes, and a football club in Leeds beats the city on a sentence about a company in Leeds. Wikidata's search rank already encodes which sense is usual, and the argmax discards it.
A floor cannot rescue the policy either. On the baseline, the eleven correct links score between 0.16 and 0.58 against their claims and the three wrong ones between 0.20 and 0.33. Seven of the eleven correct links score below the highest wrong one, so any cut that removes the Harbor link (0.30) also removes Bergen, Leeds, Python, Go, Kafka, Signal and Notion. The description-to-claim cosine does not separate right from wrong on this set.
The literal "POET" case behaves the same way. Probed by hand with the claim
"POET Technologies is shipping its optical interposer platform to two
data-center customers.", the occupation (Q49757, "person who writes poetry")
is both Wikidata's top hit and the best-scoring candidate at 0.34, so either
selector links it. The company's own item (Q30298339) is not among the
candidates a search for "POET" returns.
Decision¶
Embedding-only selection is not enough, and on this measurement it is a
regression, so it is not the default. It ships as the non-default value of
subjects.wikidata_candidate_selection so this comparison can be re-run and a
later selector can be measured against both. The default, top_hit, is the
baseline resolver, unchanged (re-measured with the setting in place: identical
rows).
The four baseline failures point at three different fixes, none of them a better cosine:
- Harbor and Lantern need a judgement that a payroll company is not a body of water when no candidate is right. That is the LLM-assisted step, a separate change with its own cost per name.
- Postgres is a defect in the prefix-expansion filter, which discards the correct top hit because "PostgreSQL" continues "Postgres" with a letter.
- Rust is the abstention floor's recall cost, a correct top hit scored below 0.15 against a claim that mentions it in passing.
Claim metrics from the same suite version¶
One full particles extractor benchmark general-extractor --suite
prose-article-seed-001 run on 2026-09-30 (claude-sonnet-4-6, extractor
0.16.0, embedding judge at 0.80) scored recall 0.94, precision 0.83 and
calibration error 0.09 over the six cases. It is a single run, and v0.3.0 has
two more cases than the v0.2.0 the provider-survey pages measured, so it is
comparable with neither.