Predicate canonicalisation: how far the free-text predicates collapse (2026-10-01)¶
A dated snapshot, as of 2026-10-01. This page measures the first step a
predicate profile needs: turning the free-text predicates of the
owner's live store into a smaller set of canonical predicates. It compares two
methods that cost no LLM call. Both were run read-only, on a copy of the store
taken with SQLite's backup command, over the listing that particles query
--predicates prints.
What was measured¶
The store has grown since the 2026-09-30 count the register row quotes (6,358 predicates over 11,074 claims).
| 2026-10-01 | |
|---|---|
| ACTIVE particles | 37,382 |
| Particles carrying a structured claim | 13,817 |
| Claims the listing counts | 13,387 |
| Distinct predicates | 7,504 |
| Kind | all TOKEN, none URI |
| Predicates that occur once | 6,084 (81%) |
| Share of claims under the 20 commonest predicates | 13.0% |
Two methods:
- Normalisation. Lowercase the predicate, drop articles, drop leading
auxiliaries and modals (
is,was,can, andhaswhen a participle follows it), then lemmatise the head verb.movedbecomesmoveandlivesbecomeslive. Negation and prepositions are kept, sodoes not supportstays apart fromsupports, andmoved tostays apart frommoved from. Lemmas are taken from the listing's own head verbs where one fits, with an irregular-verb table and suffix rules as the fallback. - Embedding clustering. Encode each predicate with the store's own local
encoder (
all-MiniLM-L6-v2) and cluster with average-linkage agglomerative clustering on cosine distance. The cut is made at cosine similarity 0.85 and at 0.75.
The script is scripts/measure_predicate_clusters.py. It reads the listing,
not the store:
uv run particles query --predicates > predicates.txt
uv run python scripts/measure_predicate_clusters.py predicates.txt \
--threshold 0.85 --threshold 0.75 --json clusters.json
Cluster counts¶
| Method | Clusters | Reduction | Clusters with 2+ terms | Claims in those | Largest cluster |
|---|---|---|---|---|---|
| None (as stored) | 7,504 | ||||
| Normalisation | 6,792 | 9.5% | 520 | 5,030 (37.6%) | 12 terms |
| Embedding at 0.85 | 6,336 | 15.6% | 869 | 5,192 (38.8%) | 14 terms |
| Embedding at 0.75 | 5,191 | 30.8% | 1,303 | 8,406 (62.8%) | 26 terms |
No method comes close to a vocabulary a reviewer could confirm term by term. The long tail is the reason: four predicates in five occur once, and most of those are one-off phrasings ("passed with", "has passing unit test count") that no surface method can tie to a commoner form without knowing what they mean.
Where the methods go wrong¶
A merge is only useful to a profile if every member means the same relation
in the same direction. A cardinality rule applied to a cluster that mixes
renamed to with renamed from would treat the old name and the new name as
one value. The check below counts clusters with 2+ terms that mix a negated
form with a plain one, or a to / into form with a from form.
| Method | Mix negated and plain | Claims | Mix to and from |
Claims |
|---|---|---|---|---|
| Normalisation | 0 | 0 | 0 | 0 |
| Embedding at 0.85 | 23 | 151 | 26 | 221 |
| Embedding at 0.75 | 73 | 1,007 | 27 | 282 |
The script flagged two normalisation clusters as mixed polarity. Both are
false alarms, since every member is a negative form (lacks / lack /
lacked, and the bare is not / was not / does not), so the table counts
them as zero. The embedding errors are real. At 0.85 the encoder merges
was renumbered from with was renumbered to, renamed to with renamed
from, version bumped from with version bumped to, and isolate with do
not isolate. At 0.75 it also merges is with is not, includes with does
not include, and returns with never returns.
Normalisation makes the opposite trade. Every merge it makes is a tense,
number or modal variant of one verb with the same preposition: has status /
had status / have status, contains / contained / contain. It cannot
see a synonym (is located at and lives in stay apart), and it merges
modals that change meaning: passes and must pass land together, though
one is a fact and the other a requirement. Its largest cluster is the bare
copula (is, was, are, is a, can be, must be…), which is not a
relation at all.
The residence family, worked through¶
The register row cites lives in (4), moved to (7), was moved to (8) and
moved from (6) as the residence slot spread over four phrasings, the 4 to 9
way variance measured when the update rung was designed. On
today's copy those predicates land as follows:
| Predicate | Normalisation | Embedding at 0.85 | Embedding at 0.75 |
|---|---|---|---|
lives in (4) |
with live in, lived in |
alone | with live in, lived in |
moved to (7) |
with was moved to, must be moved to |
with moved from, moved |
one cluster of 9: to, from, into, by, under |
was moved to (11) |
with moved to, must be moved to |
with was moved from, was moved into |
the same cluster of 9 |
moved from (6) |
with was moved from |
with moved to, moved |
the same cluster of 9 |
Reading the claims behind them changes the example. In this store none of
the four is about where a person lives. moved to and was moved to record
a decision record's file moving from proposed/ to active/, or a register
row moving to its closed section. moved from names the directory it left.
lives in says where a module or setting is kept ("the scoring package lives
in the Client layer"). The residence slot comes from persona stores such as
the rot and LongMemEval ones, where the subject is a person.
Two conclusions follow for the design:
- The slot is a predicate on a class, not a predicate.
move toon a decision record is one value at a time (a record sits in one directory), and on a person it is residence, also one at a time but a different slot.lives inon a module andlives inon a person are two slots. A canonical predicate has to be keyed by the subject's class, which is the per-class co-occurrence the vocabulary report would carry. - Direction is part of the predicate.
move tonames the new value andmove fromnames the old one. Normalisation keeps them apart, which is right. A profile can then say thatmove togives a slot its current value andmove fromrecords a past one, which no merged cluster can say.
Reading¶
- Normalisation is the safe first step. On this store it merges nothing it should not, and it is deterministic, free and explainable. It is also weak: it cuts the vocabulary by a tenth and leaves 6,272 singletons.
- Embedding similarity is a suggestion source, never a merge. At 0.85 it
finds real aliases normalisation misses (
has passing test countand its eight variants,was pushed to/pushed to/were pushed to), and in the same pass it joins opposite directions and opposite polarities. Its candidates are worth ranking for a reviewer by the claims they would cover. Applied without review, it would hand a cardinality rule the wrong values. - Canonicalisation will not cover the tail. Even at 0.75, where the errors above are frequent, 3,888 clusters are singletons. A profile can only ever govern the head of the distribution, and every claim outside it keeps the probe path. That is why the probe itself was changed first (see the slot check's kind).
Added 2026-10-02: the core form, and how far a profile reaches¶
The predicate profiles proposal needs a canonical form computed
from one predicate alone. The normalisation above picked each lemma by
checking which base forms the listing contains, so adding one predicate to
the store could change another's form. The SDK's function instead reduces
the head verb to an inflectional stem (moved, moves, moving all
become mov). On the same listing:
| Normalisation above | Core stem | |
|---|---|---|
| Clusters | 6,792 | 6,759 |
| Merges in common | 1,039 | 1,039 |
| Merges only this one makes | 3 | 50 |
Clusters mixing negated and plain, or to and from |
0 | 0 |
The stem lost three merges, had run for with its variants, which treating
run as its own participle fixes. The 50 it adds are inflections the
attestation rule missed, such as created and creates.
How far could a profile reach on this store? On a read-only copy taken 2026-10-02, the update sweep's own candidate search, with no LLM call, finds 1,985 same-subject pairs it would consider.
| Filter | Pairs |
|---|---|
| Both claims carry a triple | 578 |
| Both triples about the same resolved subject | 450 |
| That subject has a class | 62 (3.1%) |
| Both predicates share a canonical form | 11 (0.55%), on 4 keys |
2,363 of 9,850 subjects have a class, nearly all from the code-artifact authority (files, records, symbols, commands). A profile can only act on a classed pair, so on this store it could settle at most 62 of the pairs the sweep considers, and 11 without reviewed aliases.
Added 2026-10-02: the vocabulary proposal step¶
The SDK now stores reviewed predicate rulings in a vocabulary document
(particles vocab), and its proposal step suggests alias and profile
candidates from a store's own predicates. It follows the design above: a
canonical predicate is a normalised form on a subject class, so only claims
whose triple is about a classed subject take part, and an alias candidate is an
embedding cluster at cosine 0.85 within one class, split so that no candidate
mixes a negated form with a plain one, or a to form with a from form. A
candidate is only a suggestion. Nothing enters the document until a reviewer
confirms it, and nothing was confirmed here.
It was run on a fresh read-only copy of the owner's store, taken with SQLite's backup command and migrated, with no LLM call:
particles vocab create owner --prefix own --namespace https://vocab.example.org/owner/
particles vocab propose owner --top 30 --json proposals.json
| 2026-10-02 | |
|---|---|
| Structured claims | 13,817 |
| Distinct predicates | 7,718 |
| Claims about a classed subject | 3,740 (27%) |
| Subject classes those claims fall in | 4 (artifact: record, file, symbol, command) |
| Predicate and class pairs | 2,293 |
| Normalised forms on a class | 2,173 |
| Alias candidates | 96 (79 of two forms, 17 of three to five) |
| Distinct predicates in an alias candidate | 236 (3.1% of 7,718) |
| Claims an alias candidate covers | 569 |
| Profile candidates | 2,173, of which 87 cover 5 or more claims |
| Profile candidates with a past-value partner | 7 (renumber to beside renumber from, mov to beside mov from) |
The store grew by 214 predicates since 2026-10-01. The run took 29 seconds, most of it loading the encoder.
Three readings:
- The class requirement is what limits reach. Seven claims in ten have no canonical predicate, because the subject their triple is about has no class: 7,137 claims are about a bare local subject, 2,486 name no resolved subject at all, and 427 are about a subject linked to Wikidata. Class-free clustering of the same predicates at 0.85 found 869 multi-term clusters on 2026-10-01; within a class it finds 96. Recording a Wikidata subject's class would reach only the 427. The larger lever is a class for the bare local subjects, which nothing assigns today.
- The split by direction and polarity holds. No candidate below joins
renamed towithrenamed from, or a negated form with a plain one, which were the embedding errors measured on 2026-10-01. Direction partners surface instead as a profile candidate's past-value partner. - Many candidates still need a no. Rows 10 (
passedwithpassed after), 12 (was closed aswithwas closed in), 19 (shippedwithshipped as) and 26 (states in §3.2withstates in §2.4) join forms whose preposition, or a value written into the predicate, changes the relation. A reviewer declines those, and a declined candidate is never proposed again.
The top 30 alias candidates¶
By claims covered. Members are surface predicates with their claim counts; the similarity is the lowest pairwise cosine in the candidate (average linkage can admit a pair below 0.85).
| # | Claims | Predicates | Min sim. | Class | Members (claims) |
|---|---|---|---|---|---|
| 1 | 140 | 3 | 0.85 | record | has status (136), had status (3), has status/version (1) |
| 2 | 38 | 2 | 0.89 | file | implements (37), correctly implements (1) |
| 3 | 25 | 2 | 0.89 | record | specifies (24), specifies that (1) |
| 4 | 19 | 2 | 0.91 | record | status (13), status is (6) |
| 5 | 14 | 6 | 0.88 | record | was minted for (4), was minted (3), was minted as (3), minted (2), was minted because (1), was minted in (1) |
| 6 | 13 | 3 | 0.90 | record | has spec_impact (10), has spec_impact of (2), has likely spec_impact of (1) |
| 7 | 13 | 5 | 0.92 | record | was promoted to (7), promoted to (2), was promoted as (2), is promoted to (1), was promoted on (1) |
| 8 | 10 | 2 | 0.88 | record | tracks (9), tracks question (1) |
| 9 | 8 | 3 | 0.96 | file | contains section (6), contains sections (1), will contain section (1) |
| 10 | 8 | 4 | 0.87 | file | passed after (3), passed (2), passes (2), passed with (1) |
| 11 | 8 | 3 | 0.88 | record | became active at version (4), is active at version (3), active at version (1) |
| 12 | 8 | 5 | 0.81 | record | was closed as (4), is closed as (1), was closed in (1), was closed instead of (1), was closed to (1) |
| 13 | 8 | 6 | 0.95 | record | is marked (2), marked (2), is marked as (1), was marked (1), was marked as (1), will be marked (1) |
| 14 | 7 | 3 | 0.93 | file | declares prefix (5), declares only prefixes (1), declares prefixes (1) |
| 15 | 7 | 2 | 0.89 | file | was edited (4), was edited in (3) |
| 16 | 7 | 3 | 0.85 | record | landed on (5), landed at (1), was landed on (1) |
| 17 | 7 | 2 | 0.92 | symbol | returns result block type (5), produces result block type (2) |
| 18 | 6 | 4 | 0.87 | file | has passing test count (3), has passing test count of (1), has passing tests count (1), passing test count (1) |
| 19 | 6 | 3 | 0.87 | record | shipped (4), shipped as (1), ships as (1) |
| 20 | 6 | 5 | 0.89 | record | shipped as version (2), shipped at version (1), shipped at versions (1), shipped in version (1), was shipped as version (1) |
| 21 | 6 | 2 | 0.94 | symbol | has field (4), has fields (2) |
| 22 | 5 | 2 | 0.96 | file | asserts (3), asserts on (2) |
| 23 | 5 | 4 | 0.87 | file | auto-merged (2), auto-merged cleanly during (1), auto-merged cleanly with (1), auto-merged with (1) |
| 24 | 5 | 2 | 0.95 | record | establishes rule (3), establishes rule that (2) |
| 25 | 5 | 2 | 0.93 | record | has trigger status (4), trigger status (1) |
| 26 | 5 | 4 | 0.87 | record | states in §3.2 (2), states in §2.4 (1), states in §2.9 (1), states in §7 (1) |
| 27 | 5 | 2 | 0.91 | record | trigger fired on (4), trigger fired due to (1) |
| 28 | 4 | 2 | 0.96 | file | version bumped to (2), was bumped to version (2) |
| 29 | 4 | 3 | 0.90 | file | was committed as (2), committed as (1), was committed with (1) |
| 30 | 4 | 2 | 0.85 | file | enforces (3), enforces policy (1) |
Added 2026-10-03: the reviewed profiles meet no sweep pair¶
The owner reviewed the 60 proposals on their own store (a document personal
with 16 aliases and 28 profiles). On a fresh read-only copy with that document
adopted, a measurement script (since removed with the proposal) lined the profiles up against the
1,985 same-subject pairs the update sweep would consider, before any LLM call:
| Why a pair cannot reach a profile | Pairs |
|---|---|
| One side has no structured claim | 1,407 |
| Subject has no class | 388 |
| The two claims are about different or unresolved subjects | 128 |
| Classed, but the predicate has no profile | 57 |
| One side profiled, the other a different predicate | 5 |
| Both sides under one profile | 0 |
On this store the profiles would therefore save no check and prevent no
retirement, whatever the checks answer. The predicates with the most claims, which
the proposal step ranks first, rarely appear on both sides of a sweep pair:
the sweep mostly pairs two different relations about one record or file. The
only classed pairs that share a predicate are restates on files (6) and
self-certifies on commands (3).
As a ceiling test, the only two shared predicates were profiled on the copy
too, and the update checks were run on all 1,985 pairs (claude-haiku-4-5,
about US$3.2). The contradiction check confirmed 71 pairs, so saving 4 checks
would have met the 5% bar. Every one of the 9 pairs the two profiles cover is
the same claim worded twice, which the contradiction check does not confirm.
The profiles saved no check and prevented no retirement.
The same bound was then taken on the owner's older development store, whose claims come from Wikipedia articles, Reddit posts and journal entries. Its claims had no structured triples and its subjects no classes, so on a copy the claims in the sweep pairs about a Wikidata-linked subject were structured and those subjects classed by their Wikidata "instance of" value. Of the 3,407 pairs the sweep would consider, 36 then had triples about one classed subject, and in 35 of them the two claims use different predicates. The 6 pairs a profile could reach, with reviewed aliases, all state one fact twice ("Tim Berners-Lee invented the web" and "is the creator of the Web"), and the contradiction check confirms none of them.
The top 30 clusters¶
By claims covered, members by claim count, at most eight shown per cluster.
normalise¶
| # | Claims | Terms | Members (claims) |
|---|---|---|---|
| 1 | 359 | 12 | is (257), was (41), are (31), is a (10), were (6), has been (4), can be (2), must be (2), +4 more |
| 2 | 239 | 3 | has status (224), had status (8), have status (7) |
| 3 | 172 | 3 | contains (153), contained (14), contain (5) |
| 4 | 149 | 3 | is located at (140), located at (8), are located at (1) |
| 5 | 141 | 4 | requires (131), require (7), required (2), would require (1) |
| 6 | 113 | 4 | passed (60), passes (49), must pass (3), pass (1) |
| 7 | 103 | 7 | uses (86), use (8), used (4), must use (2), has used (1), should use (1), will use (1) |
| 8 | 101 | 3 | covers (97), covered (3), cover (1) |
| 9 | 99 | 4 | implements (95), implemented (2), implement (1), is implemented (1) |
| 10 | 66 | 4 | includes (55), include (5), included (3), must include (3) |
| 11 | 63 | 4 | has (49), had (9), have (4), would have (1) |
| 12 | 59 | 3 | returns (51), returned (7), must return (1) |
| 13 | 53 | 2 | concerns (52), concern (1) |
| 14 | 52 | 2 | has version (45), had version (7) |
| 15 | 47 | 3 | lacks (43), lack (2), lacked (2) |
| 16 | 44 | 3 | defines (42), define (1), defined (1) |
| 17 | 44 | 6 | produced (24), produces (16), can produce (1), is producing (1), produce (1), will produce (1) |
| 18 | 41 | 2 | addresses (36), addressed (5) |
| 19 | 41 | 4 | reports (21), reported (18), must report (1), report (1) |
| 20 | 36 | 3 | added (20), adds (15), will add (1) |
| 21 | 31 | 1 | has model ID (31) |
| 22 | 31 | 1 | has title (31) |
| 23 | 29 | 2 | provides (28), provide (1) |
| 24 | 29 | 5 | runs (19), ran (5), is running (2), was run (2), will run (1) |
| 25 | 28 | 6 | carries (16), carry (6), carried (3), can carry (1), carries a (1), must carry (1) |
| 26 | 27 | 2 | result (26), results (1) |
| 27 | 27 | 1 | specifies (27) |
| 28 | 26 | 5 | remains (17), remain (5), remained (2), must remain (1), should remain (1) |
| 29 | 25 | 1 | has URL (25) |
| 30 | 25 | 4 | is located in (21), are located in (2), located in (1), was located in (1) |
embed@0.85¶
| # | Claims | Terms | Members (claims) |
|---|---|---|---|
| 1 | 257 | 1 | is (257) |
| 2 | 231 | 2 | has status (224), have status (7) |
| 3 | 155 | 2 | contains (153), contains only (2) |
| 4 | 149 | 3 | is located at (140), located at (8), are located at (1) |
| 5 | 132 | 2 | requires (131), requires only (1) |
| 6 | 98 | 2 | covers (97), cover (1) |
| 7 | 96 | 2 | implements (95), correctly implements (1) |
| 8 | 94 | 2 | uses (86), use (8) |
| 9 | 68 | 4 | passed (60), passed with (4), passed after (3), passed for (1) |
| 10 | 60 | 2 | includes (55), include (5) |
| 11 | 53 | 2 | concerns (52), concern (1) |
| 12 | 52 | 3 | passes (49), passes in (2), passes with (1) |
| 13 | 51 | 1 | returns (51) |
| 14 | 49 | 1 | has (49) |
| 15 | 47 | 3 | has version (45), has available version (1), has release version (1) |
| 16 | 45 | 6 | was renumbered from (20), was renumbered to (16), were renumbered to (6), renumbered from (1), renumbered to (1), was not renumbered to (1) |
| 17 | 43 | 1 | lacks (43) |
| 18 | 42 | 1 | defines (42) |
| 19 | 41 | 1 | was (41) |
| 20 | 40 | 2 | produced (24), produces (16) |
| 21 | 37 | 2 | status (25), status is (12) |
| 22 | 36 | 1 | addresses (36) |
| 23 | 36 | 8 | was bumped to version (11), bumped version to (6), version bumped from (6), version bumped to (5), was re-bumped to version (3), bumped to version (2), re-bumped version to (2), was bumped from version (1) |
| 24 | 31 | 1 | are (31) |
| 25 | 31 | 1 | has model ID (31) |
| 26 | 31 | 1 | has title (31) |
| 27 | 29 | 7 | was pushed to (18), pushed to (5), were pushed to (2), pushed (1), pushed with (1), was pushed (1), was pushed with (1) |
| 28 | 28 | 1 | provides (28) |
| 29 | 28 | 2 | specifies (27), specifies that (1) |
| 30 | 27 | 14 | has passing test count (8), had passing test count of (3), has passing test count of (3), passed test count (2), passing test count (2), had passing test count (1), had passing tests count (1), had passing tests count of (1), +6 more |
embed@0.75¶
| # | Claims | Terms | Members (claims) |
|---|---|---|---|
| 1 | 278 | 7 | has status (224), status (25), status is (12), had status (8), have status (7), reached status (1), status marked as (1) |
| 2 | 264 | 2 | is (257), is not (7) |
| 3 | 180 | 12 | contains (153), contains function (10), contain (5), contains only (2), contains test (2), contains tests for (2), contains check (1), contains class (1), +4 more |
| 4 | 149 | 3 | is located at (140), located at (8), are located at (1) |
| 5 | 133 | 3 | requires (131), requires no (1), requires only (1) |
| 6 | 109 | 4 | implements (95), implements verb (12), correctly implements (1), implements method (1) |
| 7 | 98 | 2 | covers (97), cover (1) |
| 8 | 94 | 2 | uses (86), use (8) |
| 9 | 72 | 6 | passed (60), passed with (4), passed after (3), passed on (3), passed for (1), passed in (1) |
| 10 | 63 | 3 | includes (55), include (5), does not include (3) |
| 11 | 62 | 3 | has (49), had (9), have (4) |
| 12 | 60 | 6 | passes (49), passes on (6), passes in (2), pass (1), passes through (1), passes with (1) |
| 13 | 56 | 4 | returns (51), always returns (3), never returns (1), returns to (1) |
| 14 | 53 | 2 | concerns (52), concern (1) |
| 15 | 52 | 3 | defines (42), is defined as (9), defined (1) |
| 16 | 51 | 26 | merged (8), is merged into (5), was merged via (5), merged into (3), merged via (3), was merged (3), was merged into (3), was merged on (3), +18 more |
| 17 | 47 | 3 | has version (45), has available version (1), has release version (1) |
| 18 | 47 | 3 | lacks (43), lack (2), lacked (2) |
| 19 | 47 | 8 | was renumbered from (20), was renumbered to (16), were renumbered to (6), renumbered from (1), renumbered to (1), version renumbered to (1), was not renumbered to (1), were renumbered, count (1) |
| 20 | 45 | 12 | was bumped to version (11), bumped version to (6), version bumped from (6), version bumped to (5), was bumped to (5), was re-bumped to version (3), bumped to version (2), re-bumped version to (2), +4 more |
| 21 | 43 | 2 | was (41), was not (2) |
| 22 | 42 | 4 | produced (24), produces (16), is produced by (1), is producing (1) |
| 23 | 40 | 23 | has passing test count (8), had passing test count of (3), has passing test count of (3), test suite passed count (3), has passing unit test count (2), has unit test count (2), passed test count (2), passing test count (2), +15 more |
| 24 | 38 | 8 | excludes (17), is excluded from (12), are excluded from (3), excluded from (2), excluded (1), excluded by (1), is excluded for (1), was excluded from (1) |
| 25 | 36 | 1 | addresses (36) |
| 26 | 35 | 5 | result (26), result is (4), result was (2), results in (2), results (1) |
| 27 | 34 | 3 | has model ID (31), has API model id (2), has api model id (1) |
| 28 | 34 | 9 | was moved to (11), moved to (7), moved from (6), was moved from (3), moved by (2), was moved into (2), moved (1), was moved by (1), +1 more |
| 29 | 33 | 4 | supports (23), supports option (5), supports only (3), support (2) |
| 30 | 33 | 9 | was committed as (18), committed as (4), was committed to (3), is committed to (2), was committed at (2), committed at (1), had committed (1), was committed on (1), +1 more |