Lint and review¶
The lint operation is the SDK's hygiene tool. Run it routinely; it surfaces problems you'd otherwise discover in queries.
Migration: lint is read-only by default (0.45.0)
As of 0.45.0, particles lint no longer applies status transitions
by default; it reports and leaves the store untouched. The
"Status transition (with --fix)" column below fires only when
you pass --fix. Any cron job or monitoring script that relied on
the old implicit auto-fix must now add --fix explicitly to keep
transitioning STALENESS / RETRACTION_CASCADE / CORPUS_LINK_INTEGRITY
particles. The same flip applies to POST /lint (fix now defaults
to false).
What lint catches¶
The headline findings (the statuses in the third column are defined in User guide → concepts → status):
| Finding | Cause | Status transition (with --fix) |
|---|---|---|
STALENESS |
A particle's valid_until has passed |
PROVENANCE_STALE (reason VALIDITY_EXPIRED); next reindex re-extracts |
RETRACTION_CASCADE |
A particle's provenance chain includes a RETRACTED / SUPERSEDED particle |
PROVENANCE_STALE (reason RETRACTED_DEPENDENCY) |
CORPUS_LINK_INTEGRITY |
A particle references a snapshot that no longer exists | PROVENANCE_STALE (reason CORPUS_ENTRY_MISSING) |
CONTRADICTION (with --semantic) |
Two ACTIVE truth-apt particles semantically contradict (LLM-judged), including claims from different sources. Candidate pairs are gated by embedding similarity (lint.contradiction_candidate_threshold, default 0.6) so the store-wide check does not pay an O(n²) LLM cost. |
Report-only; resolve via particles review. Lint never creates the INCONSISTENCY wrapper itself; the §6.6 extraction-time ladder does. |
NO_SUBJECT |
An ACTIVE CLAIM particle has zero subjects (§6.7 says it SHOULD have ≥ 1), e.g. an import or extraction that could not resolve any subject. Zero-subject claims land in the store rather than being rejected, but are unreachable by subject-filtered query. The §9 populations that legitimately have none are excluded: DOCUMENT_META claims, non-asserted (DECLINED / HYPOTHETICAL) claims, and claims marked extraction:subject_scope = SELF, i.e. journal claims about the author, whose subject the privacy gate withholds. Same predicate as the conformance subject_ids floor |
Surfaced for manual decision (re-extract, link a subject, or retract) |
RECENCY_DECAY |
An ACTIVE particle whose effective_confidence is materially discounted by content age alone: its source's recency_factor has fallen so that 1 - recency_factor ≥ lint.recency_decay_threshold (default 0.5). Sources with no decay config or no known publication date never fire. |
Report-only WARNING; never flips status (age decay is a recoverable discount, not a provenance break). Re-fetch / reindex if a fresher source version exists. |
Beyond these, lint reports coverage and quality diagnostics
(ORPHAN, PHANTOM_SUBJECT, LOW_COVERAGE_SUBJECT,
CONFIDENCE_DECAY, GRANULARITY_VIOLATION_CANDIDATE,
PENDING_EXTRACTION, SCHEMA_VERSION_MISMATCH,
WIKIDATA_LINK_MISMATCH, BARE_PROPERTIES_KEY, …), all surfaced for manual decision; use
--verbose --category <type> to inspect one category in full.
Two INFO findings concern the adjudicability default, the
assertion_modality that decides whether the write path may arbitrate a claim:
MODALITY_CLASSIFIER_STALE: one finding per classifier rule other than today's, with how many ACTIVE claims it set. A claim extracted before the stamp existed readslegacy-extraction. Runparticles modality --dry-runto size the backlog andparticles modalityto reclassify it. Claims from the journal extractor are left to re-extraction.MODALITY_GRANT_PENDING: one finding per ACTIVE claim thatparticles modalitywould have made adjudicable. Regeneration never makes that change unattended; the finding names theparticles particle reclassifycommand that would.MODALITY_LENS_DIVERGENCE: one finding per ACTIVE claim that an adopted lens'smodality_rulesread differently from its stored default. A lens never changes what the store arbitrates; the finding names theparticles particle reclassifycommand that would.
uv run particles lint # read-only: structural report, mutates nothing
uv run particles lint --semantic # adds LLM contradiction check (costs tokens)
uv run particles lint --fix # apply auto-fixable status transitions
The review workflow¶
An INCONSISTENCY particle is created when the §6.6 extraction-time
ladder finds a new candidate conflicting with an existing claim it
cannot out-rank on trust. The losing candidate is persisted alongside
it, quarantined (status PROVENANCE_STALE with
status_reason = CONFLICT_PENDING, invisible to query),
so review can recover it in full rather than from an excerpt.
uv run particles review # list pending conflicts
uv run particles review <particle-id> --action PREFER_A
uv run particles review --bulk BOTH_VALID --dry-run # preview a bulk action
Five resolution actions:
| Action | Effect |
|---|---|
PREFER_A |
The existing claim wins. The challenger is demoted (quarantined claims flip their reason to CONFLICT_RESOLVED in place); a reviewer-derived SourceTrustStatement for the preferred source is written, keyed on the corpus entry of the preferred claim's SOURCE provenance (a claim with no source provenance writes none). |
PREFER_B |
The challenger wins. The existing claim is demoted to PROVENANCE_STALE; a quarantined challenger is promoted to a new ACTIVE particle (fresh ID, provenance preserved); the trust statement is written. |
BOTH_VALID |
The contradiction is apparent, not real. Both claims stay queryable with uncertainty_nature = ALEATORY; a quarantined challenger is recovered as a new ACTIVE particle. |
DEFER |
Record a reviewer note and re-queue; the only action that leaves the conflict open. |
DISCARD |
Neither claim is worth keeping, for example two transient session-state claims from one conversation. Both claims are retracted (RETRACTED / CONFLICT_RESOLVED, a quarantined challenger included) and no trust statement is written. The retirement is not a verdict on the value, so a later restatement is not held for review; to keep a value out for good, use particles particle retract instead. |
Retired-value records. Some INCONSISTENCY records are not a
conflict between two live claims but a re-assertion of a claim you (or a
review) already retired: a source still says a value that was retracted or
superseded by judgment, and the pipeline held the new copy for you instead of
re-minting it. Their headline reads "a candidate re-asserts a claim retired by
judgment" and Particle A is the retired original. Read the actions as:
PREFER_A: the retirement stands; PREFER_B: lift it (a fresh ACTIVE
particle is minted from the held copy; the original stays retired);
DISCARD: let this copy go without ruling again (the original keeps its
retirement, so the value is still held if restated).
particles review --bulk DISCARD retracts both sides of every open conflict,
with no undo. It lists each conflict with both claims and asks before it
writes anything; --dry-run prints the list only, and --yes skips the
prompt for scripted use. Neither
writes a trust statement or triggers a cascade, because the question was
about a value, not a source. Set extraction.retired_value_quarantine.enabled:
false to restore the pre-0264 behaviour (re-assertions re-enter ACTIVE).
Every non-DEFER resolution retracts the INCONSISTENCY wrapper
itself (reason CONFLICT_RESOLVED), so resolved conflicts leave the
queue; particles review lists only what is still pending. Each
resolution also writes a REVIEW audit particle and a REVIEW_RESOLVED
event.
The SourceTrustStatements accumulated from PREFER rulings feed the
trust cascade and the query-time source-trust factor; see
Tuning → source trust rank.
One §6.6 verdict never reaches review: SUPERSEDED_BY_EXISTING (the
candidate duplicates a strictly higher-trust existing claim) drops the
candidate at extraction time. The drop is audited: a
CONFLICT_CANDIDATE_DROPPED event records the candidate excerpt, the
verdict, and the winning particle ID
(see Auditing).
Rulings on replaced claims¶
An update can retire a claim in favour of a later one: a new price, a new
employer. The re-anchor pass can too, when it restates a claim whose basis
changed. The curation queue shows each such
retirement whose replacing claim is on record as a demotion card, with
both claims quoted:
uv run particles curate --kind demotion
uv run particles curate apply affirm <key> # the replacement was right
uv run particles curate apply dismiss <key> # both claims hold
Neither gesture changes a status. A retirement records a judgment, so it is not reversible: dismissing the card leaves the retired claim retired. If the retired claim is still true, assert it again.
What the gesture adds is a labelled pair. Each affirm or dismiss appends
one line to demotion-rulings.jsonl in benchmark.runs_dir (default
~/.particles/benchmark/runs/), and the gesture's output says so ("Recorded
as a benchmark fixture."). A line holds both claims' texts and content
hashes, the subject, the demotion reason, the probe answers that produced
the retirement when they were recorded, your ruling (coexist for a
dismiss, replacement for an affirm), who made it and when. Snooze records
nothing. A later ruling on the same pair is appended, and readers take the
newest. The HTTP gestures do the same: POST /curation/affirm and a
permanent POST /curation/snooze (no snooze_days) return the disclosure in
benchmark_fixture.
particles benchmark rot --arm probe reads that file and re-asks the update
checks about every pair you ruled on, in a section of its own (see
Memory rot: real pairs).
Privacy. The file holds claim text from your own store, in plain JSON,
under your home directory. Nothing sends it anywhere: the benchmark reads it
locally, and its probe arm sends each pair's two claims to your configured
semantic_lint model, as the update sweep itself does. Recording is on by
default because the file stays local. To stop writing it:
A retirement whose replacing claim was never recorded has no card. Document supersession never records one, so a claim a document's newer version retired is not shown. Neither is an older claim that arrived after its successor and was stored already retired.
Reindex¶
When PROVENANCE_STALE particles accumulate, or when an extractor
upgrade lands:
uv run particles reindex # re-extract stale snapshots
uv run particles reindex --extractor-id <id> # re-extract particles from one extractor version
Reindex is rate-limited (default 100 extractions per minute). It respects the chunk-hash carry-forward; particles whose source chunks didn't change skip the LLM call.
Measure a version bump before sweeping it¶
An --extractor-version scope re-extracts every snapshot stamped with the old
version, at full price, though most bumps change one prompt section. Measure
first:
The estimate prints the free work plan and the sample it would take, then asks
before spending. It re-extracts a seeded sample of the snapshots the sweep
would re-run (reindex.estimate_sample_size, default 12), writes nothing, and
compares each sample's new claims with its stored ones using the same judge the
curation queue uses for duplicate pairs. It reports the share of sampled
snapshots whose claims changed, with a 95% interval, the projected number of
changed snapshots across the scope, and the cost of the full sweep projected
from what the sample measured. Re-extracting unchanged text also varies from
run to run, so read the share as an upper bound. In a script, pass --yes;
without it a non-interactive run exits 2 and spends nothing.
From there, run the sweep, narrow it with --entry-ids, or skip it. Every
extraction now records which prompt sections, chunking path, vision channel and
subject gate it reached, with a hash of each one's text, so a later release can
re-run only the snapshots a bump could change. That selection,
--only-changed-components, is refused until a version bump has run over a
store whose snapshots carry the record.
Append-only sources¶
A session transcript or an append-only archive is deposited again at every harvest, each snapshot a longer copy of the last, and extraction reads each snapshot as a delta: only the text after the point the previous extracted snapshot was read to. An extractor upgrade therefore does not reach a transcript's earlier text by itself. Reindex is how it does.
Reindex works on an append-only entry as a whole, never on one of its
snapshots. When the scope reaches such an entry, by particles reindex
--entry-ids <entry> or through --extractor-version, --extractor-id or
--provider-model, it:
- retires every ACTIVE extractor claim with a source reference to the
entry (
SUPERSEDED_BY_REINDEX), except the restatements the re-anchor pass wrote, and only once every step below has succeeded; - replays the entry's extracted snapshots in the order they were captured: the first read whole, each later one as a delta from the one before it.
Each claim then cites the snapshot, and the time, that first contained its passage. The replay costs one call per snapshot plus the new text of each, so an entry with many snapshots costs more to reindex than a single read of its latest one. The plan line counts the replayed snapshots. A replay that fails partway retires nothing; run the same reindex again and it starts the entry from the beginning.
A pending or failed snapshot the auto-discovery scope finds is extracted as the delta it is, not replayed.
The EXTRACTOR_VERSION this keys on is set by the extractor's author; see
Plugin-author guide → extractors for
when a bump is required.
Quality reports¶
Prints the extraction-quality dashboard (calibration-source distribution, corpus snapshot status, subject coverage), useful for "is the store growing healthy or accumulating staleness?" No LLM calls; instant read from the DB.
How often to run lint¶
- After every batch deposit + extract;
lintcatches structural errors immediately. - Weekly for
--semantic; it's the LLM-expensive variant. - Before every export; exports already run an implicit
lint --semantic=Falsepre-pass and splice findings into the exported output (per-article callouts in wiki / Obsidian).
You do not have to remember any of that: particles memory consolidate
runs lint alongside the other maintenance passes on a schedule; see
Scheduled consolidation.
Fixing a misjoined Subject¶
The Subject resolver is fast and usually right, but when two real-world
entities share a name or acronym it can silently land them on one
Subject. The classic failure: "AAOI", both the Accounting and
Auditing Organization for Islamic Financial Institutions
(Wikidata Q5326167) and Applied Optoelectronics (Wikidata
Q30297735). Every particle from either gets bound to the first
Subject the resolver picked, and the wiki article ends up confidently
claiming the audit organisation has a Texas manufacturing facility.
particles subjects split re-binds the wrongly-attributed particles
onto a new Subject the resolver canonicalises against the available
external KBs. The verb is metadata-only; particle confidence and
content are unchanged.
# Inspect the misjoined Subject to find the particles that belong elsewhere.
uv run particles subjects show 1a2b3c4d
# Dry-run the split; confirms the new Subject the resolver would create.
uv run particles subjects split 1a2b3c4d \
--particle pid-aaaa1111 \
--particle pid-bbbb2222 \
--new-name "Applied Optoelectronics" \
--dry-run
# Apply.
uv run particles subjects split 1a2b3c4d \
--particle pid-aaaa1111 \
--particle pid-bbbb2222 \
--new-name "Applied Optoelectronics"
If you already know the correct external identifier and want to
sidestep the resolver's search (avoiding the same wrong-match that
created the problem), pass --new-external-id instead:
uv run particles subjects split 1a2b3c4d \
--particle pid-aaaa1111 --particle pid-bbbb2222 \
--new-external-id wikidata:Q30297735
What to run next:
particles export <format>: re-render Obsidian / wiki / Logseq with the corrected attribution. The synthesis cache hits on every other Subject and re-synthesises only the source + new Subjects.particles lint(structural): surfaces any now-staleCO_EVIDENTIALorCONTRADICTSrelations that the split may have invalidated.particles query <topic>: sanity-check that the corrected subject filter returns what you expected.
particles extract --all-pending is not the next step. The
split is metadata-only; extraction had nothing to do with the
mis-binding the operator just corrected.
The source Subject is preserved with its remaining particles even if every particle was moved off it; empty Subjects survive the split for audit-trail reasons.