Skip to content

Lint and review

The lint operation is the SDK's hygiene tool. Run it routinely; it surfaces problems you'd otherwise discover in queries.

Migration: lint is read-only by default (0.45.0)

As of 0.45.0, particles lint no longer applies status transitions by default; it reports and leaves the store untouched. The "Status transition (with --fix)" column below fires only when you pass --fix. Any cron job or monitoring script that relied on the old implicit auto-fix must now add --fix explicitly to keep transitioning STALENESS / RETRACTION_CASCADE / CORPUS_LINK_INTEGRITY particles. The same flip applies to POST /lint (fix now defaults to false).

What lint catches

The headline findings (the statuses in the third column are defined in User guide → concepts → status):

Finding Cause Status transition (with --fix)
STALENESS A particle's valid_until has passed PROVENANCE_STALE (reason VALIDITY_EXPIRED); next reindex re-extracts
RETRACTION_CASCADE A particle's provenance chain includes a RETRACTED / SUPERSEDED particle PROVENANCE_STALE (reason RETRACTED_DEPENDENCY)
CORPUS_LINK_INTEGRITY A particle references a snapshot that no longer exists PROVENANCE_STALE (reason CORPUS_ENTRY_MISSING)
CONTRADICTION (with --semantic) Two ACTIVE truth-apt particles semantically contradict (LLM-judged), including claims from different sources. Candidate pairs are gated by embedding similarity (lint.contradiction_candidate_threshold, default 0.6) so the store-wide check does not pay an O(n²) LLM cost. Report-only; resolve via particles review. Lint never creates the INCONSISTENCY wrapper itself; the §6.6 extraction-time ladder does.
NO_SUBJECT An ACTIVE CLAIM particle has zero subjects (§6.7 says it SHOULD have ≥ 1), e.g. an import or extraction that could not resolve any subject. Zero-subject claims land in the store rather than being rejected, but are unreachable by subject-filtered query. The §9 populations that legitimately have none are excluded: DOCUMENT_META claims, non-asserted (DECLINED / HYPOTHETICAL) claims, and claims marked extraction:subject_scope = SELF, i.e. journal claims about the author, whose subject the privacy gate withholds. Same predicate as the conformance subject_ids floor Surfaced for manual decision (re-extract, link a subject, or retract)
RECENCY_DECAY An ACTIVE particle whose effective_confidence is materially discounted by content age alone: its source's recency_factor has fallen so that 1 - recency_factor ≥ lint.recency_decay_threshold (default 0.5). Sources with no decay config or no known publication date never fire. Report-only WARNING; never flips status (age decay is a recoverable discount, not a provenance break). Re-fetch / reindex if a fresher source version exists.

Beyond these, lint reports coverage and quality diagnostics (ORPHAN, PHANTOM_SUBJECT, LOW_COVERAGE_SUBJECT, CONFIDENCE_DECAY, GRANULARITY_VIOLATION_CANDIDATE, PENDING_EXTRACTION, SCHEMA_VERSION_MISMATCH, WIKIDATA_LINK_MISMATCH, BARE_PROPERTIES_KEY, …), all surfaced for manual decision; use --verbose --category <type> to inspect one category in full.

Two INFO findings concern the adjudicability default, the assertion_modality that decides whether the write path may arbitrate a claim:

  • MODALITY_CLASSIFIER_STALE: one finding per classifier rule other than today's, with how many ACTIVE claims it set. A claim extracted before the stamp existed reads legacy-extraction. Run particles modality --dry-run to size the backlog and particles modality to reclassify it. Claims from the journal extractor are left to re-extraction.
  • MODALITY_GRANT_PENDING: one finding per ACTIVE claim that particles modality would have made adjudicable. Regeneration never makes that change unattended; the finding names the particles particle reclassify command that would.
  • MODALITY_LENS_DIVERGENCE: one finding per ACTIVE claim that an adopted lens's modality_rules read differently from its stored default. A lens never changes what the store arbitrates; the finding names the particles particle reclassify command that would.
uv run particles lint                # read-only: structural report, mutates nothing
uv run particles lint --semantic     # adds LLM contradiction check (costs tokens)
uv run particles lint --fix          # apply auto-fixable status transitions

The review workflow

An INCONSISTENCY particle is created when the §6.6 extraction-time ladder finds a new candidate conflicting with an existing claim it cannot out-rank on trust. The losing candidate is persisted alongside it, quarantined (status PROVENANCE_STALE with status_reason = CONFLICT_PENDING, invisible to query), so review can recover it in full rather than from an excerpt.

uv run particles review                              # list pending conflicts
uv run particles review <particle-id> --action PREFER_A
uv run particles review --bulk BOTH_VALID --dry-run  # preview a bulk action

Five resolution actions:

Action Effect
PREFER_A The existing claim wins. The challenger is demoted (quarantined claims flip their reason to CONFLICT_RESOLVED in place); a reviewer-derived SourceTrustStatement for the preferred source is written, keyed on the corpus entry of the preferred claim's SOURCE provenance (a claim with no source provenance writes none).
PREFER_B The challenger wins. The existing claim is demoted to PROVENANCE_STALE; a quarantined challenger is promoted to a new ACTIVE particle (fresh ID, provenance preserved); the trust statement is written.
BOTH_VALID The contradiction is apparent, not real. Both claims stay queryable with uncertainty_nature = ALEATORY; a quarantined challenger is recovered as a new ACTIVE particle.
DEFER Record a reviewer note and re-queue; the only action that leaves the conflict open.
DISCARD Neither claim is worth keeping, for example two transient session-state claims from one conversation. Both claims are retracted (RETRACTED / CONFLICT_RESOLVED, a quarantined challenger included) and no trust statement is written. The retirement is not a verdict on the value, so a later restatement is not held for review; to keep a value out for good, use particles particle retract instead.

Retired-value records. Some INCONSISTENCY records are not a conflict between two live claims but a re-assertion of a claim you (or a review) already retired: a source still says a value that was retracted or superseded by judgment, and the pipeline held the new copy for you instead of re-minting it. Their headline reads "a candidate re-asserts a claim retired by judgment" and Particle A is the retired original. Read the actions as: PREFER_A: the retirement stands; PREFER_B: lift it (a fresh ACTIVE particle is minted from the held copy; the original stays retired); DISCARD: let this copy go without ruling again (the original keeps its retirement, so the value is still held if restated).

particles review --bulk DISCARD retracts both sides of every open conflict, with no undo. It lists each conflict with both claims and asks before it writes anything; --dry-run prints the list only, and --yes skips the prompt for scripted use. Neither writes a trust statement or triggers a cascade, because the question was about a value, not a source. Set extraction.retired_value_quarantine.enabled: false to restore the pre-0264 behaviour (re-assertions re-enter ACTIVE).

Every non-DEFER resolution retracts the INCONSISTENCY wrapper itself (reason CONFLICT_RESOLVED), so resolved conflicts leave the queue; particles review lists only what is still pending. Each resolution also writes a REVIEW audit particle and a REVIEW_RESOLVED event.

The SourceTrustStatements accumulated from PREFER rulings feed the trust cascade and the query-time source-trust factor; see Tuning → source trust rank.

One §6.6 verdict never reaches review: SUPERSEDED_BY_EXISTING (the candidate duplicates a strictly higher-trust existing claim) drops the candidate at extraction time. The drop is audited: a CONFLICT_CANDIDATE_DROPPED event records the candidate excerpt, the verdict, and the winning particle ID (see Auditing).

Rulings on replaced claims

An update can retire a claim in favour of a later one: a new price, a new employer. The re-anchor pass can too, when it restates a claim whose basis changed. The curation queue shows each such retirement whose replacing claim is on record as a demotion card, with both claims quoted:

uv run particles curate --kind demotion
uv run particles curate apply affirm <key>    # the replacement was right
uv run particles curate apply dismiss <key>   # both claims hold

Neither gesture changes a status. A retirement records a judgment, so it is not reversible: dismissing the card leaves the retired claim retired. If the retired claim is still true, assert it again.

What the gesture adds is a labelled pair. Each affirm or dismiss appends one line to demotion-rulings.jsonl in benchmark.runs_dir (default ~/.particles/benchmark/runs/), and the gesture's output says so ("Recorded as a benchmark fixture."). A line holds both claims' texts and content hashes, the subject, the demotion reason, the probe answers that produced the retirement when they were recorded, your ruling (coexist for a dismiss, replacement for an affirm), who made it and when. Snooze records nothing. A later ruling on the same pair is appended, and readers take the newest. The HTTP gestures do the same: POST /curation/affirm and a permanent POST /curation/snooze (no snooze_days) return the disclosure in benchmark_fixture.

particles benchmark rot --arm probe reads that file and re-asks the update checks about every pair you ruled on, in a section of its own (see Memory rot: real pairs).

Privacy. The file holds claim text from your own store, in plain JSON, under your home directory. Nothing sends it anywhere: the benchmark reads it locally, and its probe arm sends each pair's two claims to your configured semantic_lint model, as the update sweep itself does. Recording is on by default because the file stays local. To stop writing it:

benchmark:
  record_demotion_rulings: false

A retirement whose replacing claim was never recorded has no card. Document supersession never records one, so a claim a document's newer version retired is not shown. Neither is an older claim that arrived after its successor and was stored already retired.

Reindex

When PROVENANCE_STALE particles accumulate, or when an extractor upgrade lands:

uv run particles reindex                          # re-extract stale snapshots
uv run particles reindex --extractor-id <id>      # re-extract particles from one extractor version

Reindex is rate-limited (default 100 extractions per minute). It respects the chunk-hash carry-forward; particles whose source chunks didn't change skip the LLM call.

Measure a version bump before sweeping it

An --extractor-version scope re-extracts every snapshot stamped with the old version, at full price, though most bumps change one prompt section. Measure first:

uv run particles reindex --estimate --extractor-version 0.15.0

The estimate prints the free work plan and the sample it would take, then asks before spending. It re-extracts a seeded sample of the snapshots the sweep would re-run (reindex.estimate_sample_size, default 12), writes nothing, and compares each sample's new claims with its stored ones using the same judge the curation queue uses for duplicate pairs. It reports the share of sampled snapshots whose claims changed, with a 95% interval, the projected number of changed snapshots across the scope, and the cost of the full sweep projected from what the sample measured. Re-extracting unchanged text also varies from run to run, so read the share as an upper bound. In a script, pass --yes; without it a non-interactive run exits 2 and spends nothing.

From there, run the sweep, narrow it with --entry-ids, or skip it. Every extraction now records which prompt sections, chunking path, vision channel and subject gate it reached, with a hash of each one's text, so a later release can re-run only the snapshots a bump could change. That selection, --only-changed-components, is refused until a version bump has run over a store whose snapshots carry the record.

Append-only sources

A session transcript or an append-only archive is deposited again at every harvest, each snapshot a longer copy of the last, and extraction reads each snapshot as a delta: only the text after the point the previous extracted snapshot was read to. An extractor upgrade therefore does not reach a transcript's earlier text by itself. Reindex is how it does.

Reindex works on an append-only entry as a whole, never on one of its snapshots. When the scope reaches such an entry, by particles reindex --entry-ids <entry> or through --extractor-version, --extractor-id or --provider-model, it:

  1. retires every ACTIVE extractor claim with a source reference to the entry (SUPERSEDED_BY_REINDEX), except the restatements the re-anchor pass wrote, and only once every step below has succeeded;
  2. replays the entry's extracted snapshots in the order they were captured: the first read whole, each later one as a delta from the one before it.

Each claim then cites the snapshot, and the time, that first contained its passage. The replay costs one call per snapshot plus the new text of each, so an entry with many snapshots costs more to reindex than a single read of its latest one. The plan line counts the replayed snapshots. A replay that fails partway retires nothing; run the same reindex again and it starts the entry from the beginning.

A pending or failed snapshot the auto-discovery scope finds is extracted as the delta it is, not replayed.

The EXTRACTOR_VERSION this keys on is set by the extractor's author; see Plugin-author guide → extractors for when a bump is required.

Quality reports

uv run particles quality

Prints the extraction-quality dashboard (calibration-source distribution, corpus snapshot status, subject coverage), useful for "is the store growing healthy or accumulating staleness?" No LLM calls; instant read from the DB.

How often to run lint

  • After every batch deposit + extract; lint catches structural errors immediately.
  • Weekly for --semantic; it's the LLM-expensive variant.
  • Before every export; exports already run an implicit lint --semantic=False pre-pass and splice findings into the exported output (per-article callouts in wiki / Obsidian).

You do not have to remember any of that: particles memory consolidate runs lint alongside the other maintenance passes on a schedule; see Scheduled consolidation.

Fixing a misjoined Subject

The Subject resolver is fast and usually right, but when two real-world entities share a name or acronym it can silently land them on one Subject. The classic failure: "AAOI", both the Accounting and Auditing Organization for Islamic Financial Institutions (Wikidata Q5326167) and Applied Optoelectronics (Wikidata Q30297735). Every particle from either gets bound to the first Subject the resolver picked, and the wiki article ends up confidently claiming the audit organisation has a Texas manufacturing facility.

particles subjects split re-binds the wrongly-attributed particles onto a new Subject the resolver canonicalises against the available external KBs. The verb is metadata-only; particle confidence and content are unchanged.

# Inspect the misjoined Subject to find the particles that belong elsewhere.
uv run particles subjects show 1a2b3c4d

# Dry-run the split; confirms the new Subject the resolver would create.
uv run particles subjects split 1a2b3c4d \
    --particle pid-aaaa1111 \
    --particle pid-bbbb2222 \
    --new-name "Applied Optoelectronics" \
    --dry-run

# Apply.
uv run particles subjects split 1a2b3c4d \
    --particle pid-aaaa1111 \
    --particle pid-bbbb2222 \
    --new-name "Applied Optoelectronics"

If you already know the correct external identifier and want to sidestep the resolver's search (avoiding the same wrong-match that created the problem), pass --new-external-id instead:

uv run particles subjects split 1a2b3c4d \
    --particle pid-aaaa1111 --particle pid-bbbb2222 \
    --new-external-id wikidata:Q30297735

What to run next:

  • particles export <format>: re-render Obsidian / wiki / Logseq with the corrected attribution. The synthesis cache hits on every other Subject and re-synthesises only the source + new Subjects.
  • particles lint (structural): surfaces any now-stale CO_EVIDENTIAL or CONTRADICTS relations that the split may have invalidated.
  • particles query <topic>: sanity-check that the corrected subject filter returns what you expected.

particles extract --all-pending is not the next step. The split is metadata-only; extraction had nothing to do with the mis-binding the operator just corrected.

The source Subject is preserved with its remaining particles even if every particle was moved off it; empty Subjects survive the split for audit-trail reasons.