Configuration¶
All tuneable parameters live in particles/config.py as a Pydantic
model. The canonical sample is config.yaml.sample in the repo
root; copy it to config.yaml to override defaults. config.yaml
is gitignored.
Which sections apply to your install¶
Every section is tagged [client] or [engine] in config.yaml.sample.
Both are always loaded and both are valid to set; the tag says
which distribution acts on the section:
| Tag | Read by | Effective in |
|---|---|---|
[client] |
the Client layer, shipped in linkedparticles-core |
every install |
[engine] |
the Engine layer and the surfaces, shipped in linkedparticles |
the full install only |
If you installed linkedparticles (the ordinary case, and what every guide
here assumes), all 64 sections apply and the tags are informational. They
matter only for a linkedparticles-core-only install, which is the store-free
Client substrate: there an [engine] section still validates and still loads,
but nothing present reads it. That is a deliberate trade (the two
distributions share one import package, so they share one config model), and
the tag is how the inert surface is made visible rather than carved away.
The declaration lives in CLIENT_SECTIONS in particles/config.py; tests keep
it, the sample's tags, and the modules that actually read config in agreement.
Which config.yaml loads¶
Discovery walks upward, git-style:
PARTICLES_CONFIG, if set: absolute authority. A path that does not exist resolves to no config file rather than falling through, so this is also the explicit opt-out from everything below../config.yamlin the working directory.- The nearest
config.yamlin an ancestor directory; first match wins. The walk stops after examining the directory that holds a.gitentry (a file in a git worktree, a directory in a normal checkout), so aconfig.yamlin$HOMEis never inherited by the projects beneath it. - Otherwise: compiled-in defaults plus env-var overrides.
Step 3 is why a verb run from scripts/, from a git worktree, or from a
hook spawn now gets the config you wrote at the repo root. Before it, such
a process silently reverted every knob to its compiled default, the
cause behind blobs written somewhere nobody looks. particles config
validate and particles hook doctor both print the file they resolved;
when "which config am I loading?" matters, ask them rather than guess.
Precedence¶
Env-var overrides are registered in _ENV_OVERRIDES in
particles/config.py. The common ones:
| Env var | Config field | Purpose |
|---|---|---|
DATABASE_URL |
database_url |
SQLite path |
PARTICLES_BLOB_DIR |
blob_dir |
Where deposited blobs are stored |
PARTICLES_CONFIG |
none (bootstrap) | Path to a non-default config.yaml |
TRUST_DIFFERENTIAL_THRESHOLD |
trust.differential_threshold |
When trust differences flag inconsistencies |
RECONCILIATION_STORE_MODE |
reconciliation.store_mode |
single (default) or multi, the consensus-store reconciliation regime |
For the full env-var list and field-by-field tuning options, see
config.yaml.sample.
Secrets¶
Secrets never live in config.yaml. Application code calls helpers
in particles/secrets.py:
| Secret | Env var | Used by |
|---|---|---|
| Anthropic API key | ANTHROPIC_API_KEY |
Extractor, semantic lint, wiki synthesis |
| Numista API key | NUMISTA_API_KEY (optional) |
Numista extractor |
| GitHub API key | GITHUB_API_KEY (optional) |
GitHub extractor (gist / repo / pages) |
Migrating to a secrets manager (1Password, AWS, sops, …) only
requires editing particles/secrets.py.
LLM provider selection¶
Every chat/completion call routes through a CompletionProvider port.
The llm section picks the (provider, model) pairing per
purpose: a default plus optional overrides for extraction,
semantic_lint, query_response, synthesis, benchmark,
benchmark_answer, abstraction, verification (the memory audit's
second reading of each contradiction the semantic_lint probe flags; routing
semantic_lint to a small model for cost leaves that reading on the
default), and subject_resolution (the judge that picks among an ambiguous
name's Wikidata candidates during extraction):
llm:
default:
provider: anthropic
model: claude-sonnet-4-6
# Route a single purpose to a different model, e.g. a cheaper model for
# high-volume extraction while synthesis stays on the default.
# extraction:
# model: claude-haiku-4-5
provider is anthropic (the native adapter) or the name of any entry in
llm.providers (see the next section). max_tokens is not
set here; it is per-call and lives with each call site
(extraction.max_tokens, query.answer_max_tokens,
wiki.max_tokens).
Those per-call budgets are total output allowances on the wire, not
response-length caps: an extended-thinking model spends its thinking tokens
from the same number. A budget sized to the expected prose therefore returns a
reply with no text in it at all: extraction reports a truncated JSON array,
and query degrades to its deterministic belief listing with
answer_generation_error_cause: BUDGET. query retries that case once at
query.answer_retry_max_tokens before degrading; the others do not.
Named providers (any OpenAI-compatible endpoint)¶
Every non-Anthropic vendor, hosted (OpenAI, DeepSeek, Kimi, gateways) or
local (Ollama, llama.cpp's server, vLLM, LM Studio), speaks the OpenAI
chat/completions dialect, so all of them are named entries in
llm.providers: adding a vendor is a config block, never code.
The endpoint, resilience, and dialect policy are per-entry; only the
per-purpose model string lives in the purpose override. The local entry
(an Ollama endpoint) is compiled in, so provider: local works
with zero configuration.
llm:
default:
provider: anthropic
model: claude-sonnet-4-6
extraction: # send only extraction to a cheaper vendor
provider: openai
model: gpt-5.6-luna
providers:
openai:
base_url: https://api.openai.com/v1 # adapter appends /chat/completions
max_tokens_param: max_completion_tokens # reasoning models reject max_tokens
send_temperature: false # …and non-default temperatures
structured_output: strict # strict-dialect JSON schemas
local: # override the compiled-in Ollama entry if needed
base_url: http://localhost:11434/v1
timeout_seconds: 120
The API key is a secret named after the entry:
PARTICLES_LLM_API_KEY_<NAME> (PARTICLES_LLM_API_KEY_OPENAI, …), set in
the environment, never in config.yaml. The local entry also honours the
legacy PARTICLES_LOCAL_LLM_API_KEY. Endpoints that enforce no auth (bare
Ollama / llama.cpp) need no key and the adapter omits the Authorization
header.
The two dialect knobs are declarative statements about a known endpoint,
not runtime negotiation: max_tokens_param picks which body member carries
the length cap, and send_temperature: false drops temperature from
requests entirely. A wrong knob fails loudly with the endpoint's own
HTTP 400. structured_output: strict transforms JSON schemas to the
OpenAI-strict dialect (every key required, optionality as union-with-null).
It is required for api.openai.com; leave the default auto for tolerant
endpoints like Ollama.
Deprecation: the pre-0227
llm.localblock is honoured asllm.providers.localwith a warning for one release cycle; move it underproviderswhen convenient.
Reasoning models need a bigger token budget¶
A reasoning model (claude-sonnet-5, DeepSeek-V4, Kimi K3, the GPT-5.6
family) spends its thinking tokens from the same completion budget as the
answer, so a prompt that fits comfortably in 8192 tokens on a
non-reasoning model can exhaust that budget before the answer is finished,
or before it starts. The answer grows with the source as well, at 2 to 3.5
tokens per source character, so a long Claude Code memory file needs more
than 16384 tokens of reply on its own. The default extraction.max_tokens is
32000 for these reasons, and a reply that still comes back empty, cut short,
or unparseable is retried once at extraction.retry_max_tokens (default
64000). The Anthropic adapter streams any call whose budget is above the
Anthropic SDK's non-streaming ceiling of about 21333 tokens, which is what
lets the budgets go that high. An OpenAI-compatible provider has no such
ceiling, but its vendor may cap output below these defaults and reject the
request, and a long reply still has to finish within the entry's
timeout_seconds. Lower extraction.max_tokens for such a provider. The endpoint returns HTTP 200 with finish_reason:
length and text that stops mid-token, which the extractor's JSON parser
then reports as Failed to parse extraction response: Unterminated string.
The adapter logs a WARNING naming the pairing and the budget whenever a
reply comes back truncated, so the two are not confused; a truncated reply
with no text at all fails with a budget-shaped error rather than a
generic empty-response one.
Give these models headroom (extraction.max_tokens: 16384 cleared the
parse failures in the 2026-08 trial with DeepSeek-V4 flash/pro and Kimi K3), and
consider raising the entry's timeout_seconds for large models, since a long
thinking pass takes wall-clock time the default 120 s may not cover:
llm:
providers:
fireworks:
base_url: https://api.fireworks.ai/inference/v1
timeout_seconds: 300 # reasoning passes are slow as well as long
extraction:
max_tokens: 16384 # thinking + answer share this budget; at or under the vendor's output cap
Confidence calibration is per (extractor, model) pairing:
each particles extractor calibrate run stores a record keyed by the
extraction model it ran under, and the pipeline applies the one matching the
configured model. A newly pointed model, including any
<provider>:<model>, is therefore uncalibrated until you benchmark it (queries fall
back to the EXTRACTOR_DIRECT disclosure meanwhile), but switching back
to a model you calibrated before restores its calibration with no re-fit.
A record also applies only under the extractor version it was fitted under:
an extractor upgrade leaves it stored but not applied until you
re-fit. List the stored pairings, and which of them apply, with particles
extractor calibrations <extractor-id>.
Pick provider names before you benchmark. The calibration key is
<name>:<model>(the operator-chosen entry name, not the vendor), so renaming a provider entry orphans every calibration record made under the old name. Treat a rename as a recalibration event.Deprecation:
extraction.modelandwiki.modelmoved into this section. The old keys are migrated automatically (tollm.default.modelandllm.synthesis.model) with a warning for one release cycle; move them tollmwhen convenient.Note the scope change:
llm.default.modelis the fallback for every purpose, including semantic lint and the benchmark judge, which were previously hard-wired toclaude-sonnet-4-6. That means a non-defaultextraction.model(nowllm.default.model) will also drive lint and benchmark. To keep those on a cheaper model, setllm.semantic_lint.model/llm.benchmark.modelexplicitly.
Common knobs¶
A few config fields you'll likely want to set early:
llm.default.model(and per-purposellm.<purpose>.model): the completion model each purpose uses. See LLM provider selection above.subjects.wikidata_candidate_selection: how an extracted name with several Wikidata candidates is linked. The default,llm_judge, sends an ambiguous name to thellm.subject_resolutionmodel once, with the claim and each candidate's description, and links the candidate it names or none of them; the answer is recorded and reused for the same name, claim and candidates. A name with one well-matched candidate, or none, never reaches the model. Measured on two gold sets it linked more names correctly and lost no correct link, at about US$0.20 per 100 names on prose about well-known entities (the measurement).top_hittakes Wikidata's first search result with no model call, which is also whatllm_judgefalls back to when no LLM is reachable.exporter_common.min_particle_confidence: the cross-exporter quality threshold. Particles below thiseffective_confidenceare dropped from every export. Per-run override and the per-exporter flag lists: User guide → exporting.wiki.min_particles: minimum particles per subject for the wiki exporter to render. Default 3. See User guide → exporting → wiki articles.query.top_k: top-k truncation for the semantic search. What it does to a result list: User guide → ranking.embeddings.progress_bars: whether the embedding stack prints its tqdm progress bars (Loading weights …on model load,Batches …on each encode) to stderr. Defaultfalse(they are noise for a CLI verb likequery); settrue(orPARTICLES_EMBEDDINGS_PROGRESS_BARS=1) to restore them.obsidian.default_output_path: soparticles export obsidianworks without an argument.inbox.file_path: the iCloud-synced file theparticles inboxcommands read URLs from (inbox.poll_interval_secondstunes theinbox watchcadence). Setup walkthrough: User Guide → Depositing from your phone.-
reconciliation.store_mode:single(default) for a solo store, ormultifor a multi-contributor / consensus store. Inmultimode a confirmed cross-source contradiction is surfaced as an INCONSISTENCY (both claims stay ACTIVE, ranked per-viewer at query time) rather than one claim auto-superseding the other on trust; a contributor's claim is never dropped by another contributor's trust. -
extraction.append_only_delta: whether a snapshot of anAPPEND_ONLYsource (a session transcript, an append-only archive) is read as a delta, only the text it adds to the last snapshot the store extracted. Defaulttrue. Withfalse, every snapshot is read whole, and the earlier text is read again each time. Two knobs size the delta:extraction.append_context_chars(default 4000) is how much of the already-extracted text is shown before it, as context the model extracts nothing from, andextraction.append_chunk_chars(default 7500, minimum 1000) is the largest delta chunk sent in one call. When a new snapshot does not extend the previous one, the extraction reads it whole and says why in its quality notes.
With the delta on, a read of such a snapshot that fails partway keeps what
it already paid for. The chunks before the first failed call are
written, and so are later answered chunks the retry is certain to skip. The
snapshot stays pending, and its retry reads only the rest: from where the
kept chunks stop, or, for a snapshot read whole, the whole read again with
the written chunks skipped. While a whole read is held that way, the
entry's later snapshots wait for it, and particles extract and the
consolidation log name the wait. particles reindex <entry> releases it.
Changing append_chunk_chars, append_context_chars or
html_chunk_size while a snapshot is held partway makes its retry read
some kept chunks again, and the extraction log says so.
See Tuning for the trust / calibration / age-decay
knobs that drive effective_confidence.
Observer scope¶
Whether a Claude Code session is shown the whole store or its own project's view of it; see User Guide → Claude Code for what the view contains. Three switches, deliberately in three places:
claude_code.observer_scope(defaultstore):projectreads the session-start digest, theMEMORY.mdregion and the freshness check through the session's project. It takes effect only on a store thatparticles memory rescopehas run on; until then the surfaces stay store-wide and say so, andparticles hook doctorreports "NOT in effect".particles mcp serve --project-observer cwd: binds one MCP server process to the project of its working directory, for reads and writes. A launch flag, not config, becauseconfig.yamlis shared by every MCP client on the machine and a client started from your home directory should not be bound to it.observer_scope.harness_tags(default["claude-code"]): the corpus-entry tags that mark a harness deposit. A harvested source with no project key is unattributed: in view store-wide and for no project, never global. Add your adapter's tag here when you wire a second harness, or its keyless deposits will read as global.
Hygiene. particles memory rescope --dry-run is safe to run any time and
is the census: sources per project, sources left unattributed, and sources
whose only project no longer exists on this machine. It only ever adds tags.
Run it after deleting or moving a repository, and after importing a store from
another machine (project keys are local path names, and an import arrives with
none, so imported beliefs read as global).
What it costs. One batched join per scoped read. On a 28,000-belief store
the join took 0.2–0.4 s and the scoped digest rendered in about 1.9 s against
1.6 s store-wide, inside the 10 s claude_code.hook_deadline_seconds.
Agent-memory projection¶
The agent_memory.projection block governs the MEMORY.md projection for
the Claude Code integration; see User Guide → Claude
Code for the
walkthrough. The load-bearing knobs:
agent_memory.projection.enabled(defaulttrue): render + splice thememory-indexregion and run the session-start freshness check.falsefalls back to a plain digest push with no region writes.agent_memory.projection.fold_authored_lines(defaulttrue): move agent-authored lines outside the projected region into the append-only archive after each successful harvest (never destroyed).
Git-versioned projection history¶
Off by default. When enabled and the memory directory is inside a git repo, each render that changes files under it is committed with a structured message (run id + ranking-delta summary), giving you a diffable, rollback-able history of the view while the store stays the source of truth. Every git failure degrades silently (logged at debug, never raised): the commit is a bonus, the projection is the product.
agent_memory.projection.git.enabled(defaultfalse): master switch. Committing into your repo is opt-in; turning it on without the memory directory being a git repo is harmless (the step is simply inert).agent_memory.projection.git.sign(defaultfalse):falsepasses--no-gpg-signso an unattended session-end commit never blocks on a signing agent;truedrops the override and respects your owncommit.gpgsign. This SDK's own GPG requirement is never imposed on your memory repo, and a signing failure never fails the projection.agent_memory.projection.git.author_name/.author_email(defaultnull): passed per-commit via-c user.name/-c user.email(never written into your git config).nulluses your repo's own identity; when that is absent, the commit degrades silently.agent_memory.projection.git.max_delta_excerpts(default6): cap on the added/removed excerpt lines in the commit message. The count line always states the true totals, so a large delta is never silently truncated.
Only files under the memory directory are staged (never git add -A), and
the internal backup / snapshot / archive live outside it, so they never enter
your history.