Collections and vaults
The two axes of memory organization: what a collection is, what a vault is, and the rules that surprise people.
Two different axes, constantly confused:
- A collection is a content namespace inside a tenant:
team-notes,_valorbrain,baseline-2026-05-26. Documents carry it; searches can be scoped to it. - A vault is a sync source: a named directory of markdown the engine can ingest and keep indexed (
vault_sync,list_vaults). Several MCP tools take avaultparameter to scope reads.
Collections
Every document has collection + path, and path is unique within the collection — that pair, not the title, is the identity you write against. Re-POSTing the same path re-serves the same document instead of duplicating it.
Four behaviors worth knowing before you design namespaces:
- Dedup is by content hash, across collections too. The same bytes will not be indexed twice even in two collections — a deliberate curation stance, not a bug. If two collections need "the same" text, they need different content (different headers, dates, context).
- Coverage ≠ curation. Pulling an outside corpus in can hurt retrieval: thousands of near-clone documents pollute the candidate pool for every query. Ingest what your tenant actually asks about.
collection_settingsis where policy lives — includingexcludePatterns, the deliberate way to say "files in this collection are refused by policy" (curated personal data, experiment leftovers) so coverage reports show expected refusals instead of divergence._valorbrainis the default collection formemory_storewrites that don't name one — decisions, lessons, observations land there unless you say otherwise.
Endpoints: GET /collections (list with counts), GET /collections/health (per-collection embedding/lifecycle state), POST /api/v1/collection-settings/{collection} (policy), POST /api/v1/collection-settings/reindex (rebuild), POST /collections/{id}/lifecycle (run lifecycle for one collection).
Vaults
A vault maps a directory to an indexed, searchable state:
vault_sync— ingest/refresh a vault (long operation;run_async: truereturns a handle,operation_statuspolls it).list_vaults— configured vaults and their paths.- The
vaultparameter on read tools scopes retrieval to one vault's documents.
The vault does not isolate identity. Entity identity is (tenant_id, entity_id) — the tenant is the isolation boundary, not the vault; two vaults in one tenant share the entity graph. If you need hard isolation between two worlds, that is two tenants.
Choosing
| You want | Use |
|---|---|
| Group memories by topic/source for humans | Collections |
| Scope search to imported docs | Collection filter (collection on /search, /retrieve) |
| Keep a folder of markdowns in sync with the brain | A vault + vault_sync |
| Isolate one company/customer completely | A separate tenant |
Naming tips that age well: dates for baselines (baseline-2026-05-26), owner + purpose for living collections (team-notes), and never a name that implies a person is the only reader — tenant memory is shared by design, visibility: private is the per-user axis.