Concepts

Collections and vaults

The two axes of memory organization: what a collection is, what a vault is, and the rules that surprise people.

Two different axes, constantly confused:

  • A collection is a content namespace inside a tenant: team-notes, _valorbrain, baseline-2026-05-26. Documents carry it; searches can be scoped to it.
  • A vault is a sync source: a named directory of markdown the engine can ingest and keep indexed (vault_sync, list_vaults). Several MCP tools take a vault parameter to scope reads.

Collections

Every document has collection + path, and path is unique within the collection — that pair, not the title, is the identity you write against. Re-POSTing the same path re-serves the same document instead of duplicating it.

Four behaviors worth knowing before you design namespaces:

  1. Dedup is by content hash, across collections too. The same bytes will not be indexed twice even in two collections — a deliberate curation stance, not a bug. If two collections need "the same" text, they need different content (different headers, dates, context).
  2. Coverage ≠ curation. Pulling an outside corpus in can hurt retrieval: thousands of near-clone documents pollute the candidate pool for every query. Ingest what your tenant actually asks about.
  3. collection_settings is where policy lives — including excludePatterns, the deliberate way to say "files in this collection are refused by policy" (curated personal data, experiment leftovers) so coverage reports show expected refusals instead of divergence.
  4. _valorbrain is the default collection for memory_store writes that don't name one — decisions, lessons, observations land there unless you say otherwise.

Endpoints: GET /collections (list with counts), GET /collections/health (per-collection embedding/lifecycle state), POST /api/v1/collection-settings/{collection} (policy), POST /api/v1/collection-settings/reindex (rebuild), POST /collections/{id}/lifecycle (run lifecycle for one collection).

Vaults

A vault maps a directory to an indexed, searchable state:

  • vault_sync — ingest/refresh a vault (long operation; run_async: true returns a handle, operation_status polls it).
  • list_vaults — configured vaults and their paths.
  • The vault parameter on read tools scopes retrieval to one vault's documents.

The vault does not isolate identity. Entity identity is (tenant_id, entity_id) — the tenant is the isolation boundary, not the vault; two vaults in one tenant share the entity graph. If you need hard isolation between two worlds, that is two tenants.

Choosing

You wantUse
Group memories by topic/source for humansCollections
Scope search to imported docsCollection filter (collection on /search, /retrieve)
Keep a folder of markdowns in sync with the brainA vault + vault_sync
Isolate one company/customer completelyA separate tenant

Naming tips that age well: dates for baselines (baseline-2026-05-26), owner + purpose for living collections (team-notes), and never a name that implies a person is the only reader — tenant memory is shared by design, visibility: private is the per-user axis.

On this page