Skip to content

Corpus vs repo_id

  • Corpus-First


    A corpus is any folder you index: repo, docs, or subtree.

  • Isolation


    Each corpus has separate Postgres tables, Neo4j DB, and config.

  • Naming Migration


    API accepts repo_id but serializes as corpus_id.

Get started Configuration API

Best Practice

Use stable, lowercase slugs for corpus ids, e.g., tribrid, myapp-docs. Avoid spaces and special characters.

AliasChoices

Pydantic models specify validation_alias=AliasChoices("repo_id", "corpus_id") and serialization_alias="corpus_id" to ensure forward compatibility.

Cross-Corpus Leakage

Never mix corpus_id across requests. Isolation is enforced in storage and graph layers.

Runtime-managed corpora carry a typed internal flag

Some corpora are registered by the runtime, not by an operator: the chat Recall corpus (recall_default) and the Codex session corpora. Each carries a meta.system_kind marker, and the Corpus wire model now derives a typed internal flag from it (server/api/repos.py _corpus_from_row; see tests/unit/test_domain_models.py). These corpora index through their own path and never have an operator-run index run, so operator surfaces exclude them: the Dashboard's Recent Index Runs panel filters them out (they would otherwise read "never indexed" forever), and "delete all unindexed corpora" cleanup skips them via internal instead of a hardcoded recall_default check. See UI tour.

Deleting a corpus takes its dead runs with it

DELETE /api/corpora/{corpus_id} removes the corpus's own rows and the staging rows (__staging__<corpus>__<run>) its index runs wrote before promotion — both in one transaction, through the shared staging sweep (_delete_corpus_staging_rows in server/db/postgres.py). A run that died mid-index therefore cannot leave orphan rows behind a deleted corpus, and the sweep is prefix-scoped so a sibling corpus whose id merely starts with the same text (corpus a vs a__b) is never touched.

Conditional cleanup refuses to delete a corpus that finished indexing

DELETE /api/corpora/{corpus_id} (and /api/repos/{corpus_id}) accepts ?only_unindexed=true: the deletion is then refused with a typed 409 corpus_already_indexed — naming the corpus and the last_indexed timestamp read under the corpus write lock — when an index has completed, even if it committed after the caller listed the corpora. The UI's "delete all unindexed corpora" sends this guard, refreshes the registry before proposing any candidate (and aborts when that refresh fails), and skips a raced corpus instead of erroring, reconciling it on the final refresh. An explicit delete without the flag behaves exactly as before. See UI tour.

Models Using corpus_id

Model Fields
IndexRequest corpus_id, repo_path, force_reindex
IndexStatus corpus_id, status, progress, current_file
SearchRequest corpus_id, query, top_k
flowchart LR
    UI["UI"] --> API["API"]
    API --> Pyd["AliasChoices(repo_id, corpus_id)"]
    Pyd --> Store["Serialized as corpus_id"]

Example Requests

import httpx
req = {"corpus_id": "tribrid", "repo_path": "/code/tribrid", "force_reindex": False}
httpx.post("http://127.0.0.1:8012/api/index", json=req)
httpx.get("http://127.0.0.1:8012/api/index/tribrid/status")
curl -sS -X POST http://127.0.0.1:8012/api/index -H 'Content-Type: application/json' -d '{
  "corpus_id": "tribrid", "repo_path": "/code/tribrid", "force_reindex": false
}'
import type { IndexRequest } from "../../web/src/types/generated";
const req: IndexRequest = { corpus_id: 'tribrid', repo_path: '/code/tribrid', force_reindex: false };

Multi-Corpus UIs

Add a repo switcher bound to corpus_id. All panels (RAG, Graph, Index) should update in lockstep.