Corpus vs repo_id
-
Corpus-First
A corpus is any folder you index: repo, docs, or subtree.
-
Isolation
Each corpus has separate Postgres tables, Neo4j DB, and config.
-
Naming Migration
API accepts
repo_idbut serializes ascorpus_id.
Best Practice
Use stable, lowercase slugs for corpus ids, e.g., tribrid, myapp-docs. Avoid spaces and special characters.
AliasChoices
Pydantic models specify validation_alias=AliasChoices("repo_id", "corpus_id") and serialization_alias="corpus_id" to ensure forward compatibility.
Cross-Corpus Leakage
Never mix corpus_id across requests. Isolation is enforced in storage and graph layers.
Runtime-managed corpora carry a typed internal flag
Some corpora are registered by the runtime, not by an operator: the chat Recall corpus (recall_default) and the Codex session corpora. Each carries a meta.system_kind marker, and the Corpus wire model now derives a typed internal flag from it (server/api/repos.py _corpus_from_row; see tests/unit/test_domain_models.py). These corpora index through their own path and never have an operator-run index run, so operator surfaces exclude them: the Dashboard's Recent Index Runs panel filters them out (they would otherwise read "never indexed" forever), and "delete all unindexed corpora" cleanup skips them via internal instead of a hardcoded recall_default check. See UI tour.
Deleting a corpus takes its dead runs with it
DELETE /api/corpora/{corpus_id} removes the corpus's own rows and the staging rows (__staging__<corpus>__<run>) its index runs wrote before promotion — both in one transaction, through the shared staging sweep (_delete_corpus_staging_rows in server/db/postgres.py). A run that died mid-index therefore cannot leave orphan rows behind a deleted corpus, and the sweep is prefix-scoped so a sibling corpus whose id merely starts with the same text (corpus a vs a__b) is never touched.
Conditional cleanup refuses to delete a corpus that finished indexing
DELETE /api/corpora/{corpus_id} (and /api/repos/{corpus_id}) accepts ?only_unindexed=true: the deletion is then refused with a typed 409 corpus_already_indexed — naming the corpus and the last_indexed timestamp read under the corpus write lock — when an index has completed, even if it committed after the caller listed the corpora. The UI's "delete all unindexed corpora" sends this guard, refreshes the registry before proposing any candidate (and aborts when that refresh fails), and skips a raced corpus instead of erroring, reconciling it on the final refresh. An explicit delete without the flag behaves exactly as before. See UI tour.
Models Using corpus_id
| Model | Fields |
|---|---|
IndexRequest | corpus_id, repo_path, force_reindex |
IndexStatus | corpus_id, status, progress, current_file |
SearchRequest | corpus_id, query, top_k |
flowchart LR
UI["UI"] --> API["API"]
API --> Pyd["AliasChoices(repo_id, corpus_id)"]
Pyd --> Store["Serialized as corpus_id"] Example Requests
import httpx
req = {"corpus_id": "tribrid", "repo_path": "/code/tribrid", "force_reindex": False}
httpx.post("http://127.0.0.1:8012/api/index", json=req)
httpx.get("http://127.0.0.1:8012/api/index/tribrid/status")
curl -sS -X POST http://127.0.0.1:8012/api/index -H 'Content-Type: application/json' -d '{
"corpus_id": "tribrid", "repo_path": "/code/tribrid", "force_reindex": false
}'
import type { IndexRequest } from "../../web/src/types/generated";
const req: IndexRequest = { corpus_id: 'tribrid', repo_path: '/code/tribrid', force_reindex: false };
Multi-Corpus UIs
Add a repo switcher bound to corpus_id. All panels (RAG, Graph, Index) should update in lockstep.