Skip to content

Indexing Pipeline

  • Loader


    Git-aware discovery honoring .gitignore with root-relative patterns.

  • Chunker


    Fixed, AST-aware, or hybrid chunk strategies with line attribution.

  • Embedder


    Deterministic local or provider-backed embeddings configured in Pydantic.

  • Chunk Summaries


    Optional LLM-generated chunk_summaries to improve sparse search.

  • Graph Builder


    Entity/relationship extraction and Neo4j persistence.

Get started Configuration API

Idempotent Indexing

Use force_reindex=false for incremental updates. The indexer skips unchanged files using mtime/hash checks where available.

Storage Layout

Chunks, embeddings, and FTS are in PostgreSQL. Graph artifacts are in Neo4j. Sizes are summarized via dashboard endpoints.

The pre-run estimate is measured

POST /api/index/estimate no longer divides bytes by a constant. It samples the corpus through the configured chunker (server/indexing/estimate.py), reports token/chunk bands (estimated_*_low/high, ±estimate_relative_error), what it measured (sampled_files, sampled_bytes), and a status of ready, warming (tokenizer still loading; every measured field is null) or insufficient_sample. The estimate is also the consent gate: a failed or refused estimate blocks the run instead of letting it start unpriced. Generation tokens are measured separately from the configured tokenizer's units, and the semantic KG input is priced from the fully rendered extraction request. See Indexing API and Indexing a corpus.

Large Corpora

Configure Neo4j heap and page cache via environment for multi-million edge graphs. Monitor Postgres disk growth for pgvector indexes.

Pipeline Flow

flowchart LR
    L["FileLoader"] --> C["Chunker"]
    C --> E["Embedder"]
    E --> P["PostgreSQL"]
    C --> S["ChunkSummarizer"]
    S --> P
    C --> GB["GraphBuilder"]
    GB --> N["Neo4j"]

Chunking & Embedding Controls (Selected)

Section Field Default Notes
chunking chunk_size 1000 Target chars per chunk
chunking chunk_overlap 200 Overlap for continuity
chunking chunking_strategy ast ast \| greedy \| hybrid
chunking max_chunk_tokens 8000 Split recursively if larger
embedding embedding_type openai Provider selector
embedding embedding_model text-embedding-3-large Model id
embedding embedding_dim 3072 Must match model outputs
indexing bm25_tokenizer stemmer Tokenizer for FTS
indexing figures.* off Describe charts/drawings in Docling-converted PDFs via a vision alias (indexing.figures)

Start Indexing via API (Annotated)

import httpx
base = "http://127.0.0.1:8012/api"

req = {
    "corpus_id": "tribrid",   # (1)!
    "repo_path": "/work/src/tribrid",
    "force_reindex": False
}
httpx.post(f"{base}/index", json=req).raise_for_status()  # (2)!

status = httpx.get(f"{base}/index/tribrid/status").json()
print(status["status"], status.get("progress"))          # (3)!
  1. Create/refresh a specific corpus
  2. Start indexing
  3. Poll progress
BASE=http://127.0.0.1:8012/api
curl -sS -X POST "$BASE/index" -H 'Content-Type: application/json' -d '{
  "corpus_id":"tribrid","repo_path":"/work/src/tribrid","force_reindex":false
}'
curl -sS "$BASE/index/tribrid/status" | jq .
import type { IndexRequest, IndexStatus } from "./web/src/types/generated";

async function reindex(path: string) {
  const req: IndexRequest = { corpus_id: "tribrid", repo_path: path, force_reindex: false } as any;
  await fetch("/api/index", { method: "POST", headers: {"Content-Type":"application/json"}, body: JSON.stringify(req) }); // (2)!
  const status: IndexStatus = await (await fetch("/api/index/tribrid/status")).json(); // (3)!
  console.log(status.status, status.progress);
}

Graph Indexing (Neo4j)

Field Default Meaning
graph_indexing.enabled true Enable graph building during indexing
graph_indexing.build_lexical_graph true Build the lexical Document/Chunk graph per file
graph_indexing.build_code_graph false Select the AST code-graph policy for a code corpus
graph_indexing.semantic_kg_max_chunks 40000 Chunk ceiling for a semantic run; exceeding it fails the run instead of slicing a partial graph
graph_indexing.semantic_kg_llm_model "" Optional LiteLLM alias for GraphRAG semantic extraction; empty uses the gateway default
graph_indexing.semantic_kg_reasoning_effort medium Reasoning effort for Responses-compatible extraction models
graph_indexing.semantic_kg_llm_timeout_s 90 Per-chunk extraction timeout (seconds)

Timeout and reasoning effort now bind to the extraction LLM

Both knobs are carried by the actual GraphRAG extraction calls (semantic_extraction_llm in server/indexing/graphrag_pipeline.py): graph_indexing.semantic_kg_llm_timeout_s is the OpenAI client's per-request timeout, so every chunk extraction is bounded by it, and graph_indexing.semantic_kg_reasoning_effort is sent on every extraction request in the alias's upstream protocol (see the note below). A pipeline built with a non-positive timeout or a blank effort refuses fast with a typed error instead of promoting a graph under unbounded calls. Previously both controls were shown on the Indexing page but neither reached the pipeline — a 5-second timeout with xhigh effort still let a 13-second extraction call promote a graph. If your extraction alias is a reasoning model, budget the timeout for its thinking trace, not just its JSON reply: a high effort can spend most of a small budget before it writes anything. The binding is pinned end to end by the live writer-contract test (tests/integration/test_graphrag_pipeline_live.py), which builds the semantic pipeline with both values read from the corpus's scoped config.

One extraction call dispatches exactly once

The extraction LLM is built with client retries disabled and a no-op rate-limit handler (semantic_extraction_llm in server/indexing/graphrag_pipeline.py): the OpenAI SDK and the Neo4j GraphRAG library each retry independently by default, and stacked retry layers would silently multiply paid extraction calls after a timeout or a transient 429/5xx — one logical invocation could become several billed gateway calls while the run's request census recorded one. Now a transient failure surfaces as a single failed request in the census, and a hung call that later times out is counted as uncertain, never silently completed — so the accounting denominator matches what the provider actually saw.

The reasoning effort travels in the upstream's own protocol

The operator's effort is bound on every GraphRAG extraction and schema-proposal call (semantic_extraction_llm and reasoning_model_params in server/indexing/graphrag_pipeline.py, derive_graph_schema_proposal in server/indexing/graphrag_schema.py), but in the protocol of the alias's LiteLLM upstream: an OpenRouter upstream (most catalog aliases) receives it as OpenRouter's native reasoning object in extra_body, which passes LiteLLM untouched and is honoured by every reasoning-capable model there; an OpenAI-compatible upstream (the local serving lane) keeps the OpenAI reasoning_effort parameter. The earlier unconditional reasoning_effort parameter broke on OpenRouter: LiteLLM only maps it onto OpenRouter for models its own capability map knows, so newer aliases answered with a 400 UnsupportedParamsError and a semantic run extracted nothing. A blank effort still refuses fast with a typed error. The live graph integration tests default to the OpenRouter-routed alias openai.gpt-5.6-luna (override with GRAPH_E2E_KG_MODEL); if you point the extraction alias at another provider, verify the extraction output shape after changing the effort — the knob is the provider's to interpret, not ragweld's.

Your Semantic KG Extraction prompt is the extractor's template

The official Neo4j GraphRAG extractor now formats system_prompts.semantic_kg_extraction for every chunk (extraction_prompt_template in server/indexing/graphrag_pipeline.py); previously that prompt was shown and edited on the System Prompts page but never read, and the library's default template — which says nothing about what a name may be — ran instead, so the extraction model named persons and emails with OCR noise. The default prompt is now the official extraction template carrying ragweld's naming rules (full names, subject lines for emails, no OCR artifacts or redaction bars), and it must keep the {schema} and {text} placeholders the extractor fills ({examples} is optional). A template the extractor cannot format refuses the run before any schema gate, fence, or staged generation with a typed 422 graph_extraction_prompt_invalid whose hint points at System Prompts; a stored prompt equal to the retired default is dropped on load so the new default applies (server/config.py, server/services/config_store.py) — an operator edit is kept.

Derived graph policy: one decision per corpus

There is no separate semantic-KG toggle. server/indexing/graph_policy.py derives exactly one policy per run:

Policy When What indexing does
semantic external corpus, graph_indexing.enabled, build_code_graph off Official Neo4j GraphRAG LLM extraction (structured output against a closed-world schema) per file, plus the lexical graph
code external corpus, enabled, build_code_graph on AST entities plus the lexical graph per file, written through the same scoped GraphRAG writer
off graph_indexing.enabled false No Neo4j work; the manifest records no graph id and the graph leg truthfully returns nothing
excluded runtime-managed internal corpora (Recall, Codex sessions) Never graph-indexed, regardless of config

The UI shows the derived policy as a badge in RAG → Indexing and RAG → Retrieval, and internal corpora cannot flip the toggle.

Schema proposal and approval (semantic policy)

A semantic run cannot start without a reviewed graph schema:

  1. Propose — POST /api/index/{corpus_id}/graph-schema/proposal deterministically samples documents and positions (documents-and-positions-v2), asks the extraction alias for a closed-world GraphSchema, rejects generic catch-all labels such as OBJECT or RELATED_TO, and persists a GraphSchemaProposal keyed by a canonical schema hash plus a source-inventory fingerprint. A fixed 36-chunk budget is spread across the sampled documents and their full length, so one long report contributes evidence throughout its body instead of only at its edges; PDF sources are read through a bounded pypdfium2 page reader (_extract_schema_sample_text_for_path in server/api/index.py) — every page of a document with 36 pages or fewer, or 36 evenly spaced pages of a larger one — with each sampled page stamped # <file> page <n>. Whole-document Docling conversion never runs behind this synchronous request.
  2. Review — the RAG → Indexing proposal card leads with a one-line summary (entity types · relationships · patterns) and keeps the review human-first: node types, relationship types, connection patterns, and constraints (with each property's name, type, and required flag) sit behind a Review schema disclosure, while the model alias, GraphRAG version, schema hash, sampling recipe, sampled chunk ids with SHA-256 hashes, and raw schema JSON sit behind Technical details. The bulk cost is not shown on the card itself — the pre-run estimate prices it before approval, and indexing cannot start without that estimate.
  3. Approve — starting the run sends approved_graph_schema_hash; the server refuses with 409 graph_schema_approval_required when the hash is missing, or when any corpus file or extraction setting changed since the review (the fingerprint no longer matches).

The proposal runs under the same reasoning effort

graph_indexing.schema_proposal_reasoning_effort (default low), schema_proposal_timeout_s (default 60 s, hard-capped below the public HTTP deadline) and schema_proposal_max_output_tokens (default 16384) bound one proposal independently of the per-chunk extraction controls. The total budget also covers config loading, sampling and the persisted write: a timed-out proposal answers a typed 504 graph_schema_deadline_exceeded, an incomplete or refused gateway reply answers 502 graph_schema_generation_failed, and a corpus or setting that changed while the proposal ran answers 409 graph_schema_context_changed — in every case the previous proposal and its approval are preserved, and the attempt (including any billed provider work) is saved as a schema_proposal run record the Indexing card prices through native accounting. If a proposal's entity/relationship quality looks off, adjusting the proposal effort and regenerating is the first knob to try; changing any budget or the effort invalidates proposal reuse, so the approval you review is always tied to the settings that generated it.

Approved labels become the graph's type vocabulary

The reviewed proposal's node labels and relationship types are exactly what lands on the promoted graph — Entity.entity_type and Relationship.relation_type are open strings on the wire (server/models/tribrid_config_model.py), not a fixed enum, so Tank, LaunchSite, CONTAINS, or LOCATED_AT surface verbatim in the Graph explorer instead of being coerced to a generic kind. See Graph API.

Every node type gets a name identity; document text stays in the chunk store

Before a proposal is hashed for review, normalize_domain_schema (server/indexing/graphrag_schema.py) applies two domain rules the extractor cannot be trusted to follow:

  1. Identity. Every node type carries a STRING name property and a mandatory (EXISTENCE or KEY) constraint on it. The official exact-match entity resolver merges on name and skips nodes where it is null, and the explorer names nodes by it — so the pruner drops an extraction that omits the name instead of writing an anonymous node. A proposal whose node types lacked the property produced a generation with 2,021 of 2,112 extracted entities anonymous — present in the graph, but unaddressable by entity resolution and invisible to the explorer.
  2. No document text on the graph. Text-shaped properties (body, content, text, full_text, transcript, and similar) are stripped from node types, relationship types, and constraints. The chunk store owns the document text; a graph property that asks the extractor to copy the source text back makes the structured-output stream carry whole documents, which provider moderation cuts mid-JSON (finish_reason=content_filter) — the run then fails closed. Graph properties are attributes, never bodies.

Normalization is deterministic and idempotent — normalizing a normalized schema is a no-op — so the schema hash the operator approves is the hash of exactly the shape that lands on the graph. validate_domain_schema enforces both rules again on the validated proposal: a node type without the STRING name identity (or without a mandatory constraint on it), or any node/relationship type carrying a document-text property, refuses with a named error instead of promoting a graph of anonymous nodes or full-text copies. It also refuses any node or relationship type whose serialized additional_properties flag is not false, and closed_graph_schema (server/indexing/graphrag_schema.py) re-applies the closed-world property contract when the approved schema is loaded for execution — so the official pruner drops undeclared properties even where a dependency upgrade would reopen an empty-property type.

Duplicate extracted node ids are folded — or the run refuses

The official GraphRAG 1.19 writer CREATEs one node per extracted row, so a model response that repeats a node id inside one chunk would yield two rows with the same chunk-prefixed entity_id and abort the run on the store's (repo_id, entity_id) uniqueness constraint. Before the writer sees the graph, fold_duplicate_node_ids (server/indexing/graphrag_pipeline.py) folds duplicates that identify the same named entity (same label, same name) into the first occurrence, properties merged, first value wins; the scoped writer then logs GraphRAG extraction repeated identical entity ids for <corpus>: folded=N — a warning, and the run continues and promotes normally.

A repeated id with a conflicting identity — a different label or a different name — is no longer rescued with a #N ordinal suffix: renaming one row would silently strip its provenance, and relationships could no longer say which entity they meant. The graph is refused before anything is written and the run fails with a message naming the id; point graph_indexing.semantic_kg_llm_model at a stronger alias and re-index when you see it.

The approved hash, schema payload, extraction telemetry, entity-resolution counts, community telemetry, and any override are persisted on the generation manifest (GraphGenerationMetadata), so every promoted graph is auditable.

A textless PDF refuses the proposal synchronously

Image-only or unreadable PDFs contribute no sample text, and a corpus whose sampled documents yield none answers a typed 422 — "Graph schema proposal sampling found no embedded PDF text or other indexable text" — instead of starting unbounded whole-document OCR behind the public request window. Whole-document OCR remains an indexing operation, not a synchronous proposal request. The sampler also defers the chunker tokenizer's warm-up until the first non-empty sample, so a proposal over an image-only corpus fails fast instead of paying the tokenizer load first. See Indexing a corpus for the operator walkthrough.

A corpus with no usable domain schema answers a typed 422

When the sampled text yields no extractable domain — the proposer returns no node types or relationships, or a forbidden label — the domain validator rejects the shape and the proposal endpoint answers a typed 422 graph_schema_unusable instead of an unhandled 500. The detail carries corpus_id, the extraction alias that ran (model_alias), the proposer's own message, and an operator_hint: provide text with named entities and relationships, or point graph_indexing.semantic_kg_llm_model at another KG model alias, then generate the proposal again. A number-heavy or boilerplate-only corpus is a review outcome, not a server fault; see Indexing API for the payload shape.

Concept diagram (the bounded proposal sampling only — the full fused pipeline is on the generated retrieval-pipeline page):

flowchart LR
  subgraph s_prop["Graph schema proposal sampling (server/api/index.py)"]
    INV["Corpus inventory\n(positionally stratified entries)"]
    EX{"Source file"}
    PDFS["Bounded PDF sampler\n(_extract_schema_sample_text_for_path)\nall pages when 36 or fewer,\nelse 36 evenly spaced pages via pypdfium2"]
    TX["Generic extractor\n(extract_text_for_path,\nindexing.parquet_extract_* caps)"]
    CH["Chunker\n(tokenizer warm-up deferred\nuntil the first non-empty sample)"]
    SEL["select_schema_chunks\n(first / middle / last per document)"]
    ALIAS["Extraction alias\n(graph_indexing.semantic_kg_llm_model,\nelse the gateway default)"]
    PROP["GraphSchemaProposal\ncanonical schema hash\n+ source-inventory fingerprint"]
    REF["Typed 422:\nno embedded PDF text\nor other indexable text"]
  end
  INV --> EX
  EX -->|"PDF"| PDFS
  EX -->|"other file"| TX
  PDFS -->|"text found"| CH
  PDFS -->|"no embedded text"| REF
  TX -->|"text found"| CH
  TX -->|"empty"| REF
  CH --> SEL
  SEL -->|"no sampled chunks"| REF
  SEL -->|"sampled chunks"| ALIAS
  ALIAS --> PROP

Promotion invariants, communities, and the audited sparse override

Before the commit, the staged graph must pass every invariant in server/indexing/graph_invariants.py: exact chunk count, complete extraction (no failed or truncated chunks), nonzero entities and semantic relationships, FROM_CHUNK provenance links into the staged chunks, no cross-generation nodes or relationships, and no unresolved duplicate entities on the policy's resolution property (name for semantic graphs, entity_id for code graphs). A refusal is a typed failure (GraphPromotionRefusedError) with codes such as zero_entities, extraction_failure, or cross_generation_node; the previous generation stays active and the run record replays the failure codes.

Duplicate groups are counted on the same key the resolver merges on

The duplicate-entity invariant reads the policy's resolution property directly (resolution_property_for_policy in server/indexing/graphrag_pipeline.py, passed as a query parameter to get_graph_invariant_counts in server/db/neo4j.py). Earlier the invariant always grouped by name, so once code entities resolved on the qualified entity_id, every code graph with two same-named methods — two __init__s is the common case — was refused with unresolved_duplicate_entity even though the resolver had merged nothing (one live run: 6,237 entities, zero resolver duplicates, invariant refused anyway). Now a code graph's same-named entities pass when their entity_ids differ, and the semantic check is unchanged: two extracted entities sharing a name still refuse promotion. The live integration test pinning both behaviors is tests/integration/test_graph_promotion_invariants_live.py.

When graph_storage.include_communities is on and the staged graph passes, ragweld projects the staged entities into Neo4j Graph Data Science and runs deterministic Leiden (server/graph/communities.py, GDS 2.13), writing communityId/communityPath onto every entity. The Compose stack ships the GDS plugin, and /api/ready refuses Neo4j readiness without GDS 2.13 when graph indexing and communities are enabled.

One deliberate exception: when a fully successful approved extraction found no graph at all, an authenticated operator (the Remote-User header behind the auth proxy) may retry with graph_empty_override_reason (at least 20 visible characters). The override promotes only the chunk/vector generation — the empty graph is deleted and omitted from the manifest, so graph retrieval can never present it — and the actor, reason, timestamp, and telemetry are persisted on the manifest.

Extraction checkpoints: durable, reusable semantic chunk graphs

A semantic run checkpoints each chunk's pruned, validated graph to Postgres (graph_extraction_checkpoints) as soon as it commits, under an exact identity: the source file's SHA-256, the chunk's content and metadata digests, its line range and provenance, the rendered-prompt digest, and the whole approved recipe — schema, prompt-template hash, extraction alias, upstream and endpoint, model parameters, and the GraphRAG extractor/pruner versions. The recipe is deliberately safe to persist: only digests of source text and prompts are stored (never the text itself), the endpoint is stored without credentials, query or fragment, and the model parameters are a reviewed typed allowlist — credentials, timeouts, retries and free-form provider payloads are refused before anything is written.

What this buys the next run:

  • Reuse instead of re-paying. A re-index whose recipe and source bytes still match reuses the stored chunk graphs with no provider call; reuse adds no native request and no spend, and reuse events are counted on the run record (reused_chunks). Change the prompt, the approved schema, the alias, the reasoning effort, or a file's bytes, and exactly the affected chunks dispatch fresh.
  • Failures keep their work. A run that fails partway — a refusal on file 3 of 5, a cancellation, a worker crash — leaves every already-committed checkpoint in place; the next owner under the same recipe reuses them and dispatches only what is missing. A run's recipe partition is closed only after a complete, invariant-passing build.
  • Corruption fails closed. A checkpoint that fails validation on read is a typed error refusing the run — never a silent re-dispatch over a row the server cannot trust. Writes are fenced through the corpus row, and explicit de-index (not staging cleanup) is the removal path.
  • Ownership is serialized in Postgres. Checkpoint mutations take the same advisory + row locks as fence takeover and corpus deletion, so a checkpoint write, a delete, and a recipe rotation cannot interleave.

Two supporting pieces make the identity trustworthy: each selected source file is copied to a private snapshot and hashed before extraction, so document records and checkpoint identities describe the same bytes even if the file changes mid-run (the copy is released once that file's graph work drains); and a progress owner persists the per-chunk outcome counters into the run summary as they happen, so a cancelled or failed run still reports what completed and what remains unfinished.

Concept diagram (the extraction-checkpoint lifecycle only — the full fused pipeline is on the generated retrieval-pipeline page):

flowchart LR
  subgraph s_prep["Per-file preparation"]
    SNAP["Source snapshot\n(private copy + sha256)"]
    IDENT["Chunk identities\ncontent + metadata + prompt digests\n+ approved recipe"]
    PREP["prepare_graph_extraction_checkpoint_file\nrecipe partition under the corpus fence"]
    SNAP --> IDENT
    IDENT --> PREP
  end
  subgraph s_extract["Per-chunk extraction"]
    LOAD{"Checkpoint read\nunder the fence"}
    REUSE["Reuse stored graph\nno request, no spend"]
    CALL["Official extractor\none fresh dispatch"]
    PRUNE["Domain pruning + validation"]
    COMMIT["put_graph_extraction_checkpoint\ndurable validated graph"]
    LOAD -->|"hit + valid"| REUSE
    LOAD -->|"miss"| CALL
    CALL --> PRUNE
    PRUNE --> COMMIT
  end
  subgraph s_finish["Run completion"]
    WRITE["File graph written\nentities + lexical"]
    INV["Generation invariants"]
    FIN["finish_success\npartition complete"]
    WRITE --> INV
    INV --> FIN
  end
  PREP --> LOAD
  REUSE --> WRITE
  COMMIT --> WRITE
  FIN -->|"prunes deleted files"| DONE["Next run reuses\nthe matching identity"]

AST code graph (build_code_graph)

When graph_indexing.build_code_graph=true, indexing additionally runs a tree-sitter AST pass per source file (server/indexing/code_graph.py) for Python, TypeScript, and JavaScript:

  • one module entity per file, plus class and function/method entities carrying qualname, line range, and first-line signature
  • contains, inherits, imports, and calls relationships, weighted by the graph_indexing.ast_*_weight fields
  • each entity anchored to the chunk that defines it through the same FROM_CHUNK lexical chunk relationship, so graph retrieval can expand a hit to its callers, callees, base classes, and importing modules

Entity ids are corpus-relative: a module is its file_path and a symbol is file_path::qualname; uniqueness in Neo4j is scoped by the corpus id, so a staging id never leaks into an entity id. Relationships whose target is defined in another file are deferred: the per-file pass writes only that file's entities and intra-file edges, and the cross-file edges (imports, plus cross-file inherits and calls) are written once after every file of the run is in Neo4j, through a relationship-only upsert that MATCHes both endpoints. No placeholder node is ever created for a target — a call to an imported class is a call to the existing class node, not to a guessed function node (the earlier shape collided with the Neo4j uniqueness constraint on entity ids and failed the run). Resolution is conservative: a name that cannot be tied to a definition inside the corpus, or an import that does not resolve to a corpus file, produces no edge and is counted as unresolved rather than guessed.

The file's Document node never shares the module entity's writer id

A code corpus writes each file's lexical graph — one Document node per file plus its Chunk nodes — and its AST graph in the same per-file pass, and both halves once used the bare file path as the writer id. The official Neo4j GraphRAG writer keys nodes and relationship endpoints by writer id regardless of label, so the collision silently mis-wired the promoted graph instead of failing: every Document node carried the module entity's contains and FROM_CHUNK edges, and 3,677 chunk FROM_DOCUMENT edges in one real generation pointed at module entities instead of at the file that owns them.

The fix is two-sided (server/indexing/graphrag_pipeline.py): the Document writer id is namespaced as document::<file_path> (document_node_id, used by both the lexical and code paths through document_info), so it can never equal a module entity id, and assemble_code_file_graph refuses any assembled per-file graph in which two nodes share a writer id with a raised error — a collision can no longer be written quietly. Both behaviors are pinned in tests/unit/test_graphrag_pipeline.py and tests/integration/test_graph_communities_live.py.

Re-index code corpora indexed before this change: the repair happens at write time, and an already-promoted graph keeps the wrong wiring until it is rebuilt.

Entity identity is resolved per graph policy (resolution_property_for_policy in server/indexing/graphrag_pipeline.py). The semantic policy merges staged entities on name — equal extracted names within one generation are the same thing — while the code policy merges on the qualified entity_id (path::Qualified.symbol), which the store already keeps unique per generation. Earlier runs resolved code entities on the bare name, which collapsed every __init__ and main of the corpus into one node (a promoted generation once ended with 81 classes "containing" a single shared __init__); with the qualified-id key, two __init__ methods of different classes survive exact-match resolution untouched, and the resolution telemetry reports zero merged nodes when nothing should merge. Policies without a graph (off, excluded) refuse resolution with a typed error instead of guessing a key.

Concept diagram (the deferred cross-file write only — the full fused pipeline is on the generated retrieval-pipeline page):

flowchart LR
  F1["File pass 1"] --> U1["GraphRAG upsert\n(entities + intra-file edges)"]
  F2["File pass 2"] --> U2["GraphRAG upsert\n(entities + intra-file edges)"]
  U1 --> N["Neo4j"]
  U2 --> N
  F1 --> D["Deferred cross-file edges\n(imports / inherits / calls)"]
  F2 --> D
  D --> UF["Relationship-only upsert\nMATCH both endpoints\nafter the last file of the run"]
  UF --> N
Why deferred edges instead of placeholder nodes

Placeholder nodes forced a guessed label for cross-file targets — for example, calling an imported class created a function node. When the defining file later wrote the real class entity with the same id, the two labels violated the Neo4j uniqueness constraint on entity ids and failed the index run. Deferring the edge until every file has been upserted means both endpoints always exist under their real labels, and index order across files no longer matters.

For engineers: the deferred write runs under the neo4j_write_code_deferred Prometheus stage label in server/api/index.py, immediately after the file loop.

Re-index after enabling

The code graph is built during indexing. Toggling build_code_graph only affects new runs, so enable it per code corpus and re-index. It is off by default because it only pays for code corpora.

The full fused path (where the graph leg feeds weighted RRF fusion alongside Qdrant dense and sparse) is documented on the generated retrieval pipeline page; every graph_indexing knob is in the graph_indexing config reference.

curl -sS -X PATCH "http://127.0.0.1:58012/api/config/graph_indexing" \
  -H 'Content-Type: application/json' \
  -d '{"build_code_graph": true}' | jq .
import httpx

httpx.patch(
    "http://127.0.0.1:58012/api/config/graph_indexing",
    json={"build_code_graph": True},
).raise_for_status()  # (1)!
  1. Sectional PATCH is validated by Pydantic; re-index the corpus to rebuild the graph

Figure descriptions (indexing.figures)

For document corpora, indexing can additionally describe figures — charts, diagrams and engineering drawings inside Docling-converted PDFs — so they become retrievable chunks. It is off by default (indexing.figures.enabled=false) because each described figure is a vision-model call through the LiteLLM gateway.

  • Docling detects picture regions per page, optionally classifies them (chart, diagram, logo, photo), and sends each qualifying figure to the configured vision alias (indexing.figures.vision_model, default z-ai.glm-5.3-flash)
  • The alias must be a vision-capable, routable gateway alias: the indexing request is refused with a typed 409 (code: figure_vision_alias) before the run claims the corpus
  • The structured description becomes a chunk anchored to the figure's page and normalized bounding box, so citations box the figure in the source document viewer
  • Coverage and cost controls: min_area_fraction, skip_classes (denied only when the classifier's confident (>= 50%) prediction names the class — a listed class that only appears in the prediction long tail does not skip), max_completion_tokens (default 2500, sized to cover a reasoning alias's internal trace before its JSON reply), concurrency, timeout_s

Under the hood (server/indexing/text_extractors.py): the Docling pipeline options are built straight from the indexing.figures fields; describe=true with no resolved gateway route raises instead of silently producing no descriptions; and enrichment converters are cached per options signature — with figures off the plain shared converter is reused, and only the PDF and IMAGE formats carry the enrichment pipeline, so DOCX/PPTX/XLSX/HTML extraction is unchanged.

The description is a structured FigureAnnotation (server/models/index.py): a vision-judged kind (diagram, chart, schematic, photo, table, drawing, other), a dense prose summary — the text that gets embedded — and transcribed labels (callouts, axis labels, legend entries, part numbers), components, connections (A -> B), values (numbers with units exactly as printed) and references (sheet/figure/table/section cross-references), persisted in Chunk.metadata["figure"] and rendered into the chunk as prose-only markdown, so callouts and part numbers are searchable verbatim without JSON entering the embedded text. The prompt templates and the reply parser live in server/indexing/figure_prompts.py (profile chosen via indexing.figures.prompt_profile); parsing is deliberately forgiving — fenced replies are unwrapped, malformed JSON (trailing commas, unescaped newlines, replies truncated mid-object) is repair-parsed via json-repair, unknown kinds fall back to other, a JSON-looking but unrecoverable reply degrades to an empty annotation so raw syntax never leaks into the embedded text, and a non-JSON reply becomes the plain summary, so a malformed vision response degrades one chunk instead of failing the run.

Rendering is its own step: RagweldPictureSerializer (server/indexing/figure_serializer.py) serializes each picture as prose at the picture's position and blocks Docling's meta-serialization block, so the raw vision JSON can never leak into the markdown even though Docling carries the reply on item.meta. A picture the classifier caught but the vision alias skipped — min_area_fraction, skip_classes, or a per-figure timeout — still renders a Figure (chart)-style header plus its image placeholder, so the class name stays searchable text; and a figure whose caption and summary are both blank keeps its structured lists (Labels, Components, Connections, Values, References) instead of collapsing away. A blank vision reply ("") is not a description: the serializer treats it exactly like no description was ever attached — the picture falls back to the classified-but-undescribed header-only block and is never parsed into a FigureAnnotation. The prose block appears exactly once per picture even when Docling keeps both the newer item.meta shape and its legacy annotations.

Stamping closes the loop: stamp_provenance (server/indexing/provenance.py) maps chunk char ranges back to source spans and, when a described figure's span covers at least half a chunk's range, marks that chunk as a figure chunk — chunk.metadata["chunk_kind"] = "figure" plus the parsed FigureAnnotation under chunk.metadata["figure"] and chunk.metadata["figure_class"] when the classifier resolved one. A chunk that merely brushes a figure keeps the page region in its provenance but stays a text chunk, and figure_class is only written when a class was actually resolved, so no metadata key is backed by nothing. The extracted document triages every picture into exactly one of three counts — figures_described (non-blank description text), figures_failed (a description object exists but its text is blank: the vision call was attempted and the gateway returned nothing, e.g. an unreachable alias or a reasoning model that spent its max_completion_tokens budget on its internal trace before writing any JSON; Docling absorbs this rather than raising), and figures_skipped (no description object at all: the picture never reached the vision call) — across both the item.meta and legacy item.annotations shapes. The run emits a final Figure summary event with all three, and only when at least one picture was processed: it is a warning, not a log line, when description was enabled but nothing was described — figures_failed > 0 gets a distinct hint pointing at the gateway, the alias, and max_completion_tokens (the call was attempted and billed), while an all-skipped run points back at the area threshold and the classification deny-list. Conversion guarding is figure-aware: without figures, a Docling conversion failure degrades to "unparseable" as before; with figures on, the failure raises so one bad conversion cannot quietly drop documents from an index the operator is paying a vision model to enrich — while a bug in ragweld's serializer, source map, or figure counting always raises so regressions stay visible.

Concept diagram (figure enrichment only — the full fused pipeline is on the generated retrieval-pipeline page):

flowchart LR
  PDF["Docling-converted PDF page"] --> DET["Picture region detection"]
  DET --> CLS["Classify + filters\n(skip_classes, min_area_fraction)"]
  CLS -->|"skipped"| CAP["Caption-only text"]
  CLS -->|"passes"| VIS["Vision alias\n(indexing.figures.vision_model)"]
  VIS --> CHUNK["Figure-description chunk\n(page + bounding box)"]
  VIS --> CHUNK2["Figure chunk\n(chunk_kind=figure,\nfigure + figure_class metadata)"]
  VIS -->|"blank reply\n(figures_failed)"| HDR["Header-only block\n(never a FigureAnnotation)"]
  CHUNK2 --> QD["Qdrant dense + sparse generation"]
  CHUNK --> PG["Postgres chunk row\nwith provenance"]

Chunking is figure-aware too: server/api/index.py cuts each document through chunk_document_with_figures (server/indexing/figure_chunking.py), so a described figure block is emitted as one atomic chunk — caption, prose summary, structured lists, and trailing image placeholder together — instead of being windowed by size and risking a citation that lands on a mid-word fragment of the description. Only the text between figures is windowed by the configured chunking strategy, a document with no described figures chunks exactly as before, and an oversized figure splits only at its Labels:/Components:/Connections:/Values:/References: headings.

Figures are part of the run record

A completed run persists figures_described, figures_failed, figures_undescribed, and a figure_description_cost_usd ceiling on its IndexRunSummary (GET /api/index/{corpus_id}/runs/latest), so the counts stay auditable after the terminal stream is gone. The pre-run estimate answers in kind: IndexEstimate.estimated_seconds_figures prices the figure phase’s wall clock (~20 s per vision call, divided by indexing.figures.concurrency) alongside its cost, and the GET /api/index/status cost card adds a Figure Descriptions line when the latest committed run described any.

For the operator walkthrough (cost estimation, per-corpus tuning, troubleshooting), see Indexing a corpus; every knob is in the indexing config reference.

Failure Modes
  • File decoding errors: logged and skipped.
  • Embedding timeouts: retried with backoff; chunk remains un-embedded if persistent.
  • Graph build failures: retrieval continues with vector/sparse; flagged in logs.
  • Code graph: extraction is skipped entirely for unsupported languages (empty graph, not an error); only Python, TypeScript, and JavaScript are parsed.
  • Docling extraction is serialized process-wide: a run queued behind another run's conversion logs Waiting for the document extractor … notices with the measured elapsed wait, and a long conversion emits Converting <file>: still running (Ns elapsed) heartbeats — the first lands after ~60 seconds to rule out a wedged worker, then the interval widens (to ~5 minutes) so a 40-minute conversion narrates itself a handful of times instead of forty identical lines. See Indexing a corpus.