Skip to content

Synthetic Data Lab

  • Recipe-driven generation


    Turn an indexed corpus into grounded eval datasets, reranker triplets, semantic cards, and keywords — recipe by recipe, not one giant pipeline.

  • Judge + quality gate


    System One typed judgments (Nouls) curate rows, a verbatim evidence check rejects ungrounded ones, and a retrieval quality gate blocks publication before weak data reaches your evals.

  • Gated promotion


    Only a completed, gate-passed run can be promoted to a lineage alias — enforced server-side, including on the raw lineage endpoint, so a failed run can never become "current".

Get started Evaluation guide Config reference: synthetic

Where it lives

Open RAG → Synthetic Lab with a corpus selected. Every run is scoped to that corpus, and each run ends with artifacts, a run report, and (for full-stack recipes) a lineage bundle you can publish or promote.

Generation costs money

Generator calls route through the LiteLLM gateway using the alias you pick per run. Judging runs through System One (system_one.provider): TypeSafe Jev bills per input token, a self-hosted Laya is free but CPU-bound. A gated recipe that fails its quality gate has already spent the generation cost — the gate decides whether the output may be used, not whether it is free. Run small max_pairs first.

What one run does

A run walks the corpus's indexed chunks in bounded batches:

  1. Generate — the generator prompt (system_prompts.synthetic_generator) asks for question / expected answer / verbatim evidence-quote rows grounded in one chunk's excerpt, with self-contained questions (no "this document").
  2. Ground — a row survives only when its evidence_quote appears verbatim in the source chunk; anything else is counted as ungrounded and dropped.
  3. Judge — every grounded row is judged with typed System One Nouls: reader_question (would a real reader ask this about the document's subject — not cover-page trivia, and understandable without the source?) and answer_supported (does the located evidence quote — not the file name — support the expected answer?) must each reach their synthetic.judge minimum. A third cue words copied noul is reported but never gated.
  4. Gate — for retrieval-affecting recipes (eval_dataset, triplets), the gate retrieves the run's own generated questions against the corpus via POST /api/search and requires top-1 accuracy >= synthetic.quality_gate.top1_min over synthetic.quality_gate.sample_size samples. Entries whose expectations cannot discriminate (a whole file that is the whole corpus) are excluded from the gate's accuracy; a sample that is entirely uninformative fails the gate and says the rows need page/line locations.

The judge is System One, and every row is located

Curation no longer runs a judge prompt through the gateway. Every grounded row is judged in one System One request (server/system_one/client.py), kept only when both gated nouls reach their minimums (synthetic.judge.reader_question_min and synthetic.judge.answer_supported_min, both 0.7 by default). Each kept row also carries a typed expected_locations span — the page range for paged documents, the chunk's line span otherwise — so the published eval dataset is scored against the located evidence instead of the whole document. See System One decisions.

The quality gate is a self-consistency check, not external validation

The gate retrieves the run's own generated questions against the corpus they came from. A perfect score proves the questions are self-consistent with the index — it is not evidence of retrieval quality on real operator questions. Validate published datasets with the Evaluation guide workflows.

Recipes and artifacts

Recipe Artifact Publish action Quality-gated
Grounded QA eval_dataset_json Publish eval dataset yes
Triplets triplets_jsonl Publish triplets (reranker training) yes
Semantic summaries semantic_cards_jsonl Publish semantic summaries no
Keywords keywords file Publish keywords no

Every run also writes a human-readable report_md summary. That report is not published to a corpus store, so it has no Publish action; each artifact row carries Copy path, Preview (a bounded, read-only preview of the artifact rows), and Publish where applicable.

The recipe picker labels every lane in plain language — Eval Dataset, Semantic Summaries, Triplets, Keywords, Autotune Retrieval, and Full Stack — and the deep-link preset notice uses those labels too (Recipe preset to "Eval Dataset" … Nothing has run). Autotune Retrieval and Full Stack sit in the same picker under the same rules as the recipes above: nothing runs until you start it there, and each run writes its artifacts and report to the run record.

Publish vs promote

Publishing writes an artifact into the corpus's stores (eval datasets, triplets, cards, keywords). Promotion moves a lineage alias — baseline, canary, current, or promoted — to point at the run's bundle, which is how the rest of the platform records "this output is now the reference".

The gate is enforced server-side on both paths:

Run state Promote? What you see
completed, gate passed (or recipe has no gate), bundle attached yes alias buttons enabled
failed or still running no — typed 409 PROMOTION_BLOCKED disabled buttons + the reason
completed but the gate failed no — 409 the gate's failure reason
completed but never attached to a bundle no — 409 "not attached to a lineage bundle"

The raw lineage endpoint refuses too: POST /api/lineage/aliases/{alias} checks whether the posted bundle id belongs to a synthetic run and applies the same gate, so a direct API call cannot bypass the UI's disabled buttons. It fails closed — if a run's record no longer validates, promotion is refused rather than trusted.

flowchart LR
  subgraph s_req["Request (RAG - Synthetic Lab)"]
    REQ["POST /api/synthetic/run/start\\nprovider + recipe + models"]
  end
  subgraph s_orch["Orchestrator (server/synthetic/orchestrator.py)"]
    ORCH["Per-source chunk batches"]
    GEN["Generator LLM\\nsynthetic.generator.*\\nvia the LiteLLM gateway :54000"]
    GROUND["Grounding check\\nevidence_quote verbatim\\nin the source chunk"]
    REJ["Ungrounded + malformed rows rejected"]
    JUDGE["Judge LLM\\nsynthetic.judge.*\\nLLM-as-a-judge curation"]
    GATE["Quality gate\\nsynthetic.quality_gate.*\\nPOST /api/search on the corpus"]
    ART["Artifacts + report\\neval dataset / triplets /\\nsemantic cards / keywords"]
  end
  subgraph s_store["Run store and lineage"]
    RUNS["Run record\\nrun.json + live events"]
    BUNDLE["Lineage bundle"]
    PUBLISH["Publish endpoints\\n/synthetic/run/:id/publish/:kind"]
    PROMOTE["Promotion gate\\ncompleted + gate passed +\\nbundle attached, else 409"]
    ALIAS["POST /api/synthetic/run/:id/promote/:alias\\nand POST /api/lineage/aliases/:alias"]
    SET["Alias updated\\nbaseline / canary / current / promoted"]
  end
  REQ --> ORCH
  ORCH --> GEN
  GEN --> GROUND
  GROUND --> REJ
  GROUND --> JUDGE
  JUDGE --> GATE
  GATE --> ART
  GATE -->|"gate failed"| RUNS
  ART --> RUNS
  ART --> PUBLISH
  RUNS --> BUNDLE
  BUNDLE --> PROMOTE
  PROMOTE --> ALIAS
  ALIAS --> SET

When a run fails

A failed run is a data point, not a dead end:

  • The run detail shows the failure reason in a Run failed card; Live events below it holds the run log.
  • Retry re-launches with the exact recipe, models, and parameters the run stored — no rebuilding the request by hand.
  • Aliases stay locked until a run actually completes and passes its gate.

Open the failed run in RAG → Synthetic Lab, read the reason, fix the cause (an unindexed corpus, an unreachable gateway alias, an unreachable System One backend), then press Retry.

Start a new run with the same request body you used before (POST /api/synthetic/run/start); the run record's request block is exactly that body. Cancel a still-running run first if you need to:

curl -sS -X POST "http://127.0.0.1:58012/api/synthetic/run/<run_id>/cancel" | jq .

Read the numbers the way the run wrote them

The Grounding & Curation panel reports sources used, generated, ungrounded, malformed, judged, kept, the mean nouls (reader question, answer supported, cue words copied — probabilities, two decimals), and mined triplets. A run with many ungrounded rows is telling you the generator is reaching beyond its excerpt — narrow the per-run excerpt scope or lower pairs_per_source rather than lowering the System One thresholds.

Knobs

All knobs are generated in the synthetic config reference. The ones that matter first:

Knob Default Why it matters
synthetic.generator.max_tokens 1200 Output budget per generator call; too low truncates JSON rows
synthetic.generator.temperature 0.0 Keep at 0 for grounded, reproducible rows
synthetic.generator.concurrency 4 Parallel gateway calls; forced to 1 for the single-stream local serving row
synthetic.judge.reader_question_min 0.7 Minimum probability that a real reader would ask the question about the document's subject (not cover-page trivia)
synthetic.judge.answer_supported_min 0.7 Minimum probability that the located evidence quote — not the file name — supports the expected answer
synthetic.quality_gate.sample_size 50 Questions sampled for the gate — raise for a stronger signal
synthetic.quality_gate.top1_min 0.4 Minimum top-1 accuracy to pass; raise cautiously

If you're not sure

Start with the eval_dataset recipe and small limits, read the run report, and only promote (point an alias at) runs whose gate passed on a healthy sample. Wire the published dataset into an eval run before trusting it in any regression workflow.