Synthetic Data Lab
-
Recipe-driven generation
Turn an indexed corpus into grounded eval datasets, reranker triplets, semantic cards, and keywords — recipe by recipe, not one giant pipeline.
-
Judge + quality gate
System One typed judgments (Nouls) curate rows, a verbatim evidence check rejects ungrounded ones, and a retrieval quality gate blocks publication before weak data reaches your evals.
-
Gated promotion
Only a completed, gate-passed run can be promoted to a lineage alias — enforced server-side, including on the raw lineage endpoint, so a failed run can never become "current".
Get started Evaluation guide Config reference: synthetic
Where it lives
Open RAG → Synthetic Lab with a corpus selected. Every run is scoped to that corpus, and each run ends with artifacts, a run report, and (for full-stack recipes) a lineage bundle you can publish or promote.
Generation costs money
Generator calls route through the LiteLLM gateway using the alias you pick per run. Judging runs through System One (system_one.provider): TypeSafe Jev bills per input token, a self-hosted Laya is free but CPU-bound. A gated recipe that fails its quality gate has already spent the generation cost — the gate decides whether the output may be used, not whether it is free. Run small max_pairs first.
What one run does
A run walks the corpus's indexed chunks in bounded batches:
- Generate — the generator prompt (
system_prompts.synthetic_generator) asks for question / expected answer / verbatim evidence-quote rows grounded in one chunk's excerpt, with self-contained questions (no "this document"). - Ground — a row survives only when its
evidence_quoteappears verbatim in the source chunk; anything else is counted as ungrounded and dropped. - Judge — every grounded row is judged with typed System One Nouls: reader_question (would a real reader ask this about the document's subject — not cover-page trivia, and understandable without the source?) and answer_supported (does the located evidence quote — not the file name — support the expected answer?) must each reach their
synthetic.judgeminimum. A third cue words copied noul is reported but never gated. - Gate — for retrieval-affecting recipes (
eval_dataset,triplets), the gate retrieves the run's own generated questions against the corpus viaPOST /api/searchand requires top-1 accuracy >=synthetic.quality_gate.top1_minoversynthetic.quality_gate.sample_sizesamples. Entries whose expectations cannot discriminate (a whole file that is the whole corpus) are excluded from the gate's accuracy; a sample that is entirely uninformative fails the gate and says the rows need page/line locations.
The judge is System One, and every row is located
Curation no longer runs a judge prompt through the gateway. Every grounded row is judged in one System One request (server/system_one/client.py), kept only when both gated nouls reach their minimums (synthetic.judge.reader_question_min and synthetic.judge.answer_supported_min, both 0.7 by default). Each kept row also carries a typed expected_locations span — the page range for paged documents, the chunk's line span otherwise — so the published eval dataset is scored against the located evidence instead of the whole document. See System One decisions.
The quality gate is a self-consistency check, not external validation
The gate retrieves the run's own generated questions against the corpus they came from. A perfect score proves the questions are self-consistent with the index — it is not evidence of retrieval quality on real operator questions. Validate published datasets with the Evaluation guide workflows.
Recipes and artifacts
| Recipe | Artifact | Publish action | Quality-gated |
|---|---|---|---|
| Grounded QA | eval_dataset_json | Publish eval dataset | yes |
| Triplets | triplets_jsonl | Publish triplets (reranker training) | yes |
| Semantic summaries | semantic_cards_jsonl | Publish semantic summaries | no |
| Keywords | keywords file | Publish keywords | no |
Every run also writes a human-readable report_md summary. That report is not published to a corpus store, so it has no Publish action; each artifact row carries Copy path, Preview (a bounded, read-only preview of the artifact rows), and Publish where applicable.
The recipe picker labels every lane in plain language — Eval Dataset, Semantic Summaries, Triplets, Keywords, Autotune Retrieval, and Full Stack — and the deep-link preset notice uses those labels too (Recipe preset to "Eval Dataset" … Nothing has run). Autotune Retrieval and Full Stack sit in the same picker under the same rules as the recipes above: nothing runs until you start it there, and each run writes its artifacts and report to the run record.
Publish vs promote
Publishing writes an artifact into the corpus's stores (eval datasets, triplets, cards, keywords). Promotion moves a lineage alias — baseline, canary, current, or promoted — to point at the run's bundle, which is how the rest of the platform records "this output is now the reference".
The gate is enforced server-side on both paths:
| Run state | Promote? | What you see |
|---|---|---|
completed, gate passed (or recipe has no gate), bundle attached | yes | alias buttons enabled |
| failed or still running | no — typed 409 PROMOTION_BLOCKED | disabled buttons + the reason |
| completed but the gate failed | no — 409 | the gate's failure reason |
| completed but never attached to a bundle | no — 409 | "not attached to a lineage bundle" |
The raw lineage endpoint refuses too: POST /api/lineage/aliases/{alias} checks whether the posted bundle id belongs to a synthetic run and applies the same gate, so a direct API call cannot bypass the UI's disabled buttons. It fails closed — if a run's record no longer validates, promotion is refused rather than trusted.
flowchart LR
subgraph s_req["Request (RAG - Synthetic Lab)"]
REQ["POST /api/synthetic/run/start\\nprovider + recipe + models"]
end
subgraph s_orch["Orchestrator (server/synthetic/orchestrator.py)"]
ORCH["Per-source chunk batches"]
GEN["Generator LLM\\nsynthetic.generator.*\\nvia the LiteLLM gateway :54000"]
GROUND["Grounding check\\nevidence_quote verbatim\\nin the source chunk"]
REJ["Ungrounded + malformed rows rejected"]
JUDGE["Judge LLM\\nsynthetic.judge.*\\nLLM-as-a-judge curation"]
GATE["Quality gate\\nsynthetic.quality_gate.*\\nPOST /api/search on the corpus"]
ART["Artifacts + report\\neval dataset / triplets /\\nsemantic cards / keywords"]
end
subgraph s_store["Run store and lineage"]
RUNS["Run record\\nrun.json + live events"]
BUNDLE["Lineage bundle"]
PUBLISH["Publish endpoints\\n/synthetic/run/:id/publish/:kind"]
PROMOTE["Promotion gate\\ncompleted + gate passed +\\nbundle attached, else 409"]
ALIAS["POST /api/synthetic/run/:id/promote/:alias\\nand POST /api/lineage/aliases/:alias"]
SET["Alias updated\\nbaseline / canary / current / promoted"]
end
REQ --> ORCH
ORCH --> GEN
GEN --> GROUND
GROUND --> REJ
GROUND --> JUDGE
JUDGE --> GATE
GATE --> ART
GATE -->|"gate failed"| RUNS
ART --> RUNS
ART --> PUBLISH
RUNS --> BUNDLE
BUNDLE --> PROMOTE
PROMOTE --> ALIAS
ALIAS --> SET When a run fails
A failed run is a data point, not a dead end:
- The run detail shows the failure reason in a Run failed card; Live events below it holds the run log.
- Retry re-launches with the exact recipe, models, and parameters the run stored — no rebuilding the request by hand.
- Aliases stay locked until a run actually completes and passes its gate.
Open the failed run in RAG → Synthetic Lab, read the reason, fix the cause (an unindexed corpus, an unreachable gateway alias, an unreachable System One backend), then press Retry.
Read the numbers the way the run wrote them
The Grounding & Curation panel reports sources used, generated, ungrounded, malformed, judged, kept, the mean nouls (reader question, answer supported, cue words copied — probabilities, two decimals), and mined triplets. A run with many ungrounded rows is telling you the generator is reaching beyond its excerpt — narrow the per-run excerpt scope or lower pairs_per_source rather than lowering the System One thresholds.
Knobs
All knobs are generated in the synthetic config reference. The ones that matter first:
| Knob | Default | Why it matters |
|---|---|---|
synthetic.generator.max_tokens | 1200 | Output budget per generator call; too low truncates JSON rows |
synthetic.generator.temperature | 0.0 | Keep at 0 for grounded, reproducible rows |
synthetic.generator.concurrency | 4 | Parallel gateway calls; forced to 1 for the single-stream local serving row |
synthetic.judge.reader_question_min | 0.7 | Minimum probability that a real reader would ask the question about the document's subject (not cover-page trivia) |
synthetic.judge.answer_supported_min | 0.7 | Minimum probability that the located evidence quote — not the file name — supports the expected answer |
synthetic.quality_gate.sample_size | 50 | Questions sampled for the gate — raise for a stronger signal |
synthetic.quality_gate.top1_min | 0.4 | Minimum top-1 accuracy to pass; raise cautiously |
If you're not sure
Start with the eval_dataset recipe and small limits, read the run report, and only promote (point an alias at) runs whose gate passed on a healthy sample. Wire the published dataset into an eval run before trusting it in any regression workflow.