Evaluation Models
-
Eval Dataset
EvalDatasetItemdefines questions and expected paths. -
Metrics
EvalMetrics,EvalRun,EvalResultcapture performance. -
Comparisons
EvalComparisonResultcompares two runs.
Match Production
Tune eval_final_k and eval_multi to reflect real usage; misaligned evals mislead.
Config Snapshots
EvalRun stores both nested and flat config snapshots for reproducibility.
AI analyses are persisted, not re-charged
POST /api/eval/analyze_comparison saves the generated analysis as an EvalAnalysisArtifact (under data/eval_runs/analysis/). GET /api/eval/analysis/{run_id}?compare_run_id=... serves it back without touching the gateway; it answers 404 when nothing is cached or when the cached analysis was generated against a different baseline — a stale pair is never served. Deleting a run deletes its cached analysis. See Evaluation Guide.
Latency Budget
Track latency_p95_ms across runs to guard against regressions.
| Model | Purpose |
|---|---|
EvalDatasetItem | Single question + expected file paths, each optionally narrowed by an EvalExpectedLocation |
EvalExpectedLocation | Page or line span refining one expected path (unit page/line, 1-based inclusive start/end) |
EvalMetrics | Aggregated metrics (MRR, Recall@K, NDCG@10, MAP@5, latency percentiles) |
EvalRun | Complete run with config snapshot and results |
EvalComparisonResult | Delta between baseline and current runs |
EvalAnalysisArtifact | Persisted AI comparison analysis, keyed by (run_id, compare_run_id) |
Chunk-level scoring: locations, uninformative entries, MAP@5, and rated runs
Retrieval is scored per chunk: a path with an expected_locations span is hit only by a chunk overlapping it, and an entry whose expectations cannot fail is uninformative — excluded from the headline metrics and counted in EvalRun.uninformative_count. EvalMetrics.map_at_5 is null on runs scored before chunk-level scoring. SearchResponse now carries event_id (its run id), and POST /api/feedback accepts surface (chat/search) plus the chunk_ids the rated answer cited — it refuses (typed 409 feedback_event_not_answered) events that failed or were aborted, since they produced no answer to rate. See Evaluation Guide.
flowchart TB
Dataset["Eval Dataset"] --> Run["Eval Run"]
Run --> Metrics["Eval Metrics"]
Run --> Results["Per-Entry Results"]
Metrics --> Compare["Compare Runs"] import httpx
base = "http://localhost:8000"
print(httpx.post(f"{base}/reranker/evaluate", json={"corpus_id": "tribrid"}).json())
BASE=http://localhost:8000
curl -sS -X POST "$BASE/reranker/evaluate" -H 'Content-Type: application/json' -d '{"corpus_id":"tribrid"}' | jq .
const report = await (await fetch('/reranker/evaluate', { method:'POST', headers:{'Content-Type':'application/json'}, body: JSON.stringify({ corpus_id: 'tribrid' }) })).json();
Top-K alignment
Ensure eval_final_k >= retrieval.final_k when you want strict hit@K parity with production.
Figure eval dataset models
Page-grounded figure evaluation uses its own dataset shapes in server/models/eval_figures.py (FigureEvalDataset / FigureEvalItem with question, expected_pages (1-based), figure_ref, kind (locate/content) and tags). They are serialized to data/eval_datasets/*.json and consumed by scripts/eval_figure_grounding.py; no frontend consumes them, so they are deliberately not registered for TypeScript generation. See Figure grounding eval.
These complement — they do not replace — the EvalDatasetItem path-based shapes above.