Skip to content

Testing and Verification

  • Zero-Mocked


    Real integrations: no request interception or Python mocks.

  • Coverage by Change Type


    Components → Playwright, APIs → pytest, Retrieval → relevance.

  • Gate to Done


    You cannot return a response unless tests run and pass.

Get started Configuration API

Real Results

Validate that search returns relevant chunks, not just 200 OK.

CI Hooks

Stop hook blocks completion until validators and tests succeed.

CI provisions Docling's PDF models from a pinned manifest

Docling's real PDF tests no longer depend on mutable Hugging Face main revisions or model-download rate limits. scripts/prepare_docling_ci.py provisions the layout, figure-classifier, tableformer, and RapidOCR artifacts listed in scripts/docling_ci_models.json — immutable revisions with SHA-256-verified bytes, cross-checked against the locked package versions and Docling's configured defaults — into DOCLING_ARTIFACTS_PATH, which CI caches only after verification. A failed download never publishes an artifact, and --verify-only re-checks the cache without any network access.

No Mocks

  • No Playwright page.route(...).fulfill(...)
  • No Python unittest.mock / monkeypatch

Required Tests by Change Type

Change Required Test
New component Playwright: render, interact, verify state
Component edit Playwright: existing tests still pass + new behavior
API endpoint pytest: real request/response/data
Config field pytest: validation works, default applies
Retrieval logic pytest: search returns relevant results
Bug fix Test reproduces bug, then passes after fix

Environment hygiene in tests (enforced)

Tests that mutate os.environ must restore what they touched. A raw os.environ.pop("KEY", None) with no finally-guaranteed restore leaks across every later test in the same pytest process — the motivating bug was an unrestored os.environ.pop("LITELLM_API_KEY", None) in tests/unit/test_reranker.py that silently failed 17 unrelated tests/api tests which each passed in isolation.

  • Prefer the patching fixtures: monkeypatch.delenv("KEY", raising=False) / monkeypatch.setenv(...) — they restore automatically on teardown.
  • If you must touch os.environ directly (for example, snapshotting the whole environ for a wide seam like RAGWELD_SYNTHETIC_RUNS_ROOT), restore it in a try/finally — os.environ.clear() + os.environ.update(previous_env) in the finally counts as a wildcard restore.
  • A architecture-policy test parses every file under tests/ and fails on any raw os.environ.pop(...) / del os.environ[...] that is not paired with a restore of the same key (or a wildcard restore) in a finally body — see tests/unit/test_architecture_policy.py::test_no_unrestored_os_environ_mutations_in_tests.

The same discipline is enforced at the shell boundary: start.sh snapshots and restores the exported environment around source .env so a .env value can never clobber a caller-provided override (covered by tests/unit/test_runtime_lifecycle.py). If you add a test that needs to mutate os.environ directly, restore in finally or use monkeypatch — the policy test will reject anything else.

Examples

# RIGHT - verify real results
import httpx

def test_search_returns_relevant_chunks():
    r = httpx.post("http://127.0.0.1:8012/api/search", json={
        "query": "authentication flow",
        "corpus_id": "my-corpus",
        "top_k": 10,
    })
    r.raise_for_status()
    results = r.json()["matches"]
    assert len(results) >= 3
    assert any("auth" in m["content"].lower() for m in results)
curl -sS -X POST http://127.0.0.1:8012/api/search -H 'Content-Type: application/json' \
  -d '{"corpus_id":"my-corpus","query":"authentication flow","top_k":10}' | jq '[.matches[].file_path] | length'
// Playwright example skeleton
import { test, expect } from '@playwright/test';

test('fusion weight slider updates config', async ({ page }) => {
  await page.goto('/rag');
  const slider = page.getByTestId('vector-weight-slider');
  await slider.fill('0.6');
  await page.getByTestId('save-config').click();
  await expect(page.getByTestId('config-saved-toast')).toBeVisible();
  await page.reload();
  await expect(slider).toHaveValue('0.6');
});
  • Start full stack locally (./start.sh --with-observability)
  • Configure LLM credentials in .env
  • Convert legacy mocked tests before editing feature areas

Exhaustive e2e: real workflows, zero mocks

The Playwright suite under web/tests/e2e/exhaustive/ exercises whole operator workflows against the live stack — real corpus provisioning, real indexing, real retrieval — with no route mocking, following the same zero-mock discipline as the pytest side. Two shared helpers make these specs cheap to write and honest by construction:

Helper What it does Why it matters
corpus_fixture.ts (provisionExhaustiveCorpus) Creates a uniquely-named temp corpus, patches per-corpus config sections, optionally runs a real index (POST /api/index with force_reindex=true), and disposes everything afterward — even when a test fails Specs cannot leak corpora into the operator's registry or fight over a shared corpus id
chat_seed.ts (seedAnswerFromSearch) Runs a real POST /api/search against the indexed corpus, then seeds a chat thread in localStorage whose assistant message carries exactly those matches as sources Citation-rendering specs assert on real retrieval evidence without depending on a paid gateway model being reachable

Concept diagram (the seeded-citation mechanism only — the full fused retrieval pipeline is on the generated retrieval-pipeline page):

flowchart LR
  subgraph s_fixture["Fixture (web/tests/e2e/exhaustive)"]
    CF["corpus_fixture.ts\nprovisionExhaustiveCorpus"]
    IDX["POST /api/index\nforce_reindex=true"]
    ST["GET /api/index/{corpus_id}/status\nwait for complete"]
    CS["chat_seed.ts\nseedAnswerFromSearch"]
    CF --> IDX
    IDX --> ST
  end
  subgraph s_seed["Seeded thread (localStorage)"]
    SEARCH["POST /api/search\ncache_mode=bypass"]
    MATCHES["ChunkMatch[]\nprovenance + metadata"]
    THREAD["ragweld-chat-threads:v2\nassistant message sources = matches"]
    SEARCH --> MATCHES
    MATCHES --> THREAD
  end
  subgraph s_spec["Assertions"]
    UI["Chat citations,\nfigure badges, source viewer"]
  end
  ST --> SEARCH
  CS --> SEARCH
  THREAD --> UI

The figure workflow spec (figure_workflow.spec.ts)

A serial-mode spec that drives figure descriptions end to end over one temp corpus (the Apollo two-page fixture PDF plus the markdown acceptance fixture, so the citation list is a genuinely mixed list of figure chunks and ordinary text chunks):

  • Figures enabled from the RAG → Indexing tab (Figures & Vision card), persisted per corpus, and the operator's global config left untouched
  • POST /api/index/estimate prices the vision calls (Figures ≤ $… (~N)) before any run starts, and cancelling the dialog starts no run
  • A real index describes the figures; the replayed run log (via the indexing-show-logs and live-terminal-output test ids) reports figures_described ≥ 1
  • Figure citations carry the Figure badge, thumbnails box the figure region, and clicking one opens the page viewer with the Figure description panel
  • The badge is conditional: ordinary citations in the same list carry none
  • GUI contract: numeric fields clamp to Pydantic bounds, blur stages the edit (no per-keystroke writes), Escape abandons the staged edit without any write, and two nested indexing.figures.* edits stage into a single pending Apply

Long-running indexes need an explicit deadline

indexCorpus in corpus_fixture.ts accepts a timeoutMs override. A Docling conversion of scanned pages plus per-figure vision calls can take tens of minutes on a loaded box — figure_workflow.spec.ts passes a 30-minute deadline explicitly rather than relying on the shared EXHAUSTIVE_INDEX_TIMEOUT_MS env default (5 minutes), so the spec cannot fail for whoever forgets the env var.

The NumberField migration spec (numberfield_migration.spec.ts)

One behavior, proven across every surface family that carries a config-bound numeric input, now under the staged commit model: type a value past the field's Pydantic bound, Tab away, and (a) the box shows the clamped value, (b) blur writes nothing — the change stages and the Apply count goes up by exactly one — (c) the Apply PUT of the whole config carries the clamped value and the raw value appears in no request body, and (d) a fresh GET /api/config confirms the server persisted the clamped value, not the operator's typed one. No route mocking; the same zero-mock discipline as the rest of the exhaustive suite.

  • Data Quality — enrichment.chunk_summaries_max, probe 999999 → 1000: the exact probe that previously reached the server unclamped and came back a 422 whose only signal was a raw error string
  • Chat Settings — chat.temperature, 9 → 2: deliberately not chat.max_tokens, which is a production-scoped global that the per-corpus config would reconcile away on read, making the persistence assertion fail for a reason unrelated to NumberField
  • Reranker config — reranking.rerank_input_snippet_chars, 50000 → 2000: a field visible regardless of reranker mode, since the fixture corpus pins reranker_mode: none
  • Reranker Training Studio — training.reranker_train_epochs, 999 → 20, driven through the Inspector's "Paths + Config" tab
  • Storage Calculator — the non-config calculator inputs survive blur unchanged (no step snapping) while min/max clamping still applies

See Configuration for the full staged commit model and the guard tests behind it.

The staged commit model spec (commit_model.spec.ts)

Proves the one commit model end to end against the real app and real API: selecting a chunking-strategy card fires no write at all — no PATCH, no PUT, well past the old 300 ms debounce window — the Apply button reflects the staged count (Apply N changes with a data-dirty-count attribute), and applying an index-invalidating change (chunking / embedding / tokenization) shows a confirmation that names the index and the section before any write. Cancelling that confirmation writes nothing, so the test mutates no config.

The dirty-count-on-load spec (dirty_count_on_load.spec.ts)

Merely visiting a config surface must stage nothing. Under the staged model, a component that "self-heals" config in a mount effect would stage a permanent edit the operator never made — the footer would read "Apply 1 change" on a page they only opened, and Apply would PUT a mutation nobody intended. This drives every config-consuming surface (the RAG subtabs, Chat settings, Infrastructure → Paths, Grafana config, Admin) with no interaction and asserts the dirty count stays exactly 0 and no config write is issued on mount.

The destructive-safety spec (destructive_safety.spec.ts)

Proves the shared confirm-dialog primitive's safety contracts: the delete-index dialog focuses the typed-confirmation input (never the destructive button), the confirm stays disabled until the corpus id is typed verbatim, and cancelling issues no DELETE — nothing is destroyed. A second scenario pins the other focus branch: a danger dialog with no typed gate (the Infrastructure → Paths save) focuses Cancel, so a stray Enter declines rather than destroys.

The dock legibility spec (legibility_dock.spec.ts)

Proves legibility invariants for tabs rendered in the narrow dock pane, with real sidebar/dock clicks and no interception: the docked glossary wraps or scrolls rather than clipping mid-word, and Get Started swapped into the dock lays out by its container's width — the onboarding container keys off the #tab-start size container, not the viewport, so a step-4 heading and its warning stay readable at dock width instead of collapsing into a vertical strip with horizontal overflow. It also asserts the main-pane rendering keeps the full container width when Get Started is the main page.

Artifacts

Temporary feature tests and results go in .tests/; permanent tests go under tests/.