Skip to content

Testing

# Testing and Verification

<div class="grid chunk_summaries" markdown>

-   :material-test-tube:{ .lg .middle } **Zero-Mocked**

    ---

    Real integrations: no request interception or Python mocks.

-   :material-clipboard-text:{ .lg .middle } **Coverage by Change Type**

    ---

    Components → Playwright, APIs → pytest, Retrieval → relevance.

-   :material-shield-check:{ .lg .middle } **Gate to Done**

    ---

    You cannot return a response unless tests run and pass.

</div>

[Get started](index.md){ .md-button .md-button--primary }
[Configuration](configuration.md){ .md-button }
[API](api.md){ .md-button }

!!! tip "Real Results"
    Validate that search returns relevant chunks, not just 200 OK.

!!! note "CI Hooks"
    Stop hook blocks completion until validators and tests succeed.

!!! danger "No Mocks"
    - No Playwright `page.route(...).fulfill(...)`
    - No Python `unittest.mock` / `monkeypatch`

## Required Tests by Change Type

| Change | Required Test |
|--------|---------------|
| New component | Playwright: render, interact, verify state |
| Component edit | Playwright: existing tests still pass + new behavior |
| API endpoint | pytest: real request/response/data |
| Config field | pytest: validation works, default applies |
| Retrieval logic | pytest: search returns relevant results |
| Bug fix | Test reproduces bug, then passes after fix |

### Environment hygiene in tests (enforced)

Tests that mutate `os.environ` must restore what they touched. A raw `os.environ.pop("KEY", None)` with no `finally`-guaranteed restore leaks across **every later test in the same pytest process** — the motivating bug was an unrestored `os.environ.pop("LITELLM_API_KEY", None)` in `tests/unit/test_reranker.py` that silently failed 17 unrelated `tests/api` tests which each passed in isolation.

- [ ] Prefer the patching fixtures: `monkeypatch.delenv("KEY", raising=False)` / `monkeypatch.setenv(...)` — they restore automatically on teardown.
- [ ] If you must touch `os.environ` directly (for example, snapshotting the whole environ for a wide seam like `RAGWELD_SYNTHETIC_RUNS_ROOT`), restore it in a `try/finally``os.environ.clear()` + `os.environ.update(previous_env)` in the `finally` counts as a wildcard restore.
- [ ] A architecture-policy test parses every file under `tests/` and fails on any raw `os.environ.pop(...)` / `del os.environ[...]` that is not paired with a restore of the same key (or a wildcard restore) in a `finally` body — see `tests/unit/test_architecture_policy.py::test_no_unrestored_os_environ_mutations_in_tests`.

The same discipline is enforced at the shell boundary: `start.sh` snapshots and restores the exported environment around `source .env` so a `.env` value can never clobber a caller-provided override (covered by `tests/unit/test_runtime_lifecycle.py`). If you add a test that needs to mutate `os.environ` directly, restore in `finally` or use `monkeypatch` — the policy test will reject anything else.

### Examples

=== "Python"
```python
# RIGHT - verify real results
import httpx

def test_search_returns_relevant_chunks():
    r = httpx.post("http://127.0.0.1:8012/api/search", json={
        "query": "authentication flow",
        "corpus_id": "my-corpus",
        "top_k": 10,
    })
    r.raise_for_status()
    results = r.json()["matches"]
    assert len(results) >= 3
    assert any("auth" in m["content"].lower() for m in results)
curl -sS -X POST http://127.0.0.1:8012/api/search -H 'Content-Type: application/json' \
  -d '{"corpus_id":"my-corpus","query":"authentication flow","top_k":10}' | jq '[.matches[].file_path] | length'
// Playwright example skeleton
import { test, expect } from '@playwright/test';

test('fusion weight slider updates config', async ({ page }) => {
  await page.goto('/rag');
  const slider = page.getByTestId('vector-weight-slider');
  await slider.fill('0.6');
  await page.getByTestId('save-config').click();
  await expect(page.getByTestId('config-saved-toast')).toBeVisible();
  await page.reload();
  await expect(slider).toHaveValue('0.6');
});
  • Start full stack locally (./start.sh --with-observability)
  • Configure LLM credentials in .env
  • Convert legacy mocked tests before editing feature areas

Exhaustive e2e: real workflows, zero mocks

The Playwright suite under web/tests/e2e/exhaustive/ exercises whole operator workflows against the live stack — real corpus provisioning, real indexing, real retrieval — with no route mocking, following the same zero-mock discipline as the pytest side. Two shared helpers make these specs cheap to write and honest by construction:

Helper What it does Why it matters
corpus_fixture.ts (provisionExhaustiveCorpus) Creates a uniquely-named temp corpus, patches per-corpus config sections, optionally runs a real index (POST /api/index with force_reindex=true), and disposes everything afterward — even when a test fails Specs cannot leak corpora into the operator's registry or fight over a shared corpus id
chat_seed.ts (seedAnswerFromSearch) Runs a real POST /api/search against the indexed corpus, then seeds a chat thread in localStorage whose assistant message carries exactly those matches as sources Citation-rendering specs assert on real retrieval evidence without depending on a paid gateway model being reachable

Concept diagram (the seeded-citation mechanism only — the full fused retrieval pipeline is on the generated retrieval-pipeline page):

flowchart LR
  subgraph s_fixture["Fixture (web/tests/e2e/exhaustive)"]
    CF["corpus_fixture.ts\nprovisionExhaustiveCorpus"]
    IDX["POST /api/index\nforce_reindex=true"]
    ST["GET /api/index/{corpus_id}/status\nwait for complete"]
    CS["chat_seed.ts\nseedAnswerFromSearch"]
    CF --> IDX
    IDX --> ST
  end
  subgraph s_seed["Seeded thread (localStorage)"]
    SEARCH["POST /api/search\ncache_mode=bypass"]
    MATCHES["ChunkMatch[]\nprovenance + metadata"]
    THREAD["ragweld-chat-threads:v2\nassistant message sources = matches"]
    SEARCH --> MATCHES
    MATCHES --> THREAD
  end
  subgraph s_spec["Assertions"]
    UI["Chat citations,\nfigure badges, source viewer"]
  end
  ST --> SEARCH
  CS --> SEARCH
  THREAD --> UI

The figure workflow spec (figure_workflow.spec.ts)

A serial-mode spec that drives figure descriptions end to end over one temp corpus (the Apollo two-page fixture PDF plus the markdown acceptance fixture, so the citation list is a genuinely mixed list of figure chunks and ordinary text chunks):

  • Figures enabled from the RAG → Indexing tab (Figures & Vision card), persisted per corpus, and the operator's global config left untouched
  • POST /api/index/estimate prices the vision calls (Figures ≤ $… (~N)) before any run starts, and cancelling the dialog starts no run
  • A real index describes the figures; the replayed run log (via the indexing-show-logs and live-terminal-output test ids) reports figures_described ≥ 1
  • Figure citations carry the Figure badge, thumbnails box the figure region, and clicking one opens the page viewer with the Figure description panel
  • The badge is conditional: ordinary citations in the same list carry none
  • GUI contract: numeric fields clamp to Pydantic bounds, blur stages the edit (no per-keystroke writes), Escape abandons the staged edit without any write, and two nested indexing.figures.* edits stage into a single pending Apply

Long-running indexes need an explicit deadline

indexCorpus in corpus_fixture.ts accepts a timeoutMs override. A Docling conversion of scanned pages plus per-figure vision calls can take tens of minutes on a loaded box — figure_workflow.spec.ts passes a 30-minute deadline explicitly rather than relying on the shared EXHAUSTIVE_INDEX_TIMEOUT_MS env default (5 minutes), so the spec cannot fail for whoever forgets the env var.

The NumberField migration spec (numberfield_migration.spec.ts)

One behavior, proven across every surface family that carries a config-bound numeric input, now under the staged commit model: type a value past the field's Pydantic bound, Tab away, and (a) the box shows the clamped value, (b) blur writes nothing — the change stages and the Apply count goes up by exactly one — (c) the Apply PUT of the whole config carries the clamped value and the raw value appears in no request body, and (d) a fresh GET /api/config confirms the server persisted the clamped value, not the operator's typed one. No route mocking; the same zero-mock discipline as the rest of the exhaustive suite.

  • Data Qualityenrichment.chunk_summaries_max, probe 9999991000: the exact probe that previously reached the server unclamped and came back a 422 whose only signal was a raw error string
  • Chat Settingschat.temperature, 92: deliberately not chat.max_tokens, which is a production-scoped global that the per-corpus config would reconcile away on read, making the persistence assertion fail for a reason unrelated to NumberField
  • Reranker configreranking.rerank_input_snippet_chars, 500002000: a field visible regardless of reranker mode, since the fixture corpus pins reranker_mode: none
  • Reranker Training Studiotraining.reranker_train_epochs, 99920, driven through the Inspector's "Paths + Config" tab
  • Storage Calculator — the non-config calculator inputs survive blur unchanged (no step snapping) while min/max clamping still applies

See Configuration for the full staged commit model and the guard tests behind it.

The staged commit model spec (commit_model.spec.ts)

Proves the one commit model end to end against the real app and real API: selecting a chunking-strategy card fires no write at all — no PATCH, no PUT, well past the old 300 ms debounce window — the Apply button reflects the staged count (Apply N changes with a data-dirty-count attribute), and applying an index-invalidating change (chunking / embedding / tokenization) shows a confirmation that names the index and the section before any write. Cancelling that confirmation writes nothing, so the test mutates no config.

The dirty-count-on-load spec (dirty_count_on_load.spec.ts)

Merely visiting a config surface must stage nothing. Under the staged model, a component that "self-heals" config in a mount effect would stage a permanent edit the operator never made — the footer would read "Apply 1 change" on a page they only opened, and Apply would PUT a mutation nobody intended. This drives every config-consuming surface (the RAG subtabs, Chat settings, Infrastructure → Paths, Grafana config, Admin) with no interaction and asserts the dirty count stays exactly 0 and no config write is issued on mount.

The destructive-safety spec (destructive_safety.spec.ts)

Proves the shared confirm-dialog primitive's safety contracts: the delete-index dialog focuses the typed-confirmation input (never the destructive button), the confirm stays disabled until the corpus id is typed verbatim, and cancelling issues no DELETE — nothing is destroyed. A second scenario pins the other focus branch: a danger dialog with no typed gate (the Infrastructure → Paths save) focuses Cancel, so a stray Enter declines rather than destroys.

The dock legibility spec (legibility_dock.spec.ts)

Proves legibility invariants for tabs rendered in the narrow dock pane, with real sidebar/dock clicks and no interception: the docked glossary wraps or scrolls rather than clipping mid-word, and Get Started swapped into the dock lays out by its container's width — the onboarding container keys off the #tab-start size container, not the viewport, so a step-4 heading and its warning stay readable at dock width instead of collapsing into a vertical strip with horizontal overflow. It also asserts the main-pane rendering keeps the full container width when Get Started is the main page.

Artifacts

Temporary feature tests and results go in .tests/; permanent tests go under tests/.

```