Observability (Prometheus, Grafana, Loki)
-
Metrics
/metricsendpoint + Postgres exporter. -
Dashboards
Grafana embedded with default UID
tribrid-overview. -
Logs
Loki + Promtail for container logs and app logs.
Sampling
Adjust tracing.trace_sampling_rate to manage cost and overhead. Use 1.0 in dev and 0.1–0.2 in production.
Anonymous Grafana
The compose stack enables anonymous access for embeds. Harden in production.
Timestamps
DOCKER_LOGS_TIMESTAMPS=1 helps correlate events across services.
Components
| Service | Port | Notes |
|---|---|---|
| Prometheus | 59090 | Scrapes /api/metrics and the Postgres exporter |
| Grafana | 3301 | Embedded dashboard in UI |
| Loki | 53100 | Log aggregation |
| Promtail | — | Ships container/host logs |
flowchart LR
App["ragweld API"] --> METRICS["/api/metrics"]
METRICS --> PROM["Prometheus"]
PROM --> GRAF["Grafana"]
LOGS["Docker Logs"] --> PROMTAIL["Promtail"]
PROMTAIL --> LOKI["Loki"]
LOKI --> GRAF The provisioned dashboard family tracks the API's own metrics: a dedicated Chat dashboard (ragweld-chat) covers requests by outcome and error ratio (client disconnects excluded), time to first event and first text, duration, spend and cost per chat, tokens and reasoning share, feedback, and Recall gate decisions; the TriBrid Overview, TriBridRAG Metrics and Reranker Training boards were rebuilt around the same series; and the retired Codex Session Ingest board was removed. Index-size panels (tribrid_corpus_chunks, tribrid_corpus_graph_entities, tribrid_corpus_graph_relationships) are the one deliberate per-corpus exception to "no corpus labels": they are served from a cache that a /metrics scrape refreshes at most every 5 minutes and exclude runtime-managed corpora (Recall, Codex sessions). Every data panel names what an empty panel means, and every paging alert links the dashboard to open first.
Cost & Capacity dashboard
The provisioned Cost & Capacity dashboard (ragweld-cost-capacity, infra/grafana/provisioning/dashboards/cost-capacity.json) is the gateway-spend surface, and its current revision makes the time range honest:
- The default range is today, midnight UTC to now — calendar-day totals, not a rolling 24 hours — with the dashboard timezone pinned to UTC. Today (UTC) and Last 7 days header links restore the canonical ranges after you explore history; custom and absolute selections keep their actual boundaries, so a panel title never claims "today" over a historical range.
- Spend and token totals are native and reset-aware.
sum(increase(litellm_spend_metric_total[...]))andsum(increase(litellm_total_tokens_metric_total[...]))run over the selected range and a rolling 7 days ending at the selected end. Prometheusincreasehandles counter resets and extrapolates between scrapes, so these are approximate gateway totals, not a billing ledger. A missing series renders No data, never a fake zero; a recorded zero stays zero. Token totals count input + output once — cached and reasoning tokens are subsets, never added again. - Breakdowns are by model and lane. Two bar-gauge panels split selected-range spend and tokens by
model×metadata_lane, with pre-lane history and lane-less callers visible as unattributed rather than filtered away. No run/corpus Prometheus labels are invented; exact per-run reconciliation stays in the run accounting API and UI. - An export-error panel counts real log lines. A Loki panel counts terminal native OTLP "Failed to export ... batch" log lines from the LiteLLM service over the range. It is a batch diagnostic, not proof of Langfuse delivery: zero only means scoped gateway logs existed in the same range with no terminal errors; no scoped logs renders Unavailable, not zero. Recovered retries are excluded, and provider request failures are not callback-delivery failures.
Direct provider calls that bypass the gateway stay outside every panel here; the dashboard's operator-notes panel says so, and per-request reported cost stays in the run trace with estimates identified separately. Native OpenAI embeddings are no longer an exception: they route through the LiteLLM gateway (openai.text-embedding-3-small / openai.text-embedding-3-large) with per-request x-litellm-session-id and x-litellm-spend-logs-metadata headers (native_request_headers in server/observability/run_census.py), so embedding spend appears in the gateway panels, split by the index_embeddings / retrieval_embeddings / cache_embeddings lanes.
"Latest" ML-quality gauges are read from persisted runs
The four "Latest" series behind the Eval / Benchmark / Prompt Regressions dashboard — tribrid_eval_last_top1_accuracy, tribrid_eval_last_topk_accuracy, tribrid_promptfoo_last_pass_ratio, and tribrid_benchmark_last_avg_latency_ms — are scraped from the most recently persisted eval / Promptfoo / benchmark run at scrape time, not set from whichever request happened to complete a run inside the current process. What that means operationally:
- Restarting the API cannot zero them. A freshly started process reports the persisted truth immediately; there is no window where the dashboard shows a green 0% beside runs that are on disk.
- No persisted run means no series at all. An absent series is the only honest encoding of "no data" Prometheus has, so the Grafana panels render "No data" instead of 0. A genuine 0% run is still exported as 0 — and the dashboard's thresholds render it red, never green.
- The benchmark directory follows config. The reader resolves
chat.benchmark.results_path, the same field the benchmark writer uses, so moving where benchmark runs land does not strand the latency gauge on a permanent "No data".
Latency is the mean of successful calls, and every 'latest' tile carries its age
tribrid_benchmark_last_avg_latency_ms averages only the calls in the newest benchmark run that succeeded — a failed call's latency is how long it took to fail (a gateway timeout, a rejected key), not a model latency — and a run where every call failed exports no latency series at all. Each newest eval, Promptfoo and benchmark run also exports its own completion time (tribrid_eval_last_run_timestamp_seconds, tribrid_promptfoo_last_run_timestamp_seconds, tribrid_benchmark_last_run_timestamp_seconds), read from the run record rather than the file mtime, so a panel can show how old the headline number is instead of guessing.
Probe hysteresis on the observability status
Component readiness behind /api/observability/status and the in-app Operator Deck debounces flapping probes: an incident needs tracing.probe_failure_threshold (default 3, env PROBE_FAILURE_THRESHOLD) consecutive failed probes, each component shows its last-8 probe history, and surfaces the API cannot probe at all (an auth-protected ingress that redirects off-host) never count as failures. See Tracing.
The Langfuse component reports UI health only
In otel_langfuse mode the Operator Deck's Langfuse card is now the Langfuse UI component: it probes the Langfuse web endpoint for reachability and states plainly that native callback activation and generation delivery are unverified by this probe — generation export to Langfuse is owned by the gateway's native OTel callback, which a web-UI health check cannot see. The old card required this process's LANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY and read "not configured" without them, which described the wrong surface. See Tracing.
Dashboards
Mount your own Grafana provisioning under infra/grafana/provisioning to add/override dashboards and datasources.