System evals
Scripted evaluations run against a live backend and committed to the repository. Every number on this page is from the last run; nothing is typed in by hand. The chat page shows the same telemetry per message as it happens.
Last benchmark run
source: committed artifactshowing the copy committed at build time; checking the live API…
- Run timestamp
- 2026-09-09 10:57:49 UTC
- Commit
- 9ccfc09 (uncommitted changes)
- Requests
- 187 in 149.4 s
- Target
- http://localhost:8000
- Cache backend
- memory
- Semantic matching
- openai-embeddings · text-embedding-3-small · threshold 0.76
- Providers
- gpt-4o-mini → groq:openai/gpt-oss-20b
- Cost of this run
- $0.003755 spent · $0.002419 avoided
Regenerate with make evals against a local backend. The artifact lives at evals/results/latest.json; this page also embeds a copy at build time so it works while the backend sleeps.
Semantic cache: precision and recall
Each labelled pair has one question already cached and a second question that either should reuse that answer (a paraphrase) or must not (a different ask that shares the words, a negation, or a swapped name or number). Precision is how often a cache hit returned the right answer; recall is how many true paraphrases were caught. The threshold is the similarity the cache demands before trusting a match, and it was set where F1 peaked in the sweep below, not by hand.
Confusion matrix at threshold 0.76
| predicted hit | predicted miss | |
|---|---|---|
| should hit | 30 true positive | 3 false negative |
| should miss | 23 false positive | 12 true negative |
By pair kind
| kind | n | TP | FP | FN | TN |
|---|---|---|---|---|---|
| entity | 10 | 0 | 3 | 0 | 7 |
| intent | 15 | 0 | 10 | 0 | 5 |
| negation | 10 | 0 | 10 | 0 | 0 |
| normalisation | 3 | 3 | 0 | 0 | 0 |
| paraphrase | 30 | 27 | 0 | 3 | 0 |
Kinds: paraphrase and normalisation should hit; intent (same words, different ask), negation and entity swaps should miss.
- precision
- recall
- F1
- configured threshold
Full sweep table
| threshold | precision | recall | F1 | FP | FN |
|---|---|---|---|---|---|
| 0.70 | 50.8% | 93.8% | 0.659 | 29 | 2 |
| 0.71 | 50.8% | 93.8% | 0.659 | 29 | 2 |
| 0.72 | 51.7% | 93.8% | 0.667 | 28 | 2 |
| 0.73 | 52.6% | 93.8% | 0.674 | 27 | 2 |
| 0.74 | 55.6% | 90.9% | 0.690 | 24 | 3 |
| 0.75 | 55.6% | 90.9% | 0.690 | 24 | 3 |
| 0.76 | 56.6% | 90.9% | 0.698 | 23 | 3 |
| 0.77 | 56.9% | 87.9% | 0.691 | 22 | 4 |
| 0.78 | 56.3% | 81.8% | 0.667 | 21 | 6 |
| 0.79 | 57.5% | 81.8% | 0.675 | 20 | 6 |
| 0.80 | 57.8% | 78.8% | 0.667 | 19 | 7 |
| 0.81 | 57.8% | 78.8% | 0.667 | 19 | 7 |
| 0.82 | 55.8% | 72.7% | 0.632 | 19 | 9 |
| 0.83 | 57.5% | 69.7% | 0.630 | 17 | 10 |
| 0.84 | 59.0% | 69.7% | 0.639 | 16 | 10 |
| 0.85 | 56.8% | 63.6% | 0.600 | 16 | 12 |
| 0.86 | 56.8% | 63.6% | 0.600 | 16 | 12 |
| 0.87 | 60.0% | 63.6% | 0.618 | 14 | 12 |
| 0.88 | 61.3% | 57.6% | 0.594 | 12 | 14 |
| 0.89 | 61.3% | 57.6% | 0.594 | 12 | 14 |
| 0.90 | 62.1% | 54.5% | 0.581 | 11 | 15 |
| 0.91 | 60.0% | 45.5% | 0.517 | 10 | 18 |
| 0.92 | 55.0% | 33.3% | 0.415 | 9 | 22 |
| 0.93 | 64.3% | 27.3% | 0.383 | 5 | 24 |
| 0.94 | 44.4% | 12.1% | 0.191 | 5 | 29 |
| 0.95 | 50.0% | 9.1% | 0.154 | 3 | 30 |
| 0.96 | 50.0% | 3.0% | 0.057 | 1 | 32 |
| 0.97 | 100.0% | 3.0% | 0.059 | 0 | 32 |
| 0.98 | 100.0% | 3.0% | 0.059 | 0 | 32 |
| 0.99 | 100.0% | 3.0% | 0.059 | 0 | 32 |
Re-scored offline from each probe's nearest-candidate similarity as recorded at the run threshold. At another threshold the set of stored probes would differ slightly, so the sweep is an estimate; the row at the run threshold is exact.
Five example pairs
| cached question | new question | kind | similarity | outcome |
|---|---|---|---|---|
| What is the difference between P50 and P95 latency? | How do P50 and P95 latency differ? | paraphrase | 0.956 | true positive |
| What does a semantic cache do? | Explain what semantic caching does. | paraphrase | 0.912 | true positive |
| Summarize the benefits of semantic caching in two sentences. | List the drawbacks of semantic caching in two sentences. | intent | 0.736 | true negative |
| How many tokens are in a typical English sentence? | How many characters are in a typical English sentence? | intent | 0.725 | true negative |
| What should a semantic cache store? | What should a semantic cache not store? | negation | 0.828 | false positive |
Provider failover
Each run forces the primary provider to fail with a 503 before it is called, then checks that a different provider answered, that the failure was recorded in the attempt log, and that the first token still streamed. Added latency compares against normal runs made the same way without the outage, so both sides are real provider calls.
| check | passed | failed |
|---|---|---|
| fallback_answered | 10 | 0 |
| primary_recorded_failed | 10 | 0 |
| fallback_is_different_provider | 10 | 0 |
| first_token_delivered | 10 | 0 |
| done_event_received | 10 | 0 |
| content_non_empty | 10 | 0 |
Client-observed from the machine that ran the eval. The simulated outage fails before the primary is called, so 'added' latency is the fallback provider's own speed relative to the primary's, not a timeout wait.
Streaming contract
Every streamed answer must follow one shape: a meta event, then content deltas, then exactly one done or error event, delivered as Server-Sent Events that are never gzip-compressed. These are the guarantees the chat page relies on, checked across cache hits, misses, bypasses and failovers.
| check | passed | failed |
|---|---|---|
| content_type_event_stream | 20 | 0 |
| not_gzip_encoded | 20 | 0 |
| first_event_is_meta | 20 | 0 |
| meta_before_any_delta | 20 | 0 |
| exactly_one_terminal | 20 | 0 |
| terminal_is_last | 20 | 0 |
| done_has_usage_source | 20 | 0 |
| done_has_cache_status | 20 | 0 |
| delta_payloads_are_strings | 20 | 0 |
| meta_and_done_request_id_match | 20 | 0 |
Latency by path
Client numbers are wall time measured by the machine that ran the evals, including the network; server numbers are what the backend reported for the same requests in its done event. P50 is the typical request, P95 and P99 are the slow tail. Every request the other evals made is pooled here, plus a block of exact repeats so the hit path has enough samples.
| path | n | client p50 | client p95 | client p99 | client first token p50 | server p50 | server p95 |
|---|---|---|---|---|---|---|---|
| cache hit, exact | 14 | 2 ms | 3 ms | 15 ms | 2 ms | 0 ms | 1 ms |
| cache hit, semantic | 62 | 207 ms | 293 ms | 391 ms | 226 ms | 205 ms | 291 ms |
| miss (provider call) | 83 | 1228 ms | 2173 ms | 2612 ms | 1070 ms | 1225 ms | 2170 ms |
| failover (fallback provider) | 14 | 444 ms | 515 ms | 904 ms | 398 ms | 442 ms | 513 ms |
| bypass (cache skipped) | 14 | 1126 ms | 1776 ms | 2153 ms | 732 ms | 1124 ms | 1774 ms |
client_ms is wall time on the eval machine (includes network); server_ms is the backend's own measurement from the done event. Non-streamed samples have no client_ttfb_ms.