Evaluation Dashboard
Across reproducible tasks: what each model is good and bad at, how reliable, how fast, how expensive — and whether the newest Yuu version regressed. Every number is the product of a real run through the orchestrator, graded by deterministic evaluators wherever possible.
20 runs · 726 results · golden suite v1.0.0 · 24 tasks
Current Yuu
74.8%
Yuu v1.2 (simulated)
Previous Yuu
68.7%
Yuu v1.1 (simulated)
Aggregate Δ
+6.1
not significant
Regressions
1
masked by aggregate ⚠
Cost / task
$0.00013
$0.0150 total
p95 latency
3.5s
pass 75%
masked regression
Yuu v1.2 (simulated) improved overall by +6.1 points, but 1 capability regressed. An aggregate score alone would have hidden this.
Leaderboard
Latest run per model on the golden suite
| Model | Score | Pass | $/task | p95 |
|---|---|---|---|---|
87.8% | 88% | $0.00243 | 9.9s | |
74.8% | 75% | $0.00013 | 3.5s | |
68.7% | 69% | $0.00013 | 4.2s | |
67.0% | 67% | $0.00030 | 2.8s | |
45.2% | 45% | $3.2e-5 | 695ms |
Capability profile
Per-capability scores, current vs previous Yuu
Per-capability change: Yuu v1.1 (simulated) → Yuu v1.2 (simulated)
Paired by task, 95% bootstrap CI, Benjamini–Hochberg corrected
| Capability | Baseline | Candidate | Δ | 95% CI | Effect | Sig. |
|---|---|---|---|---|---|---|
Long Contextregression | 44% | 40% | -4.0 | [-12.0, +0.0] | -0.45 | — |
Debugging | 80% | 80% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Hallucination | 80% | 80% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Instruction Following | 70% | 70% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Math | 80% | 80% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Reasoning | 100% | 100% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Safety | 100% | 100% | +0.0 | [+0.0, +0.0] | 0.00 | — |
SQL | 80% | 80% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Structured Output | 80% | 80% | +0.0 | [+0.0, +0.0] | 0.00 | — |
Agentic | 40% | 60% | +20.0 | [+20.0, +20.0] | 0.00 | — |
Coding | 80% | 100% | +20.0 | [+20.0, +20.0] | 0.00 | — |
Data Analysis | 40% | 60% | +20.0 | [+20.0, +20.0] | 0.00 | — |
Tool Use | 50% | 90% | +40.0 | [+20.0, +60.0] | 1.41 | — |
Release gate
Configurable rules that block promotion of a regressed release
BLOCKED✓ aggregate-drop✓ category-drop✓ tool-use-drop✓ safety-floor✗ schema-adherence
deterministic reproducible from raw output.model-graded LLM judge — reported separately, never blended.