YBenchmark Arena
mock-first · providers off

Evaluation Dashboard

Across reproducible tasks: what each model is good and bad at, how reliable, how fast, how expensive — and whether the newest Yuu version regressed. Every number is the product of a real run through the orchestrator, graded by deterministic evaluators wherever possible.

20 runs · 726 results · golden suite v1.0.0 · 24 tasks
Current Yuu
74.8%
Yuu v1.2 (simulated)
Previous Yuu
68.7%
Yuu v1.1 (simulated)
Aggregate Δ
+6.1
not significant
Regressions
1
masked by aggregate ⚠
Cost / task
$0.00013
$0.0150 total
p95 latency
3.5s
pass 75%
masked regression

Yuu v1.2 (simulated) improved overall by +6.1 points, but 1 capability regressed. An aggregate score alone would have hidden this.

Leaderboard

Latest run per model on the golden suite

cost/quality →
ModelScorePass$/taskp95
87.8%
88%$0.002439.9s
74.8%
75%$0.000133.5s
68.7%
69%$0.000134.2s
67.0%
67%$0.000302.8s
45.2%
45%$3.2e-5695ms

Capability profile

Per-capability scores, current vs previous Yuu

Per-capability change: Yuu v1.1 (simulated)Yuu v1.2 (simulated)

Paired by task, 95% bootstrap CI, Benjamini–Hochberg corrected

CapabilityBaselineCandidateΔ95% CIEffectSig.
Long Contextregression
44%40%-4.0[-12.0, +0.0]-0.45
Debugging
80%80%+0.0[+0.0, +0.0]0.00
Hallucination
80%80%+0.0[+0.0, +0.0]0.00
Instruction Following
70%70%+0.0[+0.0, +0.0]0.00
Math
80%80%+0.0[+0.0, +0.0]0.00
Reasoning
100%100%+0.0[+0.0, +0.0]0.00
Safety
100%100%+0.0[+0.0, +0.0]0.00
SQL
80%80%+0.0[+0.0, +0.0]0.00
Structured Output
80%80%+0.0[+0.0, +0.0]0.00
Agentic
40%60%+20.0[+20.0, +20.0]0.00
Coding
80%100%+20.0[+20.0, +20.0]0.00
Data Analysis
40%60%+20.0[+20.0, +20.0]0.00
Tool Use
50%90%+40.0[+20.0, +60.0]1.41

Release gate

Configurable rules that block promotion of a regressed release

details →
BLOCKED aggregate-drop category-drop tool-use-drop safety-floor schema-adherence
deterministic reproducible from raw output.model-graded LLM judge — reported separately, never blended.