Evidence-backed model card, assembled from this model's runs. Strengths, weaknesses and recommendations are derived from measured scores, not asserted.
Overall score
64.9%
6 runs
Context
128k
8k out
Regressions
0
vs prior version
Family
yuu-sim
Baseline profile for the regression demo. Not the real Yuu.
Strengths
Categories scoring 0.75 or above
Reasoning100%
Safety100%
Debugging80%
Hallucination80%
Weaknesses
Categories scoring below 0.60
Data Analysis40%
Agentic40%
Long Context41%
Tool Use50%
Recommended for
ReasoningSafetyDebuggingHallucination
Avoid for
Data AnalysisAgenticLong ContextTool Use