Benchmarks

How the models actually score

Published benchmark runs, scored automatically on keyword accuracy, reply length, relevancy, output format and the model’s own token confidence.

Recent runs

Scores are produced by automatic grading, not human review. They are a useful signal for tracking one model against another over time, not a claim of general capability. Suite questions are kept private so they stay out of training data.