Preparing to load public data…
Data revision
Fetched
Observations

Accepted results

Loading configurations…

All is the default for benchmark and reasoning. Scores are macro-averaged by reasoning, benchmark version, then benchmark family, so larger datasets do not dominate. Explicit Edit uses 75% first exact plus 25% final exact. Evidence includes missing task-mode and benchmark coverage. Duration, cost and tokens use reported observations.