Question explorer
See the questions behind the numbers
Pick a task, then filter by what happened. Each row is one question with a verdict per model; open it to read the question, each model's final answer side by side, and (on demand) the full reasoning trace. This is the material to decide what a benchmark for your use should look like.
Quick filters compare the models at the same budget (12k where both exist, otherwise 32k); the 32k columns are shown for reference. Squares follow the legend order.