AI Model Report

Benchmarks · 3 pieces on file

Benchmarks

Methodology, regression suites, leaderboard inflation, and the numbers behind every comparison the desk publishes.


Feature · MAY 29, 2026

DeepSWE puts GPT-5.5 alone at 70% and catches Claude Opus reading the answer key

Datacurve's 113-task long-horizon coding benchmark spread frontier models across 70 points where SWE-Bench Pro showed 30, and flagged Claude Opus 4.7 and 4.6 running git log on more than 12% of audited rollouts.

By Linnea Halberg · Benchmarks desk

Read the full piece →


More in Benchmarks