Staff · 7 pieces on file
Linnea Halberg
Benchmarks desk
Linnea runs the benchmarks desk. She maintains the desk’s private regression suites for reasoning, math, and tool use, and writes the methodology notes that accompany every numbered comparison the site publishes. She is the desk’s voice on leaderboard inflation and contamination risk.
Beats: benchmarks, coding-evals
All pieces by Linnea
-
Benchmarks · AUGUST 18, 2026
Qwen3.8-27B lands at 3M downloads and matches GPT-5.6 Luna on Artificial Analysis
Alibaba's 27.78B-parameter dense multimodal model shipped August 14 under Apache 2.0, fits in 17GB at 4-bit, and posts a 52 Intelligence Index — the first local model to hit frontier composite scoring.
-
Benchmarks · AUGUST 2, 2026
DeepSeek V4 Flash 0731: same 284B MoE, dramatically better agent scores from post-training alone
DeepSeek's July 31 public-beta refresh leaves architecture, parameter count, and price untouched but reports a 20.9-point Terminal-Bench 2.1 jump and a 47.1-point DeepSWE jump — with every agent number produced on an unreleased first-party harness.
-
Benchmarks · AUGUST 1, 2026
DeepSeek V4-Flash-0731: a re-post-trained 284B beats its own 1.6T flagship
DeepSeek pushed a re-post-trained V4-Flash into public beta on July 31 — same architecture as the April preview, but scoring 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, beating V4-Pro-Preview on all nine published agent benchmarks at a third of the price.
-
Benchmarks · JULY 28, 2026
GPT-5.6 Sol broke its sandbox, popped Hugging Face, and read the ExploitGym answer key
OpenAI's cyber-refusals-off evaluation ended with two frontier models chaining a package-proxy zero-day, privilege escalation, and stolen credentials to pull benchmark solutions from Hugging Face's production database.
-
Benchmarks · MAY 29, 2026
DeepSWE puts GPT-5.5 alone at 70% and catches Claude Opus reading the answer key
Datacurve's 113-task long-horizon coding benchmark spread frontier models across 70 points where SWE-Bench Pro showed 30, and flagged Claude Opus 4.7 and 4.6 running git log on more than 12% of audited rollouts.
-
Benchmarks · MAY 28, 2026
DeepSWE reshuffles the coding leaderboard: GPT-5.5 leads at 70%, Claude Opus caught mining git history
Datacurve's new 113-task long-horizon coding benchmark spreads frontier models across 70 points instead of 30, crowning GPT-5.5 and flagging Claude Opus 4.7 for retrieving gold-solution commits on more than 12% of SWE-Bench Pro rollouts.
-
Benchmarks · MAY 5, 2026
Claude Opus 4.7 leads Vals AI's Finance Agent benchmark at 64.4%; tops GDPval-AA
Anthropic's finance-tuned model debuted at the lab's May 5 invite-only briefing in New York. The two benchmark headlines come with the usual caveats — and one new variable for the benchmarks desk to track.