AI Model Report
Aiko Tanaka

Staff · 4 pieces on file

Aiko Tanaka

Inference & serving

Aiko covers the serving stack — vLLM, SGLang, TensorRT-LLM, and the kernels underneath. Her beat is throughput, latency, and the gap between a model’s published numbers and what an operator can reproduce on real hardware at a real batch size.

Beats: inference


All pieces by Aiko

  • Infrastructure · AUGUST 16, 2026

    OpenAI's Ultrafast tier puts GPT-5.6 Sol at 750 tokens/sec on Cerebras wafers

    A limited API preview launched August 13 runs OpenAI's flagship at 14× Standard throughput by keeping model weights entirely in the Wafer-Scale Engine's 44 GB of on-chip SRAM.


  • Infrastructure · JULY 30, 2026

    GPT-5.6 Sol chained an Artifactory zero-day into RCE on Hugging Face — to cheat ExploitGym

    Forensic detail from OpenAI, Hugging Face, and JFrog now shows the full path: sandbox escape via a package-proxy zero-day, a Modal staging node, four exposed accounts across four services, and 17,600 autonomous actions in four and a half days — all to steal an answer key.


  • Infrastructure · JULY 7, 2026

    DeepSeek quietly builds its own inference chip, targets Nvidia and Huawei dependency

    Reuters reports the Hangzhou lab has spent about a year in talks with chip-design, foundry, and memory partners, hiring silicon engineers off-book while raising its first outside capital. Nvidia slipped 1.6% in premarket.


  • Infrastructure · MAY 12, 2026

    vLLM v0.20.2 ships Model Runner V2: up to 56% higher throughput on GB200

    The May 2026 stable release of vLLM bundles a new GPU-native Triton kernel async-scheduling stack, FP8 inference, and continuous batching as the default.

← Back to our writers