AI Model Report

Open Source · AUGUST 4, 2026

Alibaba's Qwen3.8-Max lands at 2.4T parameters, 95B active, and #2 on Vision Arena

The new flagship activates 95 billion of its 2.4 trillion parameters per token, supports a 1M-token context, and ships open weights next week alongside a 27B companion — Alibaba's largest release to date and its most direct challenge yet to Anthropic's Fable 5.

By Lars Iverson · Open source & model weights · August 4, 2026

Alibaba on Monday unveiled Qwen3.8-Max, a 2.4-trillion-parameter sparse mixture-of-experts model that activates 95 billion parameters per token and will ship open weights next week alongside a 27-billion-parameter companion, Qwen3.8-27B. It's the company's largest model to date, and its first open release explicitly pitched at the ceiling of the closed frontier.

The scale is worth sitting with. Qwen3.5, released in February, is roughly a seventh of the new model's parameter count, per SiliconANGLE. The context window stretches to one million tokens, which Alibaba's Alizila blog describes as enough to ingest more than 200 pages of text or about 100 hours of video per request, with responses capped at 131,000 output tokens. Architecturally it extends Qwen 3.5, pairing the sparse MoE routing with what Alibaba calls a hybrid attention mechanism.

Benchmarks tell the story Alibaba wants told. Qwen3.8-Max debuts at #2 in Vision Arena, sitting directly beneath Anthropic's Fable 5, and #5 in Text Arena. On Frontend Code Arena it scored 1,668 points for fourth overall, 37 points behind the top configuration of Claude Opus 5. Bloomberg reports it also ranks ahead of Moonshot's Kimi K3, a 2.8-trillion-parameter model released only two weeks earlier, on several benchmarks.

Markets read the signal quickly. Alibaba's New York-listed shares rose 4.5% and its Hong Kong shares rose 7% on Monday, per CNBC.

The demos gesture at where Alibaba wants the conversation to go. The company says Qwen3.8-Max spent 16 days of internal testing building oh-my-cli, a self-evolving agent framework it then open-sourced on GitHub, and ran a chip design optimization task across more than 500 steps. Alibaba also claims the model outperformed human participants in the WWW2025 Multimodal Dialogue Intent Recognition Challenge. These are vendor-supplied vignettes, but they're the vignettes a company chooses when it wants to be evaluated as agentic infrastructure rather than a chatbot.

The API is already live on Alibaba Cloud Model Studio. The open-weight drop next week is the more consequential moment: a 2.4T-parameter MoE downloadable under permissive terms would be, by a comfortable margin, the largest open model ever shipped, and the 95B active count keeps inference costs inside the range that mid-sized labs can actually serve.

That's the structural point Bloomberg's follow-up flagged in its "death zone" framing on Tuesday: US labs without either frontier-pushing capability or aggressive pricing are being squeezed from both sides. Qwen3.8-Max is the squeeze in one artifact. Anthropic still holds Vision Arena's top slot and Opus 5's coding lead, but the gap is now denominated in single-digit ranks and 37-point margins, and the model closing it's about to be free to download.

Sources