AI Model Report

Multimodal · SEPTEMBER 19, 2026

Qwen3.8-Omni-Flash lands with 1M-token context and a 98% cheaper audio bill

Alibaba's new native omni-modal model reads text, image, audio, and video in a single call, calls tools through function calling, and prices an hour of audio input at roughly $0.0038 — but it answers only in text.

By Lucia Castellan · Multimodal beat · September 19, 2026

Alibaba's Qwen team shipped Qwen3.8-Omni-Flash on September 18 through the Qianwen AI Platform, and the numbers it wants you to look at are the price cuts: more than 98% off per hour of audio input, and more than 93% off audio-visual, both measured against Qwen3.5-Omni-Plus. The model reads text, image, audio, and video in one call, calls tools through function calling, and holds a 1M-token context. It doesn't speak. Text-out only.

That's the shape of the release, and it's the shape of Alibaba's current omni-modal strategy. Cheap ingestion, expensive real-time voice deferred, agents front and center.

The pricing math is aggressive. International list is $0.15 per million input tokens, $0.016 per million cache-hit tokens, and $0.47 per million output tokens. At the Qwen-Omni docs' formula of 7 tokens per second of audio, an hour lands at 25,200 tokens, or roughly $0.0038 before output. A thousand-token one-page summary tacks on about $0.0005. Domestically on Bailian, multimodal input drops from roughly ¥18 per million to ¥0.8 per million. Cache-hit input runs at about a tenth of the standard rate.

Alibaba built its headline audio comparison on two minutes of source material multiplied by thirty, with video sampled at 720p and one frame per second. Real workloads will diverge, but the direction is unambiguous.

Ingestion limits are wide: up to 64 files per request, 2 GB per file by URL, two hours per file, 113 languages and dialects on audio. Maximum input runs near 991K tokens, or about 983K in thinking mode, with 131,072 tokens of output and a 262K reasoning budget. The Qwen-MM-Plugins bundle includes Video2Note, which converts hours of video into illustrated PDF notes. Plugin support is advertised for Codex, Claude Code, Qwen Code, and Gemini CLI.

The benchmark table is the marketing. Alibaba claims a 25%+ average lift across 29 evaluations; TechNode reports 26%+ across 30. WildClawBench-MM climbs 36.5 points to 71.0. AgenticVBench moves 22.3 points to 36.8. UniClawBench sits at 69.6. LongAudioSpan gains 8.3 to 82.7. OmniVideoBench gains 9.6 to 63.4. AliMeeting speaker error collapses from 88.11 to 3.35, which is the number to watch if diarization is your bottleneck.

The rough edges are visible. On launch day the Qwen-Live Harness GitHub page returned a 404. Alibaba Cloud Model Studio's model list still routed real-time conversations to qwen3.5-omni-plus-realtime. Digital Applied found no open weights or model card. The API id qwen3-8-omni-flash is live; the voice half of "omni" isn't.

Qwen2.5-Omni shipped as a 7B open-weight model in March 2025. Eighteen months on, the open-weight posture is quieter and the agent-benchmark posture is louder. The bill for listening just fell through the floor. Talking back costs extra, and hasn't arrived.

Sources

  • https://qwen.ai/blog?id=qwen3.8-omni-flash
  • https://runtimewire.com/article/alibaba-qwen3-8-omni-flash-audio-video-agents
  • https://www.orcarouter.ai/blog/qwen-3-8-omni-flash-launch
  • https://www.digitalapplied.com/blog/qwen3-8-omni-flash-omnimodal-agents-audio-video-cost
  • https://technode.com/2026/09/18/alibabas-qwen-releases-qwen3-8-omni-flash-with-1m-token-context/