Models ·
Qwen3.8-Flash-Next is Here, and it Beats Claude
Models ·
Qwen3.8-Flash-Next is Here, and it Beats Claude
AI summary · Generated from this article
Alibaba’s Qwen3.8-Flash-Next and Z.ai’s GLM-5.3-Flash brought downloadable open-weight models close to recent Claude Opus performance on coding, tool-use, and professional-work benchmarks. Released August 26, 2026, Qwen has roughly 180 billion total parameters but activates 6 billion per token; GLM has 320 billion total and activates 18 billion. Qwen scored 62.5 on SWE-bench Pro versus Claude Opus 4.6 Max’s 53.4, while GLM scored 84.3 on Terminal-Bench 2.1 versus Opus 4.8’s 85.0 and led on DeepSWE v1.1, 63.4 to 58.0. Qwen supports 262,144 native tokens, extendable to one million, while GLM supports 1,048,576. GLM uses the MIT License; Qwen uses the more restrictive Qwen Community License 1.0. These are not laptop models: Qwen occupies about 360 GB unquantized, and GLM roughly 640 GB at BF16. Independent testing must still verify vendor benchmarks, but recent-frontier capabilities are becoming deployable on private server-grade infrastructure.