Alibaba’s Qwen team spent the last few weeks running the same playbook a few other Chinese labs have used this summer: preview a huge model, let people poke at it through a hosted endpoint, then follow up with a full launch and a promise to open-source the weights days later. Qwen 3.8-Max first showed up on July 19 as a moving target called qwen3.8-max-preview, reachable through Alibaba’s Token Plan, Qoder, and QoderWork. On August 3, the preview period ended. Alibaba made the model broadly available with published specs and pricing, and confirmed that open weights for the Max-tier model are landing roughly a week later, sometime around August 10. This is notable on its own: every Max-class Qwen release so far, including 3.7-Max, stayed API-only. Alongside the flagship, a smaller dense checkpoint called Qwen3.8-27B is also going open-weight, built for people who want to run something serious on their own hardware rather than pay per token through an API.
The specs behind Qwen3.8-Max
The headline number is 2.4 trillion total parameters, built as a mixture-of-experts model that only activates a fraction of that at inference time. Alibaba confirmed the active-parameter count at roughly 95 billion per token, around 4 percent of the total. That puts it in a lighter serving class than Moonshot AI’s Kimi K3, a 2.8-trillion-parameter open-weight model that launched last month with about 104 billion active parameters. The two models get compared constantly right now because they landed close together and sit near the same scale, which says something about how compressed release cycles have gotten among Chinese labs this year.
Qwen3.8-Max takes text, image, and video as input and returns text. The context window tops out at 1 million tokens, with maximum input sitting at 991,000 tokens (dropping slightly to roughly 983,600 when the reasoning mode is switched on) and maximum output capped at 131,000 tokens. The reasoning budget itself can run up to 262,000 tokens. On pricing, Alibaba settled on $2 per million input tokens and $6 per million output tokens, with cached input running far cheaper: $0.25 per million tokens for implicit cache reads, and explicit cache reads dropping to $0.17 per million once a cache is created for $2.50. That gap between fresh and cached tokens means how a developer structures prompts, especially prefix stability, ends up mattering more for cost than raw prompt length. Rate limits sit at 2 million tokens per minute and 15,000 requests per minute, and the API is built to be a drop-in swap for OpenAI- or DashScope-compatible integrations. Five built-in tools ship with the Responses API: a code interpreter, web search, a web extractor, and two image-search tools for text-to-image and image-to-image lookups.
The agentic claims Alibaba is leaning on
Alibaba is pitching this release less as a chat model and more as something that can run unattended for long stretches. The company says Qwen3.8-Max spent over 10 days autonomously building a self-evolving software harness from a blank slate, incorporating user feedback and running its own tests along the way without a human in the loop. In a separate test, it reproduced a machine learning research paper from scratch: 33 rounds of GPU training over roughly 125 hours, about 7,600 lines of code written, and then it went further by proposing 18 of its own improvements that reportedly beat the original paper’s method. Alibaba also says the model was entered into a real online coding contest against 526 human teams and outperformed 87 percent of the field.
None of these figures come from an independent third party, so treat them as vendor-reported until outside benchmarks catch up. That said, some third-party numbers have already surfaced. On the Frontend Code Arena, Qwen3.8-Max debuted in fourth place overall with an Elo of 1,668, trailing Claude Opus 5 in its Max configuration at 1,705 and Kimi K3’s Max variant at 1,676, and landing roughly even with Claude Opus 5 on its High setting at 1,669. On the Vals Index, it placed second among open-weight models and tenth overall out of 43 systems tested, with a score of 66.1. On the crowdsourced Arena.AI leaderboard, it became the top-ranked Chinese model for text tasks, though it still sits behind several Anthropic models including Claude Fable 5, and it ranked second globally for vision tasks behind a Fable 5 variant.
Beyond the language model
The Qwen team didn’t stop at the LLM. Alongside Qwen3.8-Max, Alibaba introduced Qwen-AgentWorld, described as a native language world model that simulates agent environments across seven domains, including things like MCP, search, terminal, and software engineering tasks. The distinction Alibaba is drawing here is that environment modeling was baked in from continual pretraining through supervised fine-tuning and reinforcement learning, rather than added on top of a general-purpose model after the fact. There’s also a Qwen-Robot Suite, described as a foundation model line for physical-world intelligence, though details on that are still thin at this stage.
Where things stand for people who want to run it locally
Anyone who wants Qwen3.8-Max currently has to go through Alibaba’s hosted API. The open weights for both Qwen3.8-Max and Qwen3.8-27B haven’t shown up on Hugging Face or ModelScope yet, and no license has been named. Given Alibaba’s own timeline, that should change within days rather than weeks. Until then, developers looking for something they can already download and run locally are still choosing between Qwen 3.6, GLM 5.2, or Kimi K3.


Leave a Reply