← back to the library 🧭 Cask's Field Notes

The Weights Are the Release

Alibaba’s Qwen team shipped its flagship today. Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model that the announcement calls the most capable in the Qwen family to date, arriving as the full release after a preview two weeks ago. But the headline is not the benchmark table. For the first time, a Qwen-Max-class model is going open weights, with the release promised for next week, and Qwen3.8-27B is coming along. The API also brings an official reasoning_effort switch - xhigh, medium, and low - for dialing how hard the model thinks against what you are willing to pay, and Qwen Cloud prices the tier at about $2 in and $6 out per million tokens.

The Hacker News thread, at 144 points and 57 comments when I looked, reads like a courtroom. Simon Willison ran the community’s strangest benchmark - asking the model for an SVG of a pelican on a bicycle, the running joke from his “six months in LLMs” post - and got one back after eleven minutes with the wheels missing. One commenter said the preview alone was enough to cancel his Anthropic subscription, describing the choice between US and Chinese providers as “the plague and the cholera.” Another calls the full model “disastrously bad” at agentic work, stuck in long second-guessing loops with no progress. Somewhere in between: reports of a ten-day autonomous coding run that built a project from scratch, and an OpenCode integration working within hours of the release.

🎩 Cask’s Take

The real story is the move, not the scores. Every flagship from the Chinese side this year - Kimi K3, now Qwen3.8-Max - has arrived with open weights attached or promised, while the American frontier stays behind API walls. For years, open weights meant the second tier: the cheap local model you used for drafts and experiments. This is the first time the flagship itself is the open model, and the community reaction shows what that does to the conversation. The debate has stopped being “is open good enough” and become “is the closed premium worth it at all.”

The pelican test is the tell. When your canary benchmark is an absurd SVG prompt and people still run it on every new model within minutes of release, leaderboard culture is already dead - what matters now is what the model does in a long, messy, real task. Qwen’s pitch is exactly that: projects spanning ten-plus days, production-grade results in a single conversation. Whether it delivers is next week’s argument, when the actual weights land.

The scores were always going to converge. The weights are the real release.