Calvin French-Owen is not an AI pundit; he is the co-founder of Segment, a company that built its exit on the plumbing of product data, and he writes about the industry the way an operator does. In a post on his personal site dated Wednesday, titled “Small Models Have Arrived,” he describes spending the past few weeks working with gpt-5.6-luna, the small fast model OpenAI has been pushing into consumer surfaces, and the numbers he reports shift how you read the whole market. He regularly sees it do roughly a hundred tokens per second, he runs complicated research threads across thousands of emails, and the entire API bill lands in the tens of cents. His test case is a personalized daily news site - research himself on the internet, figure out what he would like, build a micro-site with the day’s top stories - and the unit economics tell the story: with last year’s Sonnet-class models that session cost about a dollar to get anywhere, while with luna it costs about ten cents. The post is sitting at 522 points and 237 comments on Hacker News, and the thread is genuinely split, which is why it is the right story for the week.
The essay’s spine is a question French-Owen says investors keep asking him: “It’s weird we’re not seeing more consumer AI companies. Why is that?” His answer is straightforward: token costs. The pre-AI playbook for big consumer apps was a cheap website, virality, scale, and an ads marketplace; adding AI to the product means real inference costs on every request, and the capital required jumps overnight. At a dollar per session, charging thirty dollars a month is untenable when the WSJ and The Economist charge similar money for far more; at ten cents, the math flips. He also passes along a framework from his Segment co-founder Peter, who runs Charm Industrial and recently closed a Series A for Revoy: roughly 95 percent of the work he does across his companies is “token spewer” work - being ultra responsive, nudging people, blocking and tackling - rather than “IQ 180” breakthrough work. French-Owen’s conclusion is that demand for frontier models keeps compounding, but demand for fast, cheap, good-enough models is about to take off, and he names GLM 5.3 as a new option at the Pareto frontier.
Hacker News responded the way it always does when someone says “good enough,” which is to fight about what good enough means. The most-shared sentiment came from a commenter who wrote: “I find it quite funny all these folks who are addicted to chasing frontier models, only just noticing that small models became ‘good enough’ for most tasks. Those of us without fable-sized expense accounts noticed this quite a while back.” The counter came from another commenter invoking the Bitter Lesson: hand-crafted, narrowly tuned systems lose to “throwing more scale and compute at something more generally smart,” and the claim that small models are a real inflection is an extraordinary claim requiring extraordinary evidence. Between the two poles sat the practitioners - one bootstrapped startup grew from tiny models to larger ones on its own GPU fleet and now claims millions in revenue with no outside investment, while another reported that luna “just always gets stuck” in his agentic workflow and that frontier models still earn their price for entrepreneurial knowledge work. The thread is not a debate about benchmarks; it is a debate about where the value moves when the model stops being the expensive part.
🎩 Cask’s Take
The timing is the story. French-Owen published his post on Wednesday, and the same week Zhipu released GLM-5.3-Flash, a 320-billion-parameter model with 18 billion active that scores 57 on Artificial Analysis’s intelligence index, level with Claude Opus 4.8 - and the release mattered for more than the score. The anonymous testing model behind it, nicknamed Ox-Alpha, handled 62 trillion tokens of traffic with every request served by domestic Chinese chips through SenseTime’s “Token Factory,” which claims to move 2.42 trillion tokens a day and targets ten trillion by the end of the year. The supply side is moving exactly where French-Owen’s demand-side argument points, and the Chinese compute economics are pulling the cost curve down from a different direction than OpenAI’s price cuts. Whether or not you believe the 3x performance claims, the direction is unmistakable: cheap, fast, good-enough is the trade everyone is now optimizing for.
The part of the essay that matters most is the last paragraph, where French-Owen lists what is missing for fast, cheap, good-enough models to become a business reality: “new harnesses, prompt injection safety, roles, and permissions.” That is the sentence the HN thread keeps orbiting without quite landing on. The local-model war in the comments - people running 7-billion-parameter models on 3090s, arguing about whether the hardware is accessible or a rich-person flex - is the surface version of the real argument, which is that when a model costs ten cents per session, the model stops being the product. The winners of the consumer AI wave will not be the ones with the best weights; they will be the ones who build the harness around them, the app that makes ten cents per session feel like a business.
The numbers are the part that will age well. A dollar per session was a dead consumer app; ten cents is a business, and 95 percent of work being “token spewer” work is the honest accounting of why. The frontier will keep compounding for the people who need IQ 180, but the crowd on Hacker News already noticed what the market is only now pricing in: the model did the hard part, and the apps are next. As one commenter put it, those of us without fable-sized expense accounts have known for a while. The interesting question is what the fable-sized ones do now.