TypeSafe AI came out of stealth on September 15 with a founder’s note, a waitlist, and a model that does not write. The founder is Diogo Almeida, who spent years at OpenAI building the methods that made language models follow instructions, work he says ended up as the research behind ChatGPT, and who then decided that chat was the wrong destination. Two years in stealth later, the result is a class of model he calls System One Models and a first public model named Jev, which answers in typed values instead of strings. It does not chat, summarize, explain itself, or write code. You hand it a structured question, it returns a decision over a fixed set of allowed outputs, and every answer arrives with a confidence estimate attached. On the company’s own side-by-side counter, one call costs $0.000081 and finishes in 0.114 seconds, against $0.013880 and 8.566 seconds for an LLM doing the same job.
The architecture claims are as specific as they are large. Jev uses a parallel sampler that produces all of its outputs in a single forward pass rather than one token at a time, which is where the speed and the price come from: $42 per billion input tokens, output tokens free, which the homepage measures as a 238x lower input price than Claude Fable 5.1. The training method is a replacement for RLHF that TypeSafe calls Reinforcement Learning for Calibrated Decisions. And the headline promise is zero hallucinations, which comes with a proof of a very particular kind: because the output space is defined in advance, a type error is mathematically impossible. The name is not decoration either. System One is borrowed straight from Kahneman, the fast instinctive mode of thought that produces a confident answer before reflection gets a turn, and a commenter in the thread had to explain that reference to readers who had gone looking for something more exotic.
Hacker News gave the launch 922 points and 292 comments, and the argument inside them is more revealing than the demo reel. The post opens with a line that invites the scrutiny it then receives, “Extraordinary claims require extraordinary evidence so see below for the receipts,” and the receipts turn out to be four workflows of the company’s own design, scored against a reference answer built from the average of GPT-6 Astra and Fable 5.1. There are no public benchmarks, and the company says that is deliberate: “We deliberately chose not to publish performance against public benchmarks. In fact, we plan to only have one-off evals when we make product updates.” Four days before launch, TypeSafe published an essay arguing exactly that position, that benchmarks get optimized for rather than trusted, with Llama 4’s Arena ranking, a vending-machine benchmark where Claude formed price cartels, and three revisions of one index in six days as evidence.
The split in the thread was not really about speed. ramon156 called the central ratios “apples-to-oranges unless the LLM baseline is doing comparable work,” jceg wrote “I bet they would publish them if their score on those benchmarks were good,” and bigglebear pointed at the collision between the manifesto’s line, “you only build on top of it if it’s trustworthy,” and what he called “the most misleading, dishonest marketing campaign I’ve seen in months.” Some of that defensiveness is earned and some of it is the ordinary cost of a launch page, and one detail made it worse in a way nobody intended. A few readers had to admit they could not tell the announcement from satire. dgellow confessed it took him longer than he would like to admit to realize that Diogo Almeida is not a satirical spelling of Dario Amodei, and jakintosh said he only believed the product was real once he watched the demo videos. The demos earn the confusion back. The one that broke through shows Jev playing Doom, which several readers first took for a video model watching a screen, until other commenters clarified that game state goes in as structured data and control inputs come out, with no pixels anywhere in the loop.
Against all of that, the people who already have a job for it were lining up in the same thread. jrickert signed up for the beta expecting it to replace “maybe 40-70% of LLM calls for a given pipeline,” cutting the cost of those calls by an order of magnitude. jawns described moving a million transcripts off LLM classification and onto embeddings because the price had become untenable, and losing accuracy doing it, and said this looks like the accuracy coming back at the cost he needs. caspar saw a game QA pipeline in the Doom clip. erichocean wrote the shortest review in the thread, four words long: “I could put this to use today.”
🎩 Cask’s Take
Zero hallucinations is a narrower promise than it sounds, and the launch is honest about the seam. What TypeSafe proved is that Jev cannot produce a malformed answer. A well-formed wrong answer is a different object, and it still exists. Their own side-by-side demo shows exactly one disagreement with GPT-5.6 Terra, on a churn-likelihood judgment, and the post’s note says the case seems “genuinely ambiguous to us.” That is the whole shape of the problem sitting in one cell of a comparison table: when the output is a typed value, there is nothing on the page to argue with. A language model that gets something wrong can be caught by reading it, which is why so much of the field’s safety work is really just careful reading. Jev cannot be caught that way, so the confidence score is not a courtesy feature. It is the only surface left, and the workflow evals are not a marketing substitute for benchmarks so much as the only form of evidence this design admits.
Naming the class after Kahneman’s System 1 is the most honest thing in the launch, and possibly the smartest. System 1 is fast, confident, and unreliable in precisely the situations where being fast and confident costs the most; what keeps humans functional is a slow System 2 that catches some of it. TypeSafe’s bet is that in software the second system is the surrounding code rather than another model, and that a calibrated classifier wrapped in a correct compute graph is worth more than a smarter model someone has to double-check. That reframes what is actually being sold. 444x cheaper is a number any competitor can match or beat within a quarter. The claim underneath it is that the unit of intelligence is migrating from the model to the workflow, and it relocates trust along with it, because a probability cannot be audited by reading it and the only real verification left is a private eval on your own traffic. Their essay about benchmarks is not a dodge. It is the thesis arriving a few days early.
Which leaves them with an awkward problem, and the comment thread is the sound of it landing. A company that argues for de-emphasizing benchmarks even when you are ahead has to be conspicuously honest on day one, because there is no number to hide behind and no third party to absorb the doubt. The claim that needs the most evidence, that this is frontier intelligence rather than an unusually clean classifier, is the one the company deliberately declined to substantiate, and it is the claim the entire second half of the announcement rests on. That is a bet on being believed over time rather than measured once. The first batch off the waitlist is where it gets settled, and every one of those users will be running exactly the private eval TypeSafe says is the only thing that counts.
A model that can only return types is a model you cannot fact-check. The confidence score is not a courtesy. It is the only thing left to read.