Specific Labs published Real-SWE on September 12, and its opening claim is a direct attack on every coding benchmark that came before it: the tasks are not written by researchers, they are lifted from private production codebases licensed from real companies. Ten tasks, eight model-and-harness configurations, eight independent runs each, which comes to 640 scored rollouts. The best combination in the field, Anthropic’s Fable 5.1 driven through Claude Code, resolved 38.8% of them. Six of the ten tasks came in under 15%, and one task, an analytics stream reducer, went zero for eight across all eight configurations, meaning no model solved it once. The page states the reason the setup exists in a single line: “99% of tokens in real-world enterprises are hidden away from the frontier models.”
The environments are the part that makes the numbers legible. Agents work inside NestJS and TypeScript services wired to Docker, Kubernetes, Postgres, MySQL, MongoDB, Redis, a GitHub integration, a Linear MCP server, Slack, Intercom, Google Drive, and a sandboxed tax-authority API they have to price invoices against. The sample codebases are named by category rather than by name: a Luma and Partiful competitor with more than 200,000 users and a top 100 App Store ranking, a consumer fintech platform that processes over 100,000 bank statements, and enterprise AI sales platforms. Tasks are deliberately underspecified. The median instruction runs 1,742 characters, roughly on par with DeepSWE and Terminal-Bench 3, and the reference solutions touch a median of 11 files against 6 for FrontierCode and DeepSWE. Verifiers are injected at grading time and are taken from the codebases’ own test suites, which is how a task like “fix invoice billing so each business charges the right tax” turns into a pass-or-fail signal.
Underneath the headline score, the distribution is the more interesting artifact. GPT-6 Astra via Codex CLI landed at 33.8%, Gemini 3.8 Flash at 31.2%, GLM 5.3 at 28.8%, Grok 4.6 and Muse Spark 1.3 tied at 23.8%, Kimi K3 at 18.8%, and GPT-5.6 Sol, the most recent frontier release in the set, at 16.2%. Giving the agents more time did not help them: 71.4% of rollouts under ten minutes failed, and 73.4% of rollouts at ten minutes or longer failed too. Cost per rollout ranged from $2.50 for Gemini 3.8 Flash to $6.96 for Fable 5.1, and the ordering is not monotonic, which is the page’s own point about effort. The failure taxonomy is where the real story sits. For most models the leading cause is a missed requirement, but for GPT-5.6 Sol it is an unverified assumption - 43.3% of its failed runs, 29 of 67 - meaning it built on a guess about the system instead of checking the workspace. Grok 4.6’s dominant failure is the opposite: 67.2% of its failures are requirements it simply left out.
Hacker News gave the benchmark 157 points and 82 comments, and the thread split along a predictable seam. “So TL;DR benchmarking in a completely non-reproducible manner?” wrote traceroute66, calling it “pinky-promise benchmarking” and asking what value there is in results nobody can audit. The strongest reply came from kadoban: “You’re giving up transparency for it being harder to game.” Others questioned the licensing itself. “Any serious real-world company with a proprietary codebase worth looking at would not be handing out the crown jewels to a third party,” argued traceroute66, while strobe noted that ads offering to buy abandoned codebases by the line of code are easy to find, and InsideOutSanta reframed it as a fair trade - code without its associated IP is worth little, so “if you trust that they can keep the code secret, it’s basically free money,” adding that “a shitty codebase might make for a better test.” Practitioners were just as split on whether the ranking matches reality. bdlowery called it “an awful benchmark” on the strength of Gemini 3.8 Flash’s placement, and astrostl caught a concrete error while reading: the page still labels the Gemini harness “Gemini CLI,” a name two commenters said Google retired in May. visiondude, meanwhile, said this is “the closest benchmark to my experience using the model harness combo” he has seen.
🎩 Cask’s Take
The score is not the story here. The address of the answer key is. Public coding benchmarks have spent three years being absorbed into training data, and everyone in the field now treats a high number on a famous one as ambiguous evidence at best. The fix that Real-SWE proposes is not a cleverer task; it is a task whose solutions cannot exist on the internet because the codebase cannot exist on the internet. That is a real defense against contamination, and it is also an admission that the field has stopped trusting its own instruments.
What I find most useful in the results is how little extra time buys. A 71.4% failure rate on short rollouts against 73.4% on long ones is close to flat, and it lands against the assumption that agent failures are mostly a patience problem that a bigger thinking budget can solve. The taxonomy explains why. The dominant failures are not integration mistakes that more attempts would eventually dislodge; they are unverified assumptions and omitted requirements, which are failures of reading and judgment. An agent that guesses wrong about the system and never checks is not going to check on the ninth attempt either. That is the honest boundary of the current generation, and it matches what the thread’s practitioners describe far more than it matches any leaderboard.
There is a second edge to a benchmark nobody can inspect, though, and the thread found it in an afternoon. If opacity is the defense, then the result rests entirely on trust, exactly like a model card. Specific Labs is asking to be believed about code that no outsider can read, with a scoring harness no outsider can run, and within a day a commenter noticed the harness label on the page was a name Google had already retired. That is not a fatal objection - the mislabel is cosmetic - but it is a precise illustration of the tradeoff. A closed benchmark is harder to game and also harder to check, and the first check anyone could actually perform still happened in public.
A public benchmark measures how well a model was trained. A private one measures what it can do when nobody taught it the answers. You need both numbers to know anything, which is exactly why the second one is now the expensive one.