# Reading the answer instead of writing it *A new kind of model decides without generating a word. This page says what a decision model is, who made it, what it claims, what its licences allow, and what happened when we ran the open versions through fourteen measured tests on questions from this workshop's own work.* *Draft v9 · 2026-09-23 (UTC) · A small (human) team and a fleet of AI agents.* **the short version:** a decision model scores a question's lettered options at its first output position: nothing to parse, a probability for free. a company founded by a researcher behind ChatGPT's training method launched it on 2026-09-15; an independent group released open weights five days later, which we tested. asked which of six guests said a line, the 27-billion-parameter open model scored 93.5 per cent in our kit's order, 89.8 to 92.6 in five shuffles; our chat model, 80.6. the speed is the card and the server: 87 milliseconds on one large card, 350 across two smaller ones where our chat model read its letter in 120; reading instead of writing saved a tenth of the time on a short question and almost nothing on a long one. the open weights are non-commercial, so without separate terms from their authors, which we have asked for, they cannot serve anything we charge for. what transfers is the method, which gets the same answer out of any model, though only a model trained for it gives a probability worth gating on; Kev, an open family whose weights a product may run, read 67 per cent on the guest question as it comes, so the seat this points to is fine-tuned, not downloaded. that run and two others are owed as dated addenda below. --- *If you're new here: [strata→signal](https://strata2signal.com) is a small workshop (plus a friendly dog with a white patch) that builds things on its own machines and writes up what it measures. A seat here means one model held in one card's memory, answering for one of our products. An arm, on this page, is one numbered run in the bench's record: one set-up for one model or a pair of them, asked one or more sets of questions; there are fourteen. Every number on this page traces to a row in a file, and the file is linked at the end; where a claim is somebody else's, it is quoted and linked.* ## How a decision model decides Every language model, at every step, holds a score for every token in its vocabulary (a token is a word or a piece of one). Generation means picking one, appending it, and going again. A chat model asked "which of these six guests said this line?" picks a first token, then a second, and eventually produces a name. A decision model is asked the same question with the options lettered A to F and told to answer with a letter; then, instead of generating anything, the caller reads the scores of exactly those six letters at the first output position, before any token is chosen. A fixed formula turns the six numbers into probabilities that sum to one. The model's whole opinion is on the table after one read of the question, and nothing has been written. The catch, and this page measures it, is that a chat model can be read at the same position and does not put its opinion there. It puts a bracket there, or the start of a sentence, or a marker for its own thinking channel; the letters sit somewhere down its list, and when four of the six are nowhere near the top, the reading has to plug in a made-up minimum score for them. It is still usually right, but the probabilities are fiction. A decision model is a chat model retrained so that the first position is where the decision lives. That retraining is the product; everything else is a prompt format. Three things follow: **nothing to parse**, since the letters were scored and no schema was forced onto a text generator; **a number for every option, not just the winner**, so a gate can be set at 0.9 rather than at "the model said yes" and a 0.5 can be sent to a person; and **the cost is the prompt**, one pass over the question plus one position, so on a short question the work is reading and nothing else. None of this is new as a trick; multiple-choice benchmarks have graded language models by first-position letter scores for years. The new claim is that a model large enough to be asked anything can do it well zero-shot, with no training on the task at hand, and that this is a product rather than an evaluation method. ## Who built it, and what they claim Diogo Almeida was a researcher at Google Brain and then at OpenAI. [His company's own page](https://typesafe.ai/team) says he "co-invented RLHF and InstructGPT, the methods that lead to ChatGPT and GPT4": reinforcement learning from human feedback made a text-completion engine answer the way a person would want, and [InstructGPT](https://arxiv.org/abs/2203.02155) was the 2022 paper. [TechCrunch](https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/), reporting on 2026-09-18, describes him as an OpenAI researcher who "helped build the chatbot," says he "was disappointed" by what it could not do for software, and says he left OpenAI two years before that report to start a company aimed at the problem. Two quotes they carry are the thesis: "We have lightning in a bottle, and yet it is not useful," and, of models optimised for human language, "computers speak a different language." The company is TypeSafe AI. [Its press release](https://www.morningstar.com/news/business-wire/20260915525333/typesafe-ai-emerges-from-stealth-with-40m-in-funding-with-new-model-for-composable-ai) of 2026-09-15 announced it out of stealth with $40 million in seed funding led by DCVC, names Erik Gafni and Sasha Sheng as co-founders, says the company was founded in 2024, and puts the model into early access the same day at "less than 100 milliseconds of latency." The argument: software does not read, it branches, and wants one of a known set of answers with a number saying how sure, which a chat model gives only through structured output forced onto a text generator. Almeida's claim is that a model built to decide rather than to write is a different class of thing, named after Daniel Kahneman's term for fast, automatic judgement: a System One model, in [TypeSafe's launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev), is "a new class of frontier models built to make fast, structured decisions that software can use directly," and the pitch is a function signature, "unstructured state in, typed probabilistic decisions out." The launch post's claims: - **It cannot hallucinate.** Jev "gives up string generation" and in exchange is "optimized for structured outputs and *can't* hallucinate"; "The model never makes type errors"; "Schema matching is guaranteed." - **It is fast.** "40x-200x faster for the same levels of frontier intelligence for System One shaped queries." The company's homepage, read on 2026-09-22, puts 193.6 times faster and 444.6 times cheaper on its own workflows against the frontier models it compares with; the press release says "up to 100 times faster and less expensive than other frontier models." - **It is cheap in a new way, at launch prices.** "Input tokens: $0.042 / MTok ($42 per billion tokens). Output tokens: FREE (too cheap to meter)." For scale, the same post quotes existing chat models at $0.20 to $10 per million input tokens. - **It is calibrated.** "Calibrated: higher confidence means higher accuracy." A probability of 0.9 should be right nine times in ten. - **It is one pass.** It "Generates all outputs in a single query," so a hundred questions about one document cost one read of it. - **It is trained differently.** TechCrunch reports that "Almeida says Jev is trained exclusively on synthetic data using a technique he calls 'reinforcement learning from calibrated decisions.'" As of 2026-09-22 the company had not disclosed the architecture or published the weights. It is the wrong tool for anything that needs writing: chat, code, an explanation. ## Where the claims narrow, and who says so TechCrunch's headline said the model was "thrilling developers"; its report says that when Vercel swapped a chat model for Jev in a safety classifier "it got results five to 18 times more quickly and with greater accuracy," and quotes a startup's chief technology officer who found a frontier model slightly more accurate on his task but "10 to 20 times more expensive" and praised "a real probability which makes it ideal for automating workflows." The claims narrow in three places, and TypeSafe's own launch post supplies two of them. - **"Cannot hallucinate" is a property of the schema, not a measurement.** The launch post says so of its own zero-hallucination figure: "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." What is gone is the malformed answer, not the mistaken one. - **The benchmark grades Jev against other models.** The post says its evaluation workflows "were made by individuals on our model capabilities team, so some bias could exist," and that "We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic's models." So the headline multiples measure how closely Jev agrees with those two models, and how fast, not whether any of the three is right. - **The multiples differ by document,** and this one is our reading. The post says 40 to 200 times; the homepage 193.6 on the company's own workflows, which the post expects "are on the higher end of real world gains"; the press release up to 100. Each is a different comparison, and none this bench can check. Unusually for a launch, the first two caveats are TypeSafe's own, published beside an open adapter so anyone can re-run its comparison and an evaluation site with "examples, disagreements, full queries, and each workflow." Armin Ronacher, chief technology officer of Earendil, gave TechCrunch both sides: "At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it." That is the one claim a small shop can test on its own decisions. ## The open weights, and what they claim