decision models are not new
A new company called TypeSafe AI launched a capability called Jev, named after William Stanley Jevons, as in the Jevons paradox [1]. "Jev is TypeSafe's flagship model and the first System One model. Send state and typed questions; get structured answers your code can use directly." [2] NB. System One comes from Daniel Kahneman's Thinking, Fast and Slow [3].
The basic idea is simple. A user or application sends a structured request to Jev, and Jev responds with structured output containing a decision, a confidence, and the probability of each possible answer. For example, I can specify in my request the state: "My running shoes arrived in the wrong size. Can I swap them for a size 10?" Then specify the type: Choice. Then specify the instructions: "Which team should handle this?" And finally, specify the criteria: {"returns", "shipping", "billing"}. Jev returns probabilities over the criteria and selects the most probable criterion as its choice [4].
What's the brouhaha about? Small decision problems like this are abundant in large software systems. Using a frontier model to answer all of them is expensive, wasteful, and potentially less flexible. In my example above, I've constrained the model's search space ({"returns", "shipping", "billing"}) given the state and instructions. Jev's only job is to return probabilities and confidence. It turns out doing this well is far cheaper than using a frontier model, which is great if this type of request happens tens of thousands of times per day in my application [5]. A frontier LLM can also be told to return a one-word answer, so the saving isn't just shorter output. Jev returns all its probabilities in a single parallel pass rather than generating one token at a time, and answers several questions in the same call. Last, there is built-in flexibility. Since we have the full probability distribution, I can decide whether to act at all. In TypeSafe's own example, if Jev's confidence in its answer to "Which team should handle this?" is below 0.3, the ticket isn't routed automatically; a person decides which team gets it [4].
A brief note that there are other primitives beyond Choice: Score and Noul (presumably from Bernoulli, since it returns the probability of a yes/no outcome). See the TypeSafe documentation for more [6].
It's no wonder that open source alternatives are popping up (see Laya on Hugging Face, and Ollaya, an Ollama-style local runner for it) [7]. And more recently, OpenAI released its own Decisions API [8]. Fun fact: TypeSafe's co-founder and CEO, Diogo Almeida, is ex-OpenAI, where he worked on RLHF and InstructGPT. His recent interview is titled "Why I couldn't build Jev at OpenAI" [9].
But is this novel? Not really. Machine learning systems that predate the Generative Pre-trained Transformer (GPT) wave of models used "System One" type models to give probability distributions over candidates for tasks like Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), Entity Resolution (ER), and so on.
What is new is that those classifiers were each trained for one task, whereas Jev handles instructions and candidates generically. TypeSafe also claims a new architecture and a training method for calibration [5]. Those claims deserve independent evaluation. The problem is old; a single general model that handles it at runtime may not be.
So what's interesting here is that the paradigm of "just prompt in <natural language>!" was a bit oversold. Natural language still matters: my example's state and question are both prose. What was oversold is asking one prompt to run an entire workflow. For large-scale systems, at least, decomposing a workflow into bounded decisions and composing them in code makes systems more composable, flexible, efficient, inspectable, and, on its face, cheaper. General models don't retire that engineering pattern; they make it useful in far more places. In the back-and-forth dance between researcher and engineer, this feels like an episode of Return of the Engineer.
[1] https://en.wikipedia.org/wiki/Jevons_paradox
[2] https://docs.typesafe.ai/introduction
[3] https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow
[4] https://docs.typesafe.ai/primitives/choice
[5] https://typesafe.ai/blog/introducing-system-one-models-and-jev
[6] https://docs.typesafe.ai/primitives
[7] https://huggingface.co/convaiinnovations/laya and https://huggingface.co/ollaya-dev/laya
[8] https://community.openai.com/t/decisions-api-is-now-available-in-public-beta/1403877