Breakpoint

Why your app waits seconds for an LLM to answer a yes/no question

Your code needs three typed answers about a support ticket: is the customer angry, do they get a refund, which team takes it. A chat model writes a reply…

typesafe ai··PT2M10.1S

video loads only when you press play

Your code needs three typed answers about a support ticket: is the customer angry, do they get a refund, which team takes it. A chat model writes a reply…

Your code needs three typed answers about a support ticket: is the customer angry, do they get a refund, which team takes it. A chat model writes a reply token by token, your code parses it, and sometimes the answer isn't one of the options at all.

  • Jev returns a typed answer with a probability for every allowed option in one parallel pass, instead of generating text that code has to parse.
  • Constraining output to a declared set of answers rules out type errors by construction, but a well-typed answer can still be wrong.
  • TypeSafe's speed, price and accuracy figures come from its own workflow evals, which it describes as the high end of real-world gains.

There's a hot new AI model, and it's not from OpenAI or Anthropic. It's called Jev, and it could make a lot of your app's AI calls dramatically faster, because it never writes a single word. It comes from TypeSafe, a startup that just left stealth with $40 million, and its founder, Diogo Almeida, co-wrote the research that taught ChatGPT to follow instructions. He thinks software never wanted a conversation, it wanted a decision. Say a support ticket comes in, and your app needs to know if the customer's angry, gets a refund, and which team takes it. A language model writes its reply a word at a time, which takes seconds, and then your code has to parse that text and hope nothing wandered off script. Jev skips the writing. You hand it the ticket and every answer it's allowed to give, and it fills in all the blanks at once, with a probability beside every option. It can't answer outside that list, so a type error is impossible by design, and it's trained for honest odds, so when it says 90%, it should be right 9 times in 10. TypeSafe says that takes 70 to 500 milliseconds, against three seconds to over five minutes for frontier models, and reading a million tokens costs about four cents, with the answers free. On its own tests, it puts Jev at nearly 200x faster and over 400x cheaper. That started an argument. Critics called it a really smart switch statement, since cheap, fast classifiers have been around for years, and noted that TypeSafe wrote those tests itself. Almeida conceded the fair part, that a model that can't make a type error can still be confidently wrong. So it won't replace your chat model. It's for the small calls inside your app where you already know the possible answers, and at these prices, one engineer had it playing Doom, ten decisions a second, for about seven dollars an hour. It's named after William Stanley Jevons, who saw more efficient steam engines make Britain burn more coal, not less. TypeSafe is betting cheap decisions go the same way, because software never needed a model that could talk, only one that would answer.

This explainer is based on Introducing System One Models & Jev by TypeSafe AI ↗. The original reporting and technical work belong to its publisher.