Jev vs LLMs
Should you use Jev for the decisions your app currently asks an LLM to make?

Generate text one token at a time, whether you ask for a reply, some code or JSON.
Docs for LLMsIf your app asks an LLM yes/no questions or chooses from a fixed list, I'd try Jev on one of those tasks. It was faster and cheaper in OpenRouter's benchmark, and the probabilities give your code a way to flag answers for review. Saying that, Jev is still new and doesn't generate text, so check its answers against your own examples before switching that task over.
What you send and what you get back
You can ask an LLM for JSON, but it still generates the answer as text. With Jev, you send the content to evaluate (the state) and questions with a defined answer type. Choose one of the three types below to compare the requests and responses.
{ "model": "jev-latest", "state": "Three deploys failed and production is returning 500s.", "questions": { "needs_human": { "type": "noul", "instructions": "Does this need a human right now?" } } }You get back
{ "answers": { "needs_human": { "type": "noul", "noul": 0.94 } } }
Three deploys failed and production is returning 500s.
Does this need a human right now?
Reply with JSON: {"needs_human": "yes" | "no"}You get back{"needs_human": "yes"}A noul gives you Jev's estimated probability of yes, so 0.94 means 94%. You can use that value in a rule, but check how well the probabilities match real outcomes on your own data. The LLM response shown here only contains "yes" because that's what the prompt asks for.
These examples use TypeSafe's request format and made-up answers. The responses show only the fields discussed here; the full format is in TypeSafe's docs.
What happens when you add more questions?
An LLM generates its answers one token at a time, so asking for more answers gives it more text to produce. Jev evaluates the questions in parallel within the same request. Set the question count and select Run both to see a simulation of the difference.
StateThree deploys failed in the last hour and production is returning 500s on /checkout. The error rate is 14% and rising.
{
"needs_human": "yes",
"severity": "critical",
"owning_team": "platform"
}For 3 questions, this simulation gives Jev 179 ms and the LLM 2,786 ms. Try eight questions to see how the estimated times change.
This is a simulation. The starting times come from OpenRouter's benchmark: medians of 171 ms for Jev and 1,662 ms for the LLM. The extra time per question is an estimate for both models.
Use Jev and an LLM together
For a support inbox, Jev can suggest a team, rate the urgency and estimate whether a ticket needs a person. Your code uses those answers to route the ticket, then an LLM writes a customer reply or a note for the team. Pick a ticket to see which rule applies.
if (needs_human >= 0.7) → send to a person else if (department < 0.7) → a person checks the team else → LLM writes the reply
Thanks for letting us know. Please send the dates and amounts of both charges so we can check them and confirm whether a refund is due.
The timings, probabilities and replies are made up for this example. The routing rules use 0.7 as a threshold; you'd choose yours by checking results on your own tickets.
Side by side
| Criterion | Jev | LLMs |
|---|---|---|
| What you get back | Typed answers with probabilities | Text, which can be JSON |
| How it answers | Questions evaluated in parallel | One token at a time |
| Median time in the benchmark | 171 ms | 1,662 ms |
| Benchmark cost per 1,000 judgements | $0.021 | $0.114 |
| Generates text | No | Yes |
| Inputs | Text and JSON | Text; some models also accept images, audio or video |
When to pick which
- You can define the possible answers or a scale in advance
- Your app makes the same kind of decision thousands of times a day
- You need the answer before a request can finish
- You want to use probabilities to flag answers for a person to check
- You need it to write a reply, a summary or code
- Your question is open-ended or you need a written explanation
- You need a model that accepts images, audio or video
How we tested
- Task
- OpenRouter checked 88 answers from a reviewed subset of HaluEval against their source passages.
- Evidence
- This page uses OpenRouter's published results. The demos are simulations and don't call either model.
- With
- typesafe/jev-1.13-20260917, GPT-5.6 Luna as the LLM judge