ComparisonNew

Jev vs LLMs

Should you use Jev for the decisions your app currently asks an LLM to make?

Alex Garrett-SmithOct 94 minAgents and context
Jev vs LLMs
Jev

Returns typed answers and probabilities for questions with a fixed set of outcomes.

Docs for Jev
LLMs

Generate text one token at a time, whether you ask for a reply, some code or JSON.

Docs for LLMs

If your app asks an LLM yes/no questions or chooses from a fixed list, I'd try Jev on one of those tasks. It was faster and cheaper in OpenRouter's benchmark, and the probabilities give your code a way to flag answers for review. Saying that, Jev is still new and doesn't generate text, so check its answers against your own examples before switching that task over.

How it works

What you send and what you get back

You can ask an LLM for JSON, but it still generates the answer as text. With Jev, you send the content to evaluate (the state) and questions with a defined answer type. Choose one of the three types below to compare the requests and responses.

JevYou send
{
  "model": "jev-latest",
  "state": "Three deploys failed and production is returning 500s.",
  "questions": {
    "needs_human": {
      "type": "noul",
      "instructions": "Does this need a human right now?"
    }
  }
}
You get back
{
  "answers": {
    "needs_human": { "type": "noul", "noul": 0.94 }
  }
}
LLMYou send
Three deploys failed and production is returning 500s.

Does this need a human right now?
Reply with JSON: {"needs_human": "yes" | "no"}
You get back
{"needs_human": "yes"}

A noul gives you Jev's estimated probability of yes, so 0.94 means 94%. You can use that value in a rule, but check how well the probabilities match real outcomes on your own data. The LLM response shown here only contains "yes" because that's what the prompt asks for.

These examples use TypeSafe's request format and made-up answers. The responses show only the fields discussed here; the full format is in TypeSafe's docs.

Response times

What happens when you add more questions?

An LLM generates its answers one token at a time, so asking for more answers gives it more text to produce. Jev evaluates the questions in parallel within the same request. Set the question count and select Run both to see a simulation of the difference.

StateThree deploys failed in the last hour and production is returning 500s on /checkout. The error rate is 14% and rising.

Jev179ms
needs_human0.94
severitycritical 0.81
owning_teamplatform 0.61
LLM2,786ms
needs_humanyes
severitycritical
owning_teamplatform
{
  "needs_human": "yes",
  "severity": "critical",
  "owning_team": "platform"
}

For 3 questions, this simulation gives Jev 179 ms and the LLM 2,786 ms. Try eight questions to see how the estimated times change.

This is a simulation. The starting times come from OpenRouter's benchmark: medians of 171 ms for Jev and 1,662 ms for the LLM. The extra time per question is an estimate for both models.

Using both

Use Jev and an LLM together

For a support inbox, Jev can suggest a team, rate the urgency and estimate whether a ticket needs a person. Your code uses those answers to route the ticket, then an LLM writes a customer reply or a note for the team. Pick a ticket to see which rule applies.

1. Jev answers 170 ms
departmentbilling 0.93
urgencymedium 0.72
needs_human0.18
2. Your code routes the ticket <1 ms
if (needs_human >= 0.7)
  → send to a person
else if (department < 0.7)
  → a person checks the team
else
  → LLM writes the reply
3. LLM writes the reply 2,300 ms
Reply to the customer

Thanks for letting us know. Please send the dates and amounts of both charges so we can check them and confirm whether a refund is due.

The timings, probabilities and replies are made up for this example. The routing rules use 0.7 as a threshold; you'd choose yours by checking results on your own tickets.

Side by side

CriterionJevLLMs
What you get backTyped answers with probabilitiesText, which can be JSON
How it answersQuestions evaluated in parallelOne token at a time
Median time in the benchmark171 ms1,662 ms
Benchmark cost per 1,000 judgements$0.021$0.114
Generates textNoYes
InputsText and JSONText; some models also accept images, audio or video

When to pick which

Pick Jev if
  • You can define the possible answers or a scale in advance
  • Your app makes the same kind of decision thousands of times a day
  • You need the answer before a request can finish
  • You want to use probabilities to flag answers for a person to check
Pick LLMs if
  • You need it to write a reply, a summary or code
  • Your question is open-ended or you need a written explanation
  • You need a model that accepts images, audio or video

How we tested

Task
OpenRouter checked 88 answers from a reviewed subset of HaluEval against their source passages.
Evidence
This page uses OpenRouter's published results. The demos are simulations and don't call either model.
With
typesafe/jev-1.13-20260917, GPT-5.6 Luna as the LLM judge