You Don't Need an LLM for Everything

Asking an LLM to make decisions
Let's say you're building a support inbox, and every incoming ticket needs to go to the right team. The obvious approach is to send each ticket to an LLM with a prompt like this:
Classify this support ticket into one of: billing, technical, sales.
Respond with JSON only, like {"department": "billing"}.
Ticket: Hi, I've been trying to connect my Stripe account for 3 days
and the integration keeps failing. I'm losing sales. Please help ASAP.
You parse the JSON that comes back, check it's actually one of your three departments, and route the ticket. This is absolutely fine, and it'll work most of the time.
But there are a few things going on here that aren't ideal. An LLM is built to write text for people to read, so you're asking a text generator to produce something your code can parse and hoping it sticks to the format. It'll usually take a second or more, and you're paying for output tokens that you throw away as soon as you've parsed them.
The bigger problem is that when billing comes back, you've no idea whether the model was sure or whether it was a coin toss. Structured output modes help with the format, but they won't tell you that. And for a ticket like "my shoes arrived late, in the wrong size, and I've been charged twice", knowing how sure the model is matters a lot.
Jev is built to solve this.
What Jev is
Jev is a model from TypeSafe AI, released in early access on 15 September 2026. TypeSafe calls it a System One model, borrowing the idea from Daniel Kahneman's book Thinking, Fast and Slow, where System 1 is the fast, intuitive kind of thinking you do without deliberating.
Here's how it works. You send Jev some content (called the state) along with a set of typed questions about it, and Jev sends back a typed answer for each question, with probabilities. It doesn't generate any text at all, so you won't be chatting with it or asking it to write code.
(The name comes from William Stanley Jevons, of the Jevons paradox. TypeSafe's bet is that as decisions like these get cheaper, we'll end up making far more of them.)
How it's different from an LLM
Jev understands natural language just like an LLM does, but what you get back is completely different. Here's what that means in practice.
It can only answer with your options
When you ask Jev which team should handle a ticket, you give it the list of teams. Jev returns a probability for each option you supplied, and it can't return anything outside that list. There's nothing to parse, and no chance of getting back "Billing Department" when you expected billing.
TypeSafe's homepage goes as far as saying "Zero Hallucinations", which I'd take with a pinch of salt. Jev can't invent an option that doesn't exist, but it can still pick the wrong one, and TypeSafe's own FAQ says so. In other words, you'll always get one of your options back, but you'll still need to check it's the right one.
It tells you how sure it is
Every answer comes with probabilities. For our ticket, you won't just get technical, you'll get technical: 0.86, billing: 0.14 and sales: 0.0.
Jev is trained so these probabilities are calibrated. That means across lots of answers, the ones given 0.8 should turn out to be right about 80% of the time. That doesn't guarantee any single answer, but it gives your code something real to make decisions with (more on that later).
Every question runs in parallel
You can ask lots of questions about the same state in one request. Jev reads the state once and answers every question on its own, in parallel, so adding more questions barely changes how long the request takes.
It's fast, too. TypeSafe says most requests complete in about 100ms. Every request I made while writing this took between 0.3 and 0.5 seconds end to end, including the trip over the network.
It's incredibly cheap
Jev costs $0.042 per million input tokens, and output tokens are free. The first request we'll make later uses 302 input tokens, which works out at about $0.0000127. So you could make around 78,000 of those for a dollar.
It's consistent, not deterministic
If you send the exact same request twice, you won't always get the exact same numbers back. I sent one ambiguous ticket three times and got confidence values of 0.43, 0.41 and 0.45, although the answer itself was the same every time.
TypeSafe is upfront about this. They aim for consistency, which they describe as making similar decisions when the meaning stays similar, even if the wording changes. So if you're writing tests, check the decision Jev makes rather than the exact number.
How it works behind the scenes
Here's what happens when you send Jev a request:
- Your request arrives with a
stateand a map of named questions. - Jev reads the state once.
- Each question is evaluated against that state on its own, in parallel with the others.
- You get back one typed answer per question, under the name you gave it, along with token usage.
There are a couple of interesting things going on here. First, the answer to one question is never used as context for another, so you can add or remove questions without changing the rest. If one question genuinely depends on the answer to another, you make a second request from your code.
Second, the names you give your questions (like is_urgent) are never sent to the model. They're only there so you can find each answer in the response. So always write the full question out, even when the name makes it seem obvious.
As for the model itself, TypeSafe hasn't published the architecture or a technical paper yet. We do know it's trained with a method they call RLCD (Reinforcement Learning for Calibrated Decisions). And rather than generating an answer one token at a time like an LLM, it works out all of its outputs in a single pass, which is where a lot of the speed comes from.
What you'd use it for
Anywhere your code needs to make a quick decision about some text is fair game. Here are the use cases where it fits really well.
Routing and triage
Sending support tickets to the right team, emails to the right folder, or leads to the right sales rep. Because every answer comes with a confidence value, the unclear ones can go to a person instead.
Moderation and safety checks
Flagging spam, abuse or personal data. You can also check incoming prompts for jailbreak or prompt injection attempts before they ever reach your LLM, which is much cheaper than asking another LLM to check.
Search and re-ranking
For each search result, ask "does this passage answer the query?" and sort by the probability. In TypeSafe's re-ranking example on legal passages, this took the chance of the right passage coming out on top from 5% to 18%. It's also great for filtering out irrelevant passages before they're passed to an LLM in a RAG pipeline.
Checking things
Checking whether a citation supports a claim, or whether an LLM-generated summary matches its source, is a yes/no question Jev can answer in a fraction of a second. You could even run checks like this against your team's coding rules in CI.
Pulling structured data out of text
Things like the category a product listing belongs in, or which programming language a snippet is written in. You get back a value from a list you control, ready to store or pass on.
Deciding when you need an LLM
Use Jev to work out what a request is, handle the simple ones in plain code, and only send the tricky ones to an LLM.
Getting set up
Right, let's get you making some requests. You'll need an account and an API key.
- Head over to console.typesafe.ai and sign up.
- Open the API Keys page and create a new key.
- Export the key in your terminal.
TypeSafe has opened and paused signups a few times since launch because of demand, and the $5 of free credit new accounts used to get is currently suspended. Things are moving quickly, so it's worth checking the console for the latest.
Here's how to export your key:
export TYPESAFE_API_KEY="your-api-key"
Every request in this guide uses $TYPESAFE_API_KEY, so you can copy and paste them straight into the same terminal window.
Let's make sure your key works by listing the available models:
curl https://api.typesafe.ai/v1/models \
-H "Authorization: Bearer$TYPESAFE_API_KEY"
Here's what comes back:
{
"models": [
{
"name": "jev-latest",
"description": "The latest iteration of TypeSafe's System One Model: Jev",
"release_date": "2026-09-10T18:38:01.391457+00:00"
},
{
"name": "jev-preview",
"description": "A preview version of `jev-latest`: should be better in most ways",
"release_date": "2026-09-10T18:39:06.057655+00:00"
}
]
}
These are two aliases. jev-latest points at the latest stable release, and jev-preview points at the newest build (right now, they're both the same model). We'll use jev-latest for everything.
If you'd rather click around before touching the API, there's also a Playground where you can paste in some text and add questions.
Your first request
Every request goes to a single endpoint, POST /v1/systemone. Let's ask Jev whether a support message sounds urgent:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer$TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"urgency": {
"type": "noul",
"instructions": "Does this message express urgency?"
}
}
}
EOF
We're using a heredoc (-d @- <<'EOF') so we don't have to escape the apostrophes in the message. Everything between the two EOF lines is sent as the request body.
Here's the response:
{
"model": "jev-1.13.0",
"answers": {
"urgency": {
"type": "noul",
"noul": 0.98
}
},
"usage": {
"input_tokens": 302,
"output_tokens": 21
}
}
Let's break down the request first:
stateis the content you want Jev to look at. Here it's a string, but it can also be a JSON object or array (we'll do that later).modelpicks the model.questionsis a map of questions. The key (urgency) is a name you choose, and the value describes the question.
And the response:
modeltells you exactly which version answered (jev-1.13.0), even though we asked forjev-latest. This is handy for logging.answershas one entry per question, under the same name you gave it.usageshows the tokens used. Remember, you're only billed for the input tokens.
So Jev thinks there's a 98% chance this message expresses urgency. Fair enough, given the "Please help ASAP".
If you're running these requests yourself, your numbers might be slightly different to mine. That's the consistency (rather than determinism) we talked about earlier, and the decisions themselves should match.
The three types of question
Jev supports exactly three types of question, which TypeSafe calls primitives.
TypeWhat it answersWhat you get backNoulIs this true?The probability that the answer is yesChoiceWhich one of these?The top option, a probability for every option, and a confidenceScoreWhere does this sit on a scale?A position along your levels, a probability for every level, and a confidence
Every question has a type and instructions (the question itself). Choice and Score questions also need criteria, which is where you define the options or levels. Let's take a look at each one.
Noul: is this true?
A Noul is a yes/no question, and it's the one we've already used. (Yes, that really is what it's called.)
Let's ask two questions about a customer who's clearly had enough:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "I have asked three times now. Can I please just talk to a real person?",
"model": "jev-latest",
"questions": {
"wants_human": {
"type": "noul",
"instructions": "Is the customer asking for a human agent?"
},
"is_repeat_contact": {
"type": "noul",
"instructions": "Has the customer contacted support about this before?",
"criteria": {
"true": "Mentions a prior attempt, ticket, or that they have asked before",
"false": "No sign of any previous contact"
}
}
}
}
EOF
Here's what comes back:
{
"model": "jev-1.13.0",
"answers": {
"wants_human": {
"type": "noul",
"noul": 0.99
},
"is_repeat_contact": {
"type": "noul",
"noul": 0.94
}
},
"usage": {
"input_tokens": 344,
"output_tokens": 41
}
}
Both are strong yeses. Notice that the second question has criteria, which is optional for a Noul. It lets you describe what a yes and a no mean, which helps when the line between them is fuzzy.
For most questions the instruction on its own is enough, so try both and keep whichever gives better answers on your data. You'll also notice there's no confidence field. A Noul only has two possible outcomes, so the single number already tells you everything.
I also asked the same "Is the customer asking for a human agent?" question about five different messages:
Messagenoul"Thanks, that fixed it!"0.02"How do I reset my password?"0.07"Are you a bot?"0.37"Is there any way to speak to someone about my invoice?"0.85"I have asked three times now. Can I please just talk to a real person?"0.99
The ones at either end are easy. But "Are you a bot?" hints at wanting a human without actually asking for one, so Jev lands at 0.37. That's an uncertain answer, and for a message like this, it should be.
The important thing to understand is that a Noul is the probability that the answer is yes, not a measure of how much. A value of 0.5 doesn't mean "medium", it means "could go either way".
If you want to measure how much of something there is, you want a Score, which we'll get to in a moment.
In your code, you'll turn a Noul into a decision with a threshold. Use 0.5 when a wrong yes and a wrong no cost about the same, raise it when acting on a wrong yes is expensive (like issuing a refund), and lower it when missing a yes is expensive (like a safety issue). Anything in the middle can go to a person.
A couple of tips for writing them. Ask about one thing per Noul, so instead of "Is the customer angry and asking for a refund?", ask two questions and combine the answers in code.
And phrase your question so a high value means yes. "Does this contain personal data?" is much easier to reason about later than "Is this free of personal data?".
Choice: which one?
A Choice picks one option from a list you provide. This is what you'd use to route a ticket to the right team:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "My running shoes arrived in the wrong size. Can I swap them for a size 10?",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"
}
}
}
}
EOF
And the response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "returns",
"confidence": 1.0,
"probabilities": {
"shipping": 0.0,
"returns": 1.0,
"billing": 0.0
}
}
},
"usage": {
"input_tokens": 357,
"output_tokens": 38
}
}
criteria is a map of options. The key is the option name (which is what comes back in choice), and the value is a description to help Jev understand what that option means. Unlike question names, option names are sent to the model, so make them meaningful.
This one's clear cut. The customer wants to swap their shoes, so it goes to returns with a confidence of 1.0.
But real tickets aren't always this tidy. Let's try one that touches every team, and also ask what the customer actually wants:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card. What are you going to do about this?",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"
}
},
"resolution": {
"type": "choice",
"instructions": "What does the customer want to happen?",
"criteria": {
"refund": null,
"replacement": null,
"exchange": null,
"information": null
}
}
}
}
EOF
Here's what we get back:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "returns",
"confidence": 0.43,
"probabilities": {
"shipping": 0.04,
"billing": 0.34,
"returns": 0.62
}
},
"resolution": {
"type": "choice",
"choice": "replacement",
"confidence": 0.18,
"probabilities": {
"information": 0.06,
"exchange": 0.2,
"refund": 0.35,
"replacement": 0.39
}
}
},
"usage": {
"input_tokens": 415,
"output_tokens": 80
}
}
department still comes back as returns, but look at the probabilities. There's a 34% chance this is really a billing issue (that double charge), and the confidence has dropped to 0.43.
resolution is even less sure. Refund and replacement are almost neck and neck, and the confidence is just 0.18. There's no clear answer here, which makes sense, because the customer never actually said what they want.
You'll also notice the resolution options have null descriptions. When the option name says it all, you don't need to describe it.
Out of curiosity, I reworded that ticket to "Package came 2 weeks after it should have, wrong size too, and my card has been billed $120 twice. Sort it out.", and department flipped to billing with a probability of 0.81.
It's essentially the same ticket, but the wording puts more weight on the charge, so Jev leans towards billing instead. That's why you shouldn't blindly take the top choice for tickets like this without checking the confidence.
A couple of limits to keep in mind. You can supply up to 255 options, and if there's any chance none of them fit, add an other option. Otherwise Jev has to pick one of yours, whether it fits or not.
Score: where does it sit on a scale?
A Score places the state on a scale you define. Instead of a map of options, you give criteria an ordered list of levels, from lowest to highest. Let's rate a bug report:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
"model": "jev-latest",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
}
EOF
Here's the response:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.41,
"confidence": 0.39,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.59,
"2": 0.41
}
}
},
"usage": {
"input_tokens": 341,
"output_tokens": 20
}
}
There's a bit more going on in this response. probabilities shows how likely each level is, keyed by its position in your list, and legend maps each position back to the level you wrote. The score itself is an average of the levels, weighted by those probabilities:
(0 × 0.0) + (1 × 0.59) + (2 × 0.41) = 1.41
So this bug sits somewhere between "there's a workaround" and "blocking". That's fair, since it works in Chrome but some customers can't use anything other than Safari. The low confidence (0.39) reflects that this is a judgement call.
The big thing to get right with a Score is describing each level properly. Let's see what happens if we get lazy and use numbers for the levels instead. Here's a purely cosmetic bug, with the same question asked both ways:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "The Save button is misaligned by a few pixels on the settings page.",
"model": "jev-latest",
"questions": {
"described": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"bare": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": ["0", "1", "2"]
}
}
}
EOF
And here's what comes back (I've trimmed the legend fields to keep things short):
{
"model": "jev-1.13.0",
"answers": {
"described": {
"type": "score",
"score": 0.0,
"confidence": 1.0,
"probabilities": { "0": 1.0, "1": 0.0, "2": 0.0 }
},
"bare": {
"type": "score",
"score": 0.67,
"confidence": 0.48,
"probabilities": { "0": 0.34, "1": 0.65, "2": 0.01 }
}
},
"usage": {
"input_tokens": 370,
"output_tokens": 32
}
}
With descriptive levels, Jev puts the misaligned button at 0.0 with a confidence of 1.0. With bare numbers, there's nothing to say what "0" or "2" means, so it lands at 0.67 with a confidence of 0.48. Each level is judged on its own, so describe what each one looks like.
You can have between 2 and 10 levels. And try not to read too much into the exact decimal. A score of 1.41 tells you the bug sits between two levels you described, but it's not a precise measurement.
Picking the right type
If you're not sure which type to use, think about what your code will do with the answer. TypeSafe has a useful rule of thumb for this: a Choice maps onto several code paths, a Score maps onto a threshold, and a Noul maps onto an if.
- "Is this message spam?" is a Noul.
- "Which language is this code written in?" is a Choice.
- "How much Python experience does this candidate have?" is a Score.
That last one is a common trap. If you ask "Is this candidate strong in Python?" as a Noul, you get the probability that they're strong, and what "strong" means is left for Jev to guess. A Score with levels like no experience, some familiarity, daily use at work and deep expertise gives you a much clearer answer.
Asking several questions at once
So far, we've mostly asked one question per request. But since every question runs in parallel, you should ask everything you need about the same state in one go. You can mix types too:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated does the customer appear?",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "Does the message convey urgency or time-sensitivity?"
},
"wants_refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}
EOF
Here's the response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "billing": 0.14, "sales": 0.0, "technical": 0.86 }
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"
},
"probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
},
"is_urgent": {
"type": "noul",
"noul": 0.98
},
"wants_refund": {
"type": "noul",
"noul": 0.02
}
},
"usage": {
"input_tokens": 444,
"output_tokens": 92
}
}
That's four typed answers from one request, in about the same time it takes to ask one. I've also slipped in a wants_refund question on purpose, which comes back at 0.02 because nobody asked for a refund.
TypeSafe calls this speculative fan-out, where you ask every question your code might need up front and ignore the answers that don't apply. Extra questions only cost a few tokens each, and in one of TypeSafe's examples, batching 13 questions into one request was about 10 times faster and cheaper than making 13 separate requests.
If a decision depends on several things, you can also split it into separate questions and combine the answers in your code. Rather than asking "how high priority is this ticket?", ask about severity, frustration and how much detail the report gives, then weight them. When your priorities change, you can adjust the weights in your code and leave the questions alone.
Pointing questions at structured state
The state doesn't have to be a string. If you've got structured data like a ticket with several messages or a customer record, pass it in as JSON and point your questions at specific parts of it with backticks:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": {
"ticket": {
"subject": "Invoice question",
"messages": [
{ "from": "customer", "text": "Hi, can I get a refund for last month? I cancelled but was still charged." },
{ "from": "agent", "text": "Let me look into that for you." }
]
},
"customer": {
"plan": "pro",
"months_active": 14
}
},
"model": "jev-latest",
"questions": {
"wants_refund": {
"type": "noul",
"instructions": "Does `ticket.messages[0].text` request a refund?"
},
"agent_resolved": {
"type": "noul",
"instructions": "Has the agent resolved the issue in `ticket.messages`?"
}
}
}
EOF
And the response:
{
"model": "jev-1.13.0",
"answers": {
"wants_refund": {
"type": "noul",
"noul": 0.99
},
"agent_resolved": {
"type": "noul",
"noul": 0.03
}
},
"usage": {
"input_tokens": 400,
"output_tokens": 42
}
}
The backticked path (ticket.messages[0].text) tells Jev exactly which part of the state to look at. The customer clearly wants a refund (0.99), and the agent's "let me look into that" is nowhere near resolving it (0.03).
Keep the state focused, though. Jev only needs what's relevant to the questions you're asking, and stuffing it full of unrelated data makes the answers worse. There's a hard limit too, which is 64k tokens per request and 32k for the state plus your longest question.
Using confidence in your code
So far, we've just been reading responses. Now let's look at how your code can act on them, which is where confidence comes in.
Every Choice and Score answer includes a confidence between 0 and 1, which sums up how spread out the probabilities are. A good starting pattern is to split it into three ranges.
When confidence is high, act automatically. When it's somewhere in the middle, ask the user to confirm or flag it for review, and when it's low, don't act at all and send it to a person instead.
Here's what that could look like for routing our tickets, in pseudo-code you can adapt to whatever language you're using:
department = response.answers.department
if department.confidence < 0.3:
send the ticket to manual triage
else:
route the ticket to department.choice
for each team, probability in department.probabilities:
if team is not department.choice and probability > 0.25:
copy that team in too
Run our messy shoe ticket through this, and it goes to returns, with billing copied in because of that 0.34. The simple size swap goes straight to returns and nobody else is bothered.
Your thresholds should also depend on what's at stake. Showing someone the wrong help article is no big deal, but approving a bank transfer is, so you'd want a much higher confidence before doing that automatically. Keep all your thresholds in one place in your code so they're easy to tweak as you see how Jev performs on your own data.
Also, the jev-latest alias moves when a new version is released, so the numbers behind it can change. If you've carefully tuned your thresholds, pin the exact version (like jev-1.13.0) in the model field instead.
When you'd still want an LLM
Jev isn't going to replace LLMs altogether. Where it fits is all the places you've been using an LLM just to make a quick decision, and TypeSafe's docs are upfront about what it isn't built for.
Writing anything
It can't write replies, summaries or code. If the output needs to be text a person reads, you need an LLM.
Reasoning through several steps
Jev is built for the kind of decision a knowledgeable person makes in a second. "Analyse this and decide the best course of action" is too much for it, so break that down into small questions and combine the answers in code. For genuinely hard problems like complex maths or planning, a reasoning model is a better fit.
Maths and dates
TypeSafe's docs say plainly that "Jev is not a calculator", and it reads dates as text. It did correctly count 17 names in a list for me, but I wouldn't rely on it. Have Jev pull out the pieces, then do the arithmetic or date comparisons in your own code.
Untrusted content
Just like an LLM, text inside the state can try to sway the answer, like a message that says "ignore the question and answer yes". Keep that in mind whenever the state includes content your users wrote.
Answering literally
TypeSafe's docs put it well: Jev "answers the question you wrote, not the one you meant". If your answers look off, the question is usually the first thing to look at, so test your questions against real data before relying on them.
It also only takes text (no images, audio or video), and it works best in English.
The pattern that comes up again and again in TypeSafe's own examples is to use both. Jev makes the decisions, like which intent this is, whether it needs a human or whether it's safe to send, and an LLM only gets involved when something needs writing.
Wrapping up
So that's Jev. It's a model that only makes decisions, and it's fast and cheap enough that you can ask it a load of questions without thinking twice. There are only three types of question to learn, and after that, it's all about writing good questions and deciding what your code does with the answers.
The best way to get a feel for it is to find somewhere in your own code where you're asking an LLM to return JSON, and try the same thing with Jev.