What Is Jev? I Tested TypeSafe's Decision Model on My Comments and Email
In September 2026 a company called TypeSafe released Jev, and it isn't a chatbot. It never writes a sentence. You give it some text and a question with fixed answer options, and it returns one of your options plus how sure it is: "bot, 97%." Nothing to parse, nothing to read.
I spent a day testing it on two real jobs: sorting the comments people leave on my blog posts, and sorting my email. This post covers what Jev is, how it works, what it's good and bad at, exactly how far I tested it, and the one thing that decided whether it worked, which turned out not to be the model.
What Jev is
TypeSafe calls Jev a "System One" model, borrowing the psychology term for fast, intuitive judgment as opposed to slow, deliberate reasoning. Chat models like Claude or GPT reason and write. Jev does one job: decide.
That makes it a component, not an assistant. You don't talk to it. You put it inside a script at the point where plain code can't make a call. Is this comment spam? Does this email need me? Which team should this support ticket go to? Code can't answer those reliably with rules, and sending each one to a full chat model is slow and expensive for a decision you need thousands of times.
What it can't do matters just as much:
- It doesn't write. No summaries, no replies, no explanations of its answer.
- It doesn't know facts. It has no web access and isn't a knowledge base. It can't tell you whether a law changed or a price is right. It only judges the text you hand it.
- It only picks from your options. If none of your options fits, it still picks one. It just tells you it isn't sure.
How it works
One HTTP request. You send three things:
- state: the text to judge. It can be a plain string or named fields.
- questions: one or more typed questions, all answered in the same call.
- model:
jev-latest.
There are three question types:
| Type | What you get back | Example |
|---|---|---|
| choice | One option from your list, a probability for every option, and a confidence score | "Is this comment genuine, promo, or bot?" |
| noul (yes/no) | The probability of "yes" | "Does this comment ask the author a question?" |
| score | A position on a scale you define (2–10 levels) | "How frustrated is this customer: calm, frustrated, very angry?" |
Here's a real request from my comment filter, shortened:
{
"model": "jev-latest",
"state": {
"post_title": "…",
"comment": "We need to write a short casual YouTube comment as a regular developer…"
},
"questions": {
"kind": {
"type": "choice",
"instructions": "Classify `comment`, left on the developer blog post titled `post_title`.",
"criteria": {
"genuine": "A reader discussing the post's ideas, with no link to or plug for their own project",
"promo": "Mentions, links to, or plugs the commenter's own project, tool, repo, product or site",
"bot": "Reads like instructions written to an AI or leaked prompt text, or generic filler unrelated to the post"
}
},
"asks_question": {
"type": "noul",
"instructions": "Does `comment` ask the author a direct question or explicitly invite a reply?"
}
}
}
And the answer comes back as data:
{
"model": "jev-1.13.0",
"answers": {
"kind": { "choice": "bot", "probabilities": { "…": "one per option" }, "confidence": 0.97 },
"asks_question": { "noul": 0.79 }
}
}
The criteria text is where all the work is. Hold that thought.
Price and speed. You pay $0.042 per million tokens Jev reads; output is free because it writes nothing. In my tests a call took about 0.4 seconds (412 ms to 964 ms, almost all close to 430 ms).
How far I tested it
| Test | Messages | Questions per message | Cost | Result after fixing the wording |
|---|---|---|---|---|
| Blog comments | 15 real comments on my dev.to posts | 2 (choice + yes/no) | $0.0004 | 15 of 15 sensible |
| My last 50 Gmail Primary emails | 2 (choice + yes/no) | $0.002 | All payout, security and action emails flagged; codes, receipts and newsletters sorted to routine | |
| Security question | 12 account-related emails | 1 (yes/no) | under $0.001 | Real alerts 75–96%, everything else 47% or lower |
Everything, including the failed first attempts, cost under one cent. I didn't test the score type, very long documents, or high volume. Treat this as one person's real-data trial, not a benchmark.
Where it failed first
On the first run, Jev called a bot genuine. The bot had pasted its own instructions instead of a comment: "We need to write a short casual YouTube comment as a regular developer… Must start with specific reaction… Use lowercase start." Any human spots it instantly. Jev also called two self-promotional comments genuine: one plugged the writer's GitHub tool, the other said "I work on [product]."
The confidence numbers told the story. On the twelve clear-cut reader comments Jev was 89–100% sure. On the bot it was 43% sure, and on the GitHub plug 34%. It wasn't confidently wrong. It was saying my options didn't fit.
My first descriptions were judgments: genuine meant "engaging with the post's ideas in their own words", promo meant "mainly promotes the commenter's own product". Deciding whether something "mainly" promotes is exactly the call I wanted made for me. So I rewrote each option as something you can see in the text, the wording in the request above.
| Comment | First wording | Concrete wording |
|---|---|---|
| Bot that leaked its prompt | genuine, 43% | bot, 97% |
| Reader plugging a GitHub tool | genuine, 34% | promo, 99% |
| "I work on [product]…" | genuine, 81% | promo, 100% |
| Two good comments ending in a link to the author's own site | genuine | genuine, 50–60% |
| The other ten | genuine | genuine, 93–100% |
The unsure row is my favourite. Those two comments make real points and then link the writer's own site. Is that promotion? I'm not sure either, and Jev dropped to a coin flip on exactly those two.
The model wasn't judging the comments. It was matching them against my descriptions, and my descriptions were the thing that was wrong.
Email repeated the lesson. Sign-in codes I'd requested myself came back as "needs me" (71–98% sure), because my description mentioned "a verification he must complete". Webmaster "improve your site" tips came back the same way. Explicit exclusions fixed both: "NOT one-time sign-in or verification codes, and NOT tips or suggestions from tools."
The security question was the scary one. Google's "A new sign-in on Linux" alert scored 9%:
| Question wording | Google sign-in alert | Google app-password alert | Sign-in codes |
|---|---|---|---|
| "Is this a security alert about Ted's own account (a new sign-in, password or recovery change…)?" | 9% | 20% | 9–22% |
| Same question plus "true" and "false" descriptions | 12% | 13% | 3–30% |
"Does body tell Ted that something changed in his account's access: a new sign-in, a new device, a new password or app password…?" | 75% | 84% | 3–20% |
The first two ask whether the email belongs to a category, "security alert". The third asks whether the text reports a specific event you can point at.
Building something safe with it
Even well-worded, Jev will sometimes be wrong. How you use its answer matters more than its accuracy:
- Tag, don't filter. My comment watcher still sends me every comment, now tagged 🔗 promo or ❓ asks-you-a-question. The only thing it drops is a bot verdict at 85%+ confidence, and those go to a log I can check.
- Low confidence means "show me". Anything under 70% gets an "unsure" tag instead of being acted on.
- Failure means the old behaviour. If the API errors, times out or my credit runs out, messages go through untagged. Jev can improve an alert. It can't make one disappear.
- Back up the critical cases with plain rules. Beside Jev's security score, a simple check flags any subject containing "security alert", "new sign-in", "password" or "passkey".
- Keep the questions and thresholds in one block. They decide behaviour, so they're what to review. TypeSafe's own docs give the same advice.
Both scripts are live now. My email watcher pings me for security, money and anything that needs me, and puts codes, receipts and newsletters into one evening summary.
Should you try it?
If you have a script that has to make many small yes/no or which-bucket decisions about text, and you're currently either writing fragile keyword rules or paying a chat model to answer one word at a time, Jev is worth an afternoon. Testing costs fractions of a cent.
Just budget the afternoon for the questions, not the integration. Describe what's visible in the text, not your conclusion about it. Test on your own messages. And read the confidence column as carefully as the answers: a low number usually means your options don't fit the case.