Two places to put the types
Type safety finally reached the model layer this month. That's the real news, and it's older news than the launch that carried it.
Disclosure: I build in this space — a memory layer called flmnt — so I have a stake in how you think about it. Every claim below is linked; check me rather than trust me.
Jev is a classifier with an unusually good training recipe and a very clean interface. It arrived on September 15th with a $40 million seed led by DCVC, a reported $200 million valuation, and a founder who spent four years on ChatGPT at OpenAI. It has been described since as a new class of intelligence. It isn't, and TypeSafe's own documentation is more honest about that than most of the coverage.
What it does is narrow and useful. You hand it some text and a question with a fixed set of possible answers, and it returns one of them with a probability. No prose, no parsing, no chance of getting a paragraph where you expected a boolean. It doesn't write code, hold a conversation, or choose what happens next — your code does all of that, and an if statement you wrote sits between its answer and anything downstream.
Worth being precise about one thing, because it's typed in a narrower sense than the name suggests. The input is not typed. You can send a string, an object, or an array, and nothing validates it. What's typed is the answer space: yes/no, a pick from up to 255 options you define, or a position on a scale. The model still has to interpret a prose instruction to know what you meant. The type stops it from answering out of bounds. It doesn't stop it from answering the wrong question.
That distinction — bounded answer, unbounded question — is where the interesting part of this starts.
The idea is older than the launch
Constraining the shape of an answer so a program can depend on it is ordinary engineering. Contracts, schemas, IDLs, type systems. What's new is doing it to a component that guesses.
The prior art is specific, and people are naming it. Practitioners point to GLiNER, which they run in production today, to bart-large-mnli with the zero-shot entailment trick, and to ModernBERT. Hacker News commenters disputed the can't-hallucinate framing on the grounds that type safety isn't factual correctness, and noted that grammar-constrained decoding already covers much of the same interface.
An independent benchmark ran Jev against those older approaches on four standard academic datasets — Banking77 for customer intent, Yelp for star ratings, plus emotion and phishing sets. A 22-million-parameter encoder with a logistic regression head scored 93.2% on Banking77 and ran in 8 milliseconds on a CPU, beating every zero-shot approach, including Jev. Jev scored 80.1% on the same task with no training or labeled data, compared to 48.8% for the classic zero-shot baseline.
On Yelp it reversed: Jev 67.2%, the trained model 51.9%. The gap between a three-star and a four-star review lies in tone rather than vocabulary, and that's where a small, trained classifier runs out of world knowledge.
The fairest summary I found came from a practitioner on r/LocalLLaMA rather than from any launch material: it's BERT-like, with the data, compute, and training recipe of modern LLMs, none of which existed when BERT did. Nothing here is impossible with fine-tuning. The combination of accuracy, price, speed, and zero setup makes many ideas worth trying that weren't before. That's a good product. It isn't a new kind of intelligence.
Where I come at this from
I build flmnt. It's a memory layer for teams running AI coding agents. Agents record decisions as they work — what was chosen, why, what it superseded — and when an agent asks what the team decided, flmnt returns the version that's current rather than everything that was ever said. The relationships between decisions are written down at the time, by whoever made them, not inferred later from the text.
Different product, same problem. Before a model answers anything, something has to decide what it gets to see. Jev does that by scoring the material you hand it. flmnt does it by looking up what the team already settled. Jev holds nothing between calls, so it can only judge what you remembered to send, and only as much of it as fits in one request — roughly 64k of context, about 32k of that for the material itself. flmnt holds the record back to the first day of the project. A decision you made last month, a decision someone made last year and then left the company — it's all still there, and it surfaces when it's relevant without anyone remembering it exists.
Both approaches come down to what you're willing to let the model work out for itself.
Three things you can type
When people say a system is typed, they can mean three different things, and which one they mean decides what the system can guarantee.
| What's typed | What it means | Jev | flmnt |
|---|---|---|---|
| The value | The shape of one answer | Choice, score, yes/no | Entry kinds: decision, exploration, mistake, plan |
| The frame | What's being asked, and what the options mean | A type plus a written rubric, defined at the call site | The kind of entry it is, and what that kind is allowed to mean |
| The relationships | How answers relate to each other | Not modeled. Handed to your code | Causal links, and an authored edge when one decision replaces another |
Jev covers the first two and leaves the third to you, and their documentation is explicit about it. Ask it "is the customer requesting a refund?" and "is the customer requesting something other than a refund?" as two separate questions, and the two probabilities don't add up to 1. Their published example returns 0.72 and 0.47. Each question is scored on its own, with no view of the other, so the answers aren't obliged to be consistent with each other.
That isn't a defect. It's what independent estimation means, and for most classification work it doesn't matter.
It starts mattering when the answers are about each other. Here's one from my own build log. In March we agree on one signing algorithm. In August someone reverses it, and the reversal is written in a different thread by a different person. Both decisions now sit in the record, both are well written, and neither sentence mentions the other. Ask any model, calibrated or not, whether the March decision is current, and it will read a clear, confident, well-argued decision and tell you yes. It isn't wrong about the words. The information it needs isn't in the words. Somebody must have written down that the second decision replaced the first at the moment it was made.
On a two-person project, you don't need that, because you remember. The need arrives quietly, at about the point when the person asking wasn't in the room when the decision was made, and the honest answer to "is this still what we do?" stops being something anyone can recall.
What the coverage missed
Statelessness is cheap to keep and expensive to leave. Everything good about Jev's shape follows from holding nothing between calls. Holding a customer's record means tenancy, retention, deletion, residency and audit, plus the obligation to be right about what's still true. That's not a feature a model lab adds in a sprint. It's a different business.
Governance is the story nobody is writing. Jev is already resold through Vercel's AI Gateway, OpenRouter, Cloudflare Workers AI, and Netlify. A team already on one of those adds no contract, no key, and no procurement. Enterprise AI governance assumes a model is something you onboard deliberately. A decision model bought through a gateway you already use is something nobody reviews. A model that launched this month can end up making production decisions without a single approval, which is a larger risk than anything in the launch post.
A prediction, marked as one. If Jev's value concentrates in deciding which material reaches the answering model, then holding that material is the next logical step. The tell will be a permission model. The moment you store one customer's decisions alongside another's, you have to answer who is allowed to see what, and that question drags retention, deletion, and audit in behind it. A model API needs none of that. A system of record cannot ship without it.
Reception and artifact
The artifact holds up. TypeSafe claims Jev matches a frontier model on quick judgment calls. Someone on Reddit checked it, running both against standard multiple-choice exams with the frontier model's reasoning switched off, so the comparison was like-for-like. Jev held its own.
The reception is priced on provenance. One Reddit commenter described what's happening as people meeting a category for the first time and associating the whole category with the first entrant they heard of. Another described the product as an encoder with three classification heads, pointed to open-source implementations dating back to 2020, and attributed the attention to who built it. A researcher posted that they had published the same idea in March 2025 in an arXiv paper, a model, a dataset, and a package, and then watched a lab present it as a breakthrough a year later with none of those things. In fairness, their model is narrower, so it's prior art on the idea rather than the same product. The frustration is still reasonable.
Calibration is the property the category name rests on, and it's softer than advertised. On Banking77, the benchmark showed an average reported confidence of 88% against roughly 80% actual accuracy. Temperature scaling, a technique from 2017, cut that error by about two-thirds. If you're gating automated actions on the confidence number, recalibrate on your own data first.
What shipped is a strong classifier with a published list of its own failure modes. Almost no coverage quoted the list.
Limits, in both directions
From that list: Jev answers the question you wrote rather than the one you meant, reading scoping words and negations literally. Accuracy declines as the state fills with material unrelated to the question, a phenomenon they call context rot. The state is treated as data rather than as potentially hostile, so text written to steer the model can influence the answer. Ask broadly whether an email is phishing, and you get about 60% accuracy; decompose the same judgment into eight indicators that feed a small classifier, and you get 94% accuracy. The skill is in the question.
My side has a matching weakness. Writing a decision down in advance buys authority, not accuracy. A person can author the wrong decision with complete conviction, and there's no error bar on it. A model at least reports uncertainty. My answer is that a wrong decision can be corrected and the correction is recorded. Theirs is a confidence score. Both are mitigations for the same exposure, and neither is free.
On cost, because there will be interest: both approaches reduce what the model sees; both save real money, and they do so in shapes that aren't comparable.
Jev's claim is a lower unit price. At $0.042 per million input tokens, with output free, a judgment that used to cost a frontier call now costs a rounding error, and the savings are a fixed multiple no matter how old your project is. People on Reddit report the pattern working as advertised: a Jev check in front of a long-context frontier call, seconds down to milliseconds, and you only pay for the expensive call when the check fires. One reports 70–80% off a high-volume conversation-analysis workload.
Mine isn't a multiple; it's a curve. The alternative to a memory layer is pasting history into the prompt, and that cost grows with every week the project runs — until the history stops fitting in the window, at which point there's no price that buys you the behavior and teams fall back to a stale context file, which is worse than either. Retrieval is flat. About 4% of a project's accumulated history reaches the model for any given query, and that percentage decreases as the project grows, while the cost per query doesn't change. On a two-week-old project this is worth nothing, and I won't pretend otherwise. On a two-year-old one it's the difference between a working agent and a confident wrong one.
Here are my numbers and how they're produced, so they can be argued with. Across my own workspaces, the recorded history runs from about 54,000 tokens on a young project to roughly 645,000 on the oldest. Delivered context is 4,000-5,000 tokens per query at both ends. That's 90% of accumulated history never re-sent on the small one and over 99% on the large one, at a flat $0.0073 per query either way. The method is one line: savings is delivered divided by accumulated history, subtracted from one. It measures the history I avoid, which is the part I control.
What it doesn't measure is your bill. Most of what an agent spends goes on the code and tool output it reads to do the work, and neither Jev nor I touch that. Below roughly 50,000 tokens of history, the dollar difference against pasting it all in is a wash, and prompt caching narrows it further. Past the 645,000 mark, there's nothing to compare against because the alternative doesn't fit within the window at any price.
Cost is the argument people reach for first, and it isn't the one that matters. This is. Faros measured two years of telemetry across 22,000 developers: under high AI adoption, median time in PR review is up 441.5%, bugs per PR up 54%, and 31% more pull requests merge with no review at all. LinearB, across 8.1 million pull requests, found that AI-generated PRs wait 4.6 times longer for a reviewer and are accepted 32.7% of the time, compared with 84.4% for human-written work.
Better model answers don't fix that. Knowing what your team decided, and being able to show it, does. Their work has to survive a benchmark. Mine has to survive an audit. A benchmark reports that a model was probably right. An audit asks who decided, when, and what replaced it. Only one of those standards accepts a probability as an answer.
If you got this far
flmnt keeps your team's decisions current across every agent and every session, so a decision you reversed on Tuesday doesn't ship on Thursday. Decisions are authored, not inferred, which means when someone asks why something was built a certain way, there's a record rather than a guess.
Use code JEV30 for 30 days free.
→ flmnt.ai
Get new posts in your inbox. No spam, unsubscribe anytime.