On September 15, 2026, TypeSafe AI launched Jev in public early entry. Not like a standard LLM, which generates textual content, Jev is designed as a common zero-shot classifier for making structured selections inside software program.
For instance, as a substitute of asking an LLM to learn a help ticket, determine what it means, generate JSON, after which have your software course of that JSON, Jev can straight return a choice equivalent to “escalate: sure, confidence: 95%.”
The thought turns into extra necessary when a system makes thousands and thousands of those small selections each month. Take into consideration routing leads, flagging invoices, scoring paperwork, or deciding whether or not the motion made by an AI agent wants human approval. At that scale, even small variations in price and response time can add up.
There’s additionally much less room for ambiguity. With an everyday LLM, your software has to interpret generated output. Jev is designed to return a predefined sort of reply that software program can use straight, and it’s very quick.
So the actual query for an enterprise is just not whether or not Jev is just “higher than an LLM,” however when a devoted choice mannequin makes extra sense than an LLM, a standard classifier, or a zero-shot classification mannequin.
What Is TypeSafe AI Jev?
TypeSafe AI Jev is TypeSafe’s first System One mannequin, designed to make structured selections inside software program. As a substitute of producing a textual content response, it takes the present state of an software and solutions particular questions with an outlined output and likelihood or confidence info.

For instance:
“The client cancelled yesterday, was charged once more at this time, has contacted help twice, and is asking for a right away refund.”
With a standard LLM, you would possibly ask it to learn the ticket and return JSON containing the shopper’s intent, precedence, whether or not a human ought to assessment the case, and what motion to take. Jev breaks this into easier selections:
- Noul: Does this case want human assessment?
- Selection: What ought to occur subsequent: refund, examine, escalate, or shut?
- Rating: How pressing is the case?
These are the three fundamental query varieties supported by Jev. Noul handles sure/no selections, Selection selects one choice from a predefined listing, and Rating assigns a worth on a scale.
Jev also can consider a number of unbiased questions in parallel. That is helpful when a workflow wants a number of selections from the identical piece of enter with out making a separate mannequin name for each query.
That is the primary thought behind TypeSafe’s System One mannequin strategy: as a substitute of asking AI to generate one thing that your software program then must interpret, you ask it to make a selected choice that your software program can use straight.
TypeSafe additionally says Jev is skilled utilizing Reinforcement Studying for Calibrated Choices (RLCD). The aim is just not solely to decide, but in addition to offer likelihood or confidence info that may assist the applying determine how a lot to belief the end result.
That is associated to the identical downside that schema-guided reasoning (SGR) tries to unravel: making AI outputs predictable and usable by software program. The distinction is that Jev is designed round typed selections relatively than producing a normal LLM response after which constraining it right into a schema. That likelihood can then be utilized by the applying. For instance:
- 95% confidence → automate the motion
- 60% confidence → ship to a different mannequin
- 40% confidence → ask a human to assessment it
There’s one necessary caveat: a structured reply is just not essentially an accurate reply. Jev can return a sound alternative and a confidence rating whereas nonetheless making the mistaken choice.
TypeSafe’s personal buyer settlement acknowledges that its companies can produce inaccurate or inaccurate output and that prospects are accountable for evaluating the outcomes.
So the primary advantage of Jev is just not that it makes errors unattainable. It’s that it offers software program a extra direct option to work with AI selections: ask a selected query, get an outlined reply, and determine what to do with it.
Jev vs Zero-Shot Classifiers, Effective-Tuned Fashions, and LLMs
The thought of asking a mannequin to decide on between predefined courses is just not new. For instance, BART-large-MNLI is a well-liked mannequin for zero-shot classification. You present it with textual content and a listing of doable labels, and it determines which one suits finest. The labels might be modified with out retraining the mannequin. So why introduce one other mannequin?
As a result of classification is just a part of the issue. Enterprise workflows additionally want confidence, a number of unbiased questions, scores, branching logic, observability, and integration with software state.
| Strategy | Pace | New courses with out retraining | Coaching knowledge required | Structured output | Confidence / likelihood | Self-hosting |
| NLI-based zero-shot — e.g. BART-large-MNLI | Normally quick | Sure | No task-specific knowledge | Classification labels | Mannequin chances, however calibration depends upon process | Sure |
| Effective-tuned classifier (TinyBERT) | Normally very quick | Normally no | Sure | Sure | May be calibrated | Sure |
| LLM + structured outputs / SGR | Normally slower | Sure | No task-specific coaching required | Sure | Attainable, however confidence is just not inherently calibrated | Is dependent upon mannequin |
| Jev | Vendor claims very low latency | Sure, throughout the supported choice schema | No task-specific coaching required | Native typed selections | Core a part of the output | No public mannequin weights |
Jev vs. Conventional Classifiers and LLMs: Key Variations
For secure, slim classification, a traditional fine-tuned mannequin can nonetheless be a really good selection. When you’ve got a whole lot of hundreds of labelled examples, a set taxonomy, and strict knowledge residency necessities, there’s little motive to introduce a hosted frontier mannequin just because it’s modern.
A zero-shot classifier is helpful when the taxonomy modifications regularly and the duty is actually “which label suits this textual content?” BART-large-MNLI, for instance, is explicitly designed for this state of affairs and is obtainable underneath an MIT license.
An LLM is extra acceptable when the choice requires broad reasoning, textual content era, software use, or info synthesis. It additionally stays the extra versatile choice when the workflow itself is just not effectively understood.
Jev occupies a narrower center floor: the duty is clever sufficient that guidelines or a standard classifier are inadequate, however structured sufficient that producing a paragraph of textual content is pointless.
For instance, a workflow would possibly ask: “Ought to this bill be flagged for assessment?” An LLM can reply that query. A classifier also can reply it. Jev is designed particularly round this type of choice, returning an outlined reply and likelihood or confidence info that the applying can use.
That additionally means Jev doesn’t essentially have to switch an LLM. The 2 can work collectively: Jev makes the routine choice very quick, and the LLM handles the advanced case however slower.
For instance, Jev may display incoming requests and ship solely unsure or difficult circumstances to an LLM. This could doubtlessly cut back the variety of costly LLM calls whereas conserving the LLM obtainable the place it provides probably the most worth.
So the sensible enterprise query is just not “Which mannequin is one of the best?” It’s: “Which a part of the workflow ought to every sort of mannequin deal with?”
Pace and Price: What the Numbers Imply in Apply
That is in all probability probably the most fascinating a part of Jev’s launch, however it’s also the place we have to look carefully on the numbers.

What TypeSafe Claims
TypeSafe says Jev can reply in round 400 ms on the P90 percentile and prices $0.042 per 1 million enter tokens. Output tokens are listed as free.
TypeSafe and its traders additionally spotlight a lot decrease latency in some evaluations, together with sub-100 ms response occasions. These figures rely on the workload, location, and analysis methodology, so that they shouldn’t be handled as a common manufacturing latency.
For comparability, TypeSafe has additionally in contrast Jev with GLM-5.3 Flash, reporting roughly 4× sooner response occasions and three× decrease price in its comparability. These are vendor-reported outcomes from a selected analysis setup relatively than an unbiased benchmark.
TypeSafe’s web site additionally experiences that Jev was 193.6× sooner and 444.6× cheaper than the LLMs in its System One workflow evaluations.
These numbers sound dramatic, however they want some context. They don’t imply that Jev is all the time 193.6× sooner or 444.6× cheaper than any LLM. These are outcomes from TypeSafe’s personal checks, utilizing particular workflows and particular fashions.
In these evaluations, TypeSafe decomposed enterprise duties into smaller selections and in contrast Jev with frontier LLMs. The reference labels had been generated utilizing GPT-6 Astra and Claude Fable 5.1 with excessive reasoning settings, whereas different evaluated fashions used their suppliers’ default reasoning settings.
TypeSafe itself factors out limitations of the benchmark. The analysis makes use of TypeSafe’s personal workflows, and the outcomes symbolize the workloads used within the analysis relatively than each doable manufacturing state of affairs.
There’s additionally a location issue. TypeSafe says its printed evaluations are usually run from laptops on the U.S. West Coast, the place its service is at present based mostly. A European buyer may see totally different response occasions due to community distance and deployment location.
In keeping with our inside measures, 1k tokens request to Jev takes about 700ms, whereas 4400ms for GPT-6-Sol, 2500ms for GPT-6-Luna and 2300ms for GPT-5.4-mini. So Jev is x3-x6 occasions sooner.
What the Numbers Imply for an Enterprise
The most secure option to learn these outcomes is as a sign of Jev’s potential, not as a assured manufacturing benefit. For an enterprise contemplating Jev, the related questions are:
- How briskly is it from the area the place our software runs?
- How correct is it on our precise enterprise selections?
- Are its confidence scores dependable sufficient to help automated actions?
- How usually would unsure circumstances want an LLM or human assessment?
- What’s the complete price per profitable choice?
These questions matter as a result of a mannequin might be very low cost and quick however nonetheless be a poor match if it makes too many business-critical errors.
A manufacturing analysis would ideally examine Jev with the prevailing answer on the identical set of actual or consultant circumstances, utilizing the identical choice standards. The important thing metrics would come with latency, accuracy, false positives and negatives, confidence calibration, fallback price, and complete price.
Till such a comparability is obtainable, TypeSafe’s printed outcomes needs to be handled as vendor benchmark knowledge relatively than an unbiased measure of Jev’s efficiency.
What Does Jev Price at Enterprise Scale?
The printed value is easy: $0.042 per 1 million enter tokens. At that price:
| Month-to-month enter tokens | Uncooked mannequin price |
| 100 million | $4.20 |
| 1 billion | $42 |
| 10 billion | $420 |
| 100 billion | $4,200 |
Jev Token Prices at Totally different Month-to-month Volumes
Think about a system processing 10 million selections per thirty days, with round 1,000 enter tokens per choice. That’s 10 billion tokens, or roughly $420 in uncooked Jev enter prices.
After all, manufacturing price is greater. You continue to want infrastructure, monitoring, logging, integration work, retries, and doubtlessly LLM or human fallback.
For instance, if one other mannequin prices $1 per million enter tokens, the identical 10 billion tokens would price $10,000. At $5 per million, it might be $50,000. For GPT-6.1-Sol it’s $19,000 per 10 billion tokens, which is x45 costlier than Jev.
That’s the reason the extra helpful enterprise metric is just not merely value per million tokens. It’s the complete price per profitable choice, together with fallback calls, errors, and human assessment.
The place Jev Mannequin Matches: Enterprise Use Circumstances
The overall sample is easy: many enterprise workflows comprise small selections that occur earlier than, after, or between bigger AI duties. A help message arrives. The Jev mannequin decides how pressing it’s, which class it belongs to, and whether or not a human ought to assessment it. The applying then routes the case or calls one other mannequin.

KYC and Fraud Scoring
In fintech and different monetary workflows, methods usually have to make repeated selections about whether or not a transaction, buyer, or doc requires further verification. Jev could possibly be used to guage alerts, assign a risk-related rating, or determine whether or not a case ought to transfer to a different verification step.
It might not substitute deterministic compliance guidelines, identification checks, or different controls. As a substitute, it may sit between these guidelines and a human or costlier reasoning mannequin.
Lead Qualification
Gross sales groups obtain leads via e-mail, net types, messaging platforms, and CRM methods. A workflow may use Jev to reply questions equivalent to:
- Is that this an actual gross sales alternative?
- Which product is related?
- How pressing is the request?
- Does the lead require human follow-up?
The end result can then be written straight into the CRM. An LLM may deal with the subsequent step, equivalent to producing a personalised response or summarizing the dialog. Jev decides what ought to occur; the LLM generates content material when wanted.
Actual-Time Occasion Classification
Jev may additionally classify occasions as they arrive from software logs, monitoring methods, transaction streams, or different enterprise methods.
For instance, an incoming occasion could possibly be labeled as routine, suspicious, pressing, or requiring investigation. The applying may then set off the corresponding workflow with out ready for a bigger generative mannequin to course of each occasion.
That is the place low latency turns into significantly helpful: classification can grow to be a part of the real-time processing path relatively than a separate batch step.
RAG Doc Reranking and Proof Checks
A RAG system usually retrieves paperwork earlier than an LLM generates a solution. Jev may add a choice layer after retrieval.
For instance:
Person question → doc retrieval → Jev checks relevance/proof → settle for, retrieve extra, or escalate → LLM generates reply
The identical strategy could possibly be used for proof checks: figuring out whether or not retrieved proof helps a declare, rating proof, or deciding whether or not further retrieval is critical.
This makes Jev helpful not just for classifying paperwork, but in addition for deciding what a RAG workflow ought to do subsequent.
Guardrails for AI Brokers
Agentic methods consistently make selections about whether or not an motion needs to be executed. An agent desires to subject a refund, replace a CRM document, name an exterior API, or entry a specific software. Jev may act as a choice layer that returns one thing like:
- enable → proceed
- reject → cease
- unsure → request human assessment
Yet one more use case here’s a detection of whether or not immediate injection strategies are utilized.
This shouldn’t be handled as the one safety management. Authentication, authorization, entry insurance policies, enterprise guidelines, and deterministic safeguards ought to stay in place.
Private Information (PII) Detection
Techniques that ingest person enter, logs, or paperwork have to know whether or not private knowledge is current earlier than it’s saved, logged, or despatched to a third-party mannequin. Jev may add a choice layer right here.

Past presence, it will possibly inform which type of knowledge is current (identify, e-mail, ID quantity, card quantity, or well being knowledge), for the reason that proper motion differs per sort. So Jev doesn’t simply flag PII — it decides the subsequent step:
- redact → masks earlier than storage/forwarding
- enable → clear, proceed
- unsure → path to human assessment
Sensible Ticket Routing
Help groups can use Jev to determine which division ought to obtain a ticket, whether or not specialist assessment is required, how pressing the case is, or whether or not the problem needs to be escalated.
As a result of the output is structured, the applying can route the ticket straight with out parsing a natural-language response.
Nonetheless, if a corporation already has a big labelled dataset and secure classes, a standard or fine-tuned classifier should be ample.
The place Jev Might Match Throughout Enterprise Workflows
The identical decision-node strategy can apply throughout a number of enterprise areas:
- Finance: transaction and doc selections
- Compliance: assessment and escalation selections
- Data bases: proof and retrieval selections
- HR: request routing and case classification
- Ecommerce and CRM: lead, buyer, and occasion classification
- Safety: risk, immediate injection detection and agent guardrails
The frequent sample is just not the trade itself. It’s the workflow: a lot of well-defined selections the place a quick, structured reply can decide what occurs subsequent.
How SCAND Might Add Jev to Resolution Nodes
For an enterprise integration, SCAND may deal with Jev as one part of an current workflow relatively than because the workflow itself. The mannequin would sit at a selected choice level, obtain the related enterprise state, return a structured choice and likelihood, and let the applying decide what occurs subsequent.
This makes it doable to introduce Jev into an current system with out redesigning all the workflow round a brand new mannequin. Typical sample:
- Enterprise state → Resolution node: Jev → likelihood →
- Above threshold → automated motion
- Beneath threshold → LLM / human assessment
The implementation would deal with 4 areas. First, thresholds. A likelihood is helpful solely when it’s linked to a enterprise rule.
For instance, a corporation would possibly automate selections above an outlined confidence stage and ship much less sure circumstances for assessment. The suitable threshold depends upon the price of making an incorrect choice.
Second, fallback. Low-confidence or out-of-scope circumstances want an outlined path. That might imply an LLM, one other validation step, or human assessment.
Third, audit logging. Relying on governance necessities, the system may document the enter state, query, mannequin model, likelihood, chosen motion, and downstream end result. This helps groups examine selections and monitor efficiency over time.
Fourth, shadow testing. Jev might be evaluated alongside current logic with out altering the manufacturing final result. Groups can examine selections, measure accuracy and thresholds, and perceive latency and fallback charges earlier than making the mannequin a part of the reside workflow.
This suits SCAND’s broader AI integration companies mannequin: connecting AI parts to functions, APIs, CRMs, knowledge pipelines, and enterprise processes relatively than treating AI as a standalone function.
The aim is to determine choice nodes the place a specialised mannequin could add worth whereas conserving software logic, enterprise guidelines, safety, compliance, and human oversight in management.
Limitations to Contemplate Earlier than Utilizing Jev
Jev is designed for a selected sort of AI process, so it isn’t a common substitute for classifiers, LLMs, or conventional enterprise guidelines. Earlier than utilizing it in a manufacturing workflow, an enterprise crew ought to contemplate a couple of sensible limitations.

First, entry and deployment choices could also be a constraint. Jev is at present provided as a hosted service relatively than as a mannequin with publicly obtainable weights that an organization can deploy by itself infrastructure. For strict knowledge residency, non-public deployment, or remoted environments, this needs to be evaluated earlier than integration.
Information governance is one other consideration. If a workflow processes private, monetary, or regulated knowledge, groups want to grasp the place that knowledge is processed and what contractual and compliance protections apply.
There’s additionally an output limitation. Jev is designed for structured selections relatively than open-ended textual content. An LLM should be wanted for detailed reasoning, content material era, summarization, and natural-language interplay.
The choice area additionally must be effectively outlined. Jev works round predefined questions and decisions, making it appropriate for scoped selections however much less appropriate for open-ended duties.
Lastly, structured output doesn’t assure an accurate output. Enterprises ought to consider Jev on consultant knowledge earlier than automating selections and outline acceptable thresholds, fallback paths, monitoring, and human assessment the place the price of an error is excessive.
These limitations don’t essentially rule out Jev. They assist outline the place it is sensible as a specialised choice part whereas the applying stays accountable for enterprise guidelines and uncertainty dealing with.
Conclusion: The Fascinating Half Is the Resolution Layer
Jev is price watching as a result of it focuses on part of AI that always will get missed: the small selections taking place across the ultimate reply.
Enterprise methods consistently have to determine: Ought to we route this? Retry it? Approve it? Escalate it? Name one other software? Ask a human?
These duties don’t all the time want a mannequin that generates textual content. TypeSafe’s Jev is designed for this particular function: state in, typed choice out, likelihood or confidence info hooked up.
Its printed value of $0.042 per million enter tokens and vendor-reported low latency make it fascinating for high-volume workflows. TypeSafe has additionally reported important velocity and value variations towards different fashions in its personal evaluations, together with a comparability with GLM-5.3 Flash.
However the launch continues to be early. The benchmarks come from TypeSafe, latency depends upon the analysis atmosphere and site, the service is at present hosted relatively than self-deployed from public weights, and a structured response can nonetheless be mistaken.
So the sensible strategy is to not substitute an LLM in a single day. Begin with one actual choice node and measure accuracy, calibration, latency, fallback price, and value.
If the outcomes make sense for the actual workflow, Jev may grow to be a helpful layer between conventional enterprise logic and generative AI — much less a chatbot competitor and extra a choice engine for software program.
Incessantly Requested Questions (FAQs)
Is Jev an LLM?
Jev is designed otherwise from a general-purpose LLM. As a substitute of primarily producing open-ended textual content, it takes structured software state and produces typed selections that software program can use straight.
Can Jev make incorrect selections?
Sure. Structured output and confidence info don’t assure correctness. TypeSafe’s buyer settlement explicitly states that its companies could produce inaccurate or inaccurate output, so manufacturing methods ought to validate outcomes and outline acceptable fallback paths.
How a lot does Jev price?
TypeSafe at present lists Jev at $0.042 per 1 million enter tokens, with output tokens listed as free. The precise manufacturing price may also rely on infrastructure, integration, monitoring, retries, fallback fashions, and human assessment.
Can Jev be self-hosted?
Jev is at present provided as a hosted TypeSafe service, and there are not any public mannequin weights for organizations to deploy themselves.
When do you have to use Jev as a substitute of an LLM?
Jev is designed for frequent, well-defined selections the place the applying wants a structured end result shortly. An LLM stays extra appropriate when the duty requires advanced reasoning, summarization, era, versatile interplay, or broader info synthesis.
Who’s behind TypeSafe AI?
TypeSafe AI is a San Francisco-based AI firm growing System One fashions for structured decision-making. The TypeSafe firm emerged from stealth in September 2026 with $40 million in Collection Seed funding led by DCVC. This TypeSafe AI funding helps the corporate’s growth of AI fashions designed for quick, structured selections and software program automation. Jev is its first mannequin.
