TypeSafe Jev Alternatives: Local LLMs, Trained Classifiers and Other Decision APIs Compared
Jev, Drex, Milliseconds, Kev, GLiNER2.5-Decide, local LLMs and SetFit compared: what independent tests show, which run locally, and where hosted options process data.
— Craig Major
Short answer: TypeSafe Jev selects an answer from a fixed set. Alternatives include other decision APIs, open decision models, local language models, trained classifiers and cloud classification services. Choose by testing the same labelled task, review rules and data location across options. Public rankings may not transfer to your workflow.

The TypeSafe Jev explainer covers its interface. A local LLM for text classification may help meet location needs. This comparison sits within a broader business process automation guide: first identify the decision and the person who owns it, then compare ways to run that step. No Flowgrammer comparison result is available in this method edition.
The five kinds of alternatives
Decision APIs accept text and bounded answer choices, then return a structured decision. Jev, Drex, Milliseconds and hosted Kev are in this group. They can be quick to try because there is no training step, but the provider controls the service, its versions and the data path. Inspect whether the answer probability is useful on your task.
Open decision models publish weights that a team can run on its own equipment. GLiNER2.5-Decide is an example, with a multilingual variant listed by its maker. Open weights increase control over data location but leave infrastructure and security to your team. They do not prove accuracy or French capability.
Small local language models such as Qwen or Ministral can be constrained to a fixed answer list through Ollama’s structured-output feature. That can produce Jev-shaped output in format. It does not establish equivalent speed, probabilities or accuracy. Test against human labels.
Trained classifiers such as spaCy text categorisation and SetFit learn from examples you label. They suit a stable, narrow category system when you have enough trustworthy examples and can maintain it. spaCy has French pipelines; that fact does not establish performance on Quebec French. Changing categories may require new labels and training.
Cloud classification services can supply a managed model or a training path. AWS Comprehend has a Canada (Central) endpoint, and its custom-classification language list includes French plain text. Its price was not checked in the research. Azure’s custom text classification service has a stated March 31, 2029 retirement, so a new project needs to examine the successor path.
Comparison table: what is known and what is not
This table reflects the cited research captured for September 26, 2026. A processing region marked “not confirmed” is not a Canadian residency claim. Published prices and availability need a fresh check when you procure or publish.
| Option | Type | Runs locally? | Hosted processing | French support stated | Published price in the research | Independent evidence / who ran it |
|---|---|---|---|---|---|---|
| TypeSafe Jev | Decision API | No | US on reviewed channels | English best; French needs testing | US$0.042 per million input tokens; output free | JevBench and independent task tests; board changes |
| Drex | Decision API | No; weights not released at check | US under Nace’s policy | Not published in the research | Not checked | Nace’s own scoring; one independent developer task |
| Milliseconds | Decision API | No published local route | Region not published; DPA permits Canada, US or elsewhere | English answer labels required | Not checked | JevBench snapshot; no Canadian-region proof |
| Kev 4B | Research-preview model and hosted route | Weights published | OpenRouter route; specific region not confirmed | Not published in the research | Not checked | Author fine-tune example; board snapshot |
| GLiNER2.5-Decide | Open-weight decision model | Yes, including CPU per maker | Only if you choose a host; region depends on it | Multilingual variant listed | Local operating cost not checked | No result for Decide on the cited board snapshot |
| Qwen or Ministral via Ollama | Small LLM with fixed answer format | Yes | None for a genuinely local call; other tools still matter | Model-specific; not validated here | Hardware and operation not checked | No same-task comparison in this review |
| spaCy or SetFit | Classifier trained on your examples | Yes | None for a genuinely local call | spaCy French pipelines exist; Quebec result not shown | Training and operating cost not checked | Independent spaCy task test; task-specific |
| AWS Comprehend | Cloud classification | No | Canada (Central) endpoint is listed | French plain-text custom classification listed | Not checked | No Jev-like same-task result in this review |
Compare OCR, telephony, labelling, integration, review and errors alongside model prices. Local operation also has hardware and staffing costs.
What independent tests show
JevBench is an independent English-only decision-model board, with latency measured from Germany in the cited September 24 snapshot. That snapshot put Jev 1.13.0 at 64.7 capability, decider-4b v2 at 62.2 for about half the per-task cost, Milliseconds at 54.8 and 0.17 seconds, and Kev 4B and GLiNER2.5 multi near 40. Drex and GLiNER2.5-Decide were not scored in that snapshot. The live board has since changed its listed systems, so do not quote the snapshot as a current ranking. Re-check the board before publication. It is evidence about its English tasks, not your Canadian workflow.
Independent and vendor-run task tests answer narrower questions. In an independent spaCy comparison, a trained spaCy classifier scored 100% against Jev’s 97.8% on one English task; on a Swedish task, Jev scored 88.9% against spaCy’s 84.4%. The direction reversed with task and language. Distil Labs, a fine-tuning vendor, reported 0.98 for its trained model versus 0.79 for Jev on an invoice pay-or-hold task. That is a vendor-run result, not a broad comparison of all invoice systems.
A Kev author example reported a consumer-finance fine-tune moving a score from 0.804 to 0.904. It shows that task-specific labels can matter, but it was author-run. OpenRouter reported 81.0% for Jev versus 84.4% for Claude Opus 5 on 3,080 support messages, with Jev faster and cheaper in its setup. It is OpenRouter’s test on support classification. None of these results is a Flowgrammer result or a guarantee for another category set.
Calibration is another axis. An independent benchmark reported expected calibration error of 0.161 for Jev and 0.064 for a frontier LLM on its tasks. If your workflow sends low-confidence items to people, the ordering and reliability of probabilities can matter as much as headline accuracy. Test both.
Drex was announced by Nace AI in September 2026. Its Decision Index scores in the research were Nace’s own, with official scores still pending; its weights had not been released. The one independent developer test cited in the research reported 81% for Drex against 93% for Jev on that developer’s task, with faster server-side time for Drex. One task cannot settle whether either is the right option for you. Check the current model and board before buying.
Zero-shot or train on your own examples?
A zero-shot decision API can start from a written label list without a training set. That helps when the category scheme changes frequently, volumes are low or you need to learn what the decision actually is. The trade-off is that an answer set and prompt may not capture your policy. You still need labelled examples to measure mistakes and set a human review band.
A trained classifier uses your labelled examples to learn a narrower task. It can be attractive when categories are stable and you have enough examples of both routine and hard cases. Examples need consistent labels and a refresh plan. The number needed depends on category balance, text quality and error tolerance.
Run candidates on the same held-out examples; record false routes by category and review volume. A synthetic invoice test also showed that moving rules into code changed the result from 42 to 100 of 100 in the tester’s setup. The shape of the question and surrounding code can matter as much as the model name.
Keeping a classification step in Canada
The two hard paths found in the research are running the step on your own machines or using AWS Comprehend’s Canada (Central) endpoint for a supported classification task. Other hosted options in the research either process in the US or do not confirm a region. Ask each provider to confirm the actual processing region in writing.

A local model can remove a foreign transfer for that one classification call. OCR, email, storage and voice may still use other processors. Local also brings model operation, access controls, logging and retention responsibilities; it does not settle legal duties. The Canadian privacy guide explains the Jev path and Quebec questions. A checked local classifier is described in the in-house sorting step guide. This is general information, not legal advice. Check with your own lawyer about your data and your obligations.
French and Quebec French
The cited JevBench snapshot and other public boards in the research are English-only. No independent Quebec-French benchmark for these options was found. TypeSafe says Jev is strongest in English. Milliseconds requires English answer labels even where input may be multilingual. spaCy has French pipelines; GLiNER lists a multilingual variant; AWS lists French plain-text custom classification. None of those facts measures your Quebec customer or finance text.
Create a labelled sample that reflects local vocabulary, bilingual messages, incomplete forms and the categories with the highest cost of error. Compare each option separately by language and category. Keep a person in the review path while evidence is thin.
How to choose for one workflow
- Define one decision and the allowed answers. Is this document type, AP triage, lead intent or call routing? Start with the relevant pattern: documents, invoices, leads or calls.
- Record volume, languages, data sensitivity and the actual location requirement. Decide which parts of the workflow can leave your environment and which cannot.
- Count available labelled examples, including rare exceptions. Use the same held-out set for every candidate.
- Compare accuracy by category, false automatic actions, the number sent to review, latency and complete cost. A low per-call price can be outweighed by review time.
- Set the human review band and name its owner. The human-led design guide covers that boundary. Also ask what a good AI partner should tell you not to automate.
- Check version pinning, outage behaviour, contracts and data path before allowing automatic action.
The comparison should end with a workflow decision, including the possibility that simple code and a person are enough. A model is warranted only if it improves the defined job under your constraints.
What we’re testing
We’re comparing several of these options against Jev on the same Canadian test sets; results will be added here when the work is complete and reviewed. This method edition makes no Flowgrammer performance claim.
FAQ
What are the alternatives to TypeSafe Jev, and which can run locally?
Alternatives include other hosted decision APIs, open decision models, small language models constrained to fixed outputs, classifiers trained on your examples, and cloud classification services. GLiNER2.5-Decide, local Qwen or Ministral setups, and trained spaCy or SetFit classifiers can run on your own hardware. Each needs testing on the same labelled task and an operating plan.
What is Drex from Nace AI, and is it better than Jev?
Drex is Nace AI’s hosted decision model. The research found Nace’s own scores and one independent developer task that favoured Jev on accuracy while Drex was faster server-side. Its weights had not been released at that check, and its policy said US processing. Those observations do not establish a general winner. Compare current versions on your own task.
Can I use a Jev-style decision model in French or Quebec French?
Possibly, but no independent Quebec-French benchmark for the options in this guide was found. Some products state multilingual support, spaCy has French pipelines, and AWS lists French plain-text classification. These are capability listings, not measured performance on your messages. Build local French examples, inspect errors by category and keep review for unclear cases.
Which classification models keep data in Canada?
A classification step can run on your own machines. The research also found an AWS Comprehend endpoint in Canada (Central) for supported tasks. It could not confirm Canadian processing for the other hosted options listed here, including Jev. The rest of your workflow may still send data elsewhere. This is general information, not legal advice; check with your own lawyer.
Is a fine-tuned small model more accurate than Jev for invoice approval?
On one vendor-run pay-or-hold test, Distil Labs reported 0.98 for its fine-tuned model and 0.79 for Jev. That result belongs to its task and method. A separate synthetic test showed that combining defined model answers with code checks changed the outcome substantially. Test the complete invoice workflow, including holds and human approval, before choosing a model.
Can a small Llama, Qwen or Gemma model return fixed JSON decisions?
Ollama documents structured outputs that can constrain a supported small model to a defined format, including a fixed answer list. The cited research specifically identifies Qwen and Ministral examples; it does not provide a same-task comparison for Llama or Gemma. A valid JSON answer can still be wrong. Check classification accuracy, confidence behaviour, speed and operations on your own labelled cases.
What does Jev cost compared with Drex, Milliseconds and Kev?
The checked Jev price was US$0.042 per million input tokens, with output free. The research did not record comparable current price figures for Drex, Milliseconds or Kev. JevBench’s September snapshot estimated decider-4b at about half Jev’s per-task cost under its method, which is not a vendor quote. Check each current price page and include review and hosting costs.
How do I route calls in English and French without sending text to the US?
You can evaluate a decision step on your own machines or a supported Canadian cloud region, but the live voice layer has a separate data path. The research says GPT-Live sessions process in the US or EU, so adding a local classifier would not make that full call flow Canadian-resident. Test bilingual routing and ask providers for written region details. Check with your own lawyer.
Which open decision model is best on independent benchmarks?
A single “best” model is not established for your task. In the cited September JevBench snapshot, decider-4b v2 was close to Jev’s capability score at a lower estimated cost, while other open options scored differently. The live board has changed and its tests were English-only. Re-check it, then compare candidates on your own labelled data and review rules.
When should I use zero-shot classification versus training my own classifier?
Try zero-shot when categories are changing, volume is limited or you need to learn the task before collecting a training set. A trained classifier becomes attractive when categories are stable and you can label representative examples consistently. Neither choice removes the need for held-out testing and a human review band. Compare error costs and maintenance, not only first-run accuracy.
Next step
If choosing a classifier is part of a funded operational system, contact Flowgrammer about an AI Automation System and the human review it requires.