Flowgrammer

How to Use TypeSafe Jev for Lead Qualification: Atomic Questions, Weights in Code, Human Review

Use TypeSafe Jev to answer many small fit and intent questions per lead, keep weights in code and send unsure leads to a person. Public builds and French results.

— Craig Major

Short answer: TypeSafe Jev can answer small questions about a lead’s stated fit, pain and intent. Your own rules turn those answers into a score and route. The model’s probability describes its answer to a question, not the chance of a sale. Unclear or consequential cases should reach a person before a lead is rejected or contacted.

One lead broken into small questions; the answers feed rules in code that route it to sales, nurture or a person

This is one possible decision step inside a broader lead qualification system. It is a design pattern to test, not a Flowgrammer deployment or a claim that Jev will improve conversion on your pipeline.

Why one big question fails and many small ones work

“Is this a good lead?” hides several different judgements. A company may fit your market but have no current project. A buyer may be urgent but outside the work you provide. A short form may simply omit the information needed for a fair decision. One broad label makes it hard to see which signal drove the route or why a salesperson should trust it.

Break the job into questions that can be answered from the lead’s actual words. Is the company in the service area? Did the person describe a business problem? Is there an explicit timeframe? Are they asking to talk to a person? What information is absent? TypeSafe Jev can answer a defined Choice, Score or Yes/No question about those signals. It cannot supply facts the lead never gave you.

Write the fit criteria first. A lead qualification planner and an ICP generator can help a team name the customer characteristics it actually uses. Check the criteria with sales and operations staff. If they disagree about a category, settle the rule before asking a model to apply it.

Atomic questions also make review possible. A person can inspect “urgency unclear” without reinterpreting a mysterious overall score. When a rule changes, you can update the weight or route in code and preserve the original answers for comparison. This matters when the process handles messages from forms, email and direct messages with different levels of detail.

Weights and routes belong in code

One public composite-scoring recipe asks four Score questions about budget, authority, urgency and product fit, then applies weights in code. The pattern is useful because staff can see and change what matters. The model supplies answers to questions; your team sets the relative importance and the next action. See the lead scoring guide for the broader design.

Composite lead score: your team sets the weights for each question in code, and the model only answers the questions

A practical route might have three destinations: sales follow-up, a lower-priority queue and human review. The thresholds and any automatic action are business decisions. They should be written down and tested against labelled leads. “Low score” should not automatically mean “bad person” or “never contact.” A missing budget or an unfamiliar job title may reflect an incomplete form, not an unqualified prospect.

Separate two kinds of numbers. The model’s probability expresses how strongly it selected an answer to a specific question. The business score is a weighted result of several answers and your rules. Neither is a calibrated probability that the deal will close. To make that kind of prediction, you would need separate outcome data and validation.

The confidence floor: fail closed to a person

A confidence floor is a rule: if a required answer is too uncertain, the workflow stops automatic routing and asks a person. The right floor depends on the question and the cost of an error. “Asked for a call” can use a different rule from “meets a regulated eligibility condition.” Test each question on your own labelled leads rather than borrowing a number from a public demo.

The workflow should distinguish an uncertain model answer, a missing answer, a failed model call and conflicting evidence. They are different review tasks. A person should see the lead’s original message, the question, the selected answer, the probability and the route the code proposed. They should be able to correct the category and record why. Those corrections help expose a poor question or a shifted customer pattern.

In OpenRouter’s classification test, OpenRouter warned that confidence was “not a calibrated probability.” It reported 96.3% accuracy among the 58% of its 3,080 messages at confidence at least 0.99. Those figures describe a support-message test, not lead qualification. They show why a threshold must be evaluated on the specific task and distribution.

What public lead builds show

A First Read example asked about 32 fit, pain and intent questions per lead, then used code for an ICP score, tier, route and next action. Low-confidence cases went to a person. The builder reported 65 of 65 tiers correct after tuning, up from 92.3% on the first pass, with every initial miss flagged; they estimated about US$0.22 per 1,000 leads. The data were fictional. Treat this as a demonstration of question design and exception routing, not evidence of live sales performance.

In jev-playground, its authors reported 90% correct routing on 60 hand-labelled lead cases for Jev, compared with 78% for Sonnet 5. They reported 366 milliseconds and US$0.04 per 1,000 for Jev versus 2.6 seconds and US$3.01 for Sonnet in that setup. Another model still wrote the emails. The test is small and belongs to its authors; it does not establish which model will work on your leads.

An n8n build used nine questions and fixed rules for sales, review and support. Four of five live routes matched its expected route. Only 28 of 90 planned evaluation calls completed because provider rate limits blocked 62. That operational limit matters: a workflow needs a retry or human fallback, and a test plan needs to account for calls that never return. For the surrounding process, see the lead qualification agent guide.

A Spanish WhatsApp and CRM triage example included an “unclear” option and sent uncertain cases to a person. Its author cautioned that uncertainty intervals were wide and that hundreds of labelled rows would be needed. A Vercel Labs form-router template accepted a route only at 95% confidence and sent uncertain or failed evaluations to an LLM. These are design choices, not thresholds to copy without validation.

What stays human

A person should own the customer relationship, the meaning of the qualification criteria, exceptions and any decision with a material effect on an individual. Jev does not write the reply. A person or separate writing model can draft it, with the appropriate review before sending. The human-led design guide explains why a decision step needs a named owner.

Keep a route for someone who asks for a person, even if their form is sparse. Review leads that involve unusual needs, conflicting signals or a possible error in the intake data. If the system supplies a reason for a route, base it on logged questions and rules, not an invented narrative of what the model “thought.”

French and Quebec leads: test before you promise

One independent test found Jev lost about one point from English to French on short commands: 85% versus 84% in that setup. TypeSafe says English is where accuracy is best. No public French lead test appears in the cited sources, and nothing establishes Quebec French performance, so test on your own French leads first.

A French email sorter reported 85% agreement with the tool it replaced. That is email sorting, not lead qualification. Build a sample with local phrasing, bilingual messages and the incomplete answers your team sees. Look at errors separately by language and category. Do not assume an overall English result transfers to French customer messages.

Canadian data and automated decisions

Lead forms and messages can contain personal information. Send only fields needed for the question, limit access to the responses and document the path through any gateway. The reviewed Jev channels process data in the United States, with no Canadian processing option found. See the Canadian privacy guide before sending real lead data.

Quebec's Law 25 (s.12.1) sets duties for decisions 'based exclusively on automated processing'. The law doesn't list which systems count. Whether auto-routing a lead, a document or a call is covered depends on how your system works, for example whether a person looks at or can change the result before it takes effect. That's a question for your own lawyer. Either way, we design for it: a human review path, a log of each question, answer, probability, model version and threshold, and a way for anyone affected to reach a person.

Do not auto-reject a lead without a human path. This is general information, not legal advice. Check with your own lawyer about your data and your obligations.

Can lead sorting run on a local model?

A local classifier may support an in-house sorting step for coarse categories. In a small Flowgrammer spot check supplied with this review, the tool handled English only and scored 28.7% on urgency. That is a limited observation, not a broad benchmark. Local operation also has setup, access-control and review responsibilities. Compare the in-house sorting step with the broader alternatives guide before choosing a model.

How we’re testing it

We’re running our own lead test; we’ll publish the results when the test is complete. No Flowgrammer lead result is claimed here.

FAQ

Can TypeSafe Jev do lead qualification?

TypeSafe Jev can answer defined questions about fit, stated pain, urgency and requests for a person. Your rules combine those answers into a score and route; Jev does not decide the entire relationship. Set a review path for missing or uncertain information, and check the proposed route against real labelled leads before automating any consequential step.

Is a Jev lead score the probability a lead will close?

No. Jev’s probability describes its answer to one question. A composite lead score reflects the weights your team chose for several answers. Neither number is a measured chance of closing a deal. OpenRouter also cautioned that Jev confidence was not a calibrated probability in its own test. Check both routing accuracy and later sales outcomes on your own data.

Does Jev work on French leads?

One public test reported 85% in English and 84% in France French on short, translated commands. TypeSafe says English is where Jev works best. No public French lead test in the cited sources establishes performance on Canadian or Quebec French messages. Build a labelled sample from your own inbound leads and inspect mistakes separately by language before using automatic routes.

Can Jev write the reply to a lead?

No. Jev selects from defined answers; it does not draft an email, chat reply or call script. A person or a separate writing model must produce the message. In a public lead-routing example, Jev selected the route and another model wrote the emails. Keep a review rule before sending replies that could affect a customer relationship.

Should lead scoring run on a local model?

It can support an in-house sorting step when the categories are coarse and you can operate the tool responsibly. A small Flowgrammer spot check found one local classifier weak on urgency and limited to English. That result is not a verdict on every local model. Compare alternatives on labelled leads, review error costs and account for security and operating work.

Next step

If slow or inconsistent inbound handling is a funded problem, contact Flowgrammer about a Lead Qualification AI Automation System with explicit review and ownership rules.