Flowgrammer

Build a Voice Agent That Talks While Your System Decides: GPT-Live-1 + TypeSafe Jev for Inbound Calls

Since Sept 10, 2026, GPT-Live-1 is in the API. Build voice agents that keep talking while tasks run, with TypeSafe Jev making logged routing decisions and people on the edge cases.

— Craig Major

Short answer: TypeSafe Jev can make a defined routing decision behind a GPT-Live-1 conversation. GPT-Live listens and speaks; your app sends text and fixed questions to Jev; code applies the route and a person takes uncertain or urgent cases. A caller can keep talking while a backend task runs. Neither model processes these calls in Canada.

A voice agent keeps talking to the caller while your system works on the task in the background

An AI receptionist for a small business still needs call ownership, consent and a reliable human transfer. This is a possible architecture for an inbound line, not a Flowgrammer deployment. The broader voice AI guide covers the operating questions around callers, handoffs and records. Here, the focus is the division of work between live conversation, decision rules and a human team.

What changed on September 10, 2026

OpenAI made GPT-Live-1 available in its API on September 10, 2026. It listens and speaks at the same time, so a caller can interrupt or add details without waiting for a rigid turn to end. A developer can delegate reasoning and tools to a backend model or its own service while the voice conversation continues. That makes it possible to check an order or start a booking task without making the caller wait in silence for every step.

The live layer does not become your business system. Your app still controls access to records, the state of a task, permission to take an action and what the agent may say. The live guides describe client delegation and instruction updates during a session. Background work can continue after a caller interrupts the voice. The research reports instruction appends with a 500-token limit each; confirm the current API limit before building against it.

OpenAI’s SIP guide describes phone connections, transfer and keypad tones. The researched call cap was two hours. Partners named in the research include Twilio Agent Connect, Telnyx, LiveKit and Pipecat. Those are possible connection routes, not proof that every carrier feature or handoff pattern is ready for your line. Test the actual phone path and failure behaviour before moving callers to it.

What you can build now

A receptionist that books while chatting. The caller asks for an appointment. GPT-Live gathers the needed details while your backend checks availability. Code, not the model’s conversational confidence, commits a booking after the required confirmation. A person takes ambiguous requests, complaints or requests for a human. The AI receptionist overview and receptionist workflow cover the surrounding service design.

A support line that checks an order in the background. The agent can keep listening as a backend retrieves an order record. The app should verify the caller and redact anything the voice model need not see. The record system owns the status; the spoken answer reflects the checked result. If the lookup fails, the agent can explain that and transfer or arrange follow-up.

After-hours triage. The agent can ask what happened, check a narrow urgency question and transfer an emergency to an on-call person. Write the emergency definition and fallback before launch. A failed transfer or uncertainty must not silently become a routine message. These are build possibilities, not claims that Flowgrammer has deployed them.

The pattern: Jev decides, GPT-Live talks

The useful division is: caller → phone connection → GPT-Live conversation → your app → Jev decision question → code rule → booking, transfer or message. Your app logs the decision and sends the permitted response or action back to the voice layer. The conversational voice platform guide gives the wider context.

Call flow: the live voice model talks, your app asks a decision model fixed questions, code applies your rules and books, transfers or takes a message, with every decision logged

Client delegation and sideband monitoring

With client delegation, GPT-Live hands a task to your backend while it keeps the conversation moving. The backend can validate or redact tool results before they are spoken. A sideband connection can monitor the session and coordinate state from your server. The important design choice is ownership: the app, not the voice model, must know whether a booking has been committed, an emergency transfer attempted or a caller requested a person.

A model call should receive the minimum text needed. If the caller changes their mind, the next decision should be tied to the latest authorised state. Do not let an old classification act after a newer answer makes it stale. If the decision service is unavailable, the fallback can be a safe scripted response and a human queue.

Ask small typed questions

Jev can answer fixed questions such as:

  • Intent: new lead, existing client, vendor, spam or other.
  • Urgency: emergency, same day or routine.
  • Wants a human: yes or no.

The exact labels are business rules to validate. “Other” and “unclear” need review paths. Jev reads text rather than audio, so the app supplies conversation text or structured state from the voice layer. It does not talk to the caller or own the phone connection.

Log the decision and its override

For each consequential route, record the question, answer, probability, model version, threshold, code rule, action and any human override. Keep access and retention appropriate for call data. A reviewer should be able to see why a transfer occurred without treating a fluent spoken response as evidence that the model’s decision was right.

When GPT-Live alone is enough

A simple line with low-cost mistakes may not need another decision model. The app can use GPT-Live and deterministic rules, then test the result. Adding Jev creates another integration, another data processor and another failure path. It earns that place when the team needs four things: the same logged labels and probabilities for each decision; one routing policy across phone, form and email; predictable decision-step latency and cost; or an independent check on a voice model that does not support structured outputs.

That independent check matters because a voice system can sound confident while misunderstanding a short yes/no answer or reading a value back incorrectly. A separate model is still fallible. Use it only where your own labelled call cases show it improves a decision that matters. The test should include interruptions, corrections, silence, background speech, French where relevant and requests for a person.

Latency: measure the whole call

OpenRouter reported median Jev latency of 175 milliseconds in its test. Another public receipt reported 231 milliseconds at the 95th percentile, including network time. Those are third-party numbers from their conditions, not a promise for a phone line.

ThunderPhone’s impressions covered about 12 real and 25 simulated calls and reported about 1.3 seconds median to first audio on the phone, versus about 0.7 seconds without the phone leg. It also described weak instruction following and wrong read-backs. The whole path includes telephony, live voice, backend work and any decision call. Our own test will measure that complete experience; no Flowgrammer timing result is available yet.

One public phone-line example

A Jev IVR example used Twilio ConversationRelay and a fictional clinic. Its author reported 229 of 241 labelled utterances correct, 88 of 89 scenarios, about 170 milliseconds per decision batch and 400–550 milliseconds on a live call before a keep-alive fix. The labels were the author’s own. This is a useful pattern for explicit intent and urgency descriptions, not proof of live healthcare performance or a general phone benchmark.

The author’s point that the level descriptions are the program is practical: if “urgent” is vague, neither the voice nor decision model can apply a reliable policy. Write the categories with the people who take calls, then test cases that cross their boundaries.

Where it goes wrong

Voice errors can be subtle. The agent may hear the wrong account number, follow an instruction poorly or read a date or amount back incorrectly. TypeSafe also lists numbers and dates among Jev’s weak areas. The system of record must own the value. For important identifiers, ask for keypad confirmation or another verified channel such as SMS rather than trusting a spoken repetition.

Give recording and AI disclosures through a controlled, pre-recorded or scripted step so an unexpected model response cannot skip them. Transfer to a person when the caller asks, when an emergency rule fires, or when the model and code disagree. The uncanny-valley voice guide discusses why a natural voice needs clear boundaries and honesty about what it is.

Costs: voice, backend and decisions

OpenAI’s September 2026 page lists GPT-Live-1 voice at US$0.05 per session minute, billed per second in the research. Silence and time spent waiting for a backend still count as session time. Backend models and tools have separate charges; the research also notes a 10% data-residency uplift where applicable. Re-check the current pricing and data terms before budgeting.

Two meters for a live voice call, voice minutes and backend usage, plus a small cost per decision

Jev adds a decision charge: input tokens × US$0.042 ÷ 1,000,000 at the research’s checked price, with output free. That small model price does not include the phone carrier, voice time, backend work, integration, monitoring or human transfers. There are at least two main meters, voice time and backend usage, plus the Jev decisions. A total call or monthly estimate would require your real call pattern and provider terms.

Canadian call data and human decisions

The reviewed OpenAI data controls list Canadian storage but not Canadian processing for GPT-Live sessions; the supported live processing regions in the research are the US and EU. Jev processing is in the US on the reviewed channels. Neither model offers confirmed Canadian processing for this pattern. They are separate processors, with separate data paths to assess. See the Jev Canadian privacy guide.

The OPC’s call-recording guidance says to inform callers, state the purpose, seek consent and offer an alternative. Put that information at the start of the call in a way the caller can understand. Do not assume an AI disclosure alone covers recording or every later data use.

Quebec's Law 25 (s.12.1) sets duties for decisions 'based exclusively on automated processing'. The law doesn't list which systems count. Whether auto-routing a lead, a document or a call is covered depends on how your system works, for example whether a person looks at or can change the result before it takes effect. That's a question for your own lawyer. Either way, we design for it: a human review path, a log of each question, answer, probability, model version and threshold, and a way for anyone affected to reach a person.

This is general information, not legal advice. Check with your own lawyer about your data and your obligations.

French and Quebec callers

The cited GPT-Live model page did not publish a language list or a proven Quebec French voice at the research check. TypeSafe says English is where Jev is strongest, and the cited French test was on short commands, not calls. No cited test proves the complete voice-plus-decision pattern for Quebec callers. Keep a French-speaking person available and test real, consented French calls before promising service quality.

If you use older realtime models

OpenAI’s deprecation schedule lists January 20, 2027 shutdown dates for legacy gpt-realtime, gpt-4o-realtime and gpt-realtime-mini families. If an existing line uses them, identify the exact model and plan a migration test. GPT-Live may be an option, but it should not be treated as an automatic drop-in replacement for your current call flow.

Other ways to run the decision step

Jev is one possible classifier behind a voice system. The alternatives guide compares other decision methods and the circumstances in which a separate model may add no value.

FAQ

What can businesses build with GPT-Live-1?

Businesses can build live voice agents that listen and speak while a backend works on tasks such as bookings or order checks. OpenAI released GPT-Live-1 in its API in September 2026. The app still owns access, records, actions and handoffs. Phone service can connect through SIP or a partner route, with testing on the actual carrier and workflow before callers rely on it.

Can I use TypeSafe Jev with GPT-Live?

Yes, as a separate decision step in your app. GPT-Live handles the conversation; your app sends relevant text and a fixed intent or urgency question to Jev. Code applies the answer and tells the voice layer what action is allowed. Log the result and send uncertain, urgent or human-requested calls to a person. Jev does not hear audio directly.

Should GPT-Live make routing decisions itself?

It may be enough for a simple, low-risk line after testing. A separate model can be useful when you need consistent labels and probabilities, one policy across calls and other channels, predictable decision-step cost, or an independent check. GPT-Live-1 does not support structured outputs. Adding another model also adds integration and privacy work, so compare the full workflow on labelled calls.

How much does a GPT-Live call cost?

OpenAI listed the voice layer at US$0.05 per session minute at the September 2026 check, with backend models and tools billed separately. Jev decisions add input-token charges, and a phone provider may charge too. Silence and background-task time can count as session time. Re-check all prices before budgeting; a reliable total needs your call duration, tools and handoff pattern.

Can GPT-Live keep Canadian call data in Canada?

No Canadian processing option for GPT-Live sessions was confirmed in the cited data controls. OpenAI’s Canada region covers storage rather than live-session processing, and Jev is processed in the US on the reviewed routes. Map the phone provider and both model paths before using personal information. This is general information, not legal advice. Check with your own lawyer.

Does GPT-Live speak Quebec French?

The cited model page did not establish a Quebec French voice or a tested language result for this call pattern. Jev’s public French evidence is also from short commands, not Quebec calls. Do not promise a French caller experience on that basis. Test the actual voice and decision flow with local language examples and keep a French-speaking person as the fallback.

Can Jev replace Retell?

No. Jev is a text-only decision model; it cannot answer a phone or speak. It might supply a routing answer behind a voice platform. The Retell guide covers the platform role. At the September 23 research check, a Retell community answer did not confirm GPT-Live support, so verify current integrations before planning around one.

Next step

If inbound calls create a funded service problem, contact Flowgrammer about an Inbound Voice AI Automation System with a clear transfer path to your team.