← All News & Updates

Preview: Run your actual voice agent across the stack you already use

A product preview of evaluating the agent and voice path a team intends to release, within a validated early-access scope.

Cobalt, teal, and coral voice paths converge on one agent while preserving the surrounding stack boundaries

This preview sets out a product direction: evaluate the actual voice agent a technical team intends to release, with the relevant voice path and tool behavior in scope. It does not announce universal support for every telephony provider, model, protocol, deployment, or private network.

This capability is being developed with early-access and design-partner teams; availability depends on the voice stack and evaluation scope. A team must share its stack and release requirements before Voxeval can determine whether and how the agent can be exercised in a focused engagement.

The integration path is part of the behavior

Voice-agent behavior does not come from the model alone. Audio transport, turn detection, interruption handling, speech recognition, prompt logic, tool calls, policy checks, latency, and telephony can all affect a workflow. Testing only a text transcript or a substitute bot can miss the layer that changes the actual release.

Our direction is to run calls against the agent the team is preparing, then retain evidence from the interaction that the agreed evaluation can inspect. The useful evidence depends on the stack and permissions available. It may include audio, transcript, trace, and tool-state information, but those evidence types should not be assumed for every integration.

Define a supported boundary

Early-access work begins by mapping the path that matters: how a call reaches the agent, which tools are involved, what information can be observed, and which business outcome determines success. That mapping establishes a supported boundary before a run begins.

If a required connection cannot be established safely or an essential result cannot be observed, the evaluation scope has to change or the engagement may not be a fit. Voxeval should not replace missing evidence with a confident score. A clear “not enough data” state is more useful than pretending that a partial view represents the whole workflow.

Preserve the real conditions that matter

Within an agreed scope, the evaluation plan can include caller personas, accents or languages, background audio, interruptions, IVR or DTMF paths, and tool-state conditions when those are supported and relevant. The participating team identifies which conditions matter to its release; the preview does not claim exhaustive coverage of all possible callers or environments.

Running the actual agent also requires care with policies, access, and retained evidence. Teams remain responsible for providing approved test paths and for deciding what production material, if any, may be used. Production incidents are inputs only when available and permitted; synthetic scenarios can be used without presenting them as observed production behavior.

A release-focused preview

The aim is a repeatable evaluation of one real release path, not a generic integration badge. Results should stay attached to the tested agent version, workflow, conditions, and available evidence so an engineering team knows what the conclusion covers.

Teams interested in this preview can join the early-access list with a work email. Joining requires only that email; the form does not collect agent, workflow, stack, or release details. If we follow up, we will ask for the relevant context and constraints before discussing fit, timing, or a possible evaluation boundary. Joining the list does not guarantee that a particular stack is supported.