Audio in
3CX Call Control API · 8 kHz mono
getAudioStreamProjects
Voice Agent
Voice Agent Orchestrator is a standalone on-premises platform for building, evaluating, and operating AI-powered voice agents on telephony. Organizations can run the full conversation through a single speech-to-speech model, or use pipeline mode — a composable STT → LLM → TTS stack where each stage is configured independently. Knowledge retrieval, MCP tooling, and built-in evaluation are managed through a single control plane.
Voice AI integrations are often built around specific providers, runtimes, and deployment assumptions. As organizations introduce new models, local inference, knowledge systems, or hybrid architectures, business logic becomes increasingly coupled to implementation details.
The result is slower experimentation, duplicated integrations, and growing maintenance costs whenever the underlying AI stack changes. The challenge is not supporting a single provider or model, but enabling conversational systems to evolve without forcing the surrounding platform to evolve with them.
The platform is built around two interchangeable voice runtime architectures, deployed independently from the PBX — so the AI stack can scale and ship at its own pace without coupling to the communications core.
Realtime mode runs the full conversation through a single speech-to-speech provider. Pipeline mode composes independent STT, LLM, and TTS stages across cloud, on-premises, or hybrid deployments — with a unified runtime abstraction whether providers stream natively or connect over batch APIs.
Turn-taking, barge-in, and provider orchestration live inside the voice runtime rather than in telephony integration code.
The platform integrates with 3CX through the Call Control API for real-time call handling and uses the built-in 3CX MCP server as its primary tool execution layer. Core telephony capabilities — contact lookup, extension discovery, call transfers, and routing — are exposed to agents through MCP tools, with phonebook search leveraging 3CX's fuzzy linguistic matching to compensate for speech recognition errors.
Control plane
Create multiple stacks of any architecture or deployment type, assign agents per stack — all running simultaneously in one runtime.
Agents
Configure voice, language, prompt, tooling, knowledge access, and more — per agent.
Voice processing architecture
The runtime voice path on every call — from incoming audio to spoken response.







