Voice agents in the API: custom voice is only the beginning

Xai introduced Custom Voices and Voice Library on April 30, 2026. A custom voice can be created from a short recording and used in TTS and Voice Agent APIs.
On May 7, 2026 OpenAI added the broader layer: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the API. Voice agents are moving from “can speak” to “can run a live task, call tools, translate, and transcribe continuously.”
What changes
Xai lets teams create a custom voice from reference audio, with up to 120 seconds of audio and up to 30 custom voices per team according to its docs. OpenAI adds live voice models for production interactions: reasoning voice conversations, live translation, and low-latency transcription.
What it means in practice
This matters for customer support, onboarding, product tutorials, multilingual workflows, real-time call summaries, and agents that call internal tools instead of only answering.
Where to wait
Voice cloning and live voice automation need governance: explicit consent, audit logs, retention rules for source recordings, incident handling, clear AI disclosure where appropriate, and guardrails for what the agent may promise or trigger.
OpenAI mentions active classifiers in Realtime API sessions and says developers must make AI interaction clear to users unless it is obvious from context. That is a useful baseline, not a complete company policy.
Conclusion
Custom voices from Xai and new realtime models from OpenAI point in the same direction: voice agents are moving from demo to product layer. If you work on support, onboarding, call centers, meeting transcription, or internal assistants, voice workflows will soon be part of the normal AI stack.