Voice agents: low latency and more languages in 2026
What has happened in recent months
OpenAI released GPT‑Live‑1 for developers with improvements in "turn‑taking" and interactive behavior compared to its previous Realtime version, and led tests of end‑to‑end voice agent tasks; the company also explained how it redesigned its WebRTC stack to lower latency in production and scale bidirectional real‑time audio (OpenAI, OpenAI). (openai.com)
Google has been moving Project Astra capabilities into Gemini Live and its Live API for developers, with real‑time audio, streaming video input and combined tool use; it also documents low‑latency voice‑to‑voice translation in more than 70 languages (DeepMind/Google, Google, Docs). (deepmind.google)
Apple announced "Siri AI" at WWDC in June 2026, but delayed its launch in the EU citing the DMA; Brussels replied that the decision was Apple's, and in Spain iOS 27 arrived without Siri AI on iPhone, with signs of limited testing in later betas (Apple Newsroom, AP News, Cinco Días, Cinco Días). (apple.com)
Voice quality and latency: towards natural conversations
In voice generation, ElevenLabs raised the bar with its Eleven v4 Turbo model, which reports synthesis around 100 ms and support for dozens of European languages, including Spanish and Catalan, useful for local brands with their own vocal identity (ElevenLabs Docs). (elevenlabs.io)
In recognition and "turn‑taking", OpenAI's documentation and community report significant latency reductions in recent Realtime models and in GPT‑Live‑1, with improvements in full‑duplex and coordination between reasoning and speech (OpenAI, Dev forum). (openai.com)
For those needing controlled or edge deployments, NVIDIA documents streaming ASR with minimum chunk sizes of 80 ms and intermediate results, which helps start responses before the user finishes speaking (NVIDIA, NVIDIA). (perspectives.nvidia.com)
Languages and multilingualism: less friction in calls and counters
Google details in its Live API low‑latency voice‑to‑voice translation with broad language support, designed for agents that switch tools during a conversation (useful at reception desks or for bookings) (Docs). (ai.google.dev)
In the Microsoft ecosystem, Foundry describes "GPT Realtime Translate" and updates in Azure AI Speech: multilingual transcription in a single session, later refinement and lower latency (up to 3× versus the previous version), relevant for call centers and hybrid meetings (Microsoft Learn, TechCommunity, Foundry Blog). (learn.microsoft.com)
Are SMEs getting on board?
Public figures show progress, with nuances. ONTSI reports that AI adoption in Spanish SMEs rose from 8.6% to 10.4% with 2024 data, and Eurostat places Spain at around 20% of companies using at least one AI technology in 2025; growth is real but uneven by size and sector (ONTSI, Eurostat). (ontsi.es)
In the short term, the delay of Siri AI in the EU limits its direct impact on European businesses, while API‑based options (OpenAI, Google, Microsoft) and controlled deployments (NVIDIA) are available today for real reception, booking and after‑sales use cases (Apple Newsroom, OpenAI, Google Docs, NVIDIA Riva). (apple.com)
What it means for your business
- Less waiting when speaking with the machine: reduced latency and full‑duplex make natural dialogues possible that don’t interrupt the customer; this improves the experience in contact centers and WhatsApp/voice for bookings and appointments (OpenAI). (openai.com)
- More languages without changing systems: Google and Microsoft APIs allow switching between languages and live translation in the same session, useful for clinics and restaurants with international clients (Docs Google, TechCommunity). (ai.google.dev)
- Brand voice in Spanish with competitive latency: recent TTS solutions enable consistent timbres and high‑quality voice clones, keeping response times low for phone support and kiosks (ElevenLabs Docs). (elevenlabs.io)
- Deployments with data control: if you need predictable latencies or EU data residency, you can choose cloud infrastructure in European regions or voice components at the edge with Riva/Nemotron (OpenAI Docs, NVIDIA Riva). (developers.openai.com)
Frequently asked questions
In short
At Veltim we already integrate these capabilities into our voice agents for clinics, restaurants and local businesses: 24/7 reception, automated reminders and management of appointments or reservations, always with response time and satisfaction metrics, and without inflated claims: the impact depends on your workflow, your customer base and how you measure each step with data.
Want to know how to apply it in your business? Tell us about your case.
Kevin Rojas — Fundador de Veltim. Programador e ingeniero de IA: agentes de voz, chatbots de WhatsApp e integraciones con CRM para negocios reales. LinkedIn · YouTube · Sobre Veltim
Contact: veltim.com · Phone: +34 623 05 08 56 · Madrid, Spain.