Models ·
Google Gemini 3.5 Live Translate Turns Any Voice Into 70 Languages Instantly
Models ·
Google Gemini 3.5 Live Translate Turns Any Voice Into 70 Languages Instantly
AI summary · Generated from this article
Google has released Gemini 3.5 Live Translate, a purpose-built streaming audio model that performs near real-time speech-to-speech translation across 70 or more languages. Unlike traditional systems that wait for a speaker to finish before processing, this model continuously generates translated audio as speech arrives, preserving the original speaker's intonation, pacing, and pitch throughout. The model runs as a single audio pipeline with 16kHz input and 24kHz output, available to developers via the Gemini Live API at roughly $0.037 per minute. It supports automatic language detection, noise robustness, and embeds an imperceptible SynthID watermark in all generated audio. Google Meet is expanding from five languages to over 70, unlocking more than 2,000 language combinations per meeting. Translation quality across all 70 languages remains independently unverified, and lower-resource languages may underperform compared to widely spoken ones. That gap will matter most for the real-world deployments Google is targeting in Southeast Asia and sub-Saharan Africa.