Google AI released Gemini 3.5 Live Translate, supporting 70+ languages with near real-time latency
speech-to-speech translation. This model directly processes the raw audio stream, preserving the speaker's intonation, rhythm, and pitch. Southeast Asia's super app Grab is exploring cross-language communication between drivers and passengers, with users initiating over 10 million voice calls each month. Developers can build applications by integrating LiveKit, Fishjam, Pipecat, or Vision Agents via the Gemini Live API. LiveKit has enabled multilingual real-time understanding in virtual meeting rooms; Software Mansion breaks through streaming bottlenecks through the MoQ protocol; VisionAgents AI demonstrated dynamic multilingual switching capabilities. Developers can try out and access Cookbook sample code in Google AI Studio.
Google AI 发布 Gemini 3.5 Live Translate,支持 70+ 语言、近实时延迟的语音到语音翻译。
该模型直接处理原始音频流,保留说话者语调、节奏和音高。
东南亚超级应用 Grab 正探索将其用于司机与乘客间的跨语言沟通,其用户每月发起超 1000 万次语音通话。
开发者可通过 Gemini Live API 集成 LiveKit、Fishjam、Pipecat 或 Vision Agents 构建应用。
LiveKit 已实现虚拟会议室多语言即时理解;
Software Mansion 结合 MoQ 协议突破流媒体瓶颈;
VisionAgents AI 展示了动态多语言切换能力。
开发者可在 Google AI Studio 试用并获取 Cookbook 示例代码。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.