Google announces Gemini 3.5 Transcribe speech-to-text model
Google has unveiled Gemini 3.5 Transcribe, a new speech‑to‑text model that builds on the Gemini 3.5 architecture. The model is designed to deliver faster, more accurate transcription across a wide range of languages and accents, with a particular emphasis on real‑time performance. According to the company’s blog, Gemini 3.5 Transcribe incorporates advanced acoustic modeling and contextual language understanding, enabling it to handle noisy environments and complex dialogue with higher fidelity than previous iterations.
The announcement comes as Google expands its Gemini family of multimodal models, positioning the new transcriber for integration into Workspace, Meet, and other productivity tools. Early benchmarks released by the company show a significant reduction in error rates compared to earlier Gemini models, especially for low‑resource languages. The update also includes new privacy safeguards, allowing on‑device processing for sensitive audio streams and offering users greater control over data retention.
Community reaction on Hacker News, where the announcement received 26 points and 9 comments, reflects a cautious enthusiasm. Users praised the model’s potential for accessibility and real‑time collaboration, while some expressed interest in further performance metrics and broader language coverage. As Google rolls out Gemini 3.5 Transcribe, developers and enterprises will likely evaluate its impact on workflow automation, transcription services, and multilingual communication.