{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-08-29T22:40:45.034Z","headline":"Gemini 3.5 Transcribe 完整指南：告别 ASR 转录难题","description":"Google 推出专用于语音转文字的 Gemini 3.5 Transcribe 模型，主打快速、准确且低成本的转录，原生支持说话人分离和词级毫秒时间戳。该模型支持 85+ 种语言自动识别与代码切换，可通过 custom_vocabulary 传入最多 1，000 个领域术语避免专有名词拼写错误，并提供 Smart Transcription 与 Verbatim 两种模式。","url":"https://www.aioga.com/news/cmtd00dbh09tkroq5v7kcw183/","mainEntityOfPage":"https://www.aioga.com/news/cmtd00dbh09tkroq5v7kcw183/","datePublished":"2026-08-28T13:34:29.000Z","dateModified":"2026-08-28T13:34:29.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://dev.to/googleai/stop-wrestling-with-asr-the-complete-guide-to-gemini-35-transcribe-1m6i","https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183"],"canonicalUrl":"https://www.aioga.com/news/cmtd00dbh09tkroq5v7kcw183/","directAnswer":{"@type":"Answer","text":"Google推出专用于语音转文字的Gemini 3.5 Transcribe，定位是快速、准确且具成本效益的音频转录。材料称其原生支持说话人分离、词级毫秒时间戳、85种以上语言自动识别与代码切换，并提供Smart Transcription和Verbatim模式。","url":"https://www.aioga.com/news/cmtd00dbh09tkroq5v7kcw183/","dateCreated":"2026-08-28T13:34:29.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"dev.to source article","url":"https://dev.to/googleai/stop-wrestling-with-asr-the-complete-guide-to-gemini-35-transcribe-1m6i","datePublished":"2026-08-28T13:34:29.000Z","provider":{"@type":"Organization","name":"dev.to","url":"https://dev.to/googleai/stop-wrestling-with-asr-the-complete-guide-to-gemini-35-transcribe-1m6i"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","datePublished":"2026-08-28T13:34:29.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183"}}],"aggregationSource":"Google AI：DEV 作者专属（RSS）","originalPublisher":{"name":"dev.to","url":"https://dev.to/googleai/stop-wrestling-with-asr-the-complete-guide-to-gemini-35-transcribe-1m6i"},"geoDeepAnswer":null,"article":{"id":"cmtd00dbh09tkroq5v7kcw183","slug":"cmtd00dbh09tkroq5v7kcw183","url":"https://www.aioga.com/news/cmtd00dbh09tkroq5v7kcw183/","title":"Gemini 3.5 Transcribe 完整指南：告别 ASR 转录难题","title_en":"Stop Wrestling with ASR： The Complete Guide to Gemini 3.5 Transcribe 🎙️","summary":"Google 推出专用于语音转文字的 Gemini 3.5 Transcribe 模型，主打快速、准确且低成本的转录，原生支持说话人分离和词级毫秒时间戳。该模型支持 85+ 种语言自动识别与代码切换，可通过 custom_vocabulary 传入最多 1，000 个领域术语避免专有名词拼写错误，并提供 Smart Transcription 与 Verbatim 两种模式。","source":"Google AI：DEV 作者专属（RSS）","sourceUrl":"https://dev.to/googleai/stop-wrestling-with-asr-the-complete-guide-to-gemini-35-transcribe-1m6i","aiHotUrl":"https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","publishedAt":"2026-08-28T13:34:29.000Z","category":"技巧观点","score":74,"selected":true,"articleBody":["You’ve probably used Gemini to analyze hours of video, summarize podcasts, or answer questions from recorded meetings (if you didn't you should, it's extremely useful!). But when all you need is a clean, hyper-accurate, and structured transcript from audio, spinning up a huge reasoning model with complicated prompts often feels like using a sledgehammer to crack a nut.","Enter Gemini 3.5 Transcribe ( gemini-3.5-transcribe ).","It's Google's dedicated speech-to-text model built on Gemini's audio understanding core, optimized specifically for fast, accurate, and cost-effective transcription. Whether you want an exact court-reporter transcript with millisecond timestamps, or a reading-optimized summary that removes all your awkward \"ums\" and \"uhs\" , this model handles it natively with zero prompt gymnastics.","🚀 Hands-on first: If you want to jump straight into running the code yourself, open the interactive Gemini Transcribe Colab notebook ：https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Get_started_transcribe.ipynb! It's ready to run so you can dirrectly experience how the model work. Prefer a visual UI with zero coding? You can also test speech recognition directly in Google AI Studio ：https://aistudio.google.com/prompts/new_chat?model=gemini-3.5-transcribe.","Here's what you'll find in this guide:","Before looking at the code, let's get the mental model straight. You might wonder: \"Can't I just upload an MP3 to Gemini 3.7 and say 'Transcribe this'?\"","You can, but here is why gemini-3.5-transcribe is different:","Pro tip: If you need to ask questions about what happened in an audio file (\"What was the action item for Alice?\"), use a multimodal model like Gemini 3.7. If you need the transcript itself , subtitles, or cleaned dictation notes, use Gemini Transcribe!","The Gemini 3.5 Transcribe model runs on the modern Google GenAI SDK ( google-genai v2.0+) using the Interactions API：https://ai.google.dev/gemini-api/docs/interactions-overview.","Make sure you have an API key from Google AI Studio：https://aistudio.google.com/app/apikey, set it as GEMINI_API_KEY , and let's look at how audio gets passed to the model:","Watch the demo video below to see the baseline transcription in action—handling natural speech and bilingual code-switching with ease:","When dealing with audio and video, you never want to inline raw audio bytes as base64 in your API requests—it blows up the payload size by 33%, easily hits network timeouts, and requires re-uploading the same bytes if you want to rerun a query.","The Files API：https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/File_API.ipynb solves this cleanly:","As you saw in the video above, Gemini Transcribe automatically identifies spoken languages out of the box and seamlessly handles code-switching (when someone mixes multiple languages in the same sentence—like switching between French and English mid-sentence, which happens to me all the time!).","However, if you know your audio is exclusively in a specific language or regional dialect, you can pass explicit BCP-47 language codes in transcription_config to bias recognition:","Note: Leaving language_codes=[] (or omitting it) enables full automatic detection across 85+ supported languages and locales：https://ai.google.dev/gemini-api/docs/transcribe#supported-languages. Check out the Audio Transcription Documentation：https://ai.google.dev/gemini-api/docs/transcribe for the complete list of language codes.","Every developer has suffered from an ASR model mangling proper names, confusing specialized libraries with everyday dictionary words (turning \"ScaNN\" into \"scan\" , or \"Qdrant\" into \"quadrant\" ), or inventing phonetically similar terms ( \"Sitsi\" instead of \"CitC\" , \"Thiago\" instead of \"Tiago\" ).","With custom_vocabulary , you can pass a list of up to 1,000 domain-specific terms that the model will bias towards:","Watch the side-by-side comparison video below to see how the model behaves with and without custom vocabulary biasing:","Notice how default speech recognition falls back to phonetic dictionary guesses ( Vernat , Scan , Quadrant , Syllium , Thiago , Sitsi , Spacey ). By contrast, supplying custom_vocabulary guarantees that names of team members, niche tools, internal infrastructure, and open-source libraries are transcribed with 100% precision.","Pro tip: Don't just put acronyms in your custom vocabulary. Add proper names of team members, internal service codenames, GitHub repo handles, product brand names, and niche industry terminology.","This is hands down my favorite capability of Gemini 3.5 Transcribe.","By default, speech-to-text models operate in verbatim mode: they write down everything , including every nervous stutter, throat clear, false start, and verbal tick.","When you're transcribing a speech rehearsal, interview, or voice memo, reading raw verbatim text is painful:","If you switch mode={\"type\": \"smart\"} , the model performs intelligent reading optimization:","Look at the cleaned result on that exact same rehearsal audio:","Watch the side-by-side comparison video below to see how the raw disfluencies are stripped while listening:","(If the video doesn't load, you can listen to rehearsing.wav directly：https://storage.googleapis.com/generativeai-downloads/audio/rehearsing.wav.)","Important caveat: Because Smart transcription uses language modeling to clean up disfluencies and structure the output, it might slightly rewrite, omit, or rephrase parts of what was said to make it sound natural and concise. If you are doing verbatim court reporting, medical transcription, or subtitle syncing where every exact syllable matters, stick with verbatim mode!","Also note that Smart mode is incompatible with word-level timestamps and speaker diarization (which require {\"type\": \"verbatim\"} ).","Need to know who spoke during a multi-person meeting or podcast? Enable diarization with diarization_mode=\"speaker\" :","To extract each speaker turn cleanly, iterate through the step annotations:","Watch the demo video below where two colleagues debate pain au chocolat vs. chocolatine . Notice how the waveform line dynamically changes color (Cyan for Tiago, Orange for his colleague) as each speaker takes turns:","(Direct audio link: listen to pain_au_chocolat.wav：https://storage.googleapis.com/generativeai-downloads/audio/pain_au_chocolat.wav)","When you need exact synchronization—for example, to jump to specific points in a video, build interactive transcripts, or align text with waveforms—you can request word-level millisecond start and end offsets.","Configure timestamp_granularities=[\"word\"] (and optionally combine it with diarization_mode=\"speaker\" ):","Each recognized word comes back with its exact time offsets (and speaker turn) attached in the content annotations:","Having millisecond-level offsets for every individual word unlocks huge capabilities:","💡 Behind the scenes: That's actually what I did to make the demo videos above! The word timestamps provided the exact millisecond timing to align the subtitle cards, highlight the custom terms ( \"oatmilk\" ), and trigger the color switch of the waveform line from Cyan to Orange when the speaker changed.","If you want the complete Python function to convert these word annotations into standard .srt subtitle files, you can find it directly in the interactive Cookbook Colab notebook：https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Get_started_transcribe.ipynb.","Here is a quick cheat sheet to pick the right settings for your use case:","Everything we covered above is for pre-recorded audio files (unary mode via the Files API).","Gemini also supports real-time live streaming transcription over WebSockets using gemini-3.5-transcribe-live and the Live API. It lets you stream raw 16-bit PCM chunks (100ms each) directly from a microphone and receive instantaneous interim partial hypotheses ( interim_input_transcription ) and finalized text as speech occurs.","However, streaming real-time WebSockets with asynchronous Python workers ( asyncio ), handling audio chunking, and managing ephemeral valet tokens for secure client apps is quite a bit more complex and deserves its own dedicated tutorial.","If you want to dive straight into live streaming code right now:","Gemini 3.5 Transcribe gives you the best of both worlds: strict, millisecond-accurate verbatim data when you need timestamps and diarization, and an intelligent, disfluency-stripping smart mode when you want clean text for human eyes.","Have you tried using smart mode on your own voice recordings or meetings? Drop your thoughts and edge cases in the comments below! 🚀🚀🚀","Templates let you quickly answer FAQs or store snippets for re-use.","Are you sure you want to hide this comment? It will become hidden in your post, but will still be visible via the comment's permalink：#.","For further actions, you may consider blocking this person and/or reporting abuse：/report-abuse","Google AI Studio is the fastest way to start building with Gemini. Ready to build?","DEV Community：/ — A space to discuss and keep up software development and manage your software career","Built on Forem：https://www.forem.com — the open source：https://dev.to/t/opensource software that powers DEV：https://dev.to and other inclusive communities.","Made with love and Ruby on Rails：https://dev.to/t/rails. DEV Community &copy; 2016 - 2026.","We're a place where coders share, stay up-to-date and grow their careers."],"articleImages":[{"sourceUrl":"https://media2.dev.to/dynamic/image/width=256,height=,fit=scale-down,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8j7kvp660rqzt99zui8e.png","alt":"pic","afterParagraph":46,"url":"/media/articles/cmtd00dbh09tkroq5v7kcw183/f75d1e7bc8b434f4.webp"},{"sourceUrl":"https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg","alt":"","afterParagraph":54,"url":"/media/articles/cmtd00dbh09tkroq5v7kcw183/ae2a1867ed3e47a7.jpg"},{"sourceUrl":"https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg","alt":"","afterParagraph":54,"url":"/media/articles/cmtd00dbh09tkroq5v7kcw183/c9631abcb96b6799.jpg"}],"mediaStatus":"ok","articleBodyZh":["你可能已经用过 Gemini 来分析数小时的视频、总结播客，或者回答录制会议中的问题（如果没有，你应该尝试，它非常有用！）。但是，当你只需要从音频中得到干净、超准确且结构化的文字记录时，启动一个庞大的推理模型并使用复杂的提示，往往感觉像是用大锤敲核桃。","输入Gemini 3.5 Transcribe（gemini-3.5-transcribe）。","这是 Google 专门的语音转文本模型，建立在 Gemini 的音频理解核心之上，专门优化用于快速、准确且具有成本效益的转录。无论你是想要带毫秒时间戳的精确法庭记者式文字记录，还是优化阅读的摘要，去掉所有尴尬的“嗯”“啊”，这个模型都能本地处理，无需复杂提示。","🚀 实操优先：如果你想直接运行代码，可以打开交互式 Gemini Transcribe Colab 笔记本：https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Get_started_transcribe.ipynb! 它已经准备好运行，你可以直接体验模型的工作方式。更喜欢零编码的可视化界面吗？你也可以直接在 Google AI Studio 测试语音识别：https://aistudio.google.com/prompts/new_chat?model=gemini-3.5-transcribe。","本指南中你将找到以下内容：","在查看代码之前，让我们先理清思路。你可能会想：“我不能直接上传 MP3 到 Gemini 3.7 然后说‘转录这个’吗？”","可以，但 gemini-3.5-transcribe 与众不同的原因如下：","专业提示：如果你需要对音频文件中发生的事情提问（“Alice 的行动事项是什么？”），请使用像 Gemini 3.7 这样的多模态模型。如果你只需要文字记录、字幕或整理好的听写笔记，请使用 Gemini Transcribe！","Gemini 3.5 转录模型在现代 Google GenAI SDK（google-genai v2.0+）上运行，使用 Interactions API：https://ai.google.dev/gemini-api/docs/interactions-overview。","确保你拥有 Google AI Studio 的 API 密钥：https://aistudio.google.com/app/apikey，将其设置为 GEMINI_API_KEY，然后我们来看音频是如何传递给模型的：","观看下面的演示视频，了解基线转录的实际操作——轻松处理自然语音和双语代码切换：","在处理音频和视频时，你绝不希望在 API 请求中以内联原始音频字节的形式使用 base64——这会使负载大小增加 33%，容易触发网络超时，并且如果要重新运行查询，还需要重新上传相同的字节。","Files API：https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/File_API.ipynb 可以干净利落地解决这一问题：","如你在上面的视频中所见，Gemini Transcribe 开箱即用即可自动识别语言，并无缝处理代码切换（当有人在同一句话中混用多种语言时——比如在一句话中间从法语切换到英语，这种情况我经常遇到！）。","然而，如果你知道你的音频仅使用特定语言或地区方言，你可以在 transcription_config 中传递明确的 BCP-47 语言代码来引导识别：","注意：将 language_codes=[]（或忽略它）可以启用对 85 种以上支持的语言和地区的全自动检测：https://ai.google.dev/gemini-api/docs/transcribe#supported-languages。查看音频转录文档：https://ai.google.dev/gemini-api/docs/transcribe 获取完整的语言代码列表。","每个开发者都经历过 ASR 模型篡改专有名词，将专业库与日常词典词混淆（将 \"ScaNN\" 识别为 \"scan\"，或将 \"Qdrant\" 识别为 \"quadrant\"），或杜撰发音相似的词（将 \"Sitsi\" 识别为 \"CitC\"，将 \"Thiago\" 识别为 \"Tiago\"）。","借助 custom_vocabulary，你可以传入最多 1,000 个领域特定术语的列表，模型会偏向这些术语的识别：","观看下面的并排对比视频，看看模型在有无 custom_vocabulary 偏置下的表现：","注意，默认语音识别会回退到按发音猜测的字典（Vernat、Scan、Quadrant、Syllium、Thiago、Sitsi、Spacey）。相比之下，提供 custom_vocabulary 则可保证团队成员姓名、利基工具、内部基础设施和开源库的转录达到 100% 精确。","专业提示：不要只在自定义词汇中加入缩略词。加入团队成员的正式名称、内部服务代号、GitHub仓库账号、产品品牌名称以及细分行业术语。","这是我绝对最喜欢的 Gemini 3.5 转录功能。","默认情况下，语音转文字模型以逐字模式工作：它们记录所有内容，包括每一次紧张的结巴、清嗓子、假动作和口头抽动。","当你在转录演讲排练、采访或语音备忘录时，阅读原文文本令人痛苦：","如果你切换模式={“type”： “smart”}，模型会进行智能读取优化：","看看那个排练音频的清理结果：","请观看下面的并排对比视频，了解在聆听时原始的不流率是如何被剥离的：","（如果视频无法加载，你可以直接收听 rehearsing.wav：https://storage.googleapis.com/generativeai-downloads/audio/rehearsing.wav。）","重要注意事项：由于智能转录通过语言建模来清理不流畅并构建输出，可能会稍微改写、省略或改写部分内容，使其听起来自然简洁。如果你正在进行逐字法庭记录、医疗转录或字幕同步，且每个音节都至关重要，建议坚持逐字模式！","还要注意，智能模式与字级时间戳和扬声器日记（需要{“type”： “verbatim”}）不兼容。","需要知道在多人会议或播客中谁发言了吗？启用日记功能，diarization_mode=“发言人”：","要干净地提取每个扬声器回合，可以遍历步骤注释：","请观看下面的演示视频，两位同事在讨论巧克力面包和巧克力巧克力的争论。注意波形线随着轮流的声音动态变化（Tiago是青色，同事是橙色）：","(直接音频链接：收听 pain_au_chocolat.wav：https://storage.googleapis.com/generativeai-downloads/audio/pain_au_chocolat.wav)","当你需要精确同步——例如，跳转到视频的特定点、构建互动式转录本或将文本与波形对齐——你可以请求单词级的毫秒级开始和结束偏移量。","配置时间戳_粒度=[\"word\"]（并可选择将其与 diarization_mode=\"speaker\" 组合）：","每个识别出的单词都会在内容标注中附带其精确的时间偏移（以及说话者轮次）：","拥有每个单词的毫秒级偏移量可以开启巨大的功能：","💡 背景说明：这实际上就是我用来制作上面演示视频的方法！单词时间戳提供了精确的毫秒时序，用于对齐字幕卡片、高亮自定义词汇（\"oatmilk\"），并在说话者变化时触发波形线从青色切换为橙色。","如果你想要完整的 Python 函数，将这些单词标注转换为标准 .srt 字幕文件，可以直接在交互式 Cookbook Colab 笔记本中找到：https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Get_started_transcribe.ipynb。","这里有一个快速备忘单，帮助你为你的使用场景选择正确的设置：","上面我们讲的所有内容都是针对预先录制的音频文件（通过 Files API 的单工模式）。","Gemini 还支持使用 gemini-3.5-transcribe-live 和 Live API 通过 WebSockets 实时进行流式转录。它允许你直接从麦克风流式传输原始 16 位 PCM 块（每块 100ms），并即时接收临时部分假设（interim_input_transcription）以及最终文本。","然而，通过异步 Python 工作线程（asyncio）的实时 WebSockets 流式传输、处理音频切分以及管理用于安全客户端应用的短期 valet token 相当复杂，需要单独的教程。","如果你现在想直接深入实时流式代码：","Gemini 3.5 Transcribe 为你提供了两全其美的方案：当你需要时间戳和说话者区分时，它提供严格、毫秒级精确的逐字数据；当你希望得到易于阅读的清洁文本时，它提供智能的、去除语气词的模式。","你试过在自己的语音录音或会议中使用智能模式吗？在下面的评论中分享你的想法和特殊情况吧！🚀🚀🚀","模板可以让你快速回答常见问题或存储片段以便重复使用。","你确定要隐藏这条评论吗？它将在你的帖子中变为隐藏状态，但仍可通过评论的永久链接查看：#。","对于进一步操作，你可以考虑屏蔽此人和/或举报滥用行为：/report-abuse","Google AI Studio 是开始使用 Gemini 构建的最快方式。准备好开始构建了吗？","DEV Community：/ — 一个用于讨论、跟进软件开发并管理软件职业空间","基于 Forem：https://www.forem.com — 开源软件：https://dev.to/t/opensource 支撑着 DEV：https://dev.to 和其他包容性社区。","用爱和 Ruby on Rails 制作：https://dev.to/t/rails。DEV 社区 &copy; 2016 - 2026。","我们是一个程序员分享、保持最新信息和发展职业的平台。"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Google推出专用于语音转文字的Gemini 3.5 Transcribe，定位是快速、准确且具成本效益的音频转录。材料称其原生支持说话人分离、词级毫秒时间戳、85种以上语言自动识别与代码切换，并提供Smart Transcription和Verbatim模式。","background":"原文将该模型与通用多模态模型区分：若任务是询问音频内容，可使用多模态模型；若需求是生成转录稿、字幕或清理后的口述笔记，则Gemini 3.5 Transcribe更为直接。文章还介绍了Colab和Google AI Studio两种体验方式。","viewpoint":"Aioga判断，这款产品的核心卖点不是泛化的音频理解，而是把转录流程做成专用能力。custom_vocabulary最多可传入1,000个领域术语，可能有助于减少专有名词拼写错误，但实际效果仍需结合具体音频验证。","implications":"对于会议记录、字幕制作、口述整理及需要时间戳的逐字稿场景，材料显示该模型提供了更明确的功能组合。值得关注的是，Smart Transcription与Verbatim分别面向阅读优化和保留原始表达，使用者需要按交付目标选择模式。","nextStep":"建议先依据任务类型选择模型：需要音频内容问答时考虑多模态模型，需要转录文本时测试Gemini 3.5 Transcribe。随后用包含行业术语、多人对话和中英文切换的样本，在两种转录模式下比较结果，再决定是否配置custom_vocabulary。","evidenceRefs":["title","summary","articleBody","source"],"status":"published","aiGenerated":true,"autoApproved":true,"generatedBy":"aioga-editorial:gpt-5.6-sol","reviewedBy":"aioga-editorial-review:gpt-5.6-sol","generatedAt":"2026-08-29T15:59:18.789Z","sourceHash":"84d164a0cd968595","review":{"approved":true,"groundedness":94,"clarity":92,"duplicationRisk":38,"blockingIssues":[],"notes":["候选内容整体与来源材料一致，事实陈述均可由标题、摘要或正文摘录支持。","“Aioga判断”已明确标注为观点，未将其冒充为来源事实。","“实际效果仍需结合具体音频验证”属于合理的审慎限定，不构成无来源断言。","“85种以上语言”可更贴近来源表述为“85+种语言”，但不影响事实准确性。"]},"validation":{"passed":true,"mode":"ai-auto","revisions":0,"checks":["schema","length","source-attribution","low-source-overlap","no-html","independent-ai-review"]}},"tags":["技巧观点","Google AI：DEV 作者专属（RSS）"],"translations":{"zh-CN":{"title":"Gemini 3.5 Transcribe 完整指南：告别 ASR 转录难题","summary":"Google 推出专用于语音转文字的 Gemini 3.5 Transcribe 模型，主打快速、准确且低成本的转录，原生支持说话人分离和词级毫秒时间戳。该模型支持 85+ 种语言自动识别与代码切换，可通过 custom_vocabulary 传入最多 1，000 个领域术语避免专有名词拼写错误，并提供 Smart Transcription 与 Verbatim 两种模式。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"dev.to","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe 完整指南：告别 ASR 转录难题 - Aioga AI资讯","description":"Google 推出专用于语音转文字的 Gemini 3.5 Transcribe 模型，主打快速、准确且低成本的转录，原生支持说话人分离和词级毫秒时间戳。该模型支持 85+ 种语言自动识别与代码切换，可通过 custom_vocabulary 传入最多 1，000 个领域术语避免专有名词拼写错误，并提供 Smart Transcription 与 Verba...","url":"https://www.aioga.com/news/cmtd00dbh09tkroq5v7kcw183/"},"en":{"title":"Gemini 3.5 Transcribe Complete Guide: Say Goodbye to ASR Transcription Challenges","summary":"Google has launched the Gemini 3.5 Transcribe model specifically for speech-to-text, focusing on fast, accurate, and low-cost transcription. It natively supports speaker separation and word-level millisecond timestamps. The model supports automatic recognition and code-switching for over 85 languages, allows up to 1,000 domain-specific terms to be added through custom_vocabulary to avoid misspelling proper nouns, and offers two modes: Smart Transcription and Verbatim. 🔗 Read the original via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"Industry","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Complete Guide: Say Goodbye to ASR Transcription Challenges - Aioga AI News","description":"Google has launched the Gemini 3.5 Transcribe model specifically for speech-to-text, focusing on fast, accurate, and low-cost transcription. It natively supports speaker separation...","url":"https://www.aioga.com/en/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:22:35.544Z"},"ja":{"title":"Gemini 3.5 転写 完全ガイド：ASR 転写の問題にさようなら","summary":"Google は音声の文字起こし専用の Gemini 3.5 Transcribe モデルを発表しました。これは高速で正確、かつ低コストの文字起こしを特徴としており、話者の分離や単語レベルのミリ秒タイムスタンプをネイティブにサポートしています。このモデルは 85 以上の言語の自動認識やコードスイッチをサポートしており、custom_vocabulary を通じて最大 1,000 の専門用語を入力することで固有名詞のスペルミスを防ぐことができます。また、Smart Transcription と Verbatim の 2 つのモードを提供しています。 🔗 原文を読む via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"業界動向","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 転写 完全ガイド：ASR 転写の問題にさようなら - Aioga AIニュース","description":"Google は音声の文字起こし専用の Gemini 3.5 Transcribe モデルを発表しました。これは高速で正確、かつ低コストの文字起こしを特徴としており、話者の分離や単語レベルのミリ秒タイムスタンプをネイティブにサポートしています。このモデルは 85 以上の言語の自動認識やコードスイッチをサポートしており、custom_vocabulary を通...","url":"https://www.aioga.com/ja/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:22:48.277Z"},"ko":{"title":"Gemini 3.5 Transcribe 완전 가이드: ASR 전사 문제와 작별","summary":"Google은 음성을 텍스트로 변환하는 전용 Gemini 3.5 Transcribe 모델을 출시했으며, 빠르고 정확하며 비용 효율적인 전사 기능을 강조합니다. 이 모델은 화자 분리 및 단어 수준 밀리초 타임스탬프를 기본 지원합니다. 또한 85개 이상의 언어 자동 인식 및 코드 스위칭을 지원하며, custom_vocabulary를 통해 최대 1,000개의 도메인 용어를 입력하여 고유 명사 철자 오류를 방지할 수 있고, Smart Transcription과 Verbatim 두 가지 모드를 제공합니다. 🔗 원문 읽기 via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"업계 동향","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe 완전 가이드: ASR 전사 문제와 작별 - Aioga AI 뉴스","description":"Google은 음성을 텍스트로 변환하는 전용 Gemini 3.5 Transcribe 모델을 출시했으며, 빠르고 정확하며 비용 효율적인 전사 기능을 강조합니다. 이 모델은 화자 분리 및 단어 수준 밀리초 타임스탬프를 기본 지원합니다. 또한 85개 이상의 언어 자동 인식 및 코드 스위칭을 지원하며, custom_vocabul...","url":"https://www.aioga.com/ko/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:23:49.100Z"},"es":{"title":"Guía completa de Gemini 3.5 Transcribe: Adiós a los problemas de transcripción ASR","summary":"Google lanza el modelo Gemini 3.5 Transcribe dedicado a la conversión de voz a texto, destacando por transcripciones rápidas, precisas y de bajo costo, con soporte nativo para separación de hablantes y marcas de tiempo de milisegundos a nivel de palabra. Este modelo admite el reconocimiento automático y el cambio de código en más de 85 idiomas, y permite introducir hasta 1,000 términos de vocabulario especializado mediante custom_vocabulary para evitar errores de ortografía en nombres propios, ofreciendo además dos modos: Smart Transcription y Verbatim. 🔗 Leer el artículo original via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"Industria","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Guía completa de Gemini 3.5 Transcribe: Adiós a los problemas de transcripción ASR - Aioga Noticias de IA","description":"Google lanza el modelo Gemini 3.5 Transcribe dedicado a la conversión de voz a texto, destacando por transcripciones rápidas, precisas y de bajo costo, con soporte nativo para sepa...","url":"https://www.aioga.com/es/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:23:39.983Z"},"fr":{"title":"Guide complet de Gemini 3.5 Transcribe : dites adieu aux difficultés de transcription ASR","summary":"Google a lancé le modèle Gemini 3.5 Transcribe spécialement conçu pour la conversion de la parole en texte, axé sur la transcription rapide, précise et à faible coût, avec un support natif pour la séparation des locuteurs et des index de temps au niveau du mot en millisecondes. Ce modèle prend en charge la reconnaissance automatique et le changement de code pour plus de 85 langues, permet de transmettre jusqu'à 1 000 termes spécialisés via custom_vocabulary pour éviter les fautes d'orthographe des noms propres, et offre deux modes : Smart Transcription et Verbatim. 🔗 Lire l'article original via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"Industrie","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Guide complet de Gemini 3.5 Transcribe : dites adieu aux difficultés de transcription ASR - Aioga Actualités IA","description":"Google a lancé le modèle Gemini 3.5 Transcribe spécialement conçu pour la conversion de la parole en texte, axé sur la transcription rapide, précise et à faible coût, avec un suppo...","url":"https://www.aioga.com/fr/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:24:44.507Z"},"de":{"title":"Gemini 3.5 Transcribe Komplettanleitung: Verabschieden Sie sich von ASR-Transkriptionsproblemen","summary":"Google hat das Gemini 3.5 Transcribe-Modell speziell für die Sprach-zu-Text-Übertragung vorgestellt, das sich durch schnelle, präzise und kostengünstige Transkriptionen auszeichnet und nativ Sprechertrennung sowie Wort-für-Wort-Millisekunden-Zeitstempel unterstützt. Das Modell erkennt automatisch über 85 Sprachen und Sprachwechsel, und es können bis zu 1.000 branchenspezifische Begriffe über custom_vocabulary eingegeben werden, um Schreibfehler bei Eigennamen zu vermeiden. Es bietet zwei Modi: Smart Transcription und Verbatim. 🔗 Originalartikel lesen via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Komplettanleitung: Verabschieden Sie sich von ASR-Transkriptionsproblemen - Aioga KI-News","description":"Google hat das Gemini 3.5 Transcribe-Modell speziell für die Sprach-zu-Text-Übertragung vorgestellt, das sich durch schnelle, präzise und kostengünstige Transkriptionen auszeichnet...","url":"https://www.aioga.com/de/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:24:44.313Z"},"pt-BR":{"title":"Guia completo do Gemini 3.5 Transcribe: diga adeus aos problemas de transcrição ASR","summary":"O Google lançou o modelo Gemini 3.5 Transcribe, dedicado à conversão de voz em texto, destacando-se por transcrições rápidas, precisas e de baixo custo, com suporte nativo para separação de falantes e timestamps em milissegundos ao nível de palavra. O modelo suporta reconhecimento automático de mais de 85 idiomas e alternância de código, podendo receber até 1.000 termos de domínio via custom_vocabulary para evitar erros de ortografia de nomes próprios, e oferece dois modos: Smart Transcription e Verbatim. 🔗 Leia o texto completo via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Guia completo do Gemini 3.5 Transcribe: diga adeus aos problemas de transcrição ASR - Aioga Notícias de IA","description":"O Google lançou o modelo Gemini 3.5 Transcribe, dedicado à conversão de voz em texto, destacando-se por transcrições rápidas, precisas e de baixo custo, com suporte nativo para sep...","url":"https://www.aioga.com/pt-BR/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:25:34.267Z"},"ru":{"title":"Gemini 3.5 Transcribe Полное руководство: прощай, трудности с ASR транскрипцией","summary":"Google выпустила модель Gemini 3.5 Transcribe, специально предназначенную для преобразования речи в текст, ориентированную на быструю, точную и недорогую транскрипцию, с нативной поддержкой разделения говорящих и пословных миллисекундных временных меток. Модель поддерживает автоматическое распознавание более 85 языков и переключение кодов, позволяет с помощью custom_vocabulary передавать до 1 000 терминов из конкретной области, чтобы избежать ошибок в написании собственных имен, и предлагает два режима: Smart Transcription и Verbatim. 🔗 Читать оригинал via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Полное руководство: прощай, трудности с ASR транскрипцией - Aioga Новости ИИ","description":"Google выпустила модель Gemini 3.5 Transcribe, специально предназначенную для преобразования речи в текст, ориентированную на быструю, точную и недорогую транскрипцию, с нативной п...","url":"https://www.aioga.com/ru/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:25:41.455Z"},"ar":{"title":"دليل Gemini 3.5 Transcribe الكامل: وداعًا لمشاكل النسخ الصوتي ASR","summary":"أطلقت جوجل نموذج Gemini 3.5 Transcribe المخصص لتحويل الصوت إلى نص، والذي يتميز بالتحويل السريع والدقيق وبتكلفة منخفضة، ويدعم أصلاً فصل المتحدثين والطوابع الزمنية بالمللي ثانية على مستوى الكلمات. يدعم هذا النموذج التعرف التلقائي على أكثر من 85 لغة وتبديل الأكواد، ويمكن تمرير ما يصل إلى 1000 مصطلح قطاعي عبر custom_vocabulary لتجنب أخطاء تهجئة الأسماء الخاصة، ويوفر وضعين هما Smart Transcription و Verbatim. 🔗 قراءة النص الأصلي عبر AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"دليل Gemini 3.5 Transcribe الكامل: وداعًا لمشاكل النسخ الصوتي ASR - Aioga أخبار الذكاء الاصطناعي","description":"أطلقت جوجل نموذج Gemini 3.5 Transcribe المخصص لتحويل الصوت إلى نص، والذي يتميز بالتحويل السريع والدقيق وبتكلفة منخفضة، ويدعم أصلاً فصل المتحدثين والطوابع الزمنية بالمللي ثانية على...","url":"https://www.aioga.com/ar/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:26:34.840Z"},"hi":{"title":"Gemini 3.5 प्रतिलेखन पूरी गाइड: ASR प्रतिलेखन की समस्याओं को अलविदा कहें","summary":"Google ने विशेष रूप से आवाज़ से टेक्स्ट में बदलने के लिए Gemini 3.5 Transcribe मॉडल लॉन्च किया है, जो तेज़, सटीक और कम लागत वाली ट्रांसक्रिप्शन को प्रमुखता देता है, और मूल रूप से स्पीकर अलगाव और शब्द स्तर के मिलीसेकंड टाइमस्टैम्प का समर्थन करता है। यह मॉडल 85+ भाषाओं की स्वचालित पहचान और कोड स्विचिंग का समर्थन करता है, और custom_vocabulary के माध्यम से अधिकतम 1,000 क्षेत्रीय शब्दों को पास करके विशेष नामों की वर्तनी त्रुटियों से बचा जा सकता है। यह Smart Transcription और Verbatim, दोनों मोड प्रदान करता है। 🔗 मूल लेख पढ़ें via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 प्रतिलेखन पूरी गाइड: ASR प्रतिलेखन की समस्याओं को अलविदा कहें - Aioga AI समाचार","description":"Google ने विशेष रूप से आवाज़ से टेक्स्ट में बदलने के लिए Gemini 3.5 Transcribe मॉडल लॉन्च किया है, जो तेज़, सटीक और कम लागत वाली ट्रांसक्रिप्शन को प्रमुखता देता है, और मूल रूप से स...","url":"https://www.aioga.com/hi/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:26:40.742Z"},"it":{"title":"Guida completa a Gemini 3.5 Transcribe: addio ai problemi di trascrizione ASR","summary":"Google ha lanciato il modello Gemini 3.5 Transcribe, appositamente progettato per la trascrizione vocale in testo, puntando su trascrizioni rapide, accurate e a basso costo, con supporto nativo per la separazione dei parlanti e timestamp a livello di parola in millisecondi. Il modello supporta il riconoscimento automatico di oltre 85 lingue e il cambio di codice, e tramite custom_vocabulary è possibile inserire fino a 1.000 termini settoriali per evitare errori di ortografia dei nomi propri, offrendo due modalità: Smart Transcription e Verbatim. 🔗 Leggi l'articolo completo su AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Guida completa a Gemini 3.5 Transcribe: addio ai problemi di trascrizione ASR - Aioga Notizie IA","description":"Google ha lanciato il modello Gemini 3.5 Transcribe, appositamente progettato per la trascrizione vocale in testo, puntando su trascrizioni rapide, accurate e a basso costo, con su...","url":"https://www.aioga.com/it/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:27:38.411Z"},"nl":{"title":"Gemini 3.5 Transcribe Volledige handleiding: Neem afscheid van ASR-transcriptieproblemen","summary":"Google heeft het Gemini 3.5 Transcribe-model gelanceerd, speciaal voor spraak-naar-tekst, gericht op snelle, nauwkeurige en kosteneffectieve transcriptie, met native ondersteuning voor sprekeronderbreking en woordniveau milliseconde-timestamps. Het model ondersteunt automatische herkenning en codewisseling in meer dan 85 talen, kan tot 1.000 vaktermen via custom_vocabulary invoeren om spellingfouten van eigennamen te voorkomen, en biedt twee modi: Smart Transcription en Verbatim. 🔗 Lees het originele artikel via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Volledige handleiding: Neem afscheid van ASR-transcriptieproblemen - Aioga AI-nieuws","description":"Google heeft het Gemini 3.5 Transcribe-model gelanceerd, speciaal voor spraak-naar-tekst, gericht op snelle, nauwkeurige en kosteneffectieve transcriptie, met native ondersteuning...","url":"https://www.aioga.com/nl/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:27:32.686Z"},"tr":{"title":"Gemini 3.5 Transcribe Tam Rehberi: ASR Transkripsiyon Sorunlarına Veda","summary":"Google, hızlı, doğru ve düşük maliyetli transkripsiyona odaklanan sesten metne dönüştürme için özel Gemini 3.5 Transcribe modelini tanıttı; model, konuşmacı ayrımı ve kelime düzeyinde milisaniye zaman damgalarını yerel olarak destekliyor. Bu model, 85'ten fazla dili otomatik tanıma ve kod geçişini destekliyor, custom_vocabulary ile en fazla 1.000 alan terimi girilerek özel isimlerin yazım hatalarını önleyebiliyor ve Smart Transcription ile Verbatim olmak üzere iki mod sunuyor. 🔗 Orijinal metni okuyun via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Tam Rehberi: ASR Transkripsiyon Sorunlarına Veda - Aioga AI Haberleri","description":"Google, hızlı, doğru ve düşük maliyetli transkripsiyona odaklanan sesten metne dönüştürme için özel Gemini 3.5 Transcribe modelini tanıttı; model, konuşmacı ayrımı ve kelime düzeyi...","url":"https://www.aioga.com/tr/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:28:40.376Z"},"vi":{"title":"Gemini 3.5 Transcribe Hướng dẫn đầy đủ: Nói lời tạm biệt với những khó khăn khi chuyển âm ASR","summary":"Google ra mắt mô hình Gemini 3.5 Transcribe chuyên cho chuyển giọng nói thành văn bản, nổi bật với khả năng ghi âm nhanh, chính xác và chi phí thấp, hỗ trợ sẵn tách người nói và dấu thời gian theo từ với độ chính xác mili giây. Mô hình này hỗ trợ nhận diện tự động hơn 85 ngôn ngữ và chuyển đổi mã, có thể truyền tối đa 1.000 thuật ngữ chuyên ngành qua custom_vocabulary để tránh lỗi chính tả từ riêng, và cung cấp hai chế độ: Smart Transcription và Verbatim. 🔗 Đọc bài gốc via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Hướng dẫn đầy đủ: Nói lời tạm biệt với những khó khăn khi chuyển âm ASR - Tin tức AI Aioga","description":"Google ra mắt mô hình Gemini 3.5 Transcribe chuyên cho chuyển giọng nói thành văn bản, nổi bật với khả năng ghi âm nhanh, chính xác và chi phí thấp, hỗ trợ sẵn tách người nói và dấ...","url":"https://www.aioga.com/vi/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:28:38.625Z"},"id":{"title":"Gemini 3.5 Transcribe Panduan Lengkap: Mengucapkan Selamat Tinggal pada Masalah Transkripsi ASR","summary":"Google meluncurkan model Gemini 3.5 Transcribe khusus untuk konversi suara ke teks, dengan fokus pada transkripsi yang cepat, akurat, dan biaya rendah, mendukung secara native pemisahan pembicara dan stempel waktu milidetik per kata. Model ini mendukung pengenalan otomatis dan pergantian kode lebih dari 85 bahasa, dapat menyertakan hingga 1.000 istilah domain melalui custom_vocabulary untuk menghindari kesalahan ejaan kata khusus, dan menyediakan dua mode: Smart Transcription dan Verbatim. 🔗 Baca selengkapnya via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Panduan Lengkap: Mengucapkan Selamat Tinggal pada Masalah Transkripsi ASR - Berita AI Aioga","description":"Google meluncurkan model Gemini 3.5 Transcribe khusus untuk konversi suara ke teks, dengan fokus pada transkripsi yang cepat, akurat, dan biaya rendah, mendukung secara native pemi...","url":"https://www.aioga.com/id/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:29:37.155Z"},"th":{"title":"คู่มือสมบูรณ์ของ Gemini 3.5 Transcribe: บอกลาอุปสรรคการถอดความ ASR","summary":"Google เปิดตัวโมเดล Gemini 3.5 Transcribe สำหรับการแปลงเสียงเป็นข้อความโดยเฉพาะ เน้นการถอดความที่รวดเร็ว ถูกต้อง และต้นทุนต่ำ รองรับการแยกผู้พูดและการประทับเวลาหน่วยมิลลิวินาทีบนคำโดยตรง โมเดลนี้รองรับการรู้จำอัตโนมัติและการสลับโค้ดได้มากกว่า 85 ภาษา สามารถป้อนคำศัพท์เฉพาะสูงสุด 1,000 คำผ่าน custom_vocabulary เพื่อป้องกันการสะกดชื่อเฉพาะผิด และมีโหมด Smart Transcription และ Verbatim 🔗 อ่านต้นฉบับ via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"คู่มือสมบูรณ์ของ Gemini 3.5 Transcribe: บอกลาอุปสรรคการถอดความ ASR - ข่าว AI Aioga","description":"Google เปิดตัวโมเดล Gemini 3.5 Transcribe สำหรับการแปลงเสียงเป็นข้อความโดยเฉพาะ เน้นการถอดความที่รวดเร็ว ถูกต้อง และต้นทุนต่ำ รองรับการแยกผู้พูดและการประทับเวลาหน่วยมิลลิวินาทีบนคำ...","url":"https://www.aioga.com/th/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:29:47.546Z"},"pl":{"title":"Gemini 3.5 Transcribe Kompletny przewodnik: pożegnanie z problemami transkrypcji ASR","summary":"Google wprowadza model Gemini 3.5 Transcribe, specjalnie przeznaczony do przekształcania mowy w tekst, z naciskiem na szybkie, dokładne i niskokosztowe transkrypcje, z natywnym wsparciem dla rozdzielania mówców i znaczników czasowych na poziomie milisekundowym słów. Model obsługuje automatyczne rozpoznawanie ponad 85 języków i przełączanie kodów, można przekazać maksymalnie 1 000 terminów branżowych za pomocą custom_vocabulary, aby uniknąć błędów w pisowni nazw własnych, oraz oferuje dwa tryby: Smart Transcription i Verbatim. 🔗 Przeczytaj oryginał via AIHOT · https://aihot.virxact.com/items/cmtd00dbh09tkroq5v7kcw183","category":"行业动态","source":"Google AI：DEV 作者专属（RSS）","aggregationSource":"Google AI：DEV 作者专属（RSS）","pageTitle":"Gemini 3.5 Transcribe Kompletny przewodnik: pożegnanie z problemami transkrypcji ASR - Aioga Wiadomości AI","description":"Google wprowadza model Gemini 3.5 Transcribe, specjalnie przeznaczony do przekształcania mowy w tekst, z naciskiem na szybkie, dokładne i niskokosztowe transkrypcje, z natywnym wsp...","url":"https://www.aioga.com/pl/news/cmtd00dbh09tkroq5v7kcw183/","contentTranslated":true,"sourceHash":"35c14a5a3b67241f","translatedAt":"2026-08-29T10:30:58.751Z"}}}}