Back to list
Google Launches Gemini 3.5 Transcribe Featuring Automatic Filler Word Removal and Support for Over 85 Languages
Product LaunchGoogle GeminiAI TranscriptionAudio Intelligence

Google Launches Gemini 3.5 Transcribe Featuring Automatic Filler Word Removal and Support for Over 85 Languages

Google has officially expanded its Gemini Audio suite with the introduction of Gemini 3.5 Transcribe. This new AI-powered tool is designed to significantly enhance the quality of audio-to-text conversions by automatically filtering out common speech disfluencies, such as "ums" and "ahs." Beyond simply cleaning up speech, the model features advanced capabilities for detecting specialized technical jargon across various industries and offers robust support for more than 85 languages. This release follows the recent debut of Gemini 3.5 Live Translate, marking another step in Google's rollout of its 3.5-generation models. While the tech community continues to wait for the flagship Gemini 3.5 Pro model, Gemini 3.5 Transcribe provides a specialized solution for users seeking polished, professional, and linguistically diverse transcription services.

The Verge

Key Takeaways

  • Automated Speech Polishing: Gemini 3.5 Transcribe automatically detects and removes filler words like "ums" and "ahs" to create cleaner text outputs.
  • Extensive Language Support: The model supports transcription in more than 85 languages, catering to a global user base.
  • Jargon Detection: Advanced capabilities allow the AI to identify and correctly transcribe specialized industry-specific jargon.
  • Product Roadmap: This release follows Gemini 3.5 Live Translate, though the highly anticipated Gemini 3.5 Pro model is still pending release.

In-Depth Analysis

Refining Audio Intelligence with Gemini 3.5 Transcribe

Google's latest update to the Gemini Audio ecosystem, Gemini 3.5 Transcribe, represents a focused effort to move beyond literal transcription toward intelligent content refinement. One of the most prominent features of this new model is its ability to automatically edit out speech disfluencies. By identifying and removing "ums," "ahs," and other common filler words, the AI ensures that the resulting transcript is not just a record of what was said, but a professional document ready for immediate use. This functionality addresses a long-standing pain point in automated transcription, where raw text often requires extensive manual editing to remove the natural hesitations of human speech.

Furthermore, the model's ability to detect specialized jargon suggests a higher level of contextual awareness. In professional environments—ranging from medical and legal fields to high-tech engineering—standard transcription tools often struggle with niche terminology. By integrating jargon detection, Gemini 3.5 Transcribe aims to provide higher accuracy for specialized users, ensuring that technical terms are captured correctly without the need for constant manual correction.

Global Accessibility and Language Diversity

The scale of Gemini 3.5 Transcribe’s language support is a significant highlight of this release. With support for more than 85 languages, Google is positioning this tool as a comprehensive solution for international organizations and multilingual users. This broad support, combined with the model's ability to handle specialized terminology, indicates that Google is prioritizing the versatility of its AI tools across different cultural and professional contexts.

This release is strategically positioned within the Gemini 3.5 family. It follows the launch of Gemini 3.5 Live Translate, suggesting a modular approach to Google's AI rollout. By releasing specialized tools like Live Translate and Transcribe, Google is providing immediate utility to users in specific niches while the broader, more powerful Gemini 3.5 Pro model remains in development. This phased rollout allows Google to refine specific audio and linguistic capabilities before integrating them into a more comprehensive flagship model.

Industry Impact

The introduction of Gemini 3.5 Transcribe has several implications for the AI and transcription industries. First, it raises the bar for "out-of-the-box" transcript quality. As AI models become more adept at editing speech for clarity rather than just transcribing it verbatim, the demand for manual transcription and editing services may decrease. For businesses, this translates to faster turnaround times for meeting notes, interviews, and content creation.

Second, the focus on jargon and multilingual support highlights the growing importance of specialized AI. General-purpose models are increasingly being supplemented or replaced by tools that understand the specific nuances of different languages and professional sectors. Google's move to include these features in the Gemini 3.5 Transcribe model underscores a trend toward AI that is not only conversational but also technically precise and globally accessible. As the industry awaits Gemini 3.5 Pro, these incremental updates to the Gemini Audio suite demonstrate Google's commitment to maintaining a competitive edge in the rapidly evolving landscape of audio intelligence.

Frequently Asked Questions

Question: What are the main features of Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is designed to automatically remove filler words such as "ums" and "ahs" from audio transcriptions. It also features the ability to detect specialized industry jargon and supports more than 85 different languages, providing a more polished and accurate text output compared to standard transcription tools.

Question: How does this model fit into the Gemini 3.5 product lineup?

Gemini 3.5 Transcribe is a new addition to the Gemini family, following the launch of Gemini 3.5 Live Translate. It is part of Google's updated Gemini Audio capabilities. While these specialized tools are now available, Google has not yet released the Gemini 3.5 Pro model.

Question: Can Gemini 3.5 Transcribe handle technical or professional terminology?

Yes, one of the key features of Gemini 3.5 Transcribe is its ability to automatically detect and accurately transcribe specialized jargon, making it suitable for use in professional and technical fields where specific terminology is common.

Related News

Nolla Health Launches AI System in Utah to Scan Faces and Autonomously Prescribe Acne Treatment
Product Launch

Nolla Health Launches AI System in Utah to Scan Faces and Autonomously Prescribe Acne Treatment

Healthcare startup Nolla Health has officially announced the launch of an artificial intelligence-powered application in Utah that allows residents to receive prescriptions for acne treatment without human doctor intervention. By scanning their faces directly through the startup's mobile application, users enable an AI system to analyze the severity of their acne and autonomously generate a medical prescription. The service, which was earlier reported by Bloomberg, marks a significant milestone in automated clinical care and digital health, bringing algorithmic assessment and direct prescribing capabilities into consumers' hands within the state of Utah.

Product Launch

HyperFrames Studio Desktop Launches on Product Hunt as an Agent-Native Video Editing Workspace

HyperFrames Studio (Desktop) has officially launched on Product Hunt, introduced as the first video editor specifically engineered for AI coding agents and human creators. Developed by the team behind HeyGen's open-source HyperFrames project, the desktop application bridges the gap between agentic code generation and visual video editing. While AI agents like Claude Code and OpenAI Codex can generate video sequences by writing code as HTML and rendering to MP4, fine-tuning visual details and timing purely through chat prompts has historically been challenging. HyperFrames Studio solves this friction by providing a shared desktop workspace where creators remain in the director's seat while collaborating directly with their coding agents. Available for macOS and Linux, the release represents a significant shift toward agent-driven multimedia production workflows.

Product Launch

Spira Maxima Launches on Product Hunt: An End-to-End AI Video Model Converting Scripts into Viral Social Clips

Spira AI has officially unveiled Spira Maxima on Product Hunt, introducing an advanced social video model engineered to transform plain scripts into fully edited, viral-ready video content in a single pass. Designed by a team with roots at TikTok, CapCut, Meta, Snap, Midjourney, and Creatify AI, Spira Maxima addresses the industry-wide bottleneck of video post-production. Instead of requiring creators to manually cut B-roll, sync voiceovers, design captions, and select background tracks, the system automates the entire finishing workflow. Creators can deploy AI presenters, generate personalized clones with custom voice samples, and integrate native product footage post-trained on real-world social engagement data. By eliminating the manual friction between raw generation and final publishing, Spira Maxima sets a new benchmark for automated social media marketing and automated content pipelines.