Code-switching stops breaking the transcript: Gemini 3.5 Transcribe logs 2.6% WER across 85 languages — MarkTechPost
Google's speech model transcribes recorded audio at 2.6% word error rate and streams at 4.0%, detects languages on its own, and follows mid-sentence switches. Batch pricing runs about $0.005 a minute, API only, no open weights.
Positive-sum angle: Cheap accurate transcription lifts everything downstream of speech: contact centres, clinics, courts, and the small teams building vertical tools on top. When the transcript stops mangling Singlish and Taglish, every product above it inherits a market it could not serve before.
What's the impact: Operators running Manila and KL contact centres should reprice quality assurance per minute of audio rather than per auditor. The API is public; the region's code-switching evaluation sets are not, so build the benchmark before a vendor sells one back.