Gemini 3.5 Transcribe now available on AI Gateway

Gemini 3.5 Transcribe from Google is now available on AI Gateway. It takes audio and returns text, in two variants:

  • google/gemini-3.5-transcribe transcribes a complete recording in a single request.
  • google/gemini-3.5-transcribe-live transcribes audio over a WebSocket, returning a transcript that updates while the recording is still going.

The model detects the language on its own, covers 85+, and follows a speaker who switches language partway through. You can also supply custom vocabulary so it recognizes names, jargon, and spellings.

Streaming transcription is new in AI SDK V7:

Live transcription

streamTranscribe opens the socket and takes a ReadableStream of raw audio chunks, so you can pass a microphone straight through. Tell it the format you are sending with inputAudioFormat:

Complete recordings

For audio you already have on disk, transcribe sends it in one request and returns the text:

You can also try the model without writing any code. Open Gemini 3.5 Transcribe Live and send audio to read the transcript in the browser.

AI Gateway provides a unified API for calling models, tracking usage and cost, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, budgets for API keys, routing rules, and more.

AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests.

You can view all transcription models available on AI Gateway, or start from the speech quickstart.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论