Skip to main content
Dictation streams on GET /ws/transcribe after you create a transcription session. Handshake failures usually mean another speech stream is still finishing, auth is wrong, or you sent Ambient framing by mistake. Dictation does not use START_TIME, ambient-style AUDIO with a data field, or the ambient end marker RU9G.

Common causes

  • Opening /ws/transcribe before create returns a transcription_session_id, or while another speech stream on that ID is still finishing.
  • Reconnecting too soon after AUDIO_END and inbound EOF.
  • Wrong browser Sec-WebSocket-Protocol string (for example SukiTranscriptionAuth, or token and session ID in the wrong order).
  • Sending Ambient START_TIME, data, or RU9G on the Dictation socket.

Fix

1

Create the Dictation Session First

Create or reuse a Dictation session, then open GET /ws/transcribe. Only one speech stream can be open on that transcription_session_id at a time.
2

Authenticate with SukiAmbientAuth

In the browser, use Sec-WebSocket-Protocol: SukiAmbientAuth,<sdp_suki_token>,<transcription_session_id>. Put the token before the session ID. Do not use SukiTranscriptionAuth. Only the session ID value differs from Ambient (transcription_session_id vs ambient_session_id).
3

Use the Dictation Wire Format

Send UTF-8 JSON text frames. Put audio in audioData as Base64 PCM_S16LE. After the last chunk, send {"type":"EVENT","event":"AUDIO_END"}. See Dictation streaming wire format.
4

Wait Before the Next Speech WebSocket

There is no partner status API for Dictation (READY, IDLE, RUNNING). Do not poll those states. After AUDIO_END and EOF, wait about 5 seconds before you open the next WebSocket on the same session. If another stream is still finishing, the handshake fails with FailedPrecondition.
5

Close the Socket, Then REST End

After you finish reading frames (including terminal EOF), close the WebSocket, then call End Dictation session. The End response body is empty. Keep transcript text from frames where is_final is true.
Do not end Dictation audio with {"type":"AUDIO","data":"RU9G"}. That is the ambient /ws/stream pattern. Do not send binary WebSocket frames for audio on this JSON protocol.

Next steps

Build a Dictation streaming client - Create, stream, read frames, and end End Dictation with AUDIO_END - Use AUDIO_END, not ambient RU9G WAV header streamed as audio - Strip RIFF before PCM_S16LE Ambient WebSocket disconnects - Ambient /ws/stream keep-alives and RU9G Wrong staging vs production endpoints - Match hosts and tokens to the environment
Last modified on September 29, 2026