> ## Documentation Index
> Fetch the complete documentation index at: https://developer.suki.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Dictation WebSocket Handshake Fails

> Fix Dictation /ws/transcribe open failures from session state, auth, or Ambient framing

Dictation streams on **`GET /ws/transcribe`** after you create a transcription session. Handshake failures usually mean another speech stream is still finishing, auth is wrong, or you sent Ambient framing by mistake.

Dictation does **not** use **`START_TIME`**, ambient-style **`AUDIO`** with a **`data`** field, or the ambient end marker **`RU9G`**.

## Common causes

* Opening `/ws/transcribe` before create returns a `transcription_session_id`, or while another speech stream on that ID is still finishing.
* Reconnecting too soon after **`AUDIO_END`** and inbound **`EOF`**.
* Wrong browser `Sec-WebSocket-Protocol` string (for example `SukiTranscriptionAuth`, or token and session ID in the wrong order).
* Sending Ambient **`START_TIME`**, **`data`**, or **`RU9G`** on the Dictation socket.

## Fix

<Steps>
  <Step title="Create the Dictation Session First">
    Create or reuse a Dictation session, then open **`GET /ws/transcribe`**. Only one speech stream can be open on that `transcription_session_id` at a time.
  </Step>

  <Step title="Authenticate with SukiAmbientAuth">
    In the browser, use `Sec-WebSocket-Protocol: SukiAmbientAuth,<sdp_suki_token>,<transcription_session_id>`. Put the token before the session ID. Do not use `SukiTranscriptionAuth`. Only the session ID value differs from Ambient (`transcription_session_id` vs `ambient_session_id`).
  </Step>

  <Step title="Use the Dictation Wire Format">
    Send UTF-8 JSON text frames. Put audio in **`audioData`** as Base64 **PCM\_S16LE**. After the last chunk, send `{"type":"EVENT","event":"AUDIO_END"}`. See [Dictation streaming wire format](/documentation/how-to/audio-streaming/websocket-streaming-wire-format-dictation).
  </Step>

  <Step title="Wait Before the Next Speech WebSocket">
    There is **no partner status API** for Dictation (`READY`, `IDLE`, `RUNNING`). Do not poll those states. After **`AUDIO_END`** and **`EOF`**, wait about **5 seconds** before you open the next WebSocket on the same session. If another stream is still finishing, the handshake fails with **`FailedPrecondition`**.
  </Step>

  <Step title="Close the Socket, Then REST End">
    After you finish reading frames (including terminal **`EOF`**), close the WebSocket, then call [End Dictation session](/api-reference/audio-transcription/end-session). The End response body is empty. Keep transcript text from frames where **`is_final`** is **`true`**.
  </Step>
</Steps>

```json theme={"theme":{"light":"github-dark","dark":"material-theme-darker"}}
{ "type": "AUDIO", "audioData": "<base64-encoded PCM_S16LE audio>" }
{ "type": "AUDIO", "audioData": "<base64(pcm_s16le chunk 2)>" }

{ "type": "EVENT", "event": "AUDIO_END" }
```

<Warning>
  Do not end Dictation audio with `{"type":"AUDIO","data":"RU9G"}`. That is the ambient **`/ws/stream`** pattern. Do not send binary WebSocket frames for audio on this JSON protocol.
</Warning>

| Problem | Why it happens | How to fix it |
| :- | :- | :- |
| Handshake **`FailedPrecondition`** | Another speech session on that ID is still finishing, or reconnect too soon after **`EOF`** | After **`AUDIO_END`** and **`EOF`**, wait about **5 seconds** before the next WebSocket |
| **401** on handshake or immediate disconnect | Wrong `Sec-WebSocket-Protocol` string or token/session order | Use `SukiAmbientAuth,<sdp_suki_token>,<transcription_session_id>` |
| JSON parse errors | Binary frame, non-JSON payload, or multiple objects in one frame | One UTF-8 JSON object per text frame |
| Audio ignored or poor transcription | Ambient **`data`** field, non-PCM bytes, or WAV headers | Base64 **PCM\_S16LE** in **`audioData`**; strip WAV headers |
| Session never completes | Missing end-of-audio message | Send `{"type":"EVENT","event":"AUDIO_END"}` after the last **`AUDIO`** chunk |

## Next steps

<Icon icon="file-lines" iconType="solid" /> **[Build a Dictation streaming client](/documentation/tutorials/dictation-websocket-code-example)** - Create, stream, read frames, and end

<Icon icon="file-lines" iconType="solid" /> **[End Dictation with AUDIO\_END](/documentation/cookbooks/end-dictation-with-audio-end)** - Use AUDIO\_END, not ambient RU9G

<Icon icon="file-lines" iconType="solid" /> **[WAV header streamed as audio](/documentation/troubleshooting/wav-header-streamed-as-audio)** - Strip RIFF before PCM\_S16LE

<Icon icon="file-lines" iconType="solid" /> **[Ambient WebSocket disconnects](/documentation/troubleshooting/websocket-disconnects-ambient)** - Ambient `/ws/stream` keep-alives and RU9G

<Icon icon="file-lines" iconType="solid" /> **[Wrong staging vs production endpoints](/documentation/troubleshooting/wrong-environment-endpoints)** - Match hosts and tokens to the environment
