> ## Documentation Index
> Fetch the complete documentation index at: https://docs.modelstudio.console.alibabacloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Real-time audio and video translation - Qwen

> This topic introduces Qwen real-time speech and audiovisual translation, including model capabilities, supported models, and integration. The models use audio and image input to translate in real time and output text or speech in the target language for voice communication and video translation.

> Try an online demo with [one-click deployment using Function Compute](/en/model-studio/qwen3-5-livetranslate-flash-realtime#7727c7c1ed6du) .

## Features <span id="b980204d56gh6" /> <span id="06ce857a7f8ll" />

- <strong>Multi-language support</strong>: Translates between 60 languages — 29 with audio and text output, 31 with text-only output — including Chinese, English, French, German, Russian, Japanese, Korean, Spanish, Portuguese, and Arabic.
- <strong>Visual enhancement</strong>: Analyzes visual cues, such as lip movements, gestures, and on-screen text, to improve translation accuracy, especially in noisy environments or for ambiguous words.
- <strong>2.3-second latency</strong>: Delivers simultaneous interpretation with latency as low as 2.3 seconds.
- <strong>Real-time speaker diarization</strong>: Distinguishes different speakers and their speech when multiple people take turns speaking, so listeners can clearly understand who said what.
- <strong>Lossless simultaneous interpretation</strong>: Predicts semantic units to resolve cross-language word order differences, achieving quality comparable to offline translation.
- <strong>Natural voice</strong>: Matches the intonation and emotion of the source audio automatically.
- <strong>Hotword configuration</strong>: Configurable hotwords improve translation accuracy for specific terms.
- <strong>Voice cloning</strong>: Clones the speaker's voice for translated output. Supports server-side real-time cloning and pre-cloned voice profiles.

## Procedure <span id="7e1da95c25ej6" /> <span id="0719f2539546b" />

### 1. Configure the connection <span id="bdaa43cdd7hsd" /> <span id="4dbf1dc38dj77" />

Set `model` to a [supported model](#8f2355abeb4ei) when connecting. For the endpoint and a complete example using `qwen3.8-livetranslate-flash-realtime`, see [Quick start](#a36e6dc44fucp). The connection example below uses `qwen3.8-livetranslate-flash-realtime`.

The model connects over WebSocket with the following parameters:

<Note>
  In addition to WebSocket, `qwen3.5-livetranslate-flash-realtime` also supports the AOQ and WebRTC protocols. For client-side integration that prioritizes stable latency, resilience on weak networks, and built-in full-duplex noise suppression and echo cancellation, AOQ is recommended. For a protocol comparison, see [Realtime API overview](/en/model-studio/realtime-api-overview#rtov-s02h2).
</Note>

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}>
  <colgroup>
    <col style={{ width: "31.38%" }} />

    <col style={{ width: "68.62%" }} />
  </colgroup>

  <thead>
    <tr>
      <th>
        <strong>Parameter</strong>
      </th>

      <th>
        <strong>Description</strong>
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        endpoint
      </td>

      <td>
        China (Beijing) region: wss\://\{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime. Replace \{WorkspaceId} with your actual workspace ID.

        Singapore region: wss\://\{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime.

        Replace `{WorkspaceId}` with your actual [workspace ID](/en/model-studio/regions#h2_migrate_domain).
      </td>
    </tr>

    <tr>
      <td>
        query parameter
      </td>

      <td>
        The model query parameter must be set to the model name. Example: `?model=qwen3.8-livetranslate-flash-realtime`
      </td>
    </tr>

    <tr>
      <td>
        message header
      </td>

      <td>
        Use a Bearer Token for authentication: Authorization: Bearer DASHSCOPE\_API\_KEY

        > DASHSCOPE\_API\_KEY is your API key from Model Studio.
      </td>
    </tr>
  </tbody>
</table>

Sample connection code (Python):

<Accordion title="Python sample code for WebSocket connection">
  ```python expandable
  # pip install websocket-client
  import json
  import websocket
  import os

  API_KEY=os.getenv("DASHSCOPE_API_KEY")
  API_URL = "wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-livetranslate-flash-realtime"

  headers = [
      "Authorization: Bearer " + API_KEY
  ]

  def on_open(ws):
      print(f"Connected to server: {API_URL}")
  def on_message(ws, message):
      data = json.loads(message)
      print("Received event:", json.dumps(data, indent=2))
  def on_error(ws, error):
      print("Error:", error)

  ws = websocket.WebSocketApp(
      API_URL,
      header=headers,
      on_open=on_open,
      on_message=on_message,
      on_error=on_error
  )

  ws.run_forever()
  ```
</Accordion>

### 2. Configure language, modality, and voice <span id="6905316411elw" /> <span id="c547123c86b0s" />

<Tabs>
  <Tab title="qwen3.8-livetranslate-flash-realtime">
    Configure the session through [session.update](/en/model-studio/live-translator-client-events#af43722339yva):

    - **Target language**: Set `session.translation.language`, such as `en` for English.
    - **Output modalities**: Set `session.output_modalities` to `["text"]` for text only or `["text", "audio"]` for text and audio.
    - **Source transcription**: Receive increments through `conversation.item.input_audio_transcription.delta` and complete transcripts through `conversation.item.input_audio_transcription.completed`.
    - **Audio and voice**: Defaults are 16000 Hz PCM input, 24000 Hz PCM output, and voice `Tina`. See [Server events](/en/model-studio/live-translator-server-events#39689ed6e90ag) for the session structure.
  </Tab>

  <Tab title="qwen3.5-livetranslate-flash-realtime">
    Send the [session.update](/en/model-studio/live-translator-client-events#af43722339yva) client event with the following parameters:

    - <strong>Language</strong>

      - <strong>Source language:</strong> Configure using the `session.input_audio_transcription.language` parameter.

        > If not specified, the model automatically detects the source language.
      - <strong>Target language:</strong> Configure using the `session.translation.language` parameter.

        > The default value is `en` (English).

      See [Supported languages](/en/model-studio/qwen3-5-livetranslate-flash-realtime#4ffd192226f0s).
    - <strong>Output source language recognition results</strong>

      Set `session.input_audio_transcription.model` to `qwen3-asr-flash-realtime`. The server then returns both the translation and the speech recognition result (original text) for the input audio.

      The server returns these events:

      - `conversation.item.input_audio_transcription.text`: Streams the recognition results.
      - `conversation.item.input_audio_transcription.completed`: Returns the final result after the recognition is complete.
      - `conversation.item.input_audio_transcription.failed`: Returns error information when recognition fails.
    - <strong>Output modality</strong>

      Set the `session.modalities` parameter to `["text"]` (text only) or `["text","audio"]` (text and audio).
    - <strong>Voice Activity Detection (VAD) and Manual mode</strong>

      Configure how speech boundaries are detected using the `session.turn_detection` parameter:

      - <strong>VAD mode</strong> (default): Set `turn_detection` to a configuration object. The server automatically detects speech boundaries and triggers translation, suitable for scenarios where the client continuously sends audio streams.
      - <strong>Manual mode</strong>: Set `turn_detection` to `null`. The client determines speech boundaries and sends an `input_audio_buffer.commit` event to submit the audio after each utterance, suitable for push-to-talk scenarios.

      For the complete interaction steps under both modes, see [3. Input audio and images](/en/model-studio/qwen3-5-livetranslate-flash-realtime#82b7d6329836b).
    - <strong>Voice</strong>

      Configure using the `session.voice` parameter. See [Supported voices](/en/model-studio/qwen3-5-livetranslate-flash-realtime#0a5bde7593gdk).
    - <strong>Hotword</strong>

      Configure hotwords using the `session.translation.corpus.phrases` parameter. Hotwords are key-value pairs that map source terms to target translations, improving accuracy for specific terms. We recommend configuring no more than 1000 hotwords.

      Example: Map `"artificial intelligence"` to `"Artificial Intelligence"`.
    - <strong>Voice cloning</strong>

      Configure using the `session.enable_voice_clone`, `session.voice_clone_options.frequency`, and `session.voice` parameters. Supports three modes: pre-cloned voice profile (`frequency`: `never`), server-side clone once at session start (`once`), or real-time clone before each response (`always`). See [Voice cloning](/en/model-studio/qwen3-5-livetranslate-flash-realtime#kvwdr9d8k2gh4).
  </Tab>
</Tabs>

### 3. Input audio and images <span id="82b7d6329836b" /> <span id="e465eea44dhah" />

Send Base64-encoded audio and image data using the [input\_audio\_buffer.append](/en/model-studio/live-translator-client-events#8d11313f2198k) and [input\_image\_buffer.append](/en/model-studio/live-translator-client-events#e27908854eaht) events. Audio input is required; image input is optional.

> Images can be from a local file or captured in real time from a video stream.

The VAD and Manual configurations below apply to `qwen3.5-livetranslate-flash-realtime`. For `qwen3.8-livetranslate-flash-realtime`, the default turn detection configuration is `audio.input.turn_detection.type = speaker_detection`; send audio continuously and receive server-generated responses.

How the model determines that an utterance is complete depends on the VAD mode or Manual mode configured via the [turn\_detection](/en/model-studio/live-translator-client-events) parameter:

- <strong>VAD mode</strong> (default): The client continuously sends [input\_audio\_buffer.append](/en/model-studio/live-translator-client-events#8d11313f2198k) events. When the server detects speech start/end, it returns `input_audio_buffer.speech_started` and `input_audio_buffer.speech_stopped` events respectively, automatically commits the audio buffer, and triggers translation. Translation responses are generated synchronously with the streaming audio and typically begin during audio input, without waiting for the speech to end.
- <strong>Manual mode</strong>: Set `session.turn_detection` to `null`. After the client finishes sending a complete utterance, it sends an [input\_audio\_buffer.commit](/en/model-studio/live-translator-client-events#ltcommit001sec) event to commit the audio buffer. After the server returns an `input_audio_buffer.committed` event to confirm, it automatically starts generating the translation response; the client does not need to send any other event to trigger the response. To clear uncommitted audio before committing, send an [input\_audio\_buffer.clear](/en/model-studio/live-translator-client-events#ltclear001sec) event.

### 4. Receive the model response <span id="87c543412esso" /> <span id="9aed5b6ca9xnq" />

<Tabs>
  <Tab title="qwen3.8-livetranslate-flash-realtime">
    Handle responses according to the output modalities:

    - **Text only**: Concatenate the `delta` values from `response.text.delta` to obtain the translation.
    - **Text and audio**: Concatenate `delta` from `response.audio_transcript.delta` for text, and Base64-decode `delta` from `response.audio.delta` to obtain audio chunks.

    `response.done` indicates the end of a response. See [Server events](/en/model-studio/live-translator-server-events#text-delta) for event fields.
  </Tab>

  <Tab title="qwen3.5-livetranslate-flash-realtime">
    Translation responses are generated synchronously with the streaming audio and typically do not require waiting for speech to end (see the VAD/Manual mode description in the previous section). The response format depends on the output modality.

    - <strong>Text-only output</strong>

      The server streams incremental translated text (including confirmed text and tentative predicted text) through [response.text.text](/en/model-studio/live-translator-server-events#0c54be63e0c3w) events; upon completion, the full translated text is returned in a [response.text.done](/en/model-studio/live-translator-server-events#d675635a94jfb) event.
    - <strong>Text and audio output</strong>

      - <strong>Text</strong>

        The server streams incremental translated text through [response.audio\_transcript.text](/en/model-studio/live-translator-server-events#35396453cfood) events; upon completion, the full translated text is returned in a [response.audio\_transcript.done](/en/model-studio/live-translator-server-events#f4d1698567bsm) event.
      - <strong>Audio</strong>

        The server returns incremental, Base64-encoded audio data in [response.audio.delta](/en/model-studio/live-translator-server-events#a25cc50a15car) events.

    <Tip>
      `qwen3.5-livetranslate-flash-realtime` uses the `response.text.text` event for incremental text delivery, which differs from the `response.text.delta` event used by Omni (full-duplex voice conversation) models. These events have different field structures and semantics — do not use them interchangeably.
    </Tip>
  </Tab>
</Tabs>

### 5. End the session <span id="sf002endhd" /> <span id="sf002endsec" />

After sending all audio, send a [Client events](/en/model-studio/live-translator-client-events) event, then wait for the server to return a `session.finished` event before closing the WebSocket connection.

If you close the WebSocket without sending `session.finish`, the server's VAD cannot detect the end of the final speech segment. This causes translation results for that segment to be lost entirely, and the connection may hang indefinitely. Always send this event before disconnecting.

## Supported models <span id="8f2355abeb4ei" /> <span id="f2de2fadd9z3o" />

### Recommended models <span id="sm_rec_h3" />

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}>
  <colgroup>
    <col style={{ width: "45.55%" }} />

    <col style={{ width: "12.35%" }} />

    <col style={{ width: "15.38%" }} />

    <col style={{ width: "13.36%" }} />

    <col style={{ width: "13.36%" }} />
  </colgroup>

  <thead>
    <tr>
      <th rowSpan={2}>
        <strong>Model</strong>
      </th>

      <th rowSpan={2}>
        <strong>Version</strong>
      </th>

      <th>
        <strong>Context window</strong>
      </th>

      <th>
        <strong>Max input</strong>
      </th>

      <th>
        <strong>Max output</strong>
      </th>
    </tr>

    <tr>
      <th colSpan={3}>
        <strong>(tokens)</strong>
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        <strong><a href="/en/model-studio/qwen3-8-livetranslate-flash-realtime">qwen3.8-livetranslate-flash-realtime</a></strong>
      </td>

      <td>
        Stable
      </td>

      <td>
        53248
      </td>

      <td>
        49152
      </td>

      <td>
        4096
      </td>
    </tr>

    <tr>
      <td>
        <strong>qwen3.5-livetranslate-flash-realtime</strong>

        > Alias for qwen3.5-livetranslate-flash-realtime-2026-05-19
      </td>

      <td>
        Stable
      </td>

      <td rowSpan={2}>
        53248
      </td>

      <td rowSpan={2}>
        49152
      </td>

      <td rowSpan={2}>
        4096
      </td>
    </tr>

    <tr>
      <td>
        qwen3.5-livetranslate-flash-realtime-2026-05-19
      </td>

      <td>
        Snapshot
      </td>
    </tr>
  </tbody>
</table>

### Legacy models <span id="sm_lgcy_h3" />

> The following model is still available but is no longer the recommended choice. For new use cases, use the newer model above for better translation quality and cost-efficiency.

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}>
  <colgroup>
    <col style={{ width: "45.55%" }} />

    <col style={{ width: "12.35%" }} />

    <col style={{ width: "15.38%" }} />

    <col style={{ width: "13.36%" }} />

    <col style={{ width: "13.36%" }} />
  </colgroup>

  <thead>
    <tr>
      <th rowSpan={2}>
        <strong>Model</strong>
      </th>

      <th rowSpan={2}>
        <strong>Version</strong>
      </th>

      <th>
        <strong>Context window</strong>
      </th>

      <th>
        <strong>Max input</strong>
      </th>

      <th>
        <strong>Max output</strong>
      </th>
    </tr>

    <tr>
      <th colSpan={3}>
        <strong>(tokens)</strong>
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        <strong>qwen3-livetranslate-flash-realtime</strong>

        > Alias for qwen3-livetranslate-flash-realtime-2025-09-22
      </td>

      <td>
        Stable
      </td>

      <td rowSpan={2}>
        53248
      </td>

      <td rowSpan={2}>
        49152
      </td>

      <td rowSpan={2}>
        4096
      </td>
    </tr>

    <tr>
      <td>
        qwen3-livetranslate-flash-realtime-2025-09-22
      </td>

      <td>
        Snapshot
      </td>
    </tr>
  </tbody>
</table>

## Getting started <span id="a36e6dc44fucp" /> <span id="863c459c04n3t" />

<Tabs>
  <Tab title="qwen3.8-livetranslate-flash-realtime">
    ### Translate local audio

    Install `websocket-client` with `pip install websocket-client` and set these environment variables:

    - `DASHSCOPE_API_KEY`: an API key for the region and workspace.
    - `DASHSCOPE_WORKSPACE_ID`: your workspace ID.
    - `INPUT_PCM_FILE`: the local path to the audio file. This example uses mono, 16-bit, 16000 Hz PCM audio without a file header.

    The example translates audio into English, prints the translation, and saves mono, 16-bit, 24000 Hz PCM audio to `translation.pcm`. It configures the session, sends audio, requests session completion with `session.finish`, and closes the connection after receiving `session.finished`.

    ```python
    import base64
    import json
    import os
    import threading
    import time

    import websocket

    model = "qwen3.8-livetranslate-flash-realtime"
    workspace_id = os.environ["DASHSCOPE_WORKSPACE_ID"]
    api_key = os.environ["DASHSCOPE_API_KEY"]
    audio_file = os.environ["INPUT_PCM_FILE"]
    url = (
        f"wss://{workspace_id}.ap-southeast-1.maas.aliyuncs.com"
        f"/api-ws/v1/realtime?model={model}"
    )
    ws = websocket.create_connection(
        url, header={"Authorization": f"Bearer {api_key}"}, timeout=30
    )

    def receive():
        event = json.loads(ws.recv())
        if event.get("type") == "error" or event.get("code"):
            raise RuntimeError(event)
        return event

    send_errors = []

    def send_audio():
        try:
            with open(audio_file, "rb") as source:
                while chunk := source.read(3200):
                    ws.send(json.dumps({
                        "type": "input_audio_buffer.append",
                        "audio": base64.b64encode(chunk).decode("ascii"),
                    }))
                    time.sleep(0.1)
            ws.send(json.dumps({"type": "session.finish"}))
        except Exception as error:
            send_errors.append(error)

    try:
        receive()
        ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "output_modalities": ["text", "audio"],
                "translation": {"language": "en"},
            },
        }))
        while receive().get("type") != "session.updated":
            pass
        sender = threading.Thread(target=send_audio, daemon=True)
        sender.start()
        with open("translation.pcm", "wb") as output:
            while True:
                event = receive()
                event_type = event.get("type")
                if event_type in ("response.text.delta", "response.audio_transcript.delta"):
                    print(event["delta"], end="", flush=True)
                elif event_type == "response.audio.delta":
                    output.write(base64.b64decode(event["delta"]))
                elif event_type == "session.finished":
                    break
        sender.join(timeout=5)
        if send_errors:
            raise send_errors[0]
        print()
    finally:
        ws.close()
    ```
  </Tab>

  <Tab title="qwen3.5-livetranslate-flash-realtime">
    1. <strong>Prepare the environment</strong>

       Requires Python 3.10 or later.

       First, install pyaudio.

       <CodeGroup dropdown>
         ```bash macOS
         brew install portaudio && pip install pyaudio
         ```

         ```bash Debian/Ubuntu
         sudo apt-get install python3-pyaudio

         or

         pip install pyaudio
         ```

         ```bash CentOS
         sudo yum install -y portaudio portaudio-devel && pip install pyaudio
         ```

         ```powershell Windows
         pip install pyaudio
         ```
       </CodeGroup>

       Then install the WebSocket dependencies:

       ```bash
       pip install websocket-client==1.8.0 websockets
       ```

    2. <strong>Create the client</strong>

       Create a file named `livetranslate_client.py` with the following code:

       <Accordion title="Client code - livetranslate_client.py">
         ```python expandable
         import os
         import time
         import base64
         import asyncio
         import json
         import websockets
         import pyaudio
         import queue
         import threading
         import traceback

         class LiveTranslateClient:
             def __init__(self, api_key: str, target_language: str = "en", *, audio_enabled: bool = True):
                 if not api_key:
                     raise ValueError("API key cannot be empty.")

                 self.api_key = api_key
                 self.target_language = target_language
                 self.audio_enabled = audio_enabled
                 self.ws = None
                 self.api_url = "wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.5-livetranslate-flash-realtime"

                 # Audio input configuration (from microphone)
                 self.input_rate = 16000
                 self.input_chunk = 1600
                 self.input_format = pyaudio.paInt16
                 self.input_channels = 1

                 # Audio output configuration (for playback)
                 self.output_rate = 24000
                 self.output_chunk = 2400
                 self.output_format = pyaudio.paInt16
                 self.output_channels = 1

                 # State management
                 self.is_connected = False
                 self.audio_player_thread = None
                 self.audio_playback_queue = queue.Queue()
                 self.pyaudio_instance = pyaudio.PyAudio()
                 self.session_finished_event = asyncio.Event()

             async def connect(self):
                 """Establish a WebSocket connection to the translation service."""
                 headers = {"Authorization": f"Bearer {self.api_key}"}
                 try:
                     self.ws = await websockets.connect(self.api_url, additional_headers=headers)
                     self.is_connected = True
                     print(f"Successfully connected to the server: {self.api_url}")
                     await self.configure_session()
                 except Exception as e:
                     print(f"Connection failed: {e}")
                     self.is_connected = False
                     raise

             async def configure_session(self):
                 """Configure the translation session, setting the target language, voice, etc."""
                 config = {
                     "event_id": f"event_{int(time.time() * 1000)}",
                     "type": "session.update",
                     "session": {
                         # 'modalities' controls the output type.
                         # ["text", "audio"]: Returns both translated text and synthesized audio (recommended).
                         # ["text"]: Returns only the translated text.
                         "modalities": ["text", "audio"] if self.audio_enabled else ["text"],
                         "input_audio_format": "pcm",
                         "output_audio_format": "pcm",
                         # 'input_audio_transcription' configures source language recognition.
                         # Set 'model' to 'qwen3-asr-flash-realtime' to also output the source language recognition result.
                         # "input_audio_transcription": {
                         #     "model": "qwen3-asr-flash-realtime",
                         #     "language": "zh"  # source language, default 'en'
                         # },
                         "translation": {
                             "language": self.target_language,
                             # 'corpus' configures hotwords to improve the translation accuracy of specific terms.
                             # "corpus": {
                             #     "phrases": {
                             #         "Artificial Intelligence": "Artificial Intelligence",
                             #         "Machine Learning": "Machine Learning"
                             #     }
                             # }
                         }
                     }
                 }
                 print(f"Sending session configuration: {json.dumps(config, indent=2, ensure_ascii=False)}")
                 await self.ws.send(json.dumps(config))

             async def send_audio_chunk(self, audio_data: bytes):
                 """Encode and send an audio chunk to the server."""
                 if not self.is_connected:
                     return

                 event = {
                     "event_id": f"event_{int(time.time() * 1000)}",
                     "type": "input_audio_buffer.append",
                     "audio": base64.b64encode(audio_data).decode()
                 }
                 await self.ws.send(json.dumps(event))

             async def send_image_frame(self, image_bytes: bytes, *, event_id: str | None = None):
                 # Send an image frame to the server.
                 if not self.is_connected:
                     return

                 if not image_bytes:
                     raise ValueError("image_bytes cannot be empty.")

                 # Encode to Base64
                 image_b64 = base64.b64encode(image_bytes).decode()

                 event = {
                     "event_id": event_id or f"event_{int(time.time() * 1000)}",
                     "type": "input_image_buffer.append",
                     "image": image_b64,
                 }

                 await self.ws.send(json.dumps(event))

             def _audio_player_task(self):
                 stream = self.pyaudio_instance.open(
                     format=self.output_format,
                     channels=self.output_channels,
                     rate=self.output_rate,
                     output=True,
                     frames_per_buffer=self.output_chunk,
                 )
                 try:
                     while self.is_connected or not self.audio_playback_queue.empty():
                         try:
                             audio_chunk = self.audio_playback_queue.get(timeout=0.1)
                             if audio_chunk is None: # Termination signal
                                 break
                             stream.write(audio_chunk)
                             self.audio_playback_queue.task_done()
                         except queue.Empty:
                             continue
                 finally:
                     stream.stop_stream()
                     stream.close()

             def start_audio_player(self):
                 """Start the audio player thread (only when audio output is enabled)."""
                 if not self.audio_enabled:
                     return
                 if self.audio_player_thread is None or not self.audio_player_thread.is_alive():
                     self.audio_player_thread = threading.Thread(target=self._audio_player_task, daemon=True)
                     self.audio_player_thread.start()

             async def handle_server_messages(self, on_text_received):
                 """Handle incoming messages from the server in a loop."""
                 try:
                     async for message in self.ws:
                         event = json.loads(message)
                         event_type = event.get("type")
                         if event_type == "response.audio.delta" and self.audio_enabled:
                             audio_b64 = event.get("delta", "")
                             if audio_b64:
                                 audio_data = base64.b64decode(audio_b64)
                                 self.audio_playback_queue.put(audio_data)

                         elif event_type == "response.done":
                             print("\n[INFO] Response round complete.")
                             usage = event.get("response", {}).get("usage", {})
                             if usage:
                                 print(f"[INFO] token usage: {json.dumps(usage, indent=2, ensure_ascii=False)}")
                         elif event_type == "session.finished":
                             print("[INFO] Session finished.")
                             self.session_finished_event.set()
                         # Process source language recognition results (requires enabling input_audio_transcription.model)
                         # elif event_type == "conversation.item.input_audio_transcription.text":
                         #     stash = event.get("stash", "")  # Pending recognition text
                         #     print(f"[Recognizing] {stash}")
                         # elif event_type == "conversation.item.input_audio_transcription.completed":
                         #     transcript = event.get("transcript", "")  # Complete recognition result
                         #     print(f"[Source language] {transcript}")
                         elif event_type == "response.text.text":
                             # Streaming translated text in text-only modality
                             text = event.get("text", "")
                             stash = event.get("stash", "")
                             print(f"\r[Translating] {text}{stash}", end="", flush=True)
                         elif event_type == "response.audio_transcript.done":
                             print("\n[INFO] Translation complete.")
                             text = event.get("transcript", "")
                             if text:
                                 print(f"[INFO] Translated text: {text}")
                         elif event_type == "response.text.done":
                             print("\n[INFO] Translation complete.")
                             text = event.get("text", "")
                             if text:
                                 print(f"[INFO] Translated text: {text}")

                 except websockets.exceptions.ConnectionClosed as e:
                     print(f"[WARNING] Connection closed: {e}")
                     self.is_connected = False
                 except Exception as e:
                     print(f"[ERROR] An unexpected error occurred while processing messages: {e}")
                     traceback.print_exc()
                     self.is_connected = False

             async def start_microphone_streaming(self):
                 """Capture audio from the microphone and stream it to the server."""
                 stream = self.pyaudio_instance.open(
                     format=self.input_format,
                     channels=self.input_channels,
                     rate=self.input_rate,
                     input=True,
                     frames_per_buffer=self.input_chunk
                 )
                 print("Microphone is on. Start speaking...")
                 try:
                     while self.is_connected:
                         audio_chunk = await asyncio.get_event_loop().run_in_executor(
                             None, stream.read, self.input_chunk
                         )
                         await self.send_audio_chunk(audio_chunk)
                 finally:
                     stream.stop_stream()
                     stream.close()

             async def close(self):
                 """Gracefully close the connection and release resources."""
                 # Send session.finish to ensure the server completes translation of the final speech segment
                 if self.is_connected and self.ws:
                     finish_event = {
                         "event_id": f"event_{int(time.time() * 1000)}",
                         "type": "session.finish",
                     }
                     await self.ws.send(json.dumps(finish_event))
                     print("Sent session.finish, waiting for server to finish processing...")
                     try:
                         await asyncio.wait_for(self.session_finished_event.wait(), timeout=15)
                         print("Server processing complete.")
                     except asyncio.TimeoutError:
                         print("Timed out waiting for session.finished.")
                 self.is_connected = False
                 if self.ws:
                     await self.ws.close()
                     print("WebSocket connection closed.")

                 if self.audio_player_thread:
                     self.audio_playback_queue.put(None) # Send termination signal
                     self.audio_player_thread.join(timeout=1)
                     print("Audio player thread stopped.")

                 self.pyaudio_instance.terminate()
                 print("PyAudio instance released.")
         ```
       </Accordion>

    3. <strong>Interact with the model</strong>

       In the same directory, create a file named `main.py` with the following code:

       <Accordion title="main.py">
         ```python expandable
         import os
         import asyncio
         from livetranslate_client import LiveTranslateClient

         def print_banner():
             print("=" * 60)
             print("  Powered by Qwen qwen3.5-livetranslate-flash-realtime")
             print("=" * 60 + "\n")

         def get_user_config():
             """Get user configuration."""
             print("Select a mode:")
             print("1. Voice + Text [Default] | 2. Text Only")
             mode_choice = input("Enter your choice (press Enter for Voice + Text): ").strip()
             audio_enabled = (mode_choice != "2")

             if audio_enabled:
                 lang_map = {
                     "1": "en", "2": "zh", "3": "ru", "4": "fr", "5": "de", "6": "pt",
                     "7": "es", "8": "it", "9": "ko", "10": "ja"
                 }
                 print("Select the target language (Voice + Text mode):")
                 print("1. English | 2. Chinese | 3. Russian | 4. French | 5. German | 6. Portuguese | 7. Spanish | 8. Italian | 9. Korean | 10. Japanese")
             else:
                 lang_map = {
                     "1": "en", "2": "zh", "3": "ru", "4": "fr", "5": "de", "6": "pt", "7": "es", "8": "it",
                     "9": "id", "10": "ko", "11": "ja", "12": "vi", "13": "th", "14": "ar",
                     "15": "yue", "16": "hi", "17": "el", "18": "tr"
                 }
                 print("Select the target language (Text Only mode):")
                 print("1. English | 2. Chinese | 3. Russian | 4. French | 5. German | 6. Portuguese | 7. Spanish | 8. Italian | 9. Indonesian | 10. Korean | 11. Japanese | 12. Vietnamese | 13. Thai | 14. Arabic | 15. Cantonese | 16. Hindi | 17. Greek | 18. Turkish")

             choice = input("Enter your choice (defaults to the first option): ").strip()
             target_language = lang_map.get(choice, next(iter(lang_map.values())))

             return target_language, audio_enabled

         async def main():
             """Main program entry point."""
             print_banner()

             api_key = os.environ.get("DASHSCOPE_API_KEY")
             if not api_key:
                 print("[ERROR] Please set the DASHSCOPE_API_KEY environment variable.")
                 print("  For example: export DASHSCOPE_API_KEY='your_api_key_here'")
                 return

             target_language, audio_enabled = get_user_config()
             print("\nConfiguration complete:")
             print(f"  - Target language: {target_language}")
             if not audio_enabled:
                 print("  - Output mode: Text Only")

             client = LiveTranslateClient(api_key=api_key, target_language=target_language, audio_enabled=audio_enabled)

             # Define the callback function.
             def on_translation_text(text):
                 print(text, end="", flush=True)

             try:
                 print("Connecting to the translation service...")
                 await client.connect()

                 # Start audio playback based on the mode.
                 client.start_audio_player()

                 print("\n" + "-" * 60)
                 print("Connection successful! Speak into the microphone.")
                 print("The program will translate your speech in real time and play the translated audio. Press Ctrl+C to exit.")
                 print("-" * 60 + "\n")

                 # Run message handling and microphone recording concurrently.
                 message_handler = asyncio.create_task(client.handle_server_messages(on_translation_text))
                 tasks = [message_handler]
                 # Capture audio from the microphone for translation, regardless of whether audio output is enabled.
                 microphone_streamer = asyncio.create_task(client.start_microphone_streaming())
                 tasks.append(microphone_streamer)

                 await asyncio.gather(*tasks)

             except KeyboardInterrupt:
                 print("\n\nUser interrupted. Exiting...")
             except Exception as e:
                 print(f"\nA critical error occurred: {e}")
             finally:
                 print("\nCleaning up resources...")
                 await client.close()
                 print("Program exited.")

         if __name__ == "__main__":
             asyncio.run(main())
         ```
       </Accordion>

       Run `main.py` and speak into your microphone. The model translates your speech and outputs audio and text in real time.
  </Tab>
</Tabs>

## Voice cloning <span id="kvwdr9d8k2gh4" /> <span id="0uiaftzm7a2zi" />

The model clones the speaker's voice from input audio and uses it for translated output. Use a pre-cloned voice profile or let the server clone in real time. Useful for conference interpreting, live streaming, and video dubbing.

Set the following parameters in [session.update](/en/model-studio/live-translator-client-events#af43722339yva) to enable voice cloning:

- `session.enable_voice_clone`: Set to `true` to enable voice cloning.
- `session.voice_clone_options.frequency`: Controls when voice cloning occurs. Accepted values:

  - `never`: Does not clone on the server. Uses a pre-cloned voice profile instead. Set `session.voice` to your custom cloned voice ID.
  - `once`: Clones the voice from the input audio once at session start, then reuses it for all subsequent output. Best for single-speaker scenarios. Set `session.voice` to `default`.
  - `always`: Clones the voice before each response, dynamically adapting to speaker changes. Best for multi-speaker conversations. Set `session.voice` to `default`.
- `session.voice`: Specifies the output voice. The value depends on the `frequency` setting:

  - Set to `default`: Use with `frequency` set to `once` or `always`. The server clones the speaker's voice from the input audio. A default voice is used until cloning completes.
  - Set to a custom cloned voice ID (for example, `qwen-translate-vc-xxx-yyy-zzz`): Use with `frequency` set to `never`. You must prepare the voice in advance using the [Voice Cloning API](/en/model-studio/qwen3-5-livetranslate-flash-realtime) with `targetModel` set to the translation model you use.

> When `frequency` is set to `once` or `always` , the `voice` parameter must be set to `default` . Any other value causes the server to return an error.

### Voice cloning configuration examples <span id="vc003examplehd" /> <span id="vc003examplesec" />

<strong>Pre-cloned voice profile</strong> (consistent quality; recommended when a stable voice identity is required):

```json
{
    "type": "session.update",
    "session": {
        "modalities": ["text","audio"],
        "voice": "qwen-translate-vc-xxx-yyy-zzz",
        "translation": {
            "language": "en"
        },
        "enable_voice_clone": true,
        "voice_clone_options": {
            "frequency": "never"
        }
    }
}
```

<strong>Server-side cloning, once per session</strong> (best for single-speaker scenarios):

```json
{
    "type": "session.update",
    "session": {
        "modalities": ["text","audio"],
        "voice": "default",
        "translation": {
            "language": "en"
        },
        "enable_voice_clone": true,
        "voice_clone_options": {
            "frequency": "once"
        }
    }
}
```

<strong>Server-side cloning, every response</strong> (best for multi-speaker conversations):

```json
{
    "type": "session.update",
    "session": {
        "modalities": ["text","audio"],
        "voice": "default",
        "translation": {
            "language": "en"
        },
        "enable_voice_clone": true,
        "voice_clone_options": {
            "frequency": "always"
        }
    }
}
```

## Improve translation with images <span id="de2138642c40d" /> <span id="8ff6c64d6b99i" />

Image input helps disambiguate homonyms and recognize uncommon proper nouns during translation. Send no more than 2 images per second.

Download the following sample images: [medical mask.png](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/en-US/20250923/tjpeys/%E5%8F%A3%E7%BD%A9.png), [masquerade mask.png](https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/en-US/20250923/ifqttq/%E9%9D%A2%E5%85%B7.png)

Download the code below to the same directory as `livetranslate_client.py` and run it. Say `"What is mask?"` into your microphone. The model uses the image to disambiguate: `medical mask.png` yields "What is a medical mask?" and `masquerade mask.png` yields "What is a masquerade mask?".

```python expandable
import os
import time
import json
import asyncio
import contextlib
import functools

from livetranslate_client import LiveTranslateClient

IMAGE_PATH = "medical mask.png"
# IMAGE_PATH = "masquerade mask.png"

def print_banner():
    print("=" * 60)
    print("  Powered by Qwen qwen3.5-livetranslate-flash-realtime — single-turn interaction example (mask)")
    print("=" * 60 + "\n")

async def stream_microphone_once(client: LiveTranslateClient, image_bytes: bytes):
    pa = client.pyaudio_instance
    stream = pa.open(
        format=client.input_format,
        channels=client.input_channels,
        rate=client.input_rate,
        input=True,
        frames_per_buffer=client.input_chunk,
    )
    print(f"[INFO] Recording started. Please speak...")
    loop = asyncio.get_event_loop()
    last_img_time = 0.0
    frame_interval = 0.5  # 2 fps
    try:
        while client.is_connected:
            data = await loop.run_in_executor(None, stream.read, client.input_chunk)
            await client.send_audio_chunk(data)

            # Append an image frame every 0.5 seconds
            now = time.time()
            if now - last_img_time >= frame_interval:
                await client.send_image_frame(image_bytes)
                last_img_time = now
    finally:
        stream.stop_stream()
        stream.close()

async def main():
    print_banner()
    api_key = os.environ.get("DASHSCOPE_API_KEY")
    if not api_key:
        print("[ERROR] Please set the DASHSCOPE_API_KEY environment variable.")
        return

    client = LiveTranslateClient(api_key=api_key, target_language="zh", audio_enabled=True)

    def on_text(text: str):
        print(text, end="", flush=True)

    try:
        await client.connect()
        client.start_audio_player()
        message_task = asyncio.create_task(client.handle_server_messages(on_text))
        with open(IMAGE_PATH, "rb") as f:
            img_bytes = f.read()
        await stream_microphone_once(client, img_bytes)
        await asyncio.sleep(15)
    finally:
        await client.close()
        if not message_task.done():
            message_task.cancel()
            with contextlib.suppress(asyncio.CancelledError):
                await message_task

if __name__ == "__main__":
    asyncio.run(main())
```

## One-click Function Compute deployment <span id="7727c7c1ed6du" /> <span id="78113540b3h6i" />

To deploy the application:

1. Open the [Function Compute template](https://fc.console.alibabacloud.com/applications/create?template=qwen-livetranslate-flash-realtime-intl@dev), enter your API key, and click <strong>Create and Deploy Default Environment</strong> to test the application.
2. Wait for about a minute. In <strong>Environment Details > Environment Context</strong>, retrieve the endpoint, change the protocol from <strong>http</strong> to <strong>https</strong> (for example, [https://qwen-livetranslate-flash-realtime-intl.fcv3.xxx.ap-southeast-1.fc.devsapp.net/](https://qwen-livetranslate-flash-realtime-intl.fcv3.xxx.ap-southeast-1.fc.devsapp.net/)), and open the URL in a browser to interact with the model.

   <Tip>
     This endpoint uses a self-signed certificate and is for temporary testing only. Your browser will display a security warning on your first visit. This is expected behavior. <strong>Do not use this endpoint in a production environment</strong>. To proceed, follow the on-screen instructions (for example, click Advanced → Proceed to (unsafe)).
   </Tip>

> If you are prompted to configure Resource Access Management permissions, follow the on-screen instructions.

> To view the project source code, go to <strong>Resource Information</strong> > <strong>Function Resources</strong> .

> Both [Function Compute](https://www.alibabacloud.com/help/en/functioncompute/trial-quota-1) and [Model Studio](/en/model-studio/new-free-quota) provide a free quota for new users, sufficient for basic debugging. After the free quota is used up, pay-as-you-go billing applies.

## Interaction flow <span id="b93454ed9e636" /> <span id="1e5b3ecf042cx" />

<Tabs>
  <Tab title="qwen3.8-livetranslate-flash-realtime">
    The server detects turns and generates responses by default. The following table summarizes the main interaction stages. See [Server events](/en/model-studio/live-translator-server-events) for event fields.

    <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "20%" }} /><col style={{ width: "35%" }} /><col style={{ width: "45%" }} /></colgroup><thead><tr><th><p><strong>Stage</strong></p></th><th><p><strong>Client action</strong></p></th><th><p><strong>Server events</strong></p></th></tr></thead><tbody><tr><td><p>Create and configure the session</p></td><td><p>Connect and send <code>session.update</code></p></td><td><p><code>session.created</code>, <code>session.updated</code></p></td></tr><tr><td><p>Send audio</p></td><td><p><code>input\_audio\_buffer.append</code></p></td><td><p>Source transcription is streamed through <code>conversation.item.input\_audio\_transcription.delta</code>, followed by <code>conversation.item.input\_audio\_transcription.completed</code>.</p></td></tr><tr><td><p>Receive translation and audio</p></td><td><p>Keep receiving server events</p></td><td><p>Translation is returned through <code>response.text.delta</code> (text only) or <code>response.audio\_transcript.delta</code> (text and audio). Audio is returned through <code>response.audio.delta</code>. <code>response.done</code> marks the end of a response.</p></td></tr><tr><td><p>End the session</p></td><td><p><code>session.finish</code></p></td><td><p>Close the connection after receiving <code>session.finished</code>.</p></td></tr></tbody></table>
  </Tab>

  <Tab title="qwen3.5-livetranslate-flash-realtime">
    Translation uses an event-driven WebSocket model. How speech boundaries are determined depends on VAD mode or Manual mode (see [3. Input audio and images](/en/model-studio/qwen3-5-livetranslate-flash-realtime#82b7d6329836b)). The table below is based on VAD mode (default) and annotates the different server events in Manual mode.

    <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}>
      <colgroup>
        <col style={{ width: "17.66%" }} />

        <col style={{ width: "33.8%" }} />

        <col style={{ width: "48.54%" }} />
      </colgroup>

      <thead>
        <tr>
          <th>
            <strong>Lifecycle</strong>
          </th>

          <th>
            <strong>Client event</strong>
          </th>

          <th>
            <strong>Server event</strong>
          </th>
        </tr>
      </thead>

      <tbody>
        <tr>
          <td>
            Session initialization
          </td>

          <td>
            session.update

            > Session configuration
          </td>

          <td>
            session.created

            > Session created

            session.updated

            > Session configuration updated
          </td>
        </tr>

        <tr>
          <td>
            User audio input
          </td>

          <td>
            input\_audio\_buffer.append

            > Append audio to the buffer

            input\_image\_buffer.append

            > Append image to the buffer

            input\_audio\_buffer.commit

            > (Manual mode only) Commit the audio buffer
          </td>

          <td>
            <strong>VAD mode</strong>:

            input\_audio\_buffer.speech\_started

            > Speech start detected

            input\_audio\_buffer.speech\_stopped

            > Speech end detected; server automatically commits the audio buffer

            <strong>Manual mode</strong>:

            input\_audio\_buffer.committed

            > Returned after the client sends input\_audio\_buffer.commit, confirming the audio buffer has been committed
          </td>
        </tr>

        <tr>
          <td>
            Server audio output
          </td>

          <td>
            None
          </td>

          <td>
            response.created

            > Signals that the server starts generating a response.

            response.output\_item.added

            > Signals that a new output item is available.

            conversation.item.created

            > A new message item is created in the conversation.

            response.content\_part.added

            > Signals that a new content part has been added to the assistant message.

            response.text.text

            > Incremental translated text in text-only modality

            response.audio\_transcript.text

            > Incremental translated text in audio+text modality

            response.audio.delta

            > Contains an incremental chunk of the synthesized audio.

            response.text.done

            > Translation text complete in text-only modality

            response.audio\_transcript.done

            > Translation text complete in audio+text modality

            response.audio.done

            > Signals that the synthesized audio is complete.

            response.content\_part.done

            > Signals that a text or audio content part for the assistant message is complete.

            response.output\_item.done

            > Signals that the entire output item for the assistant message is complete.

            response.done

            > Signals that the entire response is complete.
          </td>
        </tr>

        <tr>
          <td>
            Session termination
          </td>

          <td>
            session.finish

            > Notifies the server that audio input is complete
          </td>

          <td>
            session.finished

            > Server processing complete; session ended
          </td>
        </tr>
      </tbody>
    </table>
  </Tab>
</Tabs>

After sending all audio, send a `session.finish` event and wait for `session.finished` before closing the WebSocket. If you close the connection without sending `session.finish`, the server cannot know that audio input has ended, and the recognition and translation results for the last speech segment will be lost.

## API <span id="46fcc43673t0z" /> <span id="d6f3ba031di77" />

For the connection workflow and examples, see [AOQ integration](/en/model-studio/realtime-aoq-access). For supported models and versions, see [Model and protocol support](/en/model-studio/realtime-api-overview#rtov-s02h2).

- [Qwen-Livetranslate-Realtime](/en/model-studio/live-translator-api).
- [Realtime API overview](/en/model-studio/realtime-api-overview) (WebRTC protocol description)

## Billing <span id="e02a82e668y2c" /> <span id="6a95f2fc38za0" />

<strong>Qwen3.8-LiveTranslate-Flash-Realtime and Qwen3.5-LiveTranslate-Flash-Realtime</strong>

- <strong>Audio</strong>: 7 tokens per second of input audio; 12.5 tokens per second of output audio.

- <strong>Image</strong>: Every 32×32 pixels consumes 0.5 tokens.
  <strong>Qwen3-LiveTranslate-Flash-Realtime</strong>

- <strong>Audio</strong>: Each second of audio input or output consumes 12.5 tokens.

- <strong>Image</strong>: Every 28×28 pixels consumes 0.5 tokens.

- <strong>Text</strong>: When source language speech recognition is enabled, the service returns a transcript of the input audio in addition to the translation. This transcript is billed as output text tokens.

For token prices for each model, see [Model pricing and billing](/en/model-studio/model-pricing).

## Rate limits <span id="lt-rate-limits-title" /> <span id="lt-rate-limits-section" />

For information about model rate limits, see [Rate limiting](/en/model-studio/rate-limit).

## Supported languages <span id="4ffd192226f0s" /> <span id="24e640a87fsrw" />

Use the following language codes to specify the source and target languages.

> Some target languages only support text. The legacy model qwen3-livetranslate-flash-realtime supports only the following 18 languages: en, zh, ru, fr, de, pt, es, it, id, ko, ja, vi, th, ar, yue, hi, el, tr.

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "33.33%" }} /><col style={{ width: "33.33%" }} /><col style={{ width: "33.34%" }} /></colgroup><thead><tr><th><p><strong>Language code</strong></p></th><th><p><strong>Language</strong></p></th><th><p><strong>Output</strong></p></th></tr></thead><tbody><tr><td><p>zh</p></td><td><p>Chinese</p></td><td><p>Audio + text</p></td></tr><tr><td><p>en</p></td><td><p>English</p></td><td><p>Audio + text</p></td></tr><tr><td><p>ar</p></td><td><p>Arabic</p></td><td><p>Audio + text</p></td></tr><tr><td><p>de</p></td><td><p>German</p></td><td><p>Audio + text</p></td></tr><tr><td><p>fr</p></td><td><p>French</p></td><td><p>Audio + text</p></td></tr><tr><td><p>es</p></td><td><p>Spanish</p></td><td><p>Audio + text</p></td></tr><tr><td><p>pt</p></td><td><p>Portuguese</p></td><td><p>Audio + text</p></td></tr><tr><td><p>id</p></td><td><p>Indonesian</p></td><td><p>Audio + text</p></td></tr><tr><td><p>it</p></td><td><p>Italian</p></td><td><p>Audio + text</p></td></tr><tr><td><p>ko</p></td><td><p>Korean</p></td><td><p>Audio + text</p></td></tr><tr><td><p>ru</p></td><td><p>Russian</p></td><td><p>Audio + text</p></td></tr><tr><td><p>th</p></td><td><p>Thai</p></td><td><p>Audio + text</p></td></tr><tr><td><p>vi</p></td><td><p>Vietnamese</p></td><td><p>Audio + text</p></td></tr><tr><td><p>ja</p></td><td><p>Japanese</p></td><td><p>Audio + text</p></td></tr><tr><td><p>tr</p></td><td><p>Turkish</p></td><td><p>Audio + text</p></td></tr><tr><td><p>hi</p></td><td><p>Hindi</p></td><td><p>Audio + text</p></td></tr><tr><td><p>ms</p></td><td><p>Malay</p></td><td><p>Audio + text</p></td></tr><tr><td><p>nl</p></td><td><p>Dutch</p></td><td><p>Audio + text</p></td></tr><tr><td><p>ur</p></td><td><p>Urdu</p></td><td><p>Audio + text</p></td></tr><tr><td><p>nb</p></td><td><p>Norwegian Bokmål</p></td><td><p>Audio + text</p></td></tr><tr><td><p>sv</p></td><td><p>Swedish</p></td><td><p>Audio + text</p></td></tr><tr><td><p>da</p></td><td><p>Danish</p></td><td><p>Audio + text</p></td></tr><tr><td><p>he</p></td><td><p>Hebrew</p></td><td><p>Audio + text</p></td></tr><tr><td><p>fi</p></td><td><p>Finnish</p></td><td><p>Audio + text</p></td></tr><tr><td><p>pl</p></td><td><p>Polish</p></td><td><p>Audio + text</p></td></tr><tr><td><p>is</p></td><td><p>Icelandic</p></td><td><p>Audio + text</p></td></tr><tr><td><p>cs</p></td><td><p>Czech</p></td><td><p>Audio + text</p></td></tr><tr><td><p>fil</p></td><td><p>Filipino</p></td><td><p>Audio + text</p></td></tr><tr><td><p>fa</p></td><td><p>Persian</p></td><td><p>Audio + text</p></td></tr><tr><td><p>yue</p></td><td><p>Cantonese</p></td><td><p>Text</p></td></tr><tr><td><p>el</p></td><td><p>Greek</p></td><td><p>Text</p></td></tr><tr><td><p>af</p></td><td><p>Afrikaans</p></td><td><p>Text</p></td></tr><tr><td><p>ast</p></td><td><p>Asturian</p></td><td><p>Text</p></td></tr><tr><td><p>be</p></td><td><p>Belarusian</p></td><td><p>Text</p></td></tr><tr><td><p>bg</p></td><td><p>Bulgarian</p></td><td><p>Text</p></td></tr><tr><td><p>bn</p></td><td><p>Bengali</p></td><td><p>Text</p></td></tr><tr><td><p>bs</p></td><td><p>Bosnian</p></td><td><p>Text</p></td></tr><tr><td><p>ca</p></td><td><p>Catalan</p></td><td><p>Text</p></td></tr><tr><td><p>ceb</p></td><td><p>Cebuano</p></td><td><p>Text</p></td></tr><tr><td><p>et</p></td><td><p>Estonian</p></td><td><p>Text</p></td></tr><tr><td><p>gl</p></td><td><p>Galician</p></td><td><p>Text</p></td></tr><tr><td><p>gu</p></td><td><p>Gujarati</p></td><td><p>Text</p></td></tr><tr><td><p>hr</p></td><td><p>Croatian</p></td><td><p>Text</p></td></tr><tr><td><p>hu</p></td><td><p>Hungarian</p></td><td><p>Text</p></td></tr><tr><td><p>jv</p></td><td><p>Javanese</p></td><td><p>Text</p></td></tr><tr><td><p>kk</p></td><td><p>Kazakh</p></td><td><p>Text</p></td></tr><tr><td><p>kn</p></td><td><p>Kannada</p></td><td><p>Text</p></td></tr><tr><td><p>ky</p></td><td><p>Kyrgyz</p></td><td><p>Text</p></td></tr><tr><td><p>lv</p></td><td><p>Latvian</p></td><td><p>Text</p></td></tr><tr><td><p>mk</p></td><td><p>Macedonian</p></td><td><p>Text</p></td></tr><tr><td><p>ml</p></td><td><p>Malayalam</p></td><td><p>Text</p></td></tr><tr><td><p>mr</p></td><td><p>Marathi</p></td><td><p>Text</p></td></tr><tr><td><p>pa</p></td><td><p>Punjabi</p></td><td><p>Text</p></td></tr><tr><td><p>ro</p></td><td><p>Romanian</p></td><td><p>Text</p></td></tr><tr><td><p>sk</p></td><td><p>Slovak</p></td><td><p>Text</p></td></tr><tr><td><p>sl</p></td><td><p>Slovenian</p></td><td><p>Text</p></td></tr><tr><td><p>sw</p></td><td><p>Swahili</p></td><td><p>Text</p></td></tr><tr><td><p>tg</p></td><td><p>Tajik</p></td><td><p>Text</p></td></tr><tr><td><p>az</p></td><td><p>Azerbaijani</p></td><td><p>Text</p></td></tr><tr><td><p>uk</p></td><td><p>Ukrainian</p></td><td><p>Text</p></td></tr></tbody></table>

## Supported voices <span id="0a5bde7593gdk" /> <span id="17cae276c21f3" />

For supported voices and the corresponding `voice` parameter values, see [Voice list](/en/model-studio/omni-voice-list).
