Skip to main content
Model khusus

Pemahaman audio (Qwen3-Omni-Captioner)

Qwen3-Omni-Captioner adalah model open-source berbasis Qwen3-Omni yang menghasilkan deskripsi audio detail—meliputi ucapan, suara latar, musik, dan efek suara—tanpa memerlukan prompt. Model ini mampu mengidentifikasi emosi pembicara, elemen musikal (seperti gaya dan instrumen), serta informasi sensitif untuk keperluan analisis audio, audit keamanan, pengenalan maksud, dan pengeditan video.

Model yang didukung

  • Singapura
  • Tiongkok (Beijing)

Model

Context window

Max input

Max output

Input cost

Output cost

Free quota

(Catatan)

(tokens)

(per 1M tokens)

qwen3-omni-30b-a3b-captioner

65.536

32.768

32.768

$3,81

$3,06

1 juta token

Berlaku selama 90 hari setelah mengaktifkan Model Studio

Aturan konversi token untuk audio: Total token = Durasi audio (dalam detik) × 12,5. Jika durasi audio kurang dari satu detik, dihitung sebagai satu detik.

Mulai

Prasyarat Qwen3-Omni-Captioner hanya tersedia melalui API. Pengujian berbasis Konsol tidak didukung. Contoh kode berikut menganalisis audio online melalui URL, bukan file lokal. Pelajari cara mengirimkan file lokal dan batasan file audio.
  • Kompatibel dengan OpenAI
  • DashScope
  • Python
  • Node.js
  • curl
import os
from openai import OpenAI

client = OpenAI(
    # Kunci API berbeda berdasarkan wilayah. Untuk mendapatkan Kunci API, lihat https://www.alibabacloud.com/help/en/model-studio/get-api-key
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # URL berikut untuk wilayah Singapura. Ganti {WorkspaceId} dengan ID ruang kerja aktual Anda. Untuk Beijing, gunakan URL yang ditampilkan di tab Tiongkok (Beijing).
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

completion = client.chat.completions.create(
    model="qwen3-omni-30b-a3b-captioner",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "input_audio",
                    "input_audio": {
                        "data": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20240916/xvappi/%E8%A3%85%E4%BF%AE%E5%99%AA%E9%9F%B3.wav"
                    }
                }
            ]
        }
    ]
)
print(completion.choices[0].message.content)
The audio clip begins with a sudden, loud, metallic clanking that dominates the soundstage, immediately indicating an industrial or workshop environment. The clanking is rhythmic, consistent, and has a sharp, resonant quality, suggestive of metal tools striking metal surfaces-likely a hammer, wrench, or similar instrument being used on a hard metal object. The sound is harsh and slightly distorted, with audible clipping on each impact, likely due to the microphone’s proximity and the high volume of the sound.
As the initial clanking fades, a male voice enters, speaking in Mandarin Chinese with a tone of exasperation and complaint. His voice is clear, close-mic’d, and free from distortion. He says: “Oh my, how can I possibly work quietly like this?”. His intonation is conversational, informal, and marked by a rising, questioning inflection, typical of everyday speech rather than performance or formal address. The accent is standard Putonghua, with no strong regional markers, suggesting he is a native Mandarin speaker from the northern or central regions of China.
During the speaker’s utterance, the metallic clanking resumes, overlapping with his voice. The timing and nature of these sounds indicate the speaker is directly reacting to the ongoing noise-likely caused by another person in the same space. The environment is acoustically “dry” with minimal echo, implying a small or medium-sized room with sound-absorbing materials, further supporting the workshop or industrial setting. There are no other background noises, music, or ambient sounds, and no evidence of a public or commercial space.
The recording quality is moderate: the microphone captures both the low-end thuds and the sharp metallic transients, but the loud clanking causes digital clipping, resulting in a harsh, “crunchy” distortion during the impacts. The speaker’s voice, however, remains clear and intelligible. The overall impression is of a candid, real-world interaction-possibly a worker or office employee complaining about an interruption in a noisy environment.
In summary, the audio depicts a Mandarin-speaking man in a workshop or industrial setting, reacting with frustration to ongoing metallic clanking that disrupts his work. The recording is informal, clear, and grounded in a context of manual labor or technical work, with no evidence of scripted performance, music, or extraneous activity.

Cara kerja

  • Interaksi satu arah: Setiap permintaan bersifat independen. Percakapan multi-arah tidak didukung.
  • Tugas tetap: Hanya menghasilkan deskripsi audio dalam bahasa Inggris. Instruksi seperti pesan sistem tidak dapat mengubah perilaku, format keluaran, atau fokus konten.
  • Hanya input audio: Menerima hanya audio—tidak ada prompt teks. Format parameter message bersifat tetap.
    messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "input_audio",
                        "input_audio": {
                            "data": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20240916/xvappi/%E8%A3%85%E4%BF%AE%E5%99%AA%E9%9F%B3.wav"
                        }
                    }
                ]
            }
        ]
    

Keluaran streaming

Keluaran streaming mengembalikan hasil secara bertahap saat dihasilkan, sehingga mengurangi waktu tunggu.
  • Kompatibel dengan OpenAI
  • DashScope
Atur stream ke true untuk mengaktifkan keluaran streaming.
Python
import os
from openai import OpenAI

client = OpenAI(
    # Kunci API untuk wilayah Singapura dan Beijing berbeda. Untuk mendapatkan Kunci API, lihat https://www.alibabacloud.com/help/en/model-studio/get-api-key
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # URL berikut untuk wilayah Singapura. Ganti {WorkspaceId} dengan ID ruang kerja aktual Anda. Untuk Beijing, gunakan URL yang ditampilkan di tab Tiongkok (Beijing).
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

completion = client.chat.completions.create(
    model="qwen3-omni-30b-a3b-captioner",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "input_audio",
                    "input_audio": {
                        "data": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20240916/xvappi/%E8%A3%85%E4%BF%AE%E5%99%AA%E9%9F%B3.wav"
                    }
                }
            ]
        }
    ],
    stream=True,
    stream_options={"include_usage": True},

)
for chunk in completion:
    # Jika stream_options.include_usage bernilai True, bidang choices pada chunk terakhir adalah daftar kosong dan harus dilewati. Anda dapat memperoleh penggunaan token dari chunk.usage.
    if chunk.choices and chunk.choices[0].delta.content != "":
        print(chunk.choices[0].delta.content,end="")

Kirimkan file lokal (encoding Base64 atau jalur file)

Dua metode untuk mengunggah file lokal:
  • Gunakan encoding Base64
  • Jalur file langsung (Disarankan untuk stabilitas transmisi yang lebih baik)
Metode unggah:
  • Kirimkan melalui jalur file
  • Penerusan melalui pengodean Base64
Kirimkan jalur file secara langsung. Didukung hanya oleh SDK Python dan Java DashScope, bukan HTTP. Format jalur berbeda berdasarkan SDK dan OS.

Tentukan jalur file

Sistem

SDK

Jalur file masukan

Contoh

Linux atau macOS

Python SDK

file://{absolute_path_of_the_file}

file:///home/images/test.mp3

Java SDK

Sistem operasi Windows

Python SDK

file://{absolute_path_of_the_file}

file://D:/images/test.mp3

Java SDK

file:///{absolute_path_of_the_file}

file:///D:/images/test.mp3

Batasan:
  • Jalur file direkomendasikan untuk stabilitas transmisi. Base64 juga berfungsi untuk file di bawah 1 MB.
  • Saat mengirimkan melalui jalur file, file audio harus di bawah 10 MB.
  • Saat menggunakan Base64, string yang diencode harus di bawah 10 MB. Catatan: Base64 memperbesar ukuran file.
  • Melewatkan berdasarkan jalur file
  • Kirimkan melalui encoding Base64
Pengiriman melalui jalur file hanya didukung oleh SDK Python dan Java DashScope, bukan HTTP.
Python
import dashscope
import os

# URL berikut untuk wilayah Singapura. Ganti {WorkspaceId} dengan ID ruang kerja aktual Anda. Untuk Beijing, gunakan URL yang ditampilkan di tab Tiongkok (Beijing).
dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

# Ganti ABSOLUTE_PATH/welcome.mp3 dengan jalur mutlak file audio lokal Anda.
# Jalur lengkap file lokal harus diawali dengan file:// untuk memastikan jalur valid, contoh: file:///home/images/test.mp3
audio_file_path = "file://ABSOLUTE_PATH/welcome.mp3"
messages = [
    {
        "role": "user",
        # Kirimkan jalur file yang diawali dengan file:// dalam parameter audio.
        "content": [{"audio": audio_file_path}],
    }
]

response = dashscope.MultiModalConversation.call(
    # Kunci API berbeda berdasarkan wilayah. Untuk mendapatkan Kunci API, lihat https://www.alibabacloud.com/help/en/model-studio/get-api-key
    # Jika Anda belum mengonfigurasi variabel lingkungan, ganti baris berikut dengan Kunci API Model Studio Anda: api_key="sk-xxx"
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model="qwen3-omni-30b-a3b-captioner",
    messages=messages)

print("Output:")
print(response["output"]["choices"][0]["message"].content[0]["text"])

Referensi API

Untuk parameter Qwen3-Omni-Captioner, lihat Text Generation.

Kode error

Jika panggilan model gagal dan mengembalikan pesan error, lihat Kode error untuk penyelesaian.

FAQ

Bagaimana cara memampatkan file audio ke ukuran yang diperlukan?

# Perintah konversi dasar (templat universal)
# -i: Menentukan jalur file input. Contoh: input.mp3

# -b:a: Mengatur bitrate audio.
  # Nilai umum: 64 kbps (kualitas rendah, untuk suara dan streaming bandwidth rendah), 128k (kualitas sedang, untuk audio umum dan podcast), 192 kbps (kualitas tinggi, untuk musik dan penyiaran).
  # Bitrate yang lebih tinggi menghasilkan kualitas audio lebih baik dan ukuran file lebih besar.

# -ar: Mengatur laju sampel audio, yaitu jumlah sampel per detik.
 # Nilai umum: 8000 Hz, 22050 Hz, 44100 Hz (laju sampel standar).
 # Laju sampel yang lebih tinggi menghasilkan ukuran file lebih besar.

# -ac: Mengatur jumlah saluran audio. Nilai umum: 1 (mono), 2 (stereo). File mono lebih kecil.

# -y: Menimpa file output jika sudah ada (tidak perlu nilai). # output.mp3: Menentukan jalur file output.

ffmpeg -i input.mp3 -b:a 128k -ar 44100 -ac 1 output.mp3 -y

Batasan

Batasan file audio:
  • Durasi: Maksimal 40 menit.
  • Jumlah file: Hanya satu file audio yang didukung per permintaan.
  • Format file: AMR, WAV (CodecID: GSM_MS), WAV (PCM), 3GP, 3GPP, AAC, dan MP3.
  • Metode input file: URL publik, encoding Base64, atau jalur file lokal.
  • Ukuran file:
    • URL publik: Tidak lebih dari 1 GB.
    • Jalur file: File audio harus lebih kecil dari 10 MB.
    • Encoding Base64: String yang diencode harus di bawah 10 MB. Kirimkan file lokal.
    Untuk memampatkan file, lihat Bagaimana cara memampatkan file audio ke ukuran yang diperlukan?
Rencana Token (Team Edition)
Statistik dan Pemantauan
Asset Center
Dukungan layanan