Use the Paraformer Python SDK to transcribe audio and video files with the DashScope API.
Prerequisites
Getting started
The core class (Transcription) supports two transcription approaches:
- Async submission + sync wait: Submit a task and block until it completes and returns the result.
- Async submission + async polling: Submit a task and poll for results at any time.
Async submission + sync wait
-
Call the
async_callmethod of the core class (Transcription) and set the request parameters.- The file transcription service processes tasks submitted through the API on a best-effort basis. After submission, a task enters the queued (
PENDING) state. The queue time depends on the queue length and file duration and cannot be precisely estimated, but typically completes within a few minutes. Once processing begins, speech recognition completes at several hundred times the real-time speed. - After each task completes, the recognition result and URL download link are valid for 24 hours. After expiration, you cannot query the task or download results through the previously provided URL.
- The file transcription service processes tasks submitted through the API on a best-effort basis. After submission, a task enters the queued (
-
Call the
waitmethod of the core class (Transcription) to wait synchronously for the task to complete. Task statuses:PENDING,RUNNING,SUCCEEDED,FAILED. Thewaitcall blocks duringPENDINGorRUNNING. When the task reachesSUCCEEDEDorFAILED,waitreturns the result. Thewaitmethod returns a TranscriptionResponse.
Click to view complete example
Click to view complete example
Async submission + async polling
-
Call the
async_callmethod of the Transcription core class and set the request parameters.- The file transcription service processes tasks submitted through the API on a best-effort basis. After submission, a task enters the queued (
PENDING) state. The queue time depends on the queue length and file duration and cannot be precisely estimated, but typically completes within a few minutes. Once processing begins, speech recognition completes at several hundred times the real-time speed. - After each task completes, the recognition result and URL download link are valid for 24 hours. After expiration, you cannot query the task or download results through the previously provided URL.
- The file transcription service processes tasks submitted through the API on a best-effort basis. After submission, a task enters the queued (
-
Poll the
fetchmethod of the core class (Transcription) until the task completes. Stop polling when the status isSUCCEEDEDorFAILEDand process the result. Thefetchmethod returns a TranscriptionResponse.
Click to view complete example
Click to view complete example
Request parameters
Pass these parameters to the async_call method of the core class (Transcription).
| Parameter | Type | Default | Required | Description |
|---|---|---|---|---|
| model | str | Yes | The model name for Paraformer audio and video file transcription. Supported models. | |
| file_urls | list[str] | Yes | A list of URLs for audio and video file transcription. The HTTP and HTTPS protocols are supported. A single request supports only 1 URL.If audio files are stored in Alibaba Cloud OSS, the SDK does not support temporary URLs with the oss:// prefix. | |
| vocabulary_id | str | No | The custom vocabulary ID. Supported for v2 series models; requires language configuration. Disabled by default. Custom Vocabularies. | |
| channel_id | list[int] | [0] | No | Specifies the audio track indexes to recognize in a multi-track audio file. Indexes start from 0. For example, [0] means recognizing the first track, and [0, 1] means recognizing both the first and second tracks simultaneously. If this parameter is omitted, only the first track is processed by default. |
| disfluency_removal_enabled | bool | False | No | Filters filler words. Disabled by default. |
| timestamp_alignment_enabled | bool | False | No | Enables timestamp alignment. Disabled by default. |
| special_word_filter | str | No | Specifies sensitive words to process during speech recognition and supports setting different processing methods for different sensitive words.If this parameter is not provided, the system uses the built-in sensitive word filtering logic, and words matching the Alibaba Cloud Model Studio sensitive word list in the recognition results will be replaced with * of equal length.If this parameter is provided, the following sensitive word processing strategies can be implemented:
| |
| language_hints | list[str] | ["zh", "en"] | No | Specifies the language codes of the speech to be recognized.This parameter is only applicable to the paraformer-v2 model.Supported language codes:
|
| diarization_enabled | bool | False | No | Automatic speaker diarization. Disabled by default.Only applicable to mono audio. Multi-channel audio does not support speaker diarization.When this feature is enabled, the recognition results will include a speaker_id field to distinguish different speakers.If speaker diarization is enabled, it is recommended that the audio duration does not exceed 2 hours, otherwise recognition may fail or time out. speaker_id, see Description of recognition results. |
| speaker_count | int | No | Reference speaker count. Integer from 2 to 100.Takes effect only when diarization_enabled is true.Determined automatically by default. Setting this parameter guides the algorithm but does not guarantee the exact count. |
Response
TranscriptionResponse
A TranscriptionResponse contains task_id, task_status, and the execution result in the output property. See TranscriptionOutput.
Click to view TranscriptionResponse structure examples
Click to view TranscriptionResponse structure examples
TranscriptionResponse returned by async_call does not include submit_time or scheduled_time.submit_time and scheduled_time, use the wait() or fetch() methods instead of the async_call() return value directly. The TranscriptionResponse returned by wait() or fetch():Parameter | Description |
|---|---|
status_code | The HTTP request status code. |
code |
|
message |
|
task_id | The task ID. |
task_status | The task status. The four statuses are When a task contains multiple subtasks, if any subtask succeeds, the entire task status is marked as |
results | The recognition results of the subtasks. |
subtask_status | The subtask status. The four statuses are |
file_url | The URL of the audio file to be recognized. |
transcription_url | The URL corresponding to the audio recognition result. The recognition result is saved as a JSON file. Download the file from the URL in |
TranscriptionOutput
A TranscriptionOutput object is the output property of a TranscriptionResponse object, containing the task execution result.
Click to view TranscriptionOutput structure examples
Click to view TranscriptionOutput structure examples
- PENDING status
- RUNNING status
- SUCCEEDED status
- FAILED status
Parameter | Description |
|---|---|
code | The error code. Use with the |
message | The error message. Use with the |
task_id | The task ID. |
task_status | The task status. The four statuses are When a task contains multiple subtasks, if any subtask succeeds, the entire task status is marked as |
results | The recognition results of the subtasks. |
subtask_status | The subtask status. The four statuses are |
file_url | The URL of the audio file to be recognized. |
transcription_url | The URL corresponding to the audio recognition result. The recognition result is saved in a JSON file. Download the file from |
Recognition result description
The recognition result is saved as a JSON file.
Click to view recognition result example
Click to view recognition result example
Parameter | Type | Description |
|---|---|---|
audio_format | string | The audio format of the source file. |
channels | array[integer] | The audio track index information of the source file. Returns [0] for mono audio, [0, 1] for dual-track audio, and so on. |
original_sampling_rate | integer | The sampling rate (Hz) of the audio in the source file. |
original_duration | integer | The original audio duration (ms) of the source file. |
channel_id | integer | The audio track index of the transcription result, starting from 0. |
content_duration | integer | The duration (ms) of content identified as speech in the audio track. The Paraformer speech recognition model service only transcribes and meters content identified as speech in the audio track, and bills accordingly. Non-speech content is neither metered nor billed. Typically, the speech content duration is shorter than the original audio duration. Since the determination of whether speech content exists is made by an AI model, there may be some deviation from the actual situation. |
transcript | string | The paragraph-level speech transcription result. |
sentences | array | The sentence-level speech transcription result. |
words | array | The word-level speech transcription result. |
begin_time | integer | The start timestamp (ms). |
end_time | integer | The end timestamp (ms). |
text | string | The speech transcription result. |
speaker_id | integer | The index of the current speaker, starting from 0, used to distinguish different speakers. This field is only displayed in the recognition results when speaker diarization is enabled. |
punctuation | string | The predicted punctuation after the word (if any). |
API reference
Core class (Transcription)
Import the Transcription class: from dashscope.audio.asr import Transcription.
| Member method | Method signature | Description |
|---|---|---|
| async_call | Asynchronously submits a speech recognition task.This method returns TranscriptionResponse. | |
| wait | Blocks the current thread until the asynchronous task completes (status is SUCCEEDED or FAILED).This method returns TranscriptionResponse. | |
| fetch | Asynchronously queries the task execution result.This method returns a TranscriptionResponse. |
Error codes
If you encounter errors, see Error codes for troubleshooting.
If the issue persists, join the developer community to report the issue and provide the Request ID for further investigation.
When a task contains multiple subtasks, as long as any subtask succeeds, the overall task status is marked as SUCCEEDED. You need to check the subtask_status field to determine the result of each subtask.
Error response example:
More examples
Explore more examples on GitHub.
FAQ
Features
Q: Does it support Base64-encoded audio?
No. Base64-encoded audio is not supported. Only audio accessible via publicly accessible URLs is supported. Binary streams and direct local file recognition are not supported.
Q: How to provide audio files as publicly accessible URLs?
Generally, follow these steps (this provides a general approach; specifics vary by storage product. We recommend uploading audio to Alibaba Cloud OSS):
1. Choose a storage and hosting method
1. Choose a storage and hosting method
-
Object Storage Service (recommended):
- Use a cloud provider's object storage service (such as Alibaba Cloud OSS) to upload audio files to a bucket and set them to public access.
- Advantages: High availability, CDN acceleration support, easy management.
-
Web server:
- Place audio files on a web server that supports HTTP/HTTPS access (such as Nginx or Apache).
- Advantages: Suitable for small projects or local testing.
-
Content Delivery Network (CDN):
- Host audio files on a CDN and access them through the CDN-provided URL.
- Advantages: Accelerated file delivery, suitable for high-concurrency scenarios.
2. Upload audio files
2. Upload audio files
-
Object Storage Service:
- Log in to the cloud provider's console and create a bucket.
- Upload audio files and set file permissions to "public read" or generate temporary access links.
-
Web server:
- Place audio files in the server's designated directory (such as
/var/www/html/audio/). - Ensure files are accessible via HTTP/HTTPS.
- Place audio files in the server's designated directory (such as
3. Generate a publicly accessible URL
3. Generate a publicly accessible URL
-
Object Storage Service:
- After uploading, the system automatically generates a public access URL (typically in the format
https://<bucket-name>.<region>.aliyuncs.com/<file-name>). - If you need a more user-friendly domain, you can bind a custom domain and enable HTTPS.
- After uploading, the system automatically generates a public access URL (typically in the format
-
Web server:
- The file access URL is typically the server address plus the file path (such as
https://your-domain.com/audio/file.mp3).
- The file access URL is typically the server address plus the file path (such as
-
CDN:
- After configuring CDN acceleration, use the CDN-provided URL (such as
https://cdn.your-domain.com/audio/file.mp3).
- After configuring CDN acceleration, use the CDN-provided URL (such as
4. Verify URL accessibility
4. Verify URL accessibility
- Open the URL in a browser and check whether the audio file can be played.
- Use tools (such as
curlor Postman) to verify whether the URL returns a correct HTTP response (status code 200).
oss:// prefix are not supported.
When using the RESTful API, if audio files are stored in Alibaba Cloud OSS, temporary URLs with the oss:// prefix are supported:
- The temporary URL is valid for 48 hours and cannot be used after it expires. Do not use it in a production environment.
- The API for obtaining an upload credential is limited to 100 QPS and does not support scaling out. Do not use it in production environments, high-concurrency scenarios, or stress testing scenarios.
- For production environments, use a stable storage service such as OSS to ensure long-term file availability and avoid rate limiting issues.
Q: How long does it take to get recognition results?
After submission, the task enters a queued (PENDING) state. The queue time depends on the queue length and file duration and cannot be precisely estimated, but typically completes within a few minutes. Please wait patiently. Longer audio files require more processing time.
Troubleshooting
For code errors, see Error codes.
Q: What should I do if the recognition results are out of sync with the audio playback?
Set the request parameter timestamp_alignment_enabled to true to enable timestamp calibration, which synchronizes the recognition results with the speech playback.
Q: What should I do if the task returns an InvalidFile.DownloadFailed error?
Check whether the file URL contains spaces or other non-ASCII characters (such as Chinese characters). If the file name includes spaces (for example, my audio recording.mp4), replace each space with %20 to URL-encode the file name before passing it to the file_urls parameter.
Q: Unable to get results after continuous polling?
This may be due to rate limiting. Please wait patiently. If you need capacity expansion, join the developer community to apply.Q: Why is there no recognition result (unable to recognize speech)?
- Check whether the audio meets the requirements (format, sampling rate).
- If you are using the
paraformer-v2model, check whether thelanguage_hintssetting is correct. - If none of the above resolves the issue, you can customize hot words to improve the recognition of specific words.