Skip to main content
Speech Synthesis

voice-enrollment

A large-model voice replication service used in conjunction with Cosyvoice-v3. Utilizing advanced large-model technology for feature extraction, it can replicate voices without a training process. Only a very short audio clip is required to quickly generate a highly similar and natural-sounding custom voice.

Inference Service Provider

The inference service provider for voice-enrollment is Alibaba Cloud Model Studio.

Model Capabilities

  • China (Beijing)
  • Singapore
CapabilitySupportCapabilitySupport

Input Modality

Audio

Output Modality

Audio

Model Experience

Unsupported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

Max Output Length

Context Window

Pricing

No public pricing information available.

Rate Limits

  • China (Beijing)
  • Singapore
ParameterValue

RPM (Requests Per Minute)

600

Token Plan
Model Playground
  • Audio generation
  • Music generation
Statistics and Monitoring
Asset Center