Skip to main content

Image-to-Singing-and-Acting Video – EMO

EMO generates dynamic portrait videos from a portrait image and human speech audio file. It consists of two models: EMO-detect verifies input image requirements, and EMO generates the video.

This document is only applicable to the Chinese mainland (Beijing) region. To use the model, you must use the Chinese mainland (Beijing) region's API key.

Model overview

About the models

  • EMO-detect is an image detection model that verifies whether input images meet EMO's portrait requirements.
  • EMO is a portrait video generation model that generates dynamic portrait videos from a portrait image and human speech audio file.

Example outputs

Input: portrait image + audio fileOutput: dynamic portrait video
Portrait:Sample audio: Spring Mountain songAudio: See video on the rightVideo:Style intensity: Active ("style_level": "active")
Portrait:Sample 15 - Original imageAudio: See video on the rightVideo:Style intensity: Normal ("style_level": "normal")
Portrait:娃哈哈Audio: See video on the rightVideo:Style intensity: Calm ("style_level": "calm")
The examples above were generated by the Qwen app, which integrates EMO.

Pricing and rate limits

Mode

Model name

Unit price

QPS limit for job submission

Maximum concurrent jobs

Model call

emo-detect-v1

Pay-as-you-go:

USD 0.000574 per image

5

No limit for synchronous calls

emo-v1

Pay-as-you-go:

  • 1:1 aspect ratio video: USD 0.011469 per second

  • 3:4 aspect ratio video: USD 0.022937 per second

1

Only one job runs at a time. Other jobs wait in the queue.

Prerequisites

Enable the service and obtain your API key: Obtain your API key and API host.

Call the models

  • To call the models (pay-as-you-go):
    1. Call EMO-detect to verify that your input image meets the requirements. See EMO image detection for details.
    2. Call EMO with the original image, region parameters from EMO-detect, and a clear human speech audio file to generate the video. See EMO video generation for details.
Text Generation
Image Generation
  • FAQ
Video Generation
Audio
Realtime API
Text Embedding
Model Production
Image-to-Singing-and-Acting Video – EMO - Alibaba Cloud Model Studio