Talking Avatar

Give any face your audio and it delivers the script, up to 30 seconds.

🎤Talking AvatarVideo output
Face image
Speech track
Direction (optional)0/1200
Style
Aspect ratio9:16
Duration5 to 20s
Resolution

FLUX 3 Video from 105 credits per 5s.

Lip sync runs on paid plans. Cost scales with audio length and resolution.

What is talking avatar?

A talking avatar is a video of a face delivering speech, synchronised to an audio track rather than filmed. APImage generates it from two inputs, a front-facing face image and a speech file of up to 30 seconds, using Seedance 2.0, which is the model that handles lip sync. APImage does not synthesise the voice, so the audio is your own recording or an export from a voice tool.

How to use it

  1. 1Upload one clear, front-facing face image.
  2. 2Upload the speech track, up to 30 seconds. MP3 or WAV.
  3. 3Pick a resolution and generate.
  4. 4For longer scripts, split the audio and generate several clips, then cut them together.

Specifications

Inputs
Face image plus audio file
Audio limit
30 seconds per clip
Model
Seedance 2.0, the only model that lip syncs
Voice generation
Not included. Bring your own audio.

Questions people actually ask

Make the first one

Free accounts get 3 watermarked images. Video generation starts on the paid plans.