Talking Avatar
Give any face your audio and it delivers the script, up to 30 seconds.
🎤Talking AvatarVideo output
Face image
Speech track
Direction (optional)0/1200
Style
Aspect ratio9:16
Duration5 to 20s
Resolution
FLUX 3 Video from 105 credits per 5s.
Lip sync runs on paid plans. Cost scales with audio length and resolution.
What is talking avatar?
A talking avatar is a video of a face delivering speech, synchronised to an audio track rather than filmed. APImage generates it from two inputs, a front-facing face image and a speech file of up to 30 seconds, using Seedance 2.0, which is the model that handles lip sync. APImage does not synthesise the voice, so the audio is your own recording or an export from a voice tool.
- 1Upload one clear, front-facing face image.
- 2Upload the speech track, up to 30 seconds. MP3 or WAV.
- 3Pick a resolution and generate.
- 4For longer scripts, split the audio and generate several clips, then cut them together.
- Inputs
- Face image plus audio file
- Audio limit
- 30 seconds per clip
- Model
- Seedance 2.0, the only model that lip syncs
- Voice generation
- Not included. Bring your own audio.
Where this tool earns its place in a real workflow.
Make the first one
Free accounts get 3 watermarked images. Video generation starts on the paid plans.