How Kling Avatar v2 Standard works
Audio-driven image-to-video: your clip sets timing; your portrait sets who appears on screen.
Step 1
Reference image
Upload a clear portrait — JPG, PNG, WebP, and similar formats. Forward-facing photos with visible eyes and mouth produce the most natural lip-sync.
Step 2
Speech audio
Provide the words as synthesized speech or a recording. MP3, WAV, M4A, and other common formats work. Output duration matches this audio length.
Step 3
Lip-synced MP4 out
The model generates a video where facial movement follows the audio. An optional text prompt can nudge subtle animation details; sync stays audio-driven.