How Kling Avatar v2 Standard works
Audio-driven image-to-video: your clip sets timing; your portrait sets who appears on screen.
步骤 1
Reference image
Upload a clear portrait — JPG, PNG, WebP, and similar formats. Forward-facing photos with visible eyes and mouth produce the most natural talking photo.
步骤 2
Speech audio
Provide the words as synthesized speech or a recording. MP3, WAV, M4A, and other common formats work. Output duration matches this audio length.
步骤 3
talking photo MP4 out
The model generates a video where facial movement follows the audio. An optional text prompt can nudge subtle animation details; sync stays audio-driven.