Kling AI Avatar on Kanvora

Turn one private portrait image and one private 2-300 second audio file into a speaking-person video with Kling through fal. Standard is the cost-effective tier at $0.0562 per audio second, or 10-1,405 credits. Pro is the higher-quality tier at $0.115 per audio second, or 20-2,875 credits. Kanvora does not claim a resolution that the fal endpoint contract does not publish.

Output
Source-framed portrait video
Credits
10-1,405 credits
Input
1 private portrait image + 1 private audio file
Image
JPG/PNG · ≤10 MB · ≥300px sides · 1:2.5-2.5:1
Audio
MP3/WAV/M4A/AAC · 2-300s · ≤5 MB
Duration
Audio-timed · 2-300 seconds
Expected wait
1-8 minutes
Provider path
fal Kling AI Avatar
Registry verified
Aug 22, 2026

Start with the exact private source contract, not a blank upload.

This is a direct handoff to Kanvora's real creator, not an embedded generator. The model and operation are already attached, and no text prompt will be sent. Upload both private files after opening the creator. Upload the portrait image and audio privately after opening the creator.

Open the generator
Kling AvatarSource-framed portrait video · 10-1,405 credits

No text prompt · private portrait image + private audio

Official Kanvora examples

Site-owned AI-generated proofs from this model, never a public user gallery. Open an example to carry its cataloged direction into the creator.

Official output slotAwaiting a catalog asset
Reserved proofThe media and caption keep their size when an official asset is added.
Official output slotAwaiting a catalog asset
Reserved proofThe media and caption keep their size when an official asset is added.

From model page to stored proof

One public decision path, followed by the real authenticated workflow.

  1. 1

    Prepare one supported portrait image

    Use a JPG or PNG up to 10 MB. Both dimensions must be at least 300 pixels, and the aspect ratio must stay between 1:2.5 and 2.5:1.

  2. 2

    Prepare the speaking audio

    Use an MP3, WAV, M4A, or AAC from 2 through 300 seconds and up to 5 MB. Kanvora verifies the file privately before submission.

  3. 3

    Choose Standard or Pro

    Standard costs 24 credits for a five-second estimate and 1,405 at 300 seconds. Pro costs 48 at five seconds and 2,875 at 300 seconds. Verified audio duration sets the hold.

  4. 4

    Inspect the stored speaking portrait

    Review identity, mouth timing, teeth, expression, body motion, framing, and audio alignment. Neither tier guarantees perfect identity retention or synchronization.

Compare this live operation

Choose by the live Kanvora path, not by a model-family claim from somewhere else.

ModelResolutionCreditsProvider pathUse it when
MiniMax Hailuo 2.3 Standard768p24 creditsfal image-to-videoAnimates one starting frame from a text direction; it does not create a speaking portrait from supplied audio.
Seedance 2.5 Image to Video480p or 720p74-1,156 creditsfal Seedance videoConfigurable starting-frame video with native generated audio, not supplied-avatar audio.
Seedance 2.5 Multimodal Reference480p or 720p74-1,156 creditsfal Seedance videoMultimodal reference generation rather than a dedicated talking-portrait path.
Kling AI Avatar StandardSource-framed portrait video10-1,405 creditsfal Kling AI AvatarCost-effective talking portrait at $0.0562 per verified audio second.
Kling AI Avatar ProSource-framed portrait video20-2,875 creditsfal Kling AI AvatarHigher-quality talking portrait at $0.115 per verified audio second.
Veo 3.1 Starting Frame720p or 1080p or 4K134-400 creditsGoogle direct Veo videoGoogle-direct starting-frame generation with a text prompt and native generated audio.
Veo 3.1 Reference Images720p or 1080p or 4K267-400 creditsGoogle direct Veo videoGoogle-direct subject-reference generation with one to three images.
Veo 3.1 Video Transition720p or 1080p or 4K134-400 creditsGoogle direct Veo videoGoogle-direct interpolation between private first and last frames.

Frequently asked questions

Current Kanvora behavior, stated without extending the model's capability boundary.

How many credits does Kling AI Avatar use?

fal lists Standard at $0.0562 per second and Pro at $0.115 per second. Kanvora's existing $0.012 supplier-cost coverage per credit makes Standard 10-1,405 credits and Pro 20-2,875 credits across the supported 2-300 second audio range.

What is the difference between Standard and Pro?

fal describes Standard as the cost-effective endpoint and Pro as the premium, higher-quality endpoint. Pro costs about twice as much per second. Kanvora does not add an unsupported resolution promise to that contract.

What files can I upload?

The portrait must be JPG or PNG, at most 10 MB, at least 300 pixels on both sides, and between 1:2.5 and 2.5:1. Audio may be MP3, WAV, M4A, or AAC, must be 2-300 seconds, and must be at most 5 MB.

Is this the same as Kling Lip Sync?

No. AI Avatar starts from a still portrait plus audio and creates a speaking-person video. Lip Sync starts from an existing 2-10 second video plus audio and changes visible mouth timing.

Page facts were reviewed on August 22, 2026. Model identity source: fal Kling AI Avatar Standard endpoint.

Last verified: August 22, 2026.

Take this model into the real creator.

The prepared prompt and legal model parameters go with you. Sign-in remains at the generation boundary.

Create with Kling Avatar