Kling AI Avatar on Kanvora
Turn one private portrait image and one private 2-300 second audio file into a speaking-person video with Kling through fal. Standard is the cost-effective tier at $0.0562 per audio second, or 10-1,405 credits. Pro is the higher-quality tier at $0.115 per audio second, or 20-2,875 credits. Kanvora does not claim a resolution that the fal endpoint contract does not publish.
- Output
- Source-framed portrait video
- Credits
- 10-1,405 credits
- Input
- 1 private portrait image + 1 private audio file
- Image
- JPG/PNG · ≤10 MB · ≥300px sides · 1:2.5-2.5:1
- Audio
- MP3/WAV/M4A/AAC · 2-300s · ≤5 MB
- Duration
- Audio-timed · 2-300 seconds
- Expected wait
- 1-8 minutes
- Provider path
- fal Kling AI Avatar
- Registry verified
- Aug 22, 2026
Start with the exact private source contract, not a blank upload.
This is a direct handoff to Kanvora's real creator, not an embedded generator. The model and operation are already attached, and no text prompt will be sent. Upload both private files after opening the creator. Upload the portrait image and audio privately after opening the creator.
Open the generatorNo text prompt · private portrait image + private audio
Official Kanvora examples
Site-owned AI-generated proofs from this model, never a public user gallery. Open an example to carry its cataloged direction into the creator.
From model page to stored proof
One public decision path, followed by the real authenticated workflow.
- 1
Prepare one supported portrait image
Use a JPG or PNG up to 10 MB. Both dimensions must be at least 300 pixels, and the aspect ratio must stay between 1:2.5 and 2.5:1.
- 2
Prepare the speaking audio
Use an MP3, WAV, M4A, or AAC from 2 through 300 seconds and up to 5 MB. Kanvora verifies the file privately before submission.
- 3
Choose Standard or Pro
Standard costs 24 credits for a five-second estimate and 1,405 at 300 seconds. Pro costs 48 at five seconds and 2,875 at 300 seconds. Verified audio duration sets the hold.
- 4
Inspect the stored speaking portrait
Review identity, mouth timing, teeth, expression, body motion, framing, and audio alignment. Neither tier guarantees perfect identity retention or synchronization.
Compare this live operation
Choose by the live Kanvora path, not by a model-family claim from somewhere else.
| Model | Resolution | Credits | Provider path | Use it when |
|---|---|---|---|---|
| MiniMax Hailuo 2.3 Standard | 768p | 24 credits | fal image-to-video | Animates one starting frame from a text direction; it does not create a speaking portrait from supplied audio. |
| Seedance 2.5 Image to Video | 480p or 720p | 74-1,156 credits | fal Seedance video | Configurable starting-frame video with native generated audio, not supplied-avatar audio. |
| Seedance 2.5 Multimodal Reference | 480p or 720p | 74-1,156 credits | fal Seedance video | Multimodal reference generation rather than a dedicated talking-portrait path. |
| Kling AI Avatar Standard | Source-framed portrait video | 10-1,405 credits | fal Kling AI Avatar | Cost-effective talking portrait at $0.0562 per verified audio second. |
| Kling AI Avatar Pro | Source-framed portrait video | 20-2,875 credits | fal Kling AI Avatar | Higher-quality talking portrait at $0.115 per verified audio second. |
| Veo 3.1 Starting Frame | 720p or 1080p or 4K | 134-400 credits | Google direct Veo video | Google-direct starting-frame generation with a text prompt and native generated audio. |
| Veo 3.1 Reference Images | 720p or 1080p or 4K | 267-400 credits | Google direct Veo video | Google-direct subject-reference generation with one to three images. |
| Veo 3.1 Video Transition | 720p or 1080p or 4K | 134-400 credits | Google direct Veo video | Google-direct interpolation between private first and last frames. |
Frequently asked questions
Current Kanvora behavior, stated without extending the model's capability boundary.
How many credits does Kling AI Avatar use?
fal lists Standard at $0.0562 per second and Pro at $0.115 per second. Kanvora's existing $0.012 supplier-cost coverage per credit makes Standard 10-1,405 credits and Pro 20-2,875 credits across the supported 2-300 second audio range.
What is the difference between Standard and Pro?
fal describes Standard as the cost-effective endpoint and Pro as the premium, higher-quality endpoint. Pro costs about twice as much per second. Kanvora does not add an unsupported resolution promise to that contract.
What files can I upload?
The portrait must be JPG or PNG, at most 10 MB, at least 300 pixels on both sides, and between 1:2.5 and 2.5:1. Audio may be MP3, WAV, M4A, or AAC, must be 2-300 seconds, and must be at most 5 MB.
Is this the same as Kling Lip Sync?
No. AI Avatar starts from a still portrait plus audio and creates a speaking-person video. Lip Sync starts from an existing 2-10 second video plus audio and changes visible mouth timing.
Take this model into the real creator.
The prepared prompt and legal model parameters go with you. Sign-in remains at the generation boundary.