Image to video
Animate one starting image with Kling 3.0 Turbo on Kling 4.0.
Image to video starts the generation from a picture you supply instead of from the prompt alone. It is the most reliable way to control exactly what is on screen, because the first frame is no longer a guess.
First frame
Upload one image and it becomes the opening frame of the clip. The prompt then describes what happens from there, not what things look like:
The woman turns her head towards the window and smiles.
Camera pushes in slowly. Everything else holds still.Describing the picture again ("a woman in a red coat in a cafe") wastes the prompt on information the model can already see. Spend it on motion instead.
Available model
Kling 3.0 Turbo is selected by default and accepts one starting image. There is no last-frame upload for this model. Kling 4.0 is a locked Coming soon preview and cannot be used yet.
Image requirements
- One starting image for Kling 3.0 Turbo.
- 10 MB per file.
- Use JPEG or PNG for Kling 3.0 Turbo.
Large uploads are resized to 2048 px on the long edge before they are sent, so there is no benefit to uploading a 6000 px master. Do use the cleanest source you have, though: compression artefacts, heavy grain, and watermarks in the input carry into every frame of the output.
Aspect ratio behaves differently here
When you supply a first frame, the output follows the shape of that image — the aspect-ratio control does not apply, because the model is expanding your frame rather than composing a new one. If you need a specific ratio, crop the input image to that ratio before uploading it.
Everything else
Kling 3.0 Turbo supports 3–15 seconds at 720p or 1080p, with no audio toggle. For multiple reference assets, switch to Multi Reference and choose a compatible model. See Text to video for credit and task behavior.
Continuing a shot
Any finished clip can seed the next one: take its last frame, upload it as the first frame of a new generation, and write the prompt for the next beat. Chaining short shots lets you build sequences longer than the model’s single-output limit. Each generation has its own credit cost.
When results go wrong
| Symptom | Usual cause |
|---|---|
| Barely moves | The prompt described the image instead of the motion. Name a specific action. |
| Subject warps or melts | Asked for too much movement in too few seconds. Lengthen the clip or shrink the action. |
| Output is the wrong shape | Expected — the output follows the input frame. Crop the input first. |
| Face drifts over the clip | Use a clear starting image and keep shots shorter. |
Still stuck? The contact page lists what to include so we can look at the specific generation.