Text to Clip
Describe a moment and get a still image, a 5 second video clip, and a cover frame from one call.
Text to Clip (text-to-clip) turns a one-line idea into three files: a still image of the scene, a 5 second clip that animates it, and a cover frame taken from the end of the clip. It’s the quickstart’s image-then-video flow in a single request.
Inputs
What should happen in the clip, in plain words. It’s used as the prompt for both the still and the motion, so describe the scene and what moves: “a paper boat drifting down a rainy street”, “steam rising from a coffee cup on a windowsill at sunrise”.
Aspect ratio of every output. One of 16:9 (landscape) or 9:16 (vertical). The API lists this input as required, but it has a default, so you can leave it out.
Steps
| Step | Kind | What it does | Credits |
|---|---|---|---|
still | image.generate with z-image-turbo | Generates the scene from your idea, styled as a cinematic still with natural light, at the 1k tier in your aspect ratio. | 0.08 |
frame | image.resize op | Crops and resizes the still to the size the video model expects. | free |
clip | video.generate with wan-2.2-i2v | Animates the resized still into a 5 second, 480p clip, using your idea as the motion prompt. | 4 |
cover | video.extract_frame op | Takes the last frame of the clip as a PNG. | free |
The steps run in that order, since each one uses the previous step’s output. Most of the time is the clip step, which takes a few minutes.
Outputs
| Output | Kind | Description |
|---|---|---|
video | video (MP4) | The 5 second, 480p clip. |
still | image | The generated still, before resizing. |
cover | image (PNG) | The clip’s last frame, for use as a thumbnail or the start of a follow-up clip. |
All three are saved as assets in the run’s project. The resized frame is saved too and appears in that step’s outputs, but it isn’t one of the run’s named outputs.
Price
A run costs 4.08 credits at current prices: 0.08 for the still and 4 for the clip. Both aspect ratios cost the same. Check before you run:
If the clip step fails, you still pay for the still (0.08 credits) and nothing for the clip. See Billing for how runs are charged.
Full example
Start the run, poll until it finishes, then download the three outputs.
A completed run looks like this (steps trimmed):
{
"data": {
"id": "7c1e9a40-3b2d-4f6a-8e15-0d9c4b7a2f63",
"status": "completed",
"product": "text-to-clip",
"product_version": 1,
"inputs": { "idea": "steam rising from a coffee cup on a windowsill at sunrise", "aspect": "16:9" },
"steps": [ ... ],
"outputs": {
"video": [{ "id": "e2a7...", "url": "https://api.windpaint.ai/v1/generation/assets/e2a7.../content", "content_type": "video/mp4", "duration_s": 5.0 }],
"still": [{ "id": "d77a...", "url": "https://api.windpaint.ai/v1/generation/assets/d77a.../content", "content_type": "image/png", "width": 1024, "height": 576 }],
"cover": [{ "id": "91fc...", "url": "https://api.windpaint.ai/v1/generation/assets/91fc.../content", "content_type": "image/png" }]
},
"credits": { "estimate": "4.08", "actual": "4.08" },
"error": null,
"finished_at": "2026-10-04T14:13:41.902Z"
}
}
If status is failed, error names the step that failed and why, and that step’s error has the detail. The run’s outputs stay empty, but the outputs of steps that completed (the still, for example) are in steps[].outputs and in your project’s assets.
Tips
- Write the idea as a scene with motion. The same text drives the still and the animation, so “a paper boat drifting down a rainy street” works better than “paper boat”.
- Chain clips with the cover. The
coveris the clip’s last frame. Pass its asset id asstart_frametovideo.generateto continue the shot. See Video. - Use 9:16 for vertical feeds and 16:9 for everything else. Both cost the same.
Other products
The catalog lists other products too. They depend on capabilities that don’t have a model yet, so they report available: false and can’t be run. They’ll start working when those models ship, without any change on your side. See Availability.