Models
Which capabilities have a live model, what each model supports and costs, and how to read the current catalog from the API.
Windpaint is capability-first: you ask for image.generate or video.generate, and a model implements it. Two capabilities have a live model today. The others are part of the API’s schema, show up in the catalog with no models, and return 422 if you submit to them.
The catalog endpoint is the source of truth for what’s live and what it costs. Read prices from it rather than hard-coding the numbers on this page.
curl https://api.windpaint.ai/v1/generation/capabilities \
-H "Authorization: Bearer $WINDPAINT_API_KEY"
Capabilities
| Capability | Inputs (slot: count, kind) | Output | Prompt | Model |
|---|---|---|---|---|
image.generate | none | image | Required | z-image-turbo |
video.generate | start_frame: 0–1 image, end_frame: 0–1 image | video | Required | wan-2.2-i2v (needs exactly one start_frame, no end_frame) |
image.edit | images: 1–3 image, mask: 0–1 mask | image | Required | Coming |
mask.segment | images: 1 image | mask | Required (what to segment) | Coming |
media.upscale | media: 1 image or video | same as input | Optional | Coming |
audio.speech | voice: 0–1 audio | audio | Required (the text to speak) | Coming |
video.lipsync | media: 1 image or video, audio: 1 audio | video | Optional | Coming |
text.llm | images: 0–8 image | text | Required | Coming |
“Coming” capabilities appear in GET /v1/generation/capabilities with "models": []. Submitting to one returns 422 with the message No model implements <capability> yet. They’re listed so you can see the shape of the API; there’s no date attached to any of them.
Live models
z-image-turbo
Fast text-to-image, implementing image.generate.
| Capability | image.generate |
| Resolutions | 1k (default), 2k |
| Aspect ratios | All ten (see below), default 1:1 |
| Output | One image per request |
| Licence | Apache-2.0 |
| Price | 1k: 0.08 credits, 2k: 0.30 credits |
See Images for pixel sizes and examples.
wan-2.2-i2v
Image-to-video, implementing video.generate. It animates a start frame you supply.
| Capability | video.generate |
| Inputs | start_frame: exactly 1 image (PNG, JPEG or WebP). No end_frame. |
| Resolutions | 480p only |
| Durations | 5 seconds only |
| Aspect ratios | All ten, default 16:9 |
| Output | One MP4 clip, no audio |
| Licence | Apache-2.0 |
| Price | 480p, 5 s: 4 credits per clip |
See Video for the full workflow.
Prices are in credits per output (per image, per clip) at a tier, and can change. The price you pay is the credits_estimate returned when you submit; it’s fixed for that job even if the list price changes while it runs. See Pricing for what a credit costs.
Resolution tiers
A resolution tier sets one edge of the output, and the aspect ratio sets the other.
- Image tiers set the long edge:
1k= 1024 px,2k= 2048 px,4k= 4096 px. - Video tiers set the short edge:
480p= 480 px,720p= 720 px,1080p= 1080 px.
Both edges are rounded to the nearest multiple of 16. A model only accepts the tiers it lists in resolutions; today that’s 1k and 2k for z-image-turbo and 480p for wan-2.2-i2v. Asking for any other tier returns 422.
| Request | Output size |
|---|---|
image.generate, 1k, 1:1 | 1024 × 1024 |
image.generate, 1k, 16:9 | 1024 × 576 |
image.generate, 2k, 9:16 | 1152 × 2048 |
video.generate, 480p, 16:9 | 848 × 480 |
video.generate, 480p, 9:16 | 480 × 848 |
To see the exact size for any combination before you pay, call POST /v1/generation/estimate; it returns width and height. Images has the full table for image tiers.
Aspect ratios
Every capability with an image or video output accepts the same ten ratios:
1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16, 21:9, 4:5, 5:4
The default is 16:9 for video.generate and 1:1 for everything else. Anything outside the list returns 422.
Quality and duration
quality is reserved for models with quality tiers. No live model has any, so sending quality at all returns 422.
duration is in whole seconds. For a model with durations, leaving it out uses the model’s first duration (5 for wan-2.2-i2v), and any value outside its list returns 422. Image models don’t take a duration.
Default model selection
When you leave model out of a request, Windpaint uses the capability’s first available model in catalog order. With one model per live capability today, that means image.generate runs z-image-turbo and video.generate runs wan-2.2-i2v.
If you need a specific model, name it. When more models ship for a capability, the default may change, and a request without model would follow it. The status object always reports which model ran.
No version pinning
Model ids such as z-image-turbo are stable slugs. There’s no way to pin a version of a model; if Windpaint updates what’s behind a slug, requests that name it get the update. The price is fixed per job at submit time, so a price change never affects a job you’ve already submitted.
Read the live catalog
GET /v1/generation/capabilities returns {"data": [Capability]}:
{
"data": [
{
"id": "video.generate",
"description": "Prompt, optionally with a start frame and an end frame, to a video clip.",
"output": "video",
"prompt_required": true,
"prompt_meaning": "Instruction or description.",
"inputs": [
{ "name": "start_frame", "kinds": ["image"], "min": 0, "max": 1, "description": "First frame." },
{ "name": "end_frame", "kinds": ["image"], "min": 0, "max": 1, "description": "Last frame, for first/last-frame interpolation." }
],
"aspect_ratios": ["1:1", "4:3", "3:4", "3:2", "2:3", "16:9", "9:16", "21:9", "4:5", "5:4"],
"models": [
{
"model": "wan-2.2-i2v",
"description": "Wan 2.2 image-to-video 14B. Start frame required. 480p, 5 s.",
"available": true,
"prices": [{ "resolution": "480p", "duration_s": 5, "credits": "4.00" }],
"resolutions": ["480p"],
"default_resolution": "480p",
"qualities": [],
"durations": [5],
"inputs": [
{ "name": "start_frame", "kinds": ["image"], "min": 1, "max": 1, "description": "First frame." },
{ "name": "end_frame", "kinds": ["image"], "min": 0, "max": 0, "description": "Last frame, for first/last-frame interpolation." }
],
"licence": "Apache-2.0",
"licence_restricted": false,
"licence_note": null
}
]
}
]
}
The capability’s inputs are its general slot limits; each model’s inputs repeats them with that model’s limits applied. Read the model’s list when you’re building a request: above, it’s what tells you wan-2.2-i2v needs exactly one start frame and takes no end frame.
Whether you can submit to this model right now. false when the model is temporarily unavailable or has no price; submitting then returns 503.
One entry per tier that’s for sale: resolution, duration_s (null for images) and credits as a decimal string. A tier missing from this list can’t be submitted.
Accepted resolution values. default_resolution is used when you omit it.
Accepted quality values. Empty for every model today.
Accepted duration values in seconds. Empty for models that don’t take one.
The model’s licence. Both live models are Apache-2.0.
GET /v1/generation/models returns {"data": [{"id": "<model>", "capabilities": [...]}]}, where each entry carries the same description, licence, available, resolutions, default_resolution, qualities, durations and prices for one capability the model implements.
Price a request before you send it
POST /v1/generation/estimate takes a normal request body plus capability, validates it the same way submit does, and returns the price and output size without creating a job or holding credits.
{
"data": {
"capability": "image.generate",
"model": "z-image-turbo",
"credits": "0.30",
"width": 2048,
"height": 1360,
"available": true
}
}
credits is null when the tier has no price, and available is false when a submit would be refused with 503. Because estimate validates like submit, it needs a non-empty prompt for capabilities that require one.