AI video models

Every text-to-video and image-to-video model on Cloudflare Workers AI. Open any model to see its exact input schema and an estimated cost per video.

A

HappyHorse 1.1 · Image

Alibaba

best

Animates a reference image with an optional text prompt. Smoother motion, natural skin textures and improved close-up quality.

720p1080p3–15s
$0.20/5s
A

HappyHorse 1.0 · Image

Alibaba

economy

Animates a reference image with an optional text prompt. 720P/1080P output, 3–15s. Zero data retention.

720p1080p3–15s
$0.15/5s
A

HappyHorse 1.1 · Reference

Alibaba

new

Reference-to-video. Takes 1–9 reference images (characters and scenes) plus a prompt and choreographs them into a single video, keeping each subject's identity consistent.

1–9 imagesIdentity720p
$0.25/5s
A

Wan 2.7 · Image

Alibaba

new

Generates videos from a reference image with optional text prompts. Supports 720P and 1080P output, durations 2–15s.

720p1080p2–15s
$0.20/5s
B

Seedance 2.0

ByteDance

premium

Next-generation video model with synchronized audio. Generates from text, images, video clips and audio. Native audio generation, video editing and extension.

Audio480p–4k4–12s
$0.60/5s
B

Seedance 2.0 Fast

ByteDance

fast

Faster variant of Seedance 2.0. Trades some quality for speed while sharing the same multimodal architecture.

AudioFast4–12s
$0.40/5s
B

Seedance 2.0 Mini

ByteDance

economy

Compact, cost-efficient video generation model from the Seedance family. Ideal for high-volume workloads where speed and cost matter.

Audio480p720p
$0.25/5s
B

Seedance 2.5

ByteDance

premium

Audio-video generation model for 30-second videos with reference control and editing capabilities.

AudioUp to 30sEditing
$0.75/5s
B

FLUX 3 Video

Black Forest Labs

best

Generates video from a text prompt (t2v), animates one or more reference images (i2v) or continues an existing clip (v2v). Synchronized audio, up to FHD, 5–20s.

AudioHD/FHD5–20s
$0.50/5s