AI video models
Every text-to-video and image-to-video model on Cloudflare Workers AI. Open any model to see its exact input schema and an estimated cost per video.
HappyHorse 1.1 · Image
Alibaba
Animates a reference image with an optional text prompt. Smoother motion, natural skin textures and improved close-up quality.
HappyHorse 1.0 · Image
Alibaba
Animates a reference image with an optional text prompt. 720P/1080P output, 3–15s. Zero data retention.
HappyHorse 1.1 · Reference
Alibaba
Reference-to-video. Takes 1–9 reference images (characters and scenes) plus a prompt and choreographs them into a single video, keeping each subject's identity consistent.
Wan 2.7 · Image
Alibaba
Generates videos from a reference image with optional text prompts. Supports 720P and 1080P output, durations 2–15s.
Seedance 2.0
ByteDance
Next-generation video model with synchronized audio. Generates from text, images, video clips and audio. Native audio generation, video editing and extension.
Seedance 2.0 Fast
ByteDance
Faster variant of Seedance 2.0. Trades some quality for speed while sharing the same multimodal architecture.
Seedance 2.0 Mini
ByteDance
Compact, cost-efficient video generation model from the Seedance family. Ideal for high-volume workloads where speed and cost matter.
Seedance 2.5
ByteDance
Audio-video generation model for 30-second videos with reference control and editing capabilities.
FLUX 3 Video
Black Forest Labs
Generates video from a text prompt (t2v), animates one or more reference images (i2v) or continues an existing clip (v2v). Synchronized audio, up to FHD, 5–20s.