AI video models

Every text-to-video and image-to-video model on Cloudflare Workers AI. Open any model to see its exact input schema and an estimated cost per video.

A

HappyHorse 1.1 · Text

Alibaba

best

Strong dynamic expressiveness, better visual quality and improved instruction following. Configurable resolution, aspect ratio and duration (3–15s).

720p1080p3–15s
$0.20/5s
A

HappyHorse 1.0 · Text

Alibaba

economy

HappyHorse 1.0 text-to-video. Generates videos from a text prompt with configurable resolution, aspect ratio and duration (3–15s).

720p1080p3–15s
$0.15/5s
B

Seedance 2.0

ByteDance

premium

Next-generation video model with synchronized audio. Generates from text, images, video clips and audio. Native audio generation, video editing and extension.

Audio480p–4k4–12s
$0.60/5s
B

Seedance 2.0 Fast

ByteDance

fast

Faster variant of Seedance 2.0. Trades some quality for speed while sharing the same multimodal architecture.

AudioFast4–12s
$0.40/5s
B

Seedance 2.0 Mini

ByteDance

economy

Compact, cost-efficient video generation model from the Seedance family. Ideal for high-volume workloads where speed and cost matter.

Audio480p720p
$0.25/5s
B

Seedance 2.5

ByteDance

premium

Audio-video generation model for 30-second videos with reference control and editing capabilities.

AudioUp to 30sEditing
$0.75/5s
B

FLUX 3 Video

Black Forest Labs

best

Generates video from a text prompt (t2v), animates one or more reference images (i2v) or continues an existing clip (v2v). Synchronized audio, up to FHD, 5–20s.

AudioHD/FHD5–20s
$0.50/5s
G

Veo 3.1

Google

premium

Google's latest video generation model with improved quality, motion and audio generation.

Audio720p1080p
$0.75/5s
G

Veo 3.1 Fast

Google

fast

A faster version of Veo 3.1 optimized for lower latency while maintaining high-quality video and audio output.

Audio720p1080p
$0.45/5s