Rendley docs

Model catalog

Every AI model available through the Rendley video API, with the parameters each one accepts.

58 models

Chips are the parameters each model accepts. Bold is required. A marks a parameter that takes either a publichttps URL or a library file hash, whatever the parameter is named.

Generate video28

FLUX 3

flux-3

PREMIUM multimodal video with synchronized audio. Unique strengths: storyboard keyframes — up to 10 images placed pixel-for-pixel (one = first frame, two = first+last, 3+ = evenly spaced keyframes) — and shot continuation from an existing mp4 (start_video, ≤15s ≤50MB; cannot be combined with images). 720p/1080p, 5–20s, wide aspect ratios including 21:9 and 2:1; draft mode gives a cheap 720p preview. Text/image variants cost ~2x kling-v2.6 and continuation (start_video) is very expensive (~2.5x that again). Pick this ONLY when a shot must hit exact keyframes or seamlessly continue an existing clip; for everything else use kling-v2.6 or seedance-2.0.

  • prompt
  • images
  • start_video
  • aspect_ratio
  • resolution
  • duration
  • generate_audio
  • draft

Gemini Omni 1.1

gemini-omni-1.1

Google's fast multimodal video model with native synchronized audio (360p draft / 720p / 1080p / 4K, 16:9 or 9:16). Four input modes in one model: text-to-video; image-to-video from a start frame (`image`); first→last-frame interpolation (`image` + `last_frame`); reference-guided generation (`reference_images` — identity/style, NOT literal frames, wire a character design sheet here); and video-to-video EDITING (`video` — describe the change in the prompt, e.g. 'make the sky stormy, keep everything else', output length matches the input). The clip length has no numeric parameter, but you steer it in the PROMPT (state a target, up to ~10s max); the model follows it approximately, so read the real duration after. Billed per second of output, so 360p is very cheap and 4K is costly. Pick it for reference-driven character shots with audio, quick first→last-frame motion, or editing an existing clip; for a precise/exact duration or takes longer than ~10s use kling-v2.6/seedance-2.0, and for a single photo talking-head use omni-human-1.5.

  • prompt
  • image
  • last_frame
  • video
  • reference_images
  • resolution
  • aspect_ratio

Grok Imagine Reference to Video

grok-imagine-r2v

xAI's Grok Imagine reference-to-video model. Takes 1–7 reference images used as style and content references (not starting frames) plus a prompt, and generates a 1–10s clip at 480p/720p. Pick when the user wants Grok Imagine to follow the look of supplied reference images.

  • prompt
  • reference_images
  • duration
  • aspect_ratio
  • resolution

Grok Imagine Image to Video

grok-imagine-video

xAI's Grok Imagine video model. Text-to-video, image-to-video (animate a starting image), or video editing (supply a short source video to restyle). Flexible 1–15s durations at 480p/720p. Niche pick when the user specifically wants Grok Imagine's look or its video-editing mode.

  • prompt
  • image
  • video
  • duration
  • aspect_ratio
  • resolution

Grok Imagine Video 1.5

grok-imagine-video-1.5

xAI's Grok Imagine Video 1.5 (preview) — image-to-video ONLY (`image` required) with synchronized native audio: background music, sound effects, ambience matched to the visuals, and short dialogue via an 'AUDIO:' section in the prompt. 1–15s at 480p/720p. Pick for animating a still (product showcase, portrait, character) when it should come with fitting sound in one pass; for text-to-video use grok-imagine-video.

  • prompt
  • image
  • duration
  • aspect_ratio
  • resolution

Grok Imagine Video Extension

grok-imagine-video-extension

xAI's Grok Imagine video-extension model. Takes a source MP4 (2–15s) and a prompt describing what happens next, and generates a 2–10s continuation from the last frame. Pick when the user wants to extend an existing video clip.

  • prompt
  • video
  • duration

Hailuo 2.3

hailuo-2.3

Stable, consistent motion with low artifact rates — fewer warps and flickers than larger models. Niche pick when motion stability matters more than peak fidelity.

  • prompt
  • first_frame_image
  • duration
  • resolution
  • prompt_optimizer

Kling V2.5 Turbo Pro

kling-v2.5-turbo-pro

Faster, cheaper Kling variant with first+last frame interpolation. Pick when speed or cost matters more than peak quality, or when the request specifically needs first+last frame interpolation.

  • prompt
  • start_image
  • end_image
  • duration
  • aspect_ratio

Kling V2.6

default
kling-v2.6

General-purpose video generation. Realistic motion, strong with people and natural scenes, generates audio by default. Pick this for any standard 5- or 10-second clip. For peak cinematic quality use veo-3.1 instead.

  • prompt
  • start_image
  • aspect_ratio
  • duration
  • generate_audio

Kling V2.6 Motion Control

kling-v2.6-motion-control

Specialized Kling variant for character motion transfer — takes a character image plus a motion reference video and produces a video where the character moves the same way. Pick only when the user explicitly wants to drive a character's motion from another video. Not a general-purpose video model.

  • image
  • video
  • prompt
  • character_orientation
  • mode
  • keep_original_sound

Kling 3.0 Omni

kling-v3-omni-video

Kling's unified multimodal video model: text, first+last frame (`start_image`/`end_image`), up to 7 reference images for character/product identity, elements or style (refer to them as <<<image_1>>> in the prompt), a reference video for style/camera or as an edit base, and optional native audio. 720p (standard), 1080p (pro, default) or 4K; 3–15s; 16:9, 9:16, 1:1. Premium price class — roughly 2x seedance-2.0 per second (pro) and 4K far more; audio adds to the price. The Kling-family reference-image path: pick this when the shot needs reference-image identity with Kling's motion and prompt adherence, or as the retry when a seedance call fails (same inputs); for start-frame-only shots kling-v2.6 is cheaper, for the cheapest reference-image path use seedance-2.0-mini.

  • prompt
  • start_image
  • end_image
  • reference_images
  • reference_video
  • video_reference_type
  • keep_original_sound
  • generate_audio
  • mode
  • aspect_ratio
  • duration

Kling 3.0

kling-v3-video

Kling 3.0 — text-to-video or first+last frame (`start_image`/`end_image`) with optional native audio (dialogue in double quotes, ambience, effects). 720p (standard), 1080p (pro, default) or 4K; 3–15s; 16:9, 9:16, 1:1. The leaner, cheaper sibling of kling-v3-omni-video: NO reference images and NO reference video — pick this for straight start-frame or text shots with Kling's motion and prompt adherence at a lower price; use kling-v3-omni-video when the shot needs reference-image identity or a reference video, and kling-v2.6 for the cheapest Kling start-frame path.

  • prompt
  • start_image
  • end_image
  • generate_audio
  • mode
  • aspect_ratio
  • duration
  • negative_prompt

LTX 2.5 Fast

ltx-2.5-fast

Cheapest fast video model with synchronized audio. Text-to-video or image-to-video with first-frame and optional last-frame conditioning, 16:9 or 9:16, 720p/1080p/2k/4k, 24/25/48/50 fps, durations 2–20s (over 10s only at 720p/1080p with 24/25 fps; 2k/4k and 48/50 fps cap at 10s). No reference-image identity mode. Pick this for drafts, volume work and quick previews where price and speed beat quality; for character consistency from reference images use seedance-2.0, for polished standard clips use kling-v2.6.

  • prompt
  • image
  • last_frame_image
  • resolution
  • duration
  • aspect_ratio
  • fps
  • generate_audio

MiniMax H3

minimax-h3

MiniMax H3 — multimodal video: text-to-video, first/last-frame animation (`first_frame_image`/`last_frame_image`), and reference-guided generation (up to 9 reference images, 3 reference videos, 3 reference audio clips for character/motion/voice/style). 4–15s at 768P ($0.10/s) or 2K ($0.14/s); ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, or 'adaptive' for image-to-video. Strong all-rounder when the shot needs reference identity or audio-guided rhythm at moderate cost; for peak cinematic quality use veo-3.1, for the cheapest start-frame path use kling-v2.6.

  • prompt
  • first_frame_image
  • last_frame_image
  • reference_image_urls
  • reference_video_urls
  • reference_audio_urls
  • duration
  • resolution
  • ratio

OmniHuman 1.5

omni-human-1.5

Audio-driven digital human: one image + one audio clip → the person in the image speaks/performs the audio with lip-sync, facial expression and body motion. The `image` is the PERSON (a photo of the subject), not a scene start frame; no first-frame/reference pipeline needed. Optional prompt directs scene, emotion and camera. Pick this for talking-head, UGC presenter, spokesperson or avatar shots from a single photo with pre-made speech/TTS audio. Standard price class, billed per second of output; the video is exactly as long as the audio (audio must be ≤35s) and the caller must pass the audio length in `duration`. For generic scenes without a speaking subject use seedance-2.0-mini or kling-v2.6.

  • image
  • audio
  • prompt
  • fast_mode
  • duration

P-Video

p-video

Pruna P-Video — FAST, cheap video generation: text-to-video, image-to-video (first and optional last frame), and audio-to-video (condition on a music/voice track; the video follows the audio, e.g. a character singing) in one endpoint. 1–20s at 720p ($0.02/s) or 1080p ($0.04/s), 24/48 fps, with a quarter-price `draft` mode for rapid iteration. Pick for quick, budget drafts and audio-driven shots; for peak quality use veo-3.1 or kling-v3.

  • prompt
  • image
  • last_frame_image
  • audio
  • duration
  • aspect_ratio
  • resolution
  • fps
  • draft
  • prompt_upsampling
  • save_audio

PixVerse V6

pixverse-v6

PixVerse's flagship video model — cinematic text-to-video and image-to-video with optional synchronized audio (`generate_audio_switch`: BGM, SFX, dialogue), first→last frame transitions (`image` + `last_frame_image`), and a multi-shot mode (`generate_multi_clip_switch`) that renders a storyboard-style prompt as one sequence with scene transitions. 5/8/10/15s at 360p–1080p ($0.05–$0.18/s, audio adds a little). Pick for multi-shot cinematic sequences or stylized brand-film looks from a single call; for photoreal single shots prefer kling or seedance.

  • prompt
  • image
  • last_frame_image
  • quality
  • aspect_ratio
  • duration
  • negative_prompt
  • generate_audio_switch
  • generate_multi_clip_switch

Seedance 1.5 Pro

seedance-1.5-pro

Strong on expressive motion, dynamic choreography, and dance. Niche pick for motion-driven or stylized content, or when fine motion control matters more than realism.

  • prompt
  • image
  • last_frame_image
  • duration
  • resolution
  • aspect_ratio
  • camera_fixed
  • generate_audio
  • fps

Seedance 2.0

seedance-2.0

The Seedance-family DEFAULT and the go-to reference-image video model: multimodal generation with native audio synced to dialogue/SFX, reference inputs (up to 9 images, plus reference videos and audio for lip-sync), flexible 1–15s durations, first+last frame interpolation, up to 1080p. **Slow generation — significantly higher latency than kling-v2.6 or veo-3.1.** Pick this whenever identity must be carried by reference images (character/product consistency from sheets), for motion transfer from a reference video, audio-driven lip-sync, or clips longer than 10s. For standard clips without reference images use kling-v2.6. Never upgrade to seedance-2.5 on your own — it is far more expensive and capped at 720p.

  • prompt
  • image
  • last_frame_image
  • duration
  • resolution
  • aspect_ratio
  • generate_audio
  • reference_images
  • reference_videos
  • reference_audios

Seedance 2.0 Fast

seedance-2.0-fast

Faster, cheaper Seedance 2.0 variant with the same multimodal feature set (native audio, up to 9 reference images, reference videos/audio, flexible 1–15s durations, first+last frame interpolation) but capped at 480p/720p. Pick when iterating on Seedance multimodal work or when speed and cost matter more than peak resolution; for 1080p use seedance-2.0.

  • prompt
  • image
  • last_frame_image
  • duration
  • resolution
  • aspect_ratio
  • generate_audio
  • reference_images
  • reference_videos
  • reference_audios

Seedance 2.0 Mini

seedance-2.0-mini

CHEAP Seedance — about half the price of seedance-2.0 with the same multimodal inputs (first+last frame interpolation, up to 9 reference images, up to 3 reference videos/audios, prompt-driven dialogue) and native synced audio, but 720p max (480p/720p only) and 4–15s. Pick this for volume work, drafts, iterations and any Seedance job where 720p is enough; for 1080p use seedance-2.0. Supplying reference videos moves it to a slightly higher pricing tier.

  • prompt
  • image
  • last_frame_image
  • duration
  • resolution
  • aspect_ratio
  • generate_audio
  • reference_images
  • reference_videos
  • reference_audios

Seedance 2.5

seedance-2.5

PREMIUM, EXPENSIVE, 720p-MAX variant of Seedance — NOT a default, NOT an upgrade path, NOT a fallback for seedance-2.0. Use ONLY when the user explicitly asks for seedance-2.5, or when a single shot genuinely needs what 2.0 cannot do: one native take longer than 15s (up to 30s) or more than 9 reference images (up to 30 images, 10 videos, 10 audios). Otherwise identical feature set to seedance-2.0 (native synced audio, reference inputs, first+last frame interpolation) at several times the price and lower resolution (480p/720p only). **Slow generation — significantly higher latency than kling-v2.6 or veo-3.1.** Supplying reference videos moves it onto a much more expensive pricing tier (~4x). For everything else use seedance-2.0 (reference images) or kling-v2.6 (standard clips).

  • prompt
  • image
  • last_frame_image
  • duration
  • resolution
  • aspect_ratio
  • generate_audio
  • reference_images
  • reference_videos
  • reference_audios

Veo 3.1

veo-3.1

Highest-quality video generation. Cinematic camera moves, realistic lighting, generated audio, and first+last frame interpolation. Recommended alongside kling-v2.6 — pick this when the user explicitly wants peak cinematic quality (hero shots, brand work) and the duration fits 4–8 seconds. More expensive than kling-v2.6.

  • prompt
  • image
  • last_frame
  • duration
  • aspect_ratio
  • resolution
  • generate_audio
  • reference_images

Veo 3.1 Fast

veo-3.1-fast

Faster, cheaper Veo 3.1 variant. Slightly reduced quality. Pick when the user wants the Veo aesthetic but is iterating, or as a draft pass before committing to veo-3.1.

  • prompt
  • image
  • last_frame
  • duration
  • aspect_ratio
  • resolution
  • generate_audio

Wan 2.7 I2V

wan-2.7-i2v

Image-to-video variant of Wan 2.7. Animates a starting image with optional last-frame interpolation and synchronized audio. Niche pick for longer image-to-video clips than Kling supports, or when a custom audio track is needed.

  • prompt
  • first_frame
  • last_frame
  • first_clip
  • duration
  • resolution
  • audio
  • enable_prompt_expansion

Wan 2.7 T2V

wan-2.7-t2v

Text-to-video with synchronized audio generation tightly aligned to the prompt, and longer durations than kling-v2.6 or veo-3.1. Niche pick for longer clips with synced dialogue or action audio, or when attaching a custom audio track. For image-to-video use wan-2.7-i2v.

  • prompt
  • duration
  • resolution
  • aspect_ratio
  • audio
  • enable_prompt_expansion

Wan 3.0

wan-3

Cheap text-to-video only (no image, reference or audio inputs). Cinematic motion, 480p/720p/1080p, 2–15s in 1-second steps, optional negative prompt and prompt expansion; no audio track. Cheapest 1080p text-to-video in the catalog. Pick this when a shot is prompt-only and cost matters; for a first-frame image, audio or reference-driven shots use kling-v2.6 or seedance-2.0; for a single prompt-only take longer than 15s use wan-3-prime.

  • prompt
  • negative_prompt
  • resolution
  • aspect_ratio
  • duration
  • enable_prompt_expansion

Wan 3.0 Prime

wan-3-prime

Premium text-to-video only (no image, reference or audio inputs; no audio track) — the long single-take option: one continuous prompt-only shot from 2 up to 30s at 480p/720p/1080p. Costs ~2x wan-3 per second. Pick this ONLY when a prompt-only shot genuinely must be longer than 15s in one take; for 15s or less use wan-3, for image/reference/audio-driven shots use kling-v2.6 or seedance-2.0.

  • prompt
  • negative_prompt
  • resolution
  • aspect_ratio
  • duration
  • enable_prompt_expansion

Generate image15

DALL-E 3

dalle-3

Strong prompt understanding for creative and illustrative work like ads and concept art. Niche legacy option — nano-banana covers most of the same ground.

  • prompt
  • aspect_ratio

Flux 1.1 Pro

flux-1.1-pro

Strong prompt adherence and layout control — handles complex multi-subject prompts and specific positioning. Niche pick when the prompt is unusually detailed and composition matters.

  • prompt
  • image_prompt
  • aspect_ratio
  • width
  • height
  • prompt_upsampling

Flux 2 Max

flux-2-max

Black Forest Labs' highest-fidelity image model. Exceptional detail, photoreal textures, and prompt adherence; accepts up to 8 reference images for identity/product/style consistency; output selectable from 0.5 to 4 MP (max 2048x2048) or custom width/height. Premium price class — cost scales with output megapixels AND with every reference image, so it is one of the most expensive image options. Pick this only when the user explicitly asks for maximum image quality or for Flux. For everyday generation, edits and people/product detail nano-banana-pro is the cheaper default.

  • prompt
  • image_inputs
  • aspect_ratio
  • resolution
  • width
  • height

GPT Image 2

gpt-image-2

OpenAI's top image model — sharp text, precise edits, quality knob.

  • prompt
  • image_inputs
  • aspect_ratio
  • quality
  • background

Grok Imagine Image 2

grok-imagine-image-2

xAI's Grok Imagine Image 2.0 — general-purpose text-to-image with a single-image edit mode (pass `image` and the prompt describes the change; aspect ratio is then ignored). Output at 1k or 2k, wide aspect-ratio list including phone-screen ratios (19.5:9, 20:9) and `auto`, low/medium quality knob. Standard price per image; edits cost slightly more. Pick this when the user asks for Grok/xAI specifically or wants a 2k generalist alternative; for text-heavy layouts use qwen-image-3, and for the default use nano-banana-pro.

  • prompt
  • image
  • aspect_ratio
  • resolution
  • quality

Imagen 4

imagen-4

Photorealistic image generation. Strong on natural lighting, materials, and physical accuracy. Niche pick when nano-banana's photorealism is insufficient.

  • prompt
  • aspect_ratio
  • image_size

Nano Banana

default
nano-banana

General-purpose image generation. Fast, cheap, broad subject coverage, and good at image-to-image edits (change pose, restyle, add/remove elements). Preferred default for almost all image generation requests.

  • prompt
  • image_inputs
  • aspect_ratio

Nano Banana Pro

nano-banana-pro

Higher-quality variant of nano-banana. Noticeably better at faces, hands, fine detail, in-image text rendering, and photoreal product shots. Preferred when the request specifically involves people, text rendering, or fine product detail; otherwise nano-banana is the better default.

  • prompt
  • image_inputs
  • aspect_ratio
  • resolution

Qwen Image 3

qwen-image-3

Alibaba's Qwen-Image-3. Its specialty is accurate in-image TEXT rendering and dense multi-element layouts — posters, infographics, UI mockups, packaging and ads with real copy — plus photographic detail. Text-to-image or single-image editing (optional `image`; `match_input_image` keeps the source size). Fixed aspect ratios up to 2:1/1:2, negative prompt, optional prompt expansion (turn it off for exact text). Cheap (flat price per image, one of the lowest). Pick this when the image must contain legible, correctly spelled text or a busy layout; for plain photos or general edits nano-banana-pro is the default, and for the densest text/layouts use qwen-image-3-pro.

  • prompt
  • image
  • match_input_image
  • aspect_ratio
  • negative_prompt
  • enable_prompt_expansion

Qwen Image 3 Pro

qwen-image-3-pro

Higher-tier Qwen-Image-3 for dense, accurate TEXT rendering and complex multi-element layouts — infographics with many labels, multi-panel posters, UI screens, packaging with paragraphs of copy — with photographic detail. Same inputs as qwen-image-3: text-to-image or single-image editing (optional `image`; `match_input_image`), fixed aspect ratios, negative prompt, prompt expansion. Standard price (flat per image, slightly above qwen-image-3). Pick this when a lot of text or many layout elements must all come out right; for simpler text tasks qwen-image-3 is cheaper, and for general imagery use nano-banana-pro.

  • prompt
  • image
  • match_input_image
  • aspect_ratio
  • negative_prompt
  • enable_prompt_expansion

Seedream 4

seedream-4

Stylized image generation. Strong on illustration, anime, and painterly looks. Niche pick when the user explicitly wants a non-photoreal aesthetic.

  • prompt
  • image_inputs
  • size
  • aspect_ratio
  • width
  • height
  • enhance_prompt

Seedream 4.5

seedream-4.5

Successor to seedream-4 with stronger spatial understanding, world knowledge, and in-image text rendering; keeps the stylized/illustrative strengths. Takes an optional list of up to 14 reference images for multi-reference generation and editing. Outputs 2K or 4K (or custom 1024–4096px dimensions); 1K is not supported. Standard price class (flat per image, 4K costs the same as 2K) — slightly above seedream-4. Third in the default fallback chain (nano-banana-pro → nano-banana → seedream-4.5). Pick this when the user wants the Seedream look with better layout/text accuracy or needs 4K output; for the cheapest Seedream use seedream-4; for photoreal people/products use nano-banana-pro.

  • prompt
  • image_inputs
  • size
  • aspect_ratio
  • width
  • height
  • enhance_prompt

Seedream 5 Lite

seedream-5-lite

Cheap Seedream 5 variant with built-in reasoning: interprets intent, handles example-based editing ('make it look like this reference') and domain-knowledge prompts (diagrams, infographics, product mockups) better than seedream-4/4.5. Takes an optional list of up to 14 reference images. Outputs 2K or 3K (no 1K, no 4K). Cheap price class — flat per image, below seedream-4.5 and far below seedream-5-pro. Pick this as the default Seedream model for reasoning-heavy or example-driven edits and stylized generation; for 4K use seedream-4.5; for photoreal people/products use nano-banana-pro.

  • prompt
  • image_inputs
  • size
  • aspect_ratio

Seedream 5 Pro

seedream-5-pro

Higher-cost ByteDance image generator and editor with the strongest prompt adherence and reference fidelity in the Seedream line. Takes an optional list of up to 10 reference images for multi-reference generation or precise editing. Outputs 1K or 2K only (no 4K). Premium price class — 2K costs twice as much as 1K and roughly 2–3x seedream-5-lite / seedream-4.5. Pick this only when the user explicitly asks for seedream-5-pro or when a multi-reference edit needs top-tier consistency at 2K; for general Seedream work use seedream-5-lite (cheaper) or seedream-4.5 (4K); for photoreal people/products nano-banana-pro remains the default.

  • prompt
  • image_inputs
  • size
  • aspect_ratio

Z-Image Turbo

z-image-turbo

Very fast, very cheap text-to-image model (6B, ~1 second per image). Text prompt only — NO reference or input images, no editing. Good general quality for drafts, mood boards, simple backgrounds, textures, and placeholder art; explicit width/height up to 2048x2048. Cheapest image option by a wide margin. Pick this for quick drafts, bulk variations, or simple backgrounds where cost and speed matter more than polish. For image-to-image edits or higher fidelity use nano-banana-pro.

  • prompt
  • width
  • height

Upscale video2

FLUX Video Upscale

flux-video-upscale

Premium FLUX super-resolution upscaler (1.5x–3x, output capped at ~4K). Only accepts mp4 sources of at most 20 seconds, 50MB and 2560x1440. Precise mode (creativity 0) stays faithful to the source and sharpens — use it for faces, products and real people; creative mode (creativity 1) restores and invents fine detail — use it for generated footage, textures and scenery. An optional prompt steers the added detail. Considerably more expensive than video-upscaler (billed per output megapixel-second). Pick this only for short clips where fidelity matters; otherwise use video-upscaler (the default).

  • file_hash
  • upscale_factor
  • creativity
  • prompt

Video Upscaler

default
video-upscaler

Cheap video upscaler and enhancer (BytePlus). Takes any video and outputs up to 4K at 24/30/60fps with frame interpolation when the target fps exceeds the source. Scene presets: 'aigc' for AI-generated clips (Seedance, Kling, Veo…), 'short_series' for short drama, 'ugc' for phone footage, 'old_film' for restoration, 'common' otherwise. No source length limit. Billed per second of output by resolution and fps band (4K/60fps costs ~16x 720p/30fps). Pick this by default and for any generated clip; for a faithful detail-preserving upscale of faces/products on a short clip use flux-video-upscale instead.

  • file_hash
  • target_resolution
  • target_fps
  • scene

Upscale image1

Image Upscaler

default
upscale-image

AI super-resolution upscaler (2x/4x). Enlarges images without the soft-blur look of bicubic resizing.

  • file_hash
  • scale

Remove video background1

Remove Video Background

default
remove-video-background

Removes the background from video frames, leaving only the foreground subject on a true transparent background (WebM VP9 with alpha). Useful for compositing without any chroma-key filter. Inputs up to 60 seconds.

  • file_hash

Remove image background1

Remove Image Background

default
remove-image-background

Removes the background from an image, leaving only the foreground subject with transparency. Useful for product photos and compositing.

  • file_hash

Lip sync1

Lipsync

default
lipsync

Re-syncs the mouth movement in a video clip to a separate audio track.

  • video_file_hash
  • audio_file_hash

Translate video2

ElevenLabs Dubbing

eleven-labs-dubbing

ElevenLabs dubbing: translates the speech of a video or audio file into one of 100+ languages (BCP-47 tags such as en, es-MX, pt-BR, fr-CA) while keeping every speaker's voice, emotion and timing; handles multiple and overlapping speakers and keeps background audio. Returns the dubbed AUDIO TRACK only (lossless FLAC) — not a re-muxed video: place it on the timeline over the original clip and mute the original audio. Billed per second of source media. Pick this when the user names ElevenLabs, needs a language the default model lacks, or wants the dubbed audio as a separate track; for a ready-made dubbed video use the default video-translate model.

  • file_hash
  • output_language
  • source_language
  • cloning_strength

Video Translate

default
video-translate

Translates spoken content in a video to another language with voice dubbing, preserving the original speaker's voice characteristics. Returns a new VIDEO with the dubbed track. Default for video_translate.

  • file_hash
  • output_language
  • mode

Text to speech1

ElevenLabs TTS

default
eleven-labs-tts

High-quality text-to-speech with natural-sounding voices. Choose from a large voice library for voiceover narration, dialogue, and audio content.

  • voice_id
  • prompt
  • stability
  • similarity_boost
  • style
  • use_speaker_boost
  • speed
  • seed

Voice changer1

ElevenLabs Voice Changer

default
eleven-labs-voice-changer

Changes the voice in audio/video to a target voice while preserving timing and emotional inflection. Ideal for consistent voiceover across clips.

  • file_hash
  • voice_id

Voice isolation1

ElevenLabs Voice Isolation

default
eleven-labs-voice-isolation

Isolates human voice from background noise, music, and other sounds. Use for audio cleanup and improving speech clarity in recordings.

  • file_hash

Transcribe1

ElevenLabs Speech-to-Text

default
eleven-labs-speech-to-text

Accurate speech-to-text transcription with word-level timestamps. Ideal for subtitles, captions, and transcript-guided editing.

  • file_hash
  • start_time
  • end_time

Generate music2

ElevenLabs Music

default
eleven-labs-music

Preferred default for music generation. Fixed-duration instrumental clips (5/10/15/20/30s) ideal for backgrounds, intros, outros, and short hooks. For full-length songs with vocals up to ~3 minutes use lyria-3-pro instead.

  • prompt
  • duration_seconds

Lyria 3 Pro

lyria-3-pro

Full-length song generation with vocals — strong for complete tracks up to ~3 minutes. Pick when the user wants a vocal track or a longer song; for short instrumental backgrounds use eleven-labs-music (the default).

  • prompt
  • images

Sound effects1

ElevenLabs Sound Effect

default
eleven-labs-sound-effect

Generates custom sound effects from text descriptions. Create whooshes, impacts, ambient sounds, UI sounds, risers, and any other audio effect. Duration 0.5–22 seconds.

  • prompt
  • duration_seconds

Need this list at runtime? The Models API returns the same catalog and each model's JSON Schema.