Model catalog
Every AI model available through the Rendley video API, with the parameters each one accepts.
Chips are the parameters each model accepts. Bold is required. A ↗ marks a parameter that takes either a publichttps URL or a library file hash, whatever the parameter is named.
No model matches that filter.
Generate video28
FLUX 3
flux-3PREMIUM multimodal video with synchronized audio. Unique strengths: storyboard keyframes — up to 10 images placed pixel-for-pixel (one = first frame, two = first+last, 3+ = evenly spaced keyframes) — and shot continuation from an existing mp4 (start_video, ≤15s ≤50MB; cannot be combined with images). 720p/1080p, 5–20s, wide aspect ratios including 21:9 and 2:1; draft mode gives a cheap 720p preview. Text/image variants cost ~2x kling-v2.6 and continuation (start_video) is very expensive (~2.5x that again). Pick this ONLY when a shot must hit exact keyframes or seamlessly continue an existing clip; for everything else use kling-v2.6 or seedance-2.0.
promptimages↗start_video↗aspect_ratioresolutiondurationgenerate_audiodraft
Gemini Omni 1.1
gemini-omni-1.1Google's fast multimodal video model with native synchronized audio (360p draft / 720p / 1080p / 4K, 16:9 or 9:16). Four input modes in one model: text-to-video; image-to-video from a start frame (`image`); first→last-frame interpolation (`image` + `last_frame`); reference-guided generation (`reference_images` — identity/style, NOT literal frames, wire a character design sheet here); and video-to-video EDITING (`video` — describe the change in the prompt, e.g. 'make the sky stormy, keep everything else', output length matches the input). The clip length has no numeric parameter, but you steer it in the PROMPT (state a target, up to ~10s max); the model follows it approximately, so read the real duration after. Billed per second of output, so 360p is very cheap and 4K is costly. Pick it for reference-driven character shots with audio, quick first→last-frame motion, or editing an existing clip; for a precise/exact duration or takes longer than ~10s use kling-v2.6/seedance-2.0, and for a single photo talking-head use omni-human-1.5.
promptimage↗last_frame↗video↗reference_images↗resolutionaspect_ratio
Grok Imagine Reference to Video
grok-imagine-r2vxAI's Grok Imagine reference-to-video model. Takes 1–7 reference images used as style and content references (not starting frames) plus a prompt, and generates a 1–10s clip at 480p/720p. Pick when the user wants Grok Imagine to follow the look of supplied reference images.
promptreference_images↗durationaspect_ratioresolution
Grok Imagine Image to Video
grok-imagine-videoxAI's Grok Imagine video model. Text-to-video, image-to-video (animate a starting image), or video editing (supply a short source video to restyle). Flexible 1–15s durations at 480p/720p. Niche pick when the user specifically wants Grok Imagine's look or its video-editing mode.
promptimage↗video↗durationaspect_ratioresolution
Grok Imagine Video 1.5
grok-imagine-video-1.5xAI's Grok Imagine Video 1.5 (preview) — image-to-video ONLY (`image` required) with synchronized native audio: background music, sound effects, ambience matched to the visuals, and short dialogue via an 'AUDIO:' section in the prompt. 1–15s at 480p/720p. Pick for animating a still (product showcase, portrait, character) when it should come with fitting sound in one pass; for text-to-video use grok-imagine-video.
promptimage↗durationaspect_ratioresolution
Grok Imagine Video Extension
grok-imagine-video-extensionxAI's Grok Imagine video-extension model. Takes a source MP4 (2–15s) and a prompt describing what happens next, and generates a 2–10s continuation from the last frame. Pick when the user wants to extend an existing video clip.
promptvideo↗duration
Hailuo 2.3
hailuo-2.3Stable, consistent motion with low artifact rates — fewer warps and flickers than larger models. Niche pick when motion stability matters more than peak fidelity.
promptfirst_frame_image↗durationresolutionprompt_optimizer
Kling V2.5 Turbo Pro
kling-v2.5-turbo-proFaster, cheaper Kling variant with first+last frame interpolation. Pick when speed or cost matters more than peak quality, or when the request specifically needs first+last frame interpolation.
promptstart_image↗end_image↗durationaspect_ratio
Kling V2.6
defaultkling-v2.6General-purpose video generation. Realistic motion, strong with people and natural scenes, generates audio by default. Pick this for any standard 5- or 10-second clip. For peak cinematic quality use veo-3.1 instead.
promptstart_image↗aspect_ratiodurationgenerate_audio
Kling V2.6 Motion Control
kling-v2.6-motion-controlSpecialized Kling variant for character motion transfer — takes a character image plus a motion reference video and produces a video where the character moves the same way. Pick only when the user explicitly wants to drive a character's motion from another video. Not a general-purpose video model.
image↗video↗promptcharacter_orientationmodekeep_original_sound
Kling 3.0 Omni
kling-v3-omni-videoKling's unified multimodal video model: text, first+last frame (`start_image`/`end_image`), up to 7 reference images for character/product identity, elements or style (refer to them as <<<image_1>>> in the prompt), a reference video for style/camera or as an edit base, and optional native audio. 720p (standard), 1080p (pro, default) or 4K; 3–15s; 16:9, 9:16, 1:1. Premium price class — roughly 2x seedance-2.0 per second (pro) and 4K far more; audio adds to the price. The Kling-family reference-image path: pick this when the shot needs reference-image identity with Kling's motion and prompt adherence, or as the retry when a seedance call fails (same inputs); for start-frame-only shots kling-v2.6 is cheaper, for the cheapest reference-image path use seedance-2.0-mini.
promptstart_image↗end_image↗reference_images↗reference_video↗video_reference_typekeep_original_soundgenerate_audiomodeaspect_ratioduration
Kling 3.0
kling-v3-videoKling 3.0 — text-to-video or first+last frame (`start_image`/`end_image`) with optional native audio (dialogue in double quotes, ambience, effects). 720p (standard), 1080p (pro, default) or 4K; 3–15s; 16:9, 9:16, 1:1. The leaner, cheaper sibling of kling-v3-omni-video: NO reference images and NO reference video — pick this for straight start-frame or text shots with Kling's motion and prompt adherence at a lower price; use kling-v3-omni-video when the shot needs reference-image identity or a reference video, and kling-v2.6 for the cheapest Kling start-frame path.
promptstart_image↗end_image↗generate_audiomodeaspect_ratiodurationnegative_prompt
LTX 2.5 Fast
ltx-2.5-fastCheapest fast video model with synchronized audio. Text-to-video or image-to-video with first-frame and optional last-frame conditioning, 16:9 or 9:16, 720p/1080p/2k/4k, 24/25/48/50 fps, durations 2–20s (over 10s only at 720p/1080p with 24/25 fps; 2k/4k and 48/50 fps cap at 10s). No reference-image identity mode. Pick this for drafts, volume work and quick previews where price and speed beat quality; for character consistency from reference images use seedance-2.0, for polished standard clips use kling-v2.6.
promptimage↗last_frame_image↗resolutiondurationaspect_ratiofpsgenerate_audio
MiniMax H3
minimax-h3MiniMax H3 — multimodal video: text-to-video, first/last-frame animation (`first_frame_image`/`last_frame_image`), and reference-guided generation (up to 9 reference images, 3 reference videos, 3 reference audio clips for character/motion/voice/style). 4–15s at 768P ($0.10/s) or 2K ($0.14/s); ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, or 'adaptive' for image-to-video. Strong all-rounder when the shot needs reference identity or audio-guided rhythm at moderate cost; for peak cinematic quality use veo-3.1, for the cheapest start-frame path use kling-v2.6.
promptfirst_frame_image↗last_frame_image↗reference_image_urls↗reference_video_urls↗reference_audio_urls↗durationresolutionratio
OmniHuman 1.5
omni-human-1.5Audio-driven digital human: one image + one audio clip → the person in the image speaks/performs the audio with lip-sync, facial expression and body motion. The `image` is the PERSON (a photo of the subject), not a scene start frame; no first-frame/reference pipeline needed. Optional prompt directs scene, emotion and camera. Pick this for talking-head, UGC presenter, spokesperson or avatar shots from a single photo with pre-made speech/TTS audio. Standard price class, billed per second of output; the video is exactly as long as the audio (audio must be ≤35s) and the caller must pass the audio length in `duration`. For generic scenes without a speaking subject use seedance-2.0-mini or kling-v2.6.
image↗audio↗promptfast_modeduration
P-Video
p-videoPruna P-Video — FAST, cheap video generation: text-to-video, image-to-video (first and optional last frame), and audio-to-video (condition on a music/voice track; the video follows the audio, e.g. a character singing) in one endpoint. 1–20s at 720p ($0.02/s) or 1080p ($0.04/s), 24/48 fps, with a quarter-price `draft` mode for rapid iteration. Pick for quick, budget drafts and audio-driven shots; for peak quality use veo-3.1 or kling-v3.
promptimage↗last_frame_image↗audio↗durationaspect_ratioresolutionfpsdraftprompt_upsamplingsave_audio
PixVerse V6
pixverse-v6PixVerse's flagship video model — cinematic text-to-video and image-to-video with optional synchronized audio (`generate_audio_switch`: BGM, SFX, dialogue), first→last frame transitions (`image` + `last_frame_image`), and a multi-shot mode (`generate_multi_clip_switch`) that renders a storyboard-style prompt as one sequence with scene transitions. 5/8/10/15s at 360p–1080p ($0.05–$0.18/s, audio adds a little). Pick for multi-shot cinematic sequences or stylized brand-film looks from a single call; for photoreal single shots prefer kling or seedance.
promptimage↗last_frame_image↗qualityaspect_ratiodurationnegative_promptgenerate_audio_switchgenerate_multi_clip_switch
Seedance 1.5 Pro
seedance-1.5-proStrong on expressive motion, dynamic choreography, and dance. Niche pick for motion-driven or stylized content, or when fine motion control matters more than realism.
promptimage↗last_frame_image↗durationresolutionaspect_ratiocamera_fixedgenerate_audiofps
Seedance 2.0
seedance-2.0The Seedance-family DEFAULT and the go-to reference-image video model: multimodal generation with native audio synced to dialogue/SFX, reference inputs (up to 9 images, plus reference videos and audio for lip-sync), flexible 1–15s durations, first+last frame interpolation, up to 1080p. **Slow generation — significantly higher latency than kling-v2.6 or veo-3.1.** Pick this whenever identity must be carried by reference images (character/product consistency from sheets), for motion transfer from a reference video, audio-driven lip-sync, or clips longer than 10s. For standard clips without reference images use kling-v2.6. Never upgrade to seedance-2.5 on your own — it is far more expensive and capped at 720p.
promptimage↗last_frame_image↗durationresolutionaspect_ratiogenerate_audioreference_images↗reference_videos↗reference_audios↗
Seedance 2.0 Fast
seedance-2.0-fastFaster, cheaper Seedance 2.0 variant with the same multimodal feature set (native audio, up to 9 reference images, reference videos/audio, flexible 1–15s durations, first+last frame interpolation) but capped at 480p/720p. Pick when iterating on Seedance multimodal work or when speed and cost matter more than peak resolution; for 1080p use seedance-2.0.
promptimage↗last_frame_image↗durationresolutionaspect_ratiogenerate_audioreference_images↗reference_videos↗reference_audios↗
Seedance 2.0 Mini
seedance-2.0-miniCHEAP Seedance — about half the price of seedance-2.0 with the same multimodal inputs (first+last frame interpolation, up to 9 reference images, up to 3 reference videos/audios, prompt-driven dialogue) and native synced audio, but 720p max (480p/720p only) and 4–15s. Pick this for volume work, drafts, iterations and any Seedance job where 720p is enough; for 1080p use seedance-2.0. Supplying reference videos moves it to a slightly higher pricing tier.
promptimage↗last_frame_image↗durationresolutionaspect_ratiogenerate_audioreference_images↗reference_videos↗reference_audios↗
Seedance 2.5
seedance-2.5PREMIUM, EXPENSIVE, 720p-MAX variant of Seedance — NOT a default, NOT an upgrade path, NOT a fallback for seedance-2.0. Use ONLY when the user explicitly asks for seedance-2.5, or when a single shot genuinely needs what 2.0 cannot do: one native take longer than 15s (up to 30s) or more than 9 reference images (up to 30 images, 10 videos, 10 audios). Otherwise identical feature set to seedance-2.0 (native synced audio, reference inputs, first+last frame interpolation) at several times the price and lower resolution (480p/720p only). **Slow generation — significantly higher latency than kling-v2.6 or veo-3.1.** Supplying reference videos moves it onto a much more expensive pricing tier (~4x). For everything else use seedance-2.0 (reference images) or kling-v2.6 (standard clips).
promptimage↗last_frame_image↗durationresolutionaspect_ratiogenerate_audioreference_images↗reference_videos↗reference_audios↗
Veo 3.1
veo-3.1Highest-quality video generation. Cinematic camera moves, realistic lighting, generated audio, and first+last frame interpolation. Recommended alongside kling-v2.6 — pick this when the user explicitly wants peak cinematic quality (hero shots, brand work) and the duration fits 4–8 seconds. More expensive than kling-v2.6.
promptimage↗last_frame↗durationaspect_ratioresolutiongenerate_audioreference_images↗
Veo 3.1 Fast
veo-3.1-fastFaster, cheaper Veo 3.1 variant. Slightly reduced quality. Pick when the user wants the Veo aesthetic but is iterating, or as a draft pass before committing to veo-3.1.
promptimage↗last_frame↗durationaspect_ratioresolutiongenerate_audio
Wan 2.7 I2V
wan-2.7-i2vImage-to-video variant of Wan 2.7. Animates a starting image with optional last-frame interpolation and synchronized audio. Niche pick for longer image-to-video clips than Kling supports, or when a custom audio track is needed.
promptfirst_frame↗last_frame↗first_clip↗durationresolutionaudio↗enable_prompt_expansion
Wan 2.7 T2V
wan-2.7-t2vText-to-video with synchronized audio generation tightly aligned to the prompt, and longer durations than kling-v2.6 or veo-3.1. Niche pick for longer clips with synced dialogue or action audio, or when attaching a custom audio track. For image-to-video use wan-2.7-i2v.
promptdurationresolutionaspect_ratioaudio↗enable_prompt_expansion
Wan 3.0
wan-3Cheap text-to-video only (no image, reference or audio inputs). Cinematic motion, 480p/720p/1080p, 2–15s in 1-second steps, optional negative prompt and prompt expansion; no audio track. Cheapest 1080p text-to-video in the catalog. Pick this when a shot is prompt-only and cost matters; for a first-frame image, audio or reference-driven shots use kling-v2.6 or seedance-2.0; for a single prompt-only take longer than 15s use wan-3-prime.
promptnegative_promptresolutionaspect_ratiodurationenable_prompt_expansion
Wan 3.0 Prime
wan-3-primePremium text-to-video only (no image, reference or audio inputs; no audio track) — the long single-take option: one continuous prompt-only shot from 2 up to 30s at 480p/720p/1080p. Costs ~2x wan-3 per second. Pick this ONLY when a prompt-only shot genuinely must be longer than 15s in one take; for 15s or less use wan-3, for image/reference/audio-driven shots use kling-v2.6 or seedance-2.0.
promptnegative_promptresolutionaspect_ratiodurationenable_prompt_expansion
Generate image15
DALL-E 3
dalle-3Strong prompt understanding for creative and illustrative work like ads and concept art. Niche legacy option — nano-banana covers most of the same ground.
promptaspect_ratio
Flux 1.1 Pro
flux-1.1-proStrong prompt adherence and layout control — handles complex multi-subject prompts and specific positioning. Niche pick when the prompt is unusually detailed and composition matters.
promptimage_prompt↗aspect_ratiowidthheightprompt_upsampling
Flux 2 Max
flux-2-maxBlack Forest Labs' highest-fidelity image model. Exceptional detail, photoreal textures, and prompt adherence; accepts up to 8 reference images for identity/product/style consistency; output selectable from 0.5 to 4 MP (max 2048x2048) or custom width/height. Premium price class — cost scales with output megapixels AND with every reference image, so it is one of the most expensive image options. Pick this only when the user explicitly asks for maximum image quality or for Flux. For everyday generation, edits and people/product detail nano-banana-pro is the cheaper default.
promptimage_inputs↗aspect_ratioresolutionwidthheight
GPT Image 2
gpt-image-2OpenAI's top image model — sharp text, precise edits, quality knob.
promptimage_inputs↗aspect_ratioqualitybackground
Grok Imagine Image 2
grok-imagine-image-2xAI's Grok Imagine Image 2.0 — general-purpose text-to-image with a single-image edit mode (pass `image` and the prompt describes the change; aspect ratio is then ignored). Output at 1k or 2k, wide aspect-ratio list including phone-screen ratios (19.5:9, 20:9) and `auto`, low/medium quality knob. Standard price per image; edits cost slightly more. Pick this when the user asks for Grok/xAI specifically or wants a 2k generalist alternative; for text-heavy layouts use qwen-image-3, and for the default use nano-banana-pro.
promptimage↗aspect_ratioresolutionquality
Imagen 4
imagen-4Photorealistic image generation. Strong on natural lighting, materials, and physical accuracy. Niche pick when nano-banana's photorealism is insufficient.
promptaspect_ratioimage_size
Nano Banana
defaultnano-bananaGeneral-purpose image generation. Fast, cheap, broad subject coverage, and good at image-to-image edits (change pose, restyle, add/remove elements). Preferred default for almost all image generation requests.
promptimage_inputs↗aspect_ratio
Nano Banana Pro
nano-banana-proHigher-quality variant of nano-banana. Noticeably better at faces, hands, fine detail, in-image text rendering, and photoreal product shots. Preferred when the request specifically involves people, text rendering, or fine product detail; otherwise nano-banana is the better default.
promptimage_inputs↗aspect_ratioresolution
Qwen Image 3
qwen-image-3Alibaba's Qwen-Image-3. Its specialty is accurate in-image TEXT rendering and dense multi-element layouts — posters, infographics, UI mockups, packaging and ads with real copy — plus photographic detail. Text-to-image or single-image editing (optional `image`; `match_input_image` keeps the source size). Fixed aspect ratios up to 2:1/1:2, negative prompt, optional prompt expansion (turn it off for exact text). Cheap (flat price per image, one of the lowest). Pick this when the image must contain legible, correctly spelled text or a busy layout; for plain photos or general edits nano-banana-pro is the default, and for the densest text/layouts use qwen-image-3-pro.
promptimage↗match_input_imageaspect_rationegative_promptenable_prompt_expansion
Qwen Image 3 Pro
qwen-image-3-proHigher-tier Qwen-Image-3 for dense, accurate TEXT rendering and complex multi-element layouts — infographics with many labels, multi-panel posters, UI screens, packaging with paragraphs of copy — with photographic detail. Same inputs as qwen-image-3: text-to-image or single-image editing (optional `image`; `match_input_image`), fixed aspect ratios, negative prompt, prompt expansion. Standard price (flat per image, slightly above qwen-image-3). Pick this when a lot of text or many layout elements must all come out right; for simpler text tasks qwen-image-3 is cheaper, and for general imagery use nano-banana-pro.
promptimage↗match_input_imageaspect_rationegative_promptenable_prompt_expansion
Seedream 4
seedream-4Stylized image generation. Strong on illustration, anime, and painterly looks. Niche pick when the user explicitly wants a non-photoreal aesthetic.
promptimage_inputs↗sizeaspect_ratiowidthheightenhance_prompt
Seedream 4.5
seedream-4.5Successor to seedream-4 with stronger spatial understanding, world knowledge, and in-image text rendering; keeps the stylized/illustrative strengths. Takes an optional list of up to 14 reference images for multi-reference generation and editing. Outputs 2K or 4K (or custom 1024–4096px dimensions); 1K is not supported. Standard price class (flat per image, 4K costs the same as 2K) — slightly above seedream-4. Third in the default fallback chain (nano-banana-pro → nano-banana → seedream-4.5). Pick this when the user wants the Seedream look with better layout/text accuracy or needs 4K output; for the cheapest Seedream use seedream-4; for photoreal people/products use nano-banana-pro.
promptimage_inputs↗sizeaspect_ratiowidthheightenhance_prompt
Seedream 5 Lite
seedream-5-liteCheap Seedream 5 variant with built-in reasoning: interprets intent, handles example-based editing ('make it look like this reference') and domain-knowledge prompts (diagrams, infographics, product mockups) better than seedream-4/4.5. Takes an optional list of up to 14 reference images. Outputs 2K or 3K (no 1K, no 4K). Cheap price class — flat per image, below seedream-4.5 and far below seedream-5-pro. Pick this as the default Seedream model for reasoning-heavy or example-driven edits and stylized generation; for 4K use seedream-4.5; for photoreal people/products use nano-banana-pro.
promptimage_inputs↗sizeaspect_ratio
Seedream 5 Pro
seedream-5-proHigher-cost ByteDance image generator and editor with the strongest prompt adherence and reference fidelity in the Seedream line. Takes an optional list of up to 10 reference images for multi-reference generation or precise editing. Outputs 1K or 2K only (no 4K). Premium price class — 2K costs twice as much as 1K and roughly 2–3x seedream-5-lite / seedream-4.5. Pick this only when the user explicitly asks for seedream-5-pro or when a multi-reference edit needs top-tier consistency at 2K; for general Seedream work use seedream-5-lite (cheaper) or seedream-4.5 (4K); for photoreal people/products nano-banana-pro remains the default.
promptimage_inputs↗sizeaspect_ratio
Z-Image Turbo
z-image-turboVery fast, very cheap text-to-image model (6B, ~1 second per image). Text prompt only — NO reference or input images, no editing. Good general quality for drafts, mood boards, simple backgrounds, textures, and placeholder art; explicit width/height up to 2048x2048. Cheapest image option by a wide margin. Pick this for quick drafts, bulk variations, or simple backgrounds where cost and speed matter more than polish. For image-to-image edits or higher fidelity use nano-banana-pro.
promptwidthheight
Upscale video2
FLUX Video Upscale
flux-video-upscalePremium FLUX super-resolution upscaler (1.5x–3x, output capped at ~4K). Only accepts mp4 sources of at most 20 seconds, 50MB and 2560x1440. Precise mode (creativity 0) stays faithful to the source and sharpens — use it for faces, products and real people; creative mode (creativity 1) restores and invents fine detail — use it for generated footage, textures and scenery. An optional prompt steers the added detail. Considerably more expensive than video-upscaler (billed per output megapixel-second). Pick this only for short clips where fidelity matters; otherwise use video-upscaler (the default).
file_hash↗upscale_factorcreativityprompt
Video Upscaler
defaultvideo-upscalerCheap video upscaler and enhancer (BytePlus). Takes any video and outputs up to 4K at 24/30/60fps with frame interpolation when the target fps exceeds the source. Scene presets: 'aigc' for AI-generated clips (Seedance, Kling, Veo…), 'short_series' for short drama, 'ugc' for phone footage, 'old_film' for restoration, 'common' otherwise. No source length limit. Billed per second of output by resolution and fps band (4K/60fps costs ~16x 720p/30fps). Pick this by default and for any generated clip; for a faithful detail-preserving upscale of faces/products on a short clip use flux-video-upscale instead.
file_hash↗target_resolutiontarget_fpsscene
Upscale image1
Image Upscaler
defaultupscale-imageAI super-resolution upscaler (2x/4x). Enlarges images without the soft-blur look of bicubic resizing.
file_hash↗scale
Remove video background1
Remove Video Background
defaultremove-video-backgroundRemoves the background from video frames, leaving only the foreground subject on a true transparent background (WebM VP9 with alpha). Useful for compositing without any chroma-key filter. Inputs up to 60 seconds.
file_hash↗
Remove image background1
Remove Image Background
defaultremove-image-backgroundRemoves the background from an image, leaving only the foreground subject with transparency. Useful for product photos and compositing.
file_hash↗
Lip sync1
Lipsync
defaultlipsyncRe-syncs the mouth movement in a video clip to a separate audio track.
video_file_hash↗audio_file_hash↗
Translate video2
ElevenLabs Dubbing
eleven-labs-dubbingElevenLabs dubbing: translates the speech of a video or audio file into one of 100+ languages (BCP-47 tags such as en, es-MX, pt-BR, fr-CA) while keeping every speaker's voice, emotion and timing; handles multiple and overlapping speakers and keeps background audio. Returns the dubbed AUDIO TRACK only (lossless FLAC) — not a re-muxed video: place it on the timeline over the original clip and mute the original audio. Billed per second of source media. Pick this when the user names ElevenLabs, needs a language the default model lacks, or wants the dubbed audio as a separate track; for a ready-made dubbed video use the default video-translate model.
file_hash↗output_languagesource_languagecloning_strength
Video Translate
defaultvideo-translateTranslates spoken content in a video to another language with voice dubbing, preserving the original speaker's voice characteristics. Returns a new VIDEO with the dubbed track. Default for video_translate.
file_hash↗output_languagemode
Text to speech1
ElevenLabs TTS
defaulteleven-labs-ttsHigh-quality text-to-speech with natural-sounding voices. Choose from a large voice library for voiceover narration, dialogue, and audio content.
voice_idpromptstabilitysimilarity_booststyleuse_speaker_boostspeedseed
Voice changer1
ElevenLabs Voice Changer
defaulteleven-labs-voice-changerChanges the voice in audio/video to a target voice while preserving timing and emotional inflection. Ideal for consistent voiceover across clips.
file_hash↗voice_id
Voice isolation1
ElevenLabs Voice Isolation
defaulteleven-labs-voice-isolationIsolates human voice from background noise, music, and other sounds. Use for audio cleanup and improving speech clarity in recordings.
file_hash↗
Transcribe1
ElevenLabs Speech-to-Text
defaulteleven-labs-speech-to-textAccurate speech-to-text transcription with word-level timestamps. Ideal for subtitles, captions, and transcript-guided editing.
file_hash↗start_timeend_time
Generate music2
ElevenLabs Music
defaulteleven-labs-musicPreferred default for music generation. Fixed-duration instrumental clips (5/10/15/20/30s) ideal for backgrounds, intros, outros, and short hooks. For full-length songs with vocals up to ~3 minutes use lyria-3-pro instead.
promptduration_seconds
Lyria 3 Pro
lyria-3-proFull-length song generation with vocals — strong for complete tracks up to ~3 minutes. Pick when the user wants a vocal track or a longer song; for short instrumental backgrounds use eleven-labs-music (the default).
promptimages↗
Sound effects1
ElevenLabs Sound Effect
defaulteleven-labs-sound-effectGenerates custom sound effects from text descriptions. Create whooshes, impacts, ambient sounds, UI sounds, risers, and any other audio effect. Duration 0.5–22 seconds.
promptduration_seconds
Need this list at runtime? The Models API returns the same catalog and each model's JSON Schema.