ByteDance Seedance 2.5
One-take video storytelling with native audio and timestamp-level editing
About model
Seedance 2.5 is ByteDance's new-generation video creation model, built on the unified multimodal audio-video joint-generation architecture of Seedance 2.0 and centered on turning clips into complete creative works. It generates 30-second audio-video clips in a single pass, doubled from 15, organizing multiple logically connected shots into a story with setup, development, and resolution, and supports multi-round extension that keeps characters, environments, and pacing consistent across multi-minute videos. Reference capacity grows to 30 images, 10 video clips, and 10 audio clips per pass, joined by clay render, motion, and creative referencing that let textureless 3D models drive composition, blocking, camera moves, and physically consistent lighting. Timestamp-level control directs narrative, camera, and rhythm for specific time ranges and supports targeted edits after generation, alongside green screen, camera perspective, and reference-based editing. Available on Together AI.
30s
Multi-round extensions build multi-minute stories with consistent characters and pacing
50
Up to 30 images, 10 video clips, and 10 audio clips in one pass
Timestamp
Direct narrative, camera, and pacing per time range, and edit specific moments after generation
- One-Take Storytelling: 30-second single-pass clips structured as complete stories, extendable round by round into multi-minute videos with consistent subjects
- Multimodal Referencing: Up to 30 images, 10 videos, and 10 audio clips per pass, plus clay render, motion, and creative references that control composition, blocking, and lighting
- Timestamp-Level Editing: Per-time-range control during generation and targeted edits after, with green screen, camera perspective, and reference-based editing
- Production-Ready Infrastructure: 99.9% SLA, available on Together AI serverless infrastructure
API usage
Endpoint:
Model card
Architecture Overview:
• Unified multimodal audio-video joint-generation architecture, carried forward from Seedance 2.0
• Single-pass generation of up to 30 seconds with synchronized audio, extendable across multiple rounds
• Accepts text, image, video, and audio inputs, including textureless clay renders that define spatial structure, motion paths, and camera angles
Training Methodology:
• Built by the ByteDance Seed team as the successor to Seedance 2.0, centered on foundational generation and reference-based generation
• ByteDance reports systematic optimization of object textures, skin and eye detail, lighting, and color, along with fewer uncontrolled subtitles and background music in outputs
Performance Characteristics:
• ByteDance reports smoother shot transitions and scene changes, with subjects stable across cuts and audio staying in sync through long-form videos
• Clay render referencing produces lighting that follows physical laws, including source direction, color temperature, intensity, and shadow projection
• Green screen editing preserves the subject while adapting clothing movement, hair, gait, and lighting interaction to the new environment
Prompting
Together AI API Access:
• Access Seedance 2.5 via Together AI APIs using the endpoint ByteDance/Seedance-2.5
• Authenticate using your Together AI API key in request headers
• Create a video request with a text prompt, optionally combining image, video, audio, and clay render references, then poll until the video completes
• Reference materials support up to 30 images, 10 video clips, and 10 audio clips per generation, addressable in prompts
• Direct specific time ranges in the prompt to control narrative, camera, and pacing, or extend an existing output with a follow-up request
• Available on Together AI serverless infrastructure
Applications & use cases
Film & Advertising Production:
• Produce one-take sequences with multiple shots, then extend them into multi-minute pieces
• Block complex scenes with clay render references so composition and camera match the storyboard
• Replace backgrounds with green screen editing while keeping subjects physically consistent
Education & Training Content:
• Turn lessons, historical events, and experimental procedures into vivid instructional video
• Keep recurring characters and settings consistent across a series of clips
• Localize and customize teaching materials by swapping references rather than reshooting
Simulation & Synthetic Data:
• Generate synthetic video for training perception and manipulation systems
• Simulate long-tail driving scenarios such as extreme weather and complex traffic
• Produce industrial process demonstrations and equipment walkthroughs from reference material
- TypeVideo
- Main use casesText-to-VideoImage-to-Video
- Resolution/DurationUp to 30s per generation
- DeploymentServerless
- Endpoint
- Price
$0.115 / video
- Input modalitiesTextImageVideoAudio
- Output modalitiesVideo
- ReleasedJuly 30, 2026
- CategoryVideo
