Seedance 2.5 for Product Video: Use Cases, Dos and Don'ts, and How We Use It
Seedance 2.5 is ByteDance's video model: 4 to 30 seconds in one generation, 480p or 720p with native audio and lip sync, and up to 50 reference inputs. This guide covers the specs and pricing as of September 2026, the use cases that work for online stores, the reference grammar that decides whether your product survives the render, the dos and don'ts we learned running it in production, and real examples from our own studio.
Key takeaways
- Seedance 2.5 generates a single continuous clip of any whole length from 4 to 30 seconds at 480p or 720p, with synchronized audio and lip-synced dialogue produced natively, from text, a start image, a start and end frame, or up to 50 reference inputs (30 images, 10 videos, 10 audio clips).
- Through the gateway we use it costs $0.138 per output second at 480p and $0.296 at 720p, so a 15-second 720p product ad is about $4.44 of compute before anyone edits it. ByteDance's own platform lists a 5-second 720p clip at $1.156.
- The model does not guess which reference is which. Every image has to be addressed by number in the prompt with a role and a negative ("@image1 defines the face and hair only, do not use her clothing"). Bare prose loses to pixels, and that one rule fixed our worst production incident.
- It invents a different presenter every run from the same faceless start frame, and pinning the seed does not change that. If a person has to look the same across shots, they have to be a reference image, not a description.
- Its weakest measured axis is scenery: it recomposes backgrounds it was told to keep. Its strongest is the camera move. Plan shots around that, keep on-pack text out of the clip or check it frame by frame, and add words as a separate layer after the render.
Seedance 2.5 is the video model behind most of the product clips and presenter ads our studio has rendered since August. We have paid for hundreds of its renders, judged them against the product photos they were supposed to match, and written down what worked and what did not. This is that list, with the vendor's specs and prices as of September 4, 2026, the use cases that make sense for a store, and the rules we would hand anyone about to send it their first product.
What Seedance 2.5 is
Seedance 2.5 is ByteDance's video generation model. It was announced at the Volcano Engine FORCE conference on June 23, 2026, released on July 31, and reached public API access in early August through ByteDance's BytePlus platform and a set of third-party gateways.
Three things set it apart from the previous generation. It generates a single continuous clip of any whole length from 4 to 30 seconds, where Seedance 2.0 topped out around 15 seconds and reached 30 only by stitching. It accepts up to 50 reference inputs in one request, against a ceiling of twelve files on 2.0. And every clip comes back with its audio already generated and in sync, including lip-synced dialogue when a spoken line is quoted in the prompt.
What it does not do is 4K. Native output is 480p or 720p at 24 frames per second. Some platforms offer 1080p and 4K as upscales of the 720p render. ByteDance upgraded 2.0 to native 4K with 10-bit output at the same time, so the two models now split: 2.5 for length, references and sound, 2.0 for resolution.
Specs
| Spec | Seedance 2.5 |
|---|---|
| Duration | Any whole second from 4 to 30, one generation, no stitching |
| Resolution | 480p or 720p, chosen per request (1080p and 4K on some platforms are upscales) |
| Frame rate | 24 fps |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, plus adaptive (inherits the input frame) |
| Modes | Text-to-video, image-to-video, first and last frame, reference-to-video, video edit, video extend |
| References | Up to 50 per request: 30 images, 10 video clips, 10 audio clips (video and audio each capped at 30 seconds total) |
| Audio | Generated natively and in sync; lip-synced dialogue from quoted lines; a voice reference matches timbre and accent |
| Prompt length | Up to 10,000 tokens on the gateway we use |
| Editing | Localized edits that redraw part of a frame while leaving the rest; extend from a last frame |
Pricing, and what a product ad costs
Prices below were captured on September 4, 2026.
| Where | 480p | 720p | Notes |
|---|---|---|---|
| BytePlus ModelArk (ByteDance) | not listed separately | $1.156 per 5-second 16:9 clip, about $0.23 per second | Also priced at $10.70 per million tokens without video input and $6.40 with |
| EvoLink gateway (what we use) | $0.138 per output second | $0.296 per output second | Text, image and reference modes; audio included; failed tasks not billed |
| EvoLink, video-input modes (edit, extend, video refs) | $0.084 per second | $0.180 per second | Input and output duration both billed |
So a 15-second 720p product ad is about $4.44 of compute on the gateway, or about $3.50 on ByteDance's own platform. A 30-second lifestyle film is double. A 6-second product-in-motion clip for a product page is under $2. Those are the raw model costs before any of the work around them: the references, the prompt, the quality check, the retakes, the words on top, the end card.
That framing matters because the model is not the expensive part of a good product video. The retakes are. Everything in the dos and don'ts below exists to cut retakes.
Use cases that work for a store
Product in motion. A slow push-in or a quarter turn on a hero product from a single catalogue photo. This is image-to-video: the clip begins on your exact photo and moves from there. It is the cheapest, safest use of the model and it turns a static product page into one that moves.
Presenter ads with a cast member. A person holding, wearing or using the product, speaking a line, in a location you chose. This is reference-to-video with a portrait reference, one or more product references and a spoken line in the prompt. It is the use case that replaced a studio day for the boutiques on our platform, and it is also the one with the most ways to go wrong, which is why most of the rules below are about it.
Lifestyle films of 15 to 30 seconds. Several shots in one generation, planned as a timecoded shot list. Golden-hour exteriors, a room, a table, three products across four beats. The 30-second single generation is what makes this possible without stitching.
Rooms and spaces from one still. Image-to-video on a photo of a room, with the products composited in beforehand. In our pilots a bare-room 8-second arc held the merchant's exact room through the move, and a decorated frame held the room plus three placed products.
Voice-matched narration. One audio reference makes the generated speech match that voice's timbre, pitch and accent, natively lip-synced. A founder can narrate every ad without recording every ad.
Native speech in the shopper's language. The model speaks the quoted lines in the prompt's language rather than dubbing over them, which for a store selling into several countries is the difference between one ad and five.
Examples from our studio
Every clip below was rendered with Seedance 2.5 from a merchant's own product photos, on stores that use our platform and agreed to show their work. Words and end cards are a separate layer added after the render. Tap a clip to play it; sound on for the native audio.
How each one was made, left to right:
- The boutique ad is image-to-video. Each shot begins on a frame we composed at 9:16 from the product photos, so the gown, the bag and the heels are the merchant's exact items. The headline and end card are overlaid after the render.
- The lifestyle film is one 30-second generation from a five-shot timecoded script, with the dress supplied as references and the town described in the prompt. The ambient sound is the model's own.
- The presenter ad is reference-to-video: her portrait first, then the product references, each addressed by number with a role and a negative, and the spoken line quoted in the prompt so the lip sync is native.
- The credenza ad is product in motion for a furniture store: a close-up of the doors and brass handles pulling out to the room, from the catalogue photos alone.
- The brand film is a 16:9 generation across five rooms. The products hold because they were references; the rooms are the model's, which is the trade the rest of this guide is about.
The reference grammar decides everything
If you take one thing from this guide, take this. Seedance 2.5 numbers every reference you upload and expects the prompt to address each one by that number, with a role and, ideally, a negative.
ByteDance's own guidance is explicit that the model will not reliably work out on its own that image 3 is the location and image 4 is the jacket. You have to say so in the prompt. Their own example reads, in substance: image 1 defines the courier's face and hair only, do not use her clothing or the background.
We learned this the expensive way. On August 27 a run of presenter clips came back with the cast member wearing the garment from her own portrait photo instead of the product, with perfect product references sitting in the request. The references were right. The prompt described them in ordinary prose ("the second and third images show the dress") and ordinary prose lost to pixels. The fix was to write every reference by number, in the model's grammar, with a role and a negative:
@image1 defines the face, hair and build only. Do not use her clothing or the background. @image2 through @image4 show the product to match exactly, the same item from three angles, one product, never a second copy. Product photos are not casting: if a person is visible in one, take only the item, never a face, hair, skin, body or pose.
Three details inside that block are load-bearing. The portrait goes first, because the identity clause anchors on it. A multi-view set is declared as one object, because ByteDance's guidance says a set of four views can otherwise render as four objects. And the casting ban is unconditional, because a catalogue photo with a model in it will otherwise recast your ad.
There is a ceiling. In our measurements the ordinal binding stops holding at around 20 references; with three bags in twenty images, the model read the printed words on the bags more reliably than it read the numbers in the prompt. The API accepts 30 images. We cap ourselves well below that and warn past 15.
Dos
Front-load the subject and the action. The first 20 to 30 words carry the most weight; the model locks in the subject and core action from the opening. Subject, action, environment, camera, style, constraints, in that order, in 60 to 100 words for a single shot.
Use one strong lighting word. Lighting has the largest single effect on the look of the clip. One precise lighting description beats ten adjectives.
Write a timecoded shot list for anything over 10 seconds. "0 to 3s: push in on the bag on the counter. 3 to 8s: she lifts it and turns to the window. 8 to 12s: close on the clasp." Four to six shots in 30 seconds is the practical maximum; more than that and the shots blur.
Compose the start frame at the delivery aspect before you submit. In image-to-video the clip inherits the aspect of the start image. A square catalogue photo forced into 9:16 loses the top and bottom of the product. Build the frame at 9:16 first, then animate it.
Re-read the aspect list off the gateway's own error. The accepted aspect ratios moved under us once: 4:5 rendered on August 14 and was rejected on August 26, while 4:3, 3:4 and 21:9 appeared. When a render fails with an aspect error, the error message is the current truth.
State negatives as positive instructions when the gateway has no negative channel. Some gateways expose no negative prompt field, so "no jump cuts" arrives as the tokens "jump cuts". Write what you want instead: "one continuous take, steady camera."
Respell rare words phonetically in spoken lines. The speech branch reads the line through a language model and swaps unfamiliar words for a likely neighbour. A nursery's "pawpaw" came out as "pom-pom" twice, and "paw paw" as "palm". Writing "Pah-pah" in the spoken line only, not on screen, made it say the word. Spelling tricks that can be read as a common word will fail; a spelling that cannot be lexicalized wins.
Check the clip against the reference photos, frame by frame, before it ships. Colour, print, hardware, logo position, the number of items. This is where invented details hide, and it is the check that our fidelity gate runs on every cast render before a merchant sees it.
Don'ts
Do not expect the same invented presenter twice. From one faceless start frame with byte-identical arguments, four renders produced four different women in four different rooms. Pinning the seed to the value the first run returned did not hold the identity. If someone has to look the same across shots, they are a reference image, addressed by number. A description is not an identity.
Do not send product photos that contain a person without saying what to take from them. The catalogue model in your product photo is one instruction away from becoming the presenter in your ad.
Do not send ultra-wide reference composites. References wider than roughly 2.2:1 were rejected on our gateway with a misleading "timed out fetching your input media" error. A side-by-side sheet of a square and a landscape photo failed three submits; the same content as a single near-square image passed instantly. Stack references vertically past a 2:1 aspect.
Do not trust it with small on-pack text. It held every string on two product stills in our bake-off and rewrote a "500G" weight line on a third. Treat printed text as intermittent: keep it large, keep it out of the shot, or check it in every frame.
Do not ask it to preserve scenery you care about. Scene fidelity was its weakest measured axis. A terrace still was recomposed to a wider view with the pots gone; a town became open sea. Product held; place did not. If the location matters, make it the reference and address it by number, and expect drift anyway.
Do not put words in the frame. Ask for the scene text-free and add the headline, the price and the call to action as a layer afterwards. It renders more reliably, and you can change the offer next season without paying for another render. That is how the ad above was built.
Do not exceed 30 seconds in one request. Longer stories are two generations and a cut. The video extend mode exists for continuing from a last frame, at the lower video-input rate.
Do not script customer testimonials for a synthetic presenter. A presenter-style line about the product is fine. A first-person line about a year of use is an endorsement under the FTC's guides, and a person who does not exist cannot honestly give one.
The best way to succeed: a workflow
- Write the brief first: one product, one claim, one call to action, one platform, one aspect ratio. The aspect changes every shot.
- Collect references: three to six angles of the product from your own catalogue, the portrait and full-body photo of the presenter if there is one, and a photo of the location if the location matters. Fewer, better references beat many.
- Write the reference block by number, portrait first, with a role and a negative for each, and declare multi-view sets as one object.
- Write the shot list with timecodes, four to six shots for 30 seconds, one for anything under eight.
- Render once at 480p to check blocking, product fidelity and the spoken line. It costs less than half of 720p and it will show you the problems.
- Fix the references or the prompt, not the seed. Re-render at 720p.
- Hold the final against the product photos frame by frame. Then watch it muted.
- Add words, captions and the end card as a layer. Export to the platform's published specs.
How Obsess AI uses it
Inside our platform, Seedance 2.5 is one camera in a bag of cameras, and the agent picks the camera for the job. Every camera lives in a single registry with its verified vendor facts (the aspect list, the reference caps, the duration bounds), our own measured quality on real merchant products (subject hold, camera control, scene fidelity, human fidelity, physics, small text) and its price per output second. A camera with no measured price cannot be routed to. A camera whose vendor moves an enum is re-verified from the error and the registry updated, not the other way round.
Two Seedance 2.5 routes are active. Reference-to-video is the default motion camera for our AI photographer when a clip includes a person or several product angles; the grammar above is generated by one function that every render passes through, so a fix there fixes every surface. Image-to-video drives the ads pipeline and room films from a composed start frame. Both submit at 720p for deliverable parity, both bill on the exact output seconds the gateway confirms, and every render lands in the same cost ledger as every other AI call we make.
Two more things sit around the camera. A product-fidelity gate compares cast video frames to the product references and records a verdict; a frame that fails is refunded to the merchant but still shown, with the reason, because hiding a failure is how you lose the right to run unattended. And words are never in the frame: the headline, the offer and the end card are a separate layer, so a merchant can restyle the sale text without another render. The video ad guide walks through the brief-to-export side of that; the UGC studio page shows what the merchant sees.
Seedance 2.5 versus 2.0
Choose 2.5 for anything with people, speech, several references or a shot longer than about 10 seconds. Choose 2.0 when you need a 4K master and the shot is short, silent and single-reference. Most store work is the first kind. If you are choosing an AI video tool rather than a model, our guide to AI product video covers the questions to ask before the model even comes up.
Update log
September 4, 2026: Specs and prices from the fal and EvoLink model pages and the CineD launch report on this date. Dos and don'ts from production renders between August 12 and August 28, 2026. We will update this post when ByteDance changes the reference caps, aspect list or pricing, all of which have moved once already.
Frequently Asked Questions
What is Seedance 2.5?
Seedance 2.5 is ByteDance's video generation model, announced at the Volcano Engine FORCE conference on June 23, 2026 and released on July 31, 2026, with public API access following in early August. It generates a single native clip of 4 to 30 seconds at 480p or 720p with synchronized audio, supports text-to-video, image-to-video, first-and-last-frame, reference-to-video, video edit and video extend modes, and accepts up to 50 reference inputs in one request.
How much does Seedance 2.5 cost?
It depends on the platform. ByteDance's BytePlus ModelArk lists a 16:9 five-second 720p clip at $1.156, roughly $0.23 per second, and prices the API at $10.70 per million tokens without video input and $6.40 with. Gateways resell it per output second: the one we use charges $0.138 per second at 480p and $0.296 per second at 720p for text, image and reference modes, with audio included and failed tasks not billed. Prices captured September 4, 2026.
How do references work in Seedance 2.5?
You pass up to 30 images, 10 video clips and 10 audio clips as ordered lists, then address each one in the prompt by its position: @image1, @image2 and so on. ByteDance's own guidance says the model will not reliably work out on its own that image 3 is the location and image 4 is the jacket; you have to say so, and you should say what not to take from each reference as well. A single audio reference voice-matches the generated speech.
Does Seedance 2.5 support 4K?
Not natively. Seedance 2.5 outputs 480p and 720p. Some platforms offer 1080p and 4K as upscales of the 720p render. ByteDance upgraded the older Seedance 2.0 to native 4K with 10-bit output in parallel, so if a 4K master matters more than 30-second single-shot length and 50 references, 2.0 is the one to look at.
Can Seedance 2.5 keep the same person across several clips?
Only if the person is a reference image. In our tests, four renders from the same faceless start frame with identical arguments produced four different women, and pinning the seed to the value the first run returned did not change that. Reference-to-video with a portrait reference addressed by number is the mechanism for continuity; a text description is not.
Is Seedance 2.5 good for product videos?
Yes for product-in-motion and presenter ads when the product is supplied as references and named in the prompt. It held complete products through camera moves in our tests when told what each reference was. Its weak points are scenery fidelity, which it tends to recompose, and small on-pack text, which it sometimes rewrites. Check every clip against the reference photos before it ships.
Related Articles
Keep exploring
Go deeper on the topics in this article with related guides, free tools, industry playbooks, and competitor comparisons.
Related guides
Free tools
How we compare
Sources & references
Primary documentation referenced for the technical claims on this page. We do not link out to competitor products or affiliate content; these are the standards bodies and platform docs the guidance is built against.
- fal: Seedance 2.5 model page ↗Vendor-platform listing of the model's modes, 4 to 30 second durations, 480p and 720p resolutions, seven aspect ratios, reference limits and native audio. Read September 4, 2026.
- EvoLink: Seedance 2.5 API pricing and model ids ↗The gateway our production renders run through: the five model routes, per-second pricing at 480p and 720p, reference caps, audio defaults and request limits. Captured September 4, 2026.
- CineD: ByteDance Seedance 2.5 API goes live ↗Trade-press report on the API launch: 30-second single-shot generation, 50 multimodal references, 3D white-model camera blockouts and localized editing, with the Seedance 2.0 comparison.
- FTC: Endorsement Guides ↗The rules that apply when a synthetic presenter delivers lines that read as a customer experience.
Ready to Automate Your Content Marketing?
Let Obsess AI write SEO-optimized blog posts for your Shopify store.