All posts
Model Launches

Seedance 2.5: What's Real, What Isn't, and How to Run It

ByteDance launched Seedance 2.5 on July 31, 2026 with no public API. The API opened on August 7 — 480p and 720p, 4 to 30 seconds, 50 references — and it's live on Flixly. Here's the confirmed spec sheet, and what's still unpublished.

August 3, 2026
Seedance 2.5: What's Real, What Isn't, and How to Run It

TL;DR

Seedance 2.5 is ByteDance's video model, launched July 31, 2026. Confirmed: up to 30 seconds of joint audio-video in a single generation (double Seedance 2.0's 15 seconds), and up to 50 references per job. The API opened on August 7 and settled two of the biggest open questions: output is 480p or 720p — not the 4K everyone attributed to it — and length is any whole second from 4 to 30. It is live on Flixly in text-to-video, image-to-video and reference-to-video. Still unpublished: frame rate, benchmarks, and any technical report.

ByteDance launched Seedance 2.5 on July 31, 2026, and within days the internet had assigned it a spec sheet ByteDance never wrote. The confirmed facts were narrower and more interesting than the headlines: a 30-second ceiling per generation that ByteDance describes as a single "one-take" pass, a 50-asset reference budget, and a conspicuous silence on resolution, frame rate and benchmarks.

Updated August 7, 2026: the API has since opened, and it settled the biggest open question in the least glamorous way — Seedance 2.5 outputs 480p or 720p, not the 4K that half the internet attributed to it. This post separates what ByteDance actually said from what got invented around it, and now marks which of the unknowns the API answered.

What is Seedance 2.5?

Seedance 2.5 is ByteDance's latest audio-video generation model, launched July 31, 2026 via the ByteDance Seed blog. It generates picture and sound together in a single pass rather than dubbing audio onto finished video.

It was first shown on June 23, 2026 at Volcano Engine's FORCE conference as an enterprise beta, with a public launch promised for early July. That slipped roughly five weeks. Coverage dated June 24 is next-day reporting on the FORCE unveiling, not a separate announcement. The two dates get cited interchangeably and make the timeline look messier than it is.

ByteDance frames the release around three claimed upgrades: 30-second single-segment output, up to 50 multimodal reference assets in one generation, and flexible local editing. The company positioned all three as industry firsts at FORCE. No comparison table, benchmark, or competitor survey was published to support that, and Seedance 2.5 does not appear on any independent video leaderboard as of August 3, 2026.

What are Seedance 2.5's confirmed specs?

These are the only figures traceable to a ByteDance-owned page:

  • Up to 30 seconds per generation. The Seed launch blog states "Up to 30 seconds per generation, with multi-round extensions." The model page states the same 30-second ceiling (the two pages differ on extensions, see below). Seedance 2.0's technical report documented a 4–15 second range, so the ceiling doubled.
  • Up to 50 reference assets: 30 images, 10 video clips, 10 audio clips in a single pass. The per-type breakdown is verbatim from the launch blog. The round number 50 is the sum, not a figure printed on ByteDance's English spec page.
  • Joint audio-video generation, inherited from Seedance 2.0's unified multimodal architecture. ByteDance's own wording is "building on" that architecture, so this is scale and control, not a new design.
  • Timestamp-level editing control, plus white-model control, green-screen, camera-perspective and reference-based editing. The model page also cites professional camera movement and performance blocking.
  • A staged rollout inside ByteDance's own China-market products, led by Jimeng AI and Doubao Pro, and gated behind paid tiers.

That is the complete official list from launch week. The API then added three more confirmed figures, which are facts about the shipped product rather than marketing copy:

  • Output is 480p or 720p. Two tiers, no 1080p, no 4K.
  • Duration is any whole second from 4 to 30, set per request. There is no extension parameter.
  • Billing is per token, where a job's token count is the output's pixel area times its duration times 24, divided by 1024, charged per thousand tokens — about $0.47 per second at 720p and $0.22 at 480p. Audio costs nothing extra; it is generated in the same pass.

Everything else circulating is still inference, extrapolation, or someone's SEO page.

Does Seedance 2.5 generate 4K video?

No — and this is now settled rather than merely unstated. The API offers exactly two resolutions, 480p and 720p. ByteDance's own launch blog and model page still state none, and the "native 4K" figure everywhere online belongs to a different model.

Native 4K with 10-bit color depth was announced as a Seedance 2.0 upgrade on June 23, 2026, at the same FORCE event where 2.5 was unveiled. Two announcements, one stage, and the coverage merged them. The tell is ByteDance's own "three world firsts" framing for 2.5: duration, reference count, local editing. Resolution is not on the list, which would be an odd omission if 4K were the headline.

The observable product surface argues against it too. Hands-on testing of the shipped Jimeng implementation reports only 480P and 720P available, and that review's own pricing example is quoted at 720P. Even sources asserting 4K concede this.

That prediction held. The 720p/1080p/4K ladder you'll see listed for 2.5 is Seedance 2.0's, carried over by writers who assumed continuity — and if 4K is what you need from this family, 2.0 is the version that has it.

What is Seedance 2.5's frame rate?

Unpublished, like most of the rest of the spec sheet.

Frame rate. No fps figure appears in any ByteDance source for 2.5, and the API exposes no frame-rate parameter either. There is one new piece of indirect evidence: the published billing formula multiplies by 24, the same constant ByteDance uses across the Seedance family. That is a statement about how a job is priced, not a promise about how the video is rendered, and it should not be quoted as a spec. Claims about 30 fps, 60 fps or frame interpolation remain unsupported — no frame-rate feature of any kind has been announced.

Architecture details. There is no technical report for Seedance 2.5. ByteDance Seed's research page lists it as a blog post only, with publications for 2.0 and 1.5 pro but nothing for 2.5. No parameter count, model size, training data or diffusion specifics have been disclosed. Any "720 frames via a Sparse DiT path" description is fabricated arithmetic built on the unconfirmed fps and resolution figures above.

Benchmarks. ByteDance published none. Seedance 2.5 is absent from Artificial Analysis's video leaderboards, where only Seedance 2.0 720p and 1.5 pro appear. Every quality claim in circulation is either vendor-asserted or a single reviewer's anecdote, and the demo clips are ByteDance's own selections rather than an independent sample.

How long can Seedance 2.5 videos actually be?

Thirty seconds in one pass. Anything longer is produced by chaining extensions onto existing output, not by generating a longer clip.

ByteDance's two pages disagree on how far that goes. The launch blog says "multi-round extensions." The model page says "with the option to extend twice." Neither states how many seconds each round adds — and the API, now that it exists, exposes no extension parameter at all. Whatever extension is, it is a feature of ByteDance's own apps rather than something a caller can invoke, so outside those apps a sequence longer than 30 seconds is built by composing separate generations. Reviewers have reported reaching roughly three minutes this way, and the "180-second beta long-video mode" phrasing comes from a marketing landing page that cites nothing.

The more useful finding is qualitative. One hands-on review found that extension preserves appearance but loses narrative state: the model knows what the previous segment looked like, but not where the plot had gotten to. A separate reviewer reached a three-minute sequence this way. Both can be true. Visual continuity holds, story continuity does not. If you are planning a two-minute narrative piece, that distinction decides your storyboarding approach.

How many reference images can Seedance 2.5 take?

Up to 30 images, alongside 10 video clips and 10 audio clips. The 50-reference figure is the largest documented change from 2.0, and the comparison you'll see stated is wrong in a specific way. Seedance 2.0's technical report documents 9 images, 3 video clips and 3 audio clips: 15 references, not 12. The circulating 12 counts images and video and silently drops the three audio slots. Corrected, the jump is roughly 3.3x, applied uniformly across all three modalities (9→30, 3→10, 3→10), which reads as a scaled version of the same conditioning mechanism rather than a rebuilt one.

More references does not mean better output. In the one published stress test of dense referencing, loading 30 images alongside video and audio produced a cross-modal semantic error: lemons in the reference set rendered as orange juice in the output, plus unrequested inserted content. Capacity and instruction adherence appear to trade off well before the cap. Similar guidance circulated for 2.0, where filling every slot over-constrained the model.

Two more things. The text prompt is the instruction channel and does not consume a reference slot. And ByteDance has published no per-file size limits, accepted formats, or whether the video and audio references share an aggregate duration budget. Seedance 2.0 capped reference video and audio at 15 seconds total each, so a similar limit on 2.5 is plausible but unconfirmed.

How does Seedance 2.5 compare to Sora 2, Veo 3.1, Kling 3.0 and Wan 2.7?

Partly, now. The API filled in resolution and pricing, so two of the four missing columns are answerable — and the answer on resolution is unflattering: 720p is the ceiling, where Sora 2, Veo 3.1 and Kling 3.0 all offer 1080p or better. What is still missing is any independent quality measurement: ByteDance published no benchmark and the model appears on no leaderboard, so "is it better" remains a question you settle by running the same prompt through each. The comparison that IS complete end to end from primary sources is against its own predecessor.

Spec Seedance 2.5 Seedance 2.0 Where the figure comes from
Max duration, one generation 4–30 seconds 4–15 seconds 2.5 API schema; 2.0 technical report
Extension past that ceiling "Multi-round" on the blog, "extend twice" on the model page Not compared here Two ByteDance pages, in conflict
Reference images 30 9 Same
Reference video clips 10 3 Same
Reference audio clips 10 3 Same
Total reference assets 50 15 Sum of the stated per-type caps
Output resolution 480p, 720p 480P, 720P, 1080P, 4K 2.5 API schema; 2.0 platform docs
Color depth Not published 10-bit, announced June 2026 Seedance 2.0 upgrade announcement
Frame rate Not published 24 fps in third-party docs for the 2.0 API endpoints No 2.5 source states any fps; the 2.0 figure is third-party, not ByteDance
Audio Joint audio-video generation Joint audio-video generation 2.5 launch blog, calling it inherited
Public API Live since August 7, 2026 — text-to-video, image-to-video, reference-to-video Live; 2.0, 2.0 Fast and 2.0 Mini listed Published endpoint schemas
Technical report None published Published on arXiv ByteDance Seed research page
Independent leaderboard Absent Listed on Artificial Analysis's text-to-video leaderboard, as Seedance 2.0 720p Artificial Analysis

The honest read: ByteDance positioned Seedance 2.5 as leading on the two axes it chose to compete on, duration and reference density, and on those two it does lead. It gave up resolution to get there — 720p against rivals' 1080p and 4K — and it still ships with no benchmark, no technical report and no leaderboard entry, so the quality claim remains untested by anyone outside ByteDance. Pick it when length, synchronized audio across a long take, or a large reference set is what the shot needs; pick something else when the deliverable has to be sharp at full screen. For side-by-side breakdowns of the families you can actually run, see the Seedance alternatives and Sora alternatives pages.

Can I use Seedance 2.5 right now?

Yes. The API opened on August 7, 2026 — a week after the model launched — and Seedance 2.5 is live on Flixly in all three shapes: text-to-video, image-to-video and reference-to-video.

The earlier version of this post warned that any vendor advertising a "Seedance 2.5 API" could not be reselling licensed access, because at the time none existed. That is no longer true, and the date those reseller blogs were circulating turned out to be right. Published endpoints and a real rate card now exist; the warning stands only for anything claiming 4K, a frame-rate control, or an extension parameter, none of which the API has.

What it costs. Billing is per token, and the formula is public: a job's token count is the output's pixel area times its duration in seconds times 24, divided by 1024, charged per thousand tokens. Two consequences are worth internalizing before you spend anything:

  • Length is linear. A 30-second clip costs six times a 5-second one. Iterate short.
  • Resolution is the big lever. 480p runs a little under half the price of 720p, so drafting at 480p and re-running only the prompt you like at 720p is the cheapest way to work.
  • A video reference adds its own length to the bill. Supplying reference video drops the rate to 0.6x, but the billed duration becomes your output plus the reference clip — so a 20-second reference driving a 6-second output is billed on 26 seconds, not 6. Image and audio references don't do this. It is the one place where the pricing genuinely surprises people.

On ByteDance's own surfaces, access remains metered and gated behind paid membership. One Chinese hands-on review reported a 30-second 720P generation costing around 780 Jimeng credits, on the order of ¥2 per generated second. Those figures reflect promotional launch pricing inside Jimeng and are not the API rate.

On where the model fails, the public record is still thin.

ByteDance has published no benchmarks, no error analysis and no stated limitations for Seedance 2.5. What exists is two hands-on reports: one found cross-modal errors and unrequested inserted content under dense referencing, the other found that extended segments lose narrative state.

What can you generate on Flixly right now?

Seedance 2.5 is live on Flixly, in all three modes, as of August 7, 2026 — the day its API opened. Pick it from the model selector in Text to Video, Image to Video or Reference to Video, set resolution and duration, and generate. For the practical walkthrough — prompt structure, which mode to pick, what each control does — see Seedance 2.5 on Flixly.

The Seedance 2.0 family is still there and still worth picking: 2.0, 2.0 Fast and 2.0 Mini reach 1080p and 4K, which 2.5 does not. Sora 2, Veo 3.1, Kling 3.0 and Wan 2.7 sit alongside them in the same selector.

How do you build a sequence longer than 30 seconds?

Thirty seconds is one shot, not one film — and since the API has no extension parameter, anything longer is composed rather than generated.

  1. Build your reference set first. Generate consistent character and product plates in the AI Image Generator, then feed them into Reference to Video to carry identity across shots.
  2. Fix the start and end states. First to Last Frame lets you set the first and last frame of each shot, so you decide where a shot begins and ends instead of relying on an extension to carry the story forward.
  3. Animate stills directly with Image to Video when you already have the frame you want.
  4. Finish in post. The AI Video Tools hub covers upscaling, background removal and effects, and Lip Sync Video handles dialogue when you need precise mouth timing rather than model-generated speech.
  5. Cut for distribution with the Shorts Generator and Auto Captions.

Thirty seconds in one pass is a real advance, and the reference budget is the most substantive change in the release. But a 30-second clip is still one shot. Finished content is a sequence of controlled shots plus sound and captions, and that pipeline is available today. Start in Text to Video — Seedance 2.5 is in the selector.

Frequently Asked Questions

What is Seedance 2.5?

Seedance 2.5 is ByteDance's latest audio-video generation model, launched July 31, 2026. It generates picture and sound together in a single pass rather than dubbing audio onto finished video, produces up to 30 seconds in one generation, and accepts up to 50 reference assets across images, video and audio.

Does Seedance 2.5 support 4K resolution?

No. The API offers 480p and 720p only. This was the single most-repeated claim about the model before it shipped, and it was wrong: native 4K was announced for Seedance 2.0 on June 23, 2026, at the same event where 2.5 was unveiled, and coverage merged the two. If you need 4K from this family, Seedance 2.0 is the one that has it.

How long can a Seedance 2.5 video be?

Any whole second from 4 to 30 in a single generation, double Seedance 2.0's 15-second ceiling. ByteDance's own pages also describe extending finished output — 'multi-round extensions' on the launch blog, 'the option to extend twice' on the model page — but the published API exposes no extension parameter, so that is an app-level feature on ByteDance's surfaces rather than something callable.

Is there a Seedance 2.5 API?

Yes, as of August 7, 2026. All three shapes are published — text-to-video, image-to-video and reference-to-video — and billing is per token: the pixel area of the output times its duration times 24, divided by 1024, charged per thousand tokens. In practice that works out to roughly $0.47 per second at 720p and $0.22 at 480p.

How many reference images does Seedance 2.5 accept?

Up to 50 reference inputs in total across images, video clips and audio — ByteDance's launch blog breaks that down as 30 images, 10 video and 10 audio. That is roughly a 3.3x increase over Seedance 2.0 across every modality. More references does not automatically mean better output: dense reference sets have produced cross-modal errors in testing.

Can I use Seedance 2.5 on Flixly?

Yes. It is live in all three modes — Text to Video, Image to Video and Reference to Video. Pick it from the model selector, set resolution and duration, and generate. Drafting at 480p costs less than half of 720p, so it's the sensible tier while you iterate on a prompt.

How does Seedance 2.5 compare to Sora 2 and Veo 3.1?

It leads on duration and reference density — 30 seconds in one pass and 50 references — and trails on resolution, since it tops out at 720p while its rivals offer 1080p and above. There is still no independent benchmark for it: ByteDance published none and it appears on no video leaderboard, so quality comparisons remain a matter of running the same prompt through each and judging the output yourself.

Tools mentioned in this post

seedance 2.5bytedanceai video generationseedance 2.0sora 2veo 3.1kling 3.0wan 2.7ai model news

Ready to create with Model Launches?

Jump straight into Flixly's AI studio and try model launches with 50+ models — free to start.