SUNDAY, AUGUST 2, 2026|No. 9817
Technology · AI · China

ByteDance Launches Seedance 2.5 with 30-Second Video Generation and Multi-Reference Support

ByteDance's Seedance 2.5 extends video generation to 30 seconds and adds multimodal references, with hands-on tests showing improved realism but lingering moderation and control issues.

Seedance 2.5 generates up to 30 seconds of native video with multimodal reference support.
Seedance 2.5 generates up to 30 seconds of native video with multimodal reference support.
2 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
1 countries
Related coverage

Seedance 2.5 Is Here: Can ByteDance Beat Itself?

Hands-on with Seedance 2.5: longer and more stable, but the details still fall short.

Text | AIX Finance, Author | Lei Jing, Editor | Jin Yufan

On July 31, Seedance 2.5 arrived.

The core parameter changes in this upgrade: single-generation video duration extended from 15 seconds to 30 seconds of native output, with support for multi-round extension; up to 50 reference materials per generation, covering three modalities: images, videos, and audio; enhanced local video editing capabilities, supporting local detail changes and adding or removing elements, with timestamp instruction precision controlled to within 1 second; support for directly referencing green-screen and white-model materials, and invocation through Maya and Blender plugins; coverage of more than ten languages, with lip-sync and speaking rate matching.

From the direction of the upgrade, ByteDance wants to push the Seedance model's capabilities up another level, toward more professional production scenarios such as film and television, advertising, and gaming.

On February 7, Seedance 2.0 began closed beta testing and generated widespread discussion on social media, with many users giving positive feedback. But as the hype faded, problems gradually surfaced: characters' skin had a waxy, greasy look, expressions were exaggerated and obviously AI-like, and characters looked like they all came from the same mold.

So, is Seedance 2.5 actually good? We tried Seedance 2.5 through Jimeng AI and set it six tasks for hands-on testing, covering single-person long video, multi-person fights, and multimodal material integration.

First, impressions. Seedance 2.5's realism has clearly improved: the waxy feel on faces has been reduced, skin texture is closer to real, and expressions no longer feel as overdone as before.

Prompt length also directly affects the experience. The longer and more detailed the prompt, the slower generation tends to be, and the more likely it is to trigger moderation rejection. Shorter prompts generate noticeably faster and have a higher pass rate.

One-shot output stability has also improved. Compared with the past, when you often had to “pull cards” repeatedly to get a usable video, Seedance 2.5 is indeed more controllable. In most scenes, one or two generations are enough to get a passable result.

Next, let's look at its actual performance in detail.

01. Single-person video: Better skin quality, and better acting

This round tests Seedance 2.5's most basic and core upgrade: 30-second native output and the quality of the character's performance in the video.

We set up a station farewell scene, asking for a departure with the character conveying reluctance and a restrained, suppressed emotion. This kind of static emotional performance is a real test of the model. Nearly all information is delivered through subtle facial expressions. Skin texture, slight redness around the eyes, tiny movements of the lips — if any detail is wrong, viewers are quickly pulled out of the scene.

There was an interruption during the test. The prompt was rejected by moderation on the first submission, even though the content didn't involve obvious sensitive factors. Similar situations occurred in later scenarios, and we only got through after trimming the wording. For creators who need to describe character states, camera language, and scene atmosphere in detail, this kind of moderation can disrupt the creative rhythm.

The final output did pass the quality bar. The most direct improvement is skin texture. The waxy and oily feel that was widely criticized in Seedance 2.0 is noticeably reduced in Seedance 2.5's videos. Station lights falling on the face produce natural light and shadow changes, and pores and skin texture are closer to real people.

Emotional expression has also improved. The over-the-top performances common in the Seedance 2.0 era — wide eyes, furrowed brows — are absent. The character's reluctance is conveyed mainly through the eyes and subtle facial changes, and the overall result basically matches the restrained, suppressed emotion required by the prompt.

More noteworthy than the realism is that Seedance 2.5 maintained the character's appearance, scene lighting, and emotional state across 30 seconds. Previously, generating longer content usually meant generating multiple short clips and stitching them together, and problems like face changes, lighting jumps, and emotional resets often appeared between clips. Thirty-second native output reduces the cost of cross-segment generation and post-repair. At least in single-character, low-intensity action scenes, doubling the duration didn't bring an obvious drop in stability — that is the more practical value of this upgrade.

02. Multi-person video: No face changes, and smooth action

With single-person performance passed, let's add one more person.

We set up a two-person confrontation in the Northern Song Dynasty: a swordsman in red and a blade master in black meet on a narrow path, then fight. This scenario tests two core things: first, character consistency — after camera cuts, the two people's clothing, faces, and weapons must not change; second, the plausibility of physical interaction — when weapons clash, there must be no clipping, and body movements must not violate the laws of physics.

The prompt described camera movement, scene, lighting, and character settings in considerable detail, so it was relatively long. The first submission was rejected by moderation again, and it passed only after we cut and adjusted the wording.

The generated video basically met the requirements. The two protagonists maintained good consistency across multiple shots, and skin texture was similarly closer to real. The tension of the standoff was conveyed mainly through eyes and subtle expressions, not through stylized wide-eyed anger. The fight was smooth overall, with no obvious stuttering, freezing, or stalled movements.

This round shows that in multi-character scenes, Seedance 2.5 can handle both character consistency and action plausibility: the characters don't “change face” as the camera cuts, and the fight doesn't fall apart under complex interaction.

03. Video extension: The footage connects, but the narrative doesn't

Seedance 2.5 can continue extending an existing video, so we chose the two-person fight video from the previous round and asked the model to extend it by 10 seconds, continuing the story while preserving the characters' appearance, visual style, and narrative rhythm.

Video extension tests not only whether the model can keep generating footage, but also whether it can understand the plot progress and emotional state of the previous video.

The problems in the output were clear. The original video ended with the two characters back-to-back, a few steps apart. The extended portion didn't move the action forward; instead, it held that static pose for too long, causing a clear break in rhythm at the seam. New fighting then appeared, but it lacked emotional and action transitions from the tense standoff, feeling more like a new segment stitched onto the original.

Visual continuity was better. The two characters' looks, clothing colors, and weapons remained largely consistent, and the visual style didn't change noticeably.

Video extension involves two levels of continuity: visual continuity of characters, scenes, and style, and narrative continuity of action, emotion, and causality. This round shows that Seedance 2.5 basically achieves the former but not the latter. It can recognize what the previous video “looks like,” but it still can't fully understand where the story “has gotten to.”

04. White-model reference: Accurate spatial restoration, good news for games and film

This round tests a new Seedance 2.5 capability: using a white-model video as a reference to produce a finished video.

In simple terms, a white model is a 3D previsualization video with only basic geometric structure, no materials or lighting. Seedance 2.5's job is to preserve the original spatial structure and object positions while adding materials, lighting, and environmental detail to turn it into a video.

This capability solves the problem of prompts being unable to describe spatial relationships precisely. Instead of repeatedly telling the model about camera movement and position, creators can use a white model to establish object positions and camera paths, then have the model generate a video in the desired visual style.

We uploaded a white-model preview of a city scene and asked Seedance 2.5 to generate a “ruins” video showing a future sci-fi city after a disaster. The original white model had only building masses and road outlines; the model had to fill in the atmosphere and visual details itself.

The spatial restoration turned out well. The positional relationships between buildings were largely preserved, with no obvious building misalignment or perspective errors. Seedance 2.5 gave the buildings a ruined feel: a dim, gray palette and a sense of worn-down decay.

For game and film/TV pre-production, this capability is especially valuable. Whether for level scenes, storyboard previews, or advertising shot proposals, creators can first use a white model to set the space and camera movement, then quickly generate visual previews to check shot effects and overall style.

05. Multimodal reference: More materials, but understanding still isn't accurate enough

Seedance 2.5 expands the number of reference materials per generation, supporting up to 50 materials across three modalities: images, videos, and audio.

The 50-cap is the total when different types are combined. Specifically, up to 30 images are supported, with resolution no higher than 4K and each no larger than 30MB; up to 10 videos, each no larger than 200MB and no longer than 30 seconds, with total video length capped at 30 seconds; up to 10 audio clips, each no larger than 15MB and no longer than 30 seconds, with total audio length also capped at 30 seconds. So although the number seems high, there are quite a few practical limitations.

Since we wanted to test multimodal material integration, we maxed it out. We gave Seedance 2.5 30 images plus 30 seconds of video and audio to generate a beverage ad video.

This round mainly examined whether the model could simultaneously understand the role of each material. We asked it to extract art style and product from the images, reference the scene and shots from the video, and use the provided background audio.

The output basically met the requirements. The visuals carried the illustration style of the reference images, and the scene and shots broadly matched the reference video. But there was a mismatch: the fruit being picked in the footage was a lemon, while the final product was orange juice.

This shows that Seedance 2.5 can extract art style and scene features from many materials, but it hasn't fully sorted out the relationships between different materials. With more reference materials, it still can't accurately determine which requirements each material corresponds to.

06. Industrial scenarios: Ad-quality output, but not yet industrial precision

The previous rounds focused mainly on creative content generation, but Seedance 2.5 emphasizes extending into more professional industry scenarios. In this final round, we test its industrial video generation and local editing capabilities.

We uploaded an image of an industrial robotic arm and asked Seedance 2.5 to generate a video of the arm picking up an object.

The output generally met the request. The rotation of the arm's joints and the gripping action were basically correct, but the image still had a certain “AI feel” — it looked more like an industrial product commercial than a real factory floor. In addition, the video opened with a close-up of a pen that wasn't in the prompt, which felt jarring.

We then performed a local edit, asking the model to replace the pen with a phone. It made the change quickly, leaving the rest of the original video untouched. Local editing is quite valuable for creators, because they no longer need to regenerate an entire video for a single detail error.

Based on this test, Seedance 2.5 can already generate industrial-themed visual content and perform basic local edits. But at its current level of performance, it is more suited to industrial product display, proposal demos, and marketing content. It's not yet possible to conclude that it can be used directly for industrial simulation or production-data generation.

07. Conclusion

Across the six scenarios, Seedance 2.5 shows improved character realism, maintains basic visual and performance quality in 30-second native output, and delivers more stable appearance consistency and action interaction in two-character scenes. White-model reference and local editing also give creators new ways to control space and revise footage.

The problems are mainly in the experience: prompt moderation is unstable; video extension preserves characters and style but doesn't handle narrative rhythm; and multimodal-reference generation still needs better understanding of the materials.

Returning to the question at the start: Seedance 2.5 does solve some of the problems left by Seedance 2.0. It carries forward the image-generation capabilities proven by Seedance 2.0 and focuses the upgrade on duration, stability, reference inputs, and video editing. The model's highlights are also extending from output quality toward controllability and editing efficiency in the production workflow.

The importance of this upgrade should be understood against the backdrop of ByteDance's accelerating AI commercialization. Since the second half of 2025, ByteDance has sharply increased AI computing procurement, infrastructure construction, and R&D investment, planning about 200 billion yuan in AI capital expenditure for 2026.

On the language model side, as of June 2026, Doubao's daily average Token calls exceeded 180 trillion. According to IDC data, by Token call volume, Volcano Engine held a 49.5% share of China's public cloud MaaS service market. However, Token call volume doesn't directly translate into revenue and profit. Converting technical investment into revenue is the key to ByteDance's AI strategy.

ByteDance's recent moves show its AI business shifting from expanding the C-end user base to opening up the enterprise market through Volcano Engine. Compared with general-purpose model APIs in a fiercely competitive market where prices keep falling, video generation offers higher unit prices, greater commercial value, and larger gross-margin space. Advertising agencies, manufacturers, and film/TV institutions have continuous demand for marketing materials, product videos, and film/TV content, and are more willing to pay for stable, controllable video-generation capabilities.

The value of the Seedance model family in this commercial system is thus clearer. In the Seedance 1.0 and 1.5 periods, market competitiveness was limited and customer call volumes were low. After Seedance 2.0 was released, short-drama and comic-drama production companies began buying Volcano Engine's video model services. Content companies use models to generate AI video, then distribute it through platforms like Douyin and Hongguo; traffic and revenue in turn stimulate production demand. ByteDance therefore has a chance to form a closed commercial loop spanning models, cloud services, and content platforms.

So Seedance 2.5 is an important step in ByteDance's push from consumer-grade AI products toward enterprise services. What it has to do is raise the usability rate and production efficiency in this chain, becoming a product with revenue-generation potential in the Volcano Engine system. For ByteDance, Seedance 2.5's significance is turning its technical advantages into a repeatable, deliverable production-service capability.

PAN's pipeline reviewed approximately 2 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →