A close look at each of the summer’s three frontier video models, followed by three head to head comparisons that show where each strength actually comes from.
Between late June and early August 2026, ByteDance, MiniMax, and Alibaba each shipped a frontier video generation model, and the release calendar was so compressed that two of the three launched on the same day. The temptation is to rank them. This article resists it, because a single ranking would hide the most useful fact about the trio: each model concentrates its ambition in a different place, and each is clearly ahead of the other two somewhere.

The format below reflects that. First come three model profiles, one per model, each pairing a fact file of verified specifications with a short assessment of what the numbers mean. Then come three head to head comparisons, pairing the models two at a time on the ground where their strengths collide: Seedance 2.5 against Wan 3.0 on the new 30 second long take, MiniMax H3 against Seedance 2.5 on sound and creative control, and Wan 3.0 against MiniMax H3 on inputs and openness. Final verdicts close the article. Every figure comes from official announcements and public reporting current to August 13, 2026.
Seedance 2.5
| Developer | ByteDance Seed, within the Seed foundation model family |
| Announced | June 23, 2026, Volcano Engine FORCE conference, Beijing |
| Released | July 31, 2026, worldwide on Dreamina |
| Max single pass | 30 seconds, generated natively with no stitching |
| Resolution | 4K reported in launch coverage, absent from the official headline slide |
| Audio | Co generated with the visuals and locked to the timeline |
| References | Up to 50 per generation: 30 images, 10 video clips, 10 audio files |
| Editing | Region level and timestamp level revisions, multi round extensions |
| Open weights | No |
| Pricing status | No official public rate card at launch, token rates listed on BytePlus |
Seedance 2.5 is ByteDance’s declaration that AI video has entered its production era. The company introduced it by skipping four version numbers outright, and its chief executive told the FORCE conference audience that reaching the top of the AI field is now the company’s first priority, with the model business treated as core long term infrastructure. The launch sequence backed the rhetoric: an enterprise beta through early summer, a promotional short film starring the footballer Michael Owen released through BytePlus in mid July, and then the worldwide Dreamina release on July 31.
The assessment writes itself from the fact file. Nothing else on the market combines a half minute of continuous generation with fifty simultaneous reference inputs, and that pairing exists for one audience: teams who storyboard before they generate. The open questions are equally clear. The 4K figure lives in press coverage rather than in ByteDance’s own headline material, the public rate card has not appeared, and the widest access still runs through consumer surfaces rather than a fully open developer API.
MiniMax H3
| Developer | MiniMax, successor to the Hailuo 2.3 line |
| Announced | July 31, 2026, as MiniMax H3, alias Hailuo 3.0 |
| Released | July 31, 2026, same day across apps and API partners |
| Max single pass | 15 seconds, with several shots possible inside one clip |
| Resolution | Native 2K at 24 frames per second |
| Audio | Stereo sound generated in the same pass as the picture |
| References | Up to 12 per generation: 9 images, 3 video clips, 3 audio clips |
| Editing | Instruction based revisions described in plain language |
| Open weights | Yes, available at release |
| Pricing status | About 0.073 to 0.120 dollars per second, from 0.13 on OpenRouter |
MiniMax built H3 as an argument against specialization. The company’s own release framing criticizes the industry pattern of separate expert models for image editing, reference handling, voice, and sound effects, and presents H3 as one system that understands text, images, video, and audio together before it generates anything. The lineage supports the story: Hailuo 01 established the stack, Hailuo 02 concentrated on efficiency and data quality, and H3 merges the scattered capabilities into a single omni modal model.
What the fact file cannot show is temperament. H3 is the model of the three that a solo creator can actually adopt this week: three free generations on the MiniMax Hub, a public per second price small enough to experiment against, an API with a callback option so developers are not polling for results, and weights that are genuinely downloadable. Its ceiling is equally visible, at 15 seconds and 2K, and its evidence base is young, since the technical report was still pending at release and the launch shipped without a benchmark table.
Wan 3.0
| Developer | Alibaba Tongyi Lab, Wan series |
| Announced | August 6, 2026, public beta as wan3.0-video |
| Released | In public beta on Alibaba Cloud Model Studio and Qwen platforms |
| Max single pass | 30 seconds, double the ceiling of Wan 2.7 |
| Resolution | 480p, 720p, and 1080p in the current API reference |
| Audio | Supported, with quality named as an improvement area in early reports |
| References | Text, image, audio, video, plus documents, slides, spreadsheets, web pages |
| Editing | Modification of scenes, plot, and dialogue in generated video |
| Open weights | Pledged under Apache 2.0, not shipped as of August 8 |
| Pricing status | Regional rate card published with the beta |
Wan 3.0 is the only model of the three aimed past the creator economy at the office. Alibaba’s beta announcement leans on business language, describing the transformation of static, text heavy data into video, and the input list makes the point concrete: working files in doc, xls, ppt, pdf, and md formats up to 100 megabytes or 50 pages each, additional formats including txt, key, pages, and numbers cited in Chinese press coverage, and live web pages passed in by URL. The consolidation Alibaba calls Omni Reference then holds characters, props, voices, and style steady across shots, folding what were separate 2.x models into one.
The assessment carries two asterisks. Output tops out at 1080p in the current API reference, and early testers name sound quality and on screen text accuracy as the rough edges, which matters for a model whose target output is the corporate explainer. The Apache 2.0 open weight pledge also remains a promise rather than a shipment, and the line’s history, an unshipped open release for 2.5, a closed 2.6, and a 2.7 that opened partially after sustained community pressure, is reason enough to wait for the repository link before celebrating.
Seedance 2.5 vs Wan 3.0 on Long Takes
Both models advertise the same headline number, a 30 second clip in one uninterrupted generation, and both arrived at it for the same reason: joins between short clips are where AI video breaks, as faces mutate, light sources wander, and props blink out of existence between fragments. Doubling the previous ceiling removes the joins from a large class of real work, since a standard television spot fits inside a single take for the first time.
The tiebreakers sit on either side of the generation itself. Before the render, Seedance 2.5 offers far more steering, with its 50 reference inputs against Wan 3.0’s leaner reference workflow, and its stack accepts unusual control assets such as 3D white models for spatial blocking and green screen plates for placing a character precisely in a scene. After the render, the models diverge on finishing: Seedance 2.5 revises by region and by timestamp, while Wan 3.0 counters with an intelligent duration feature that recommends the right length from the prompt, an extension tool that grows an existing timeline, and editing that reaches into scenes, plot, and dialogue. On raw output, the comparison currently favors ByteDance, whose reported ceiling is 4K against Wan 3.0’s documented 1080p, with the caveat that the 4K figure has yet to appear in ByteDance’s own specification material. Verdict: Seedance 2.5 wins the long take for cinematic work, while Wan 3.0 makes the long take cheaper to reach for everyday business output.
MiniMax H3 vs Seedance 2.5 on Sound and Control
This is the contest between the two philosophies of finish. Both models generate audio together with the picture rather than dubbing it afterwards, Seedance 2.5 by co processing sound in the same latent space as the video, H3 by rendering a stereo track in the same pass. The difference is emphasis. H3 treats the soundtrack as a first class part of the prompt, so a writer can specify the line, the moment it is delivered, and where a music cue lands, and the earliest wave of public testing kept returning to one observation, that the lip sync held up. For dialogue led formats, vertical drama, and social content, that is the whole game.
Seedance 2.5 answers with breadth of control rather than depth of sound. Fifty references against twelve is not a rounding difference, it is a different way of working, and the same is true of revision, where Seedance edits by region and timestamp against H3’s conversational instruction based editing. H3’s counterweights are pragmatic: a published price per second, day one availability everywhere from the Hailuo app to OpenRouter, and a 6 second 2K clip with sound costing about 0.78 dollars, numbers Seedance simply cannot match while its rate card remains unpublished. Verdict: H3 wins sound and speed of adoption, Seedance 2.5 wins directorial control, and the deciding question is whether your project is driven by performances or by planning.
Wan 3.0 vs MiniMax H3 on Inputs and Openness
The final comparison pits the two models that widened the doorway, in opposite directions. Wan 3.0 widened what goes in. No other model accepts a slide deck, a spreadsheet, a manual, or a URL as the seed of a video, and for organizations whose raw material is paperwork, that single capability deletes a production pipeline. H3 widened who can build. It is the only one of the three whose weights shipped openly, which converts it from a rented service into infrastructure a team can host, fine tune, and audit.
Each advantage exposes the other’s gap. H3 has nothing like document input, and its 12 file reference system, while precise, still assumes a user who thinks in shots. Wan 3.0’s openness is a pledge with a skeptical audience, given the line’s record, and its beta carries integration friction of its own, with the model, endpoint, API key, and uploaded assets all required to sit in the same cloud region. Verdict: Wan 3.0 wins the enterprise doorway, H3 wins the developer doorway, and they overlap so little that many organizations will justifiably run both.
Final Verdicts
Best for directed, cinematic production: Seedance 2.5. The longest take on the market, the deepest reference stack, and post generation editing precise to the timestamp make it the model for teams that plan shots before generating them, once its API access and official pricing catch up with its capability.
Best for creators, dialogue, and immediate adoption: MiniMax H3. Finished stereo sound, credible lip sync, open weights, and a verifiable per second price make it the model a small team can build on this week, accepting the 15 second and 2K ceilings as the cost of accessibility.
Best for turning business material into video: Wan 3.0. Document input up to 50 pages per file, a 30 second take, and story level editing make it the shortest path from a deck to a demo, provided buyers treat the resolution limit, the audio maturity, and the open weight promise with clear eyes.
The pattern across all three verdicts is the same. None of these models tries to be everything, and the summer of 2026 rewarded that restraint: three specialists shipped in six weeks, and for once the marketing language about different strengths is simply accurate.













