Three of the biggest video models shipped the same idea this summer, which usually means it wasn't a coincidence. Kling 3.0 will generate up to six camera cuts in a single pass. Seedance 2.5 runs thirty seconds continuous. Runway Gen-4.5 takes camera choreography, sequential actions, and timed beats in one prompt. The unit of AI video used to be the clip. It's quietly becoming the sequence.
That sounds like an incremental spec bump. It isn't, because the clip was never the natural unit of anything. It was the length a model could hold together before the face drifted. Every workflow built on AI video for the last two years has been a workaround for that limit: generate five seconds, generate another five, fight the continuity, stitch it in an editor. When the model holds the sequence itself, the workaround becomes the bottleneck.
Kling 3.0 plans the cuts
Kling's multi-shot mode is the most literal version of this. You get two doors. Auto multi-shot reads your scene description and decides where to cut, what angle each shot takes, and how the sequence paces. Custom multi-shot has you set shot count, duration per shot, camera angle, and on-screen action, and the model follows that storyboard.
The more useful piece is native audio running across the cuts instead of per clip. Input and output happen together (text or image in, video and sound out), so the audio bed doesn't reset every time the camera moves. Anyone who has hand-matched room tone across four generated clips knows what that saves. Kling 3.0 also handles multi-character coreference for three or more characters, which is exactly the thing that used to fall apart the instant a scene had a crowd.
Seedance 2.5 removes the seam
ByteDance took a different route: rather than cutting inside a generation, extend the generation. Seedance 2.5 launched globally on Dreamina at the end of July with up to thirty seconds of continuous output, support for as many as fifty multimodal reference assets in one workflow, and targeted editing of the result instead of a full regenerate.
Fifty references is the number worth staring at. That isn't a style prompt, it's an asset library of characters, locations, props, and lighting plates handed to the model as context for a half-minute of footage. ByteDance's own pitch is about reducing clip stitching, visual drift, and rework, which is an unusually honest description of what AI video actually costs today.
Runway Gen-4.5 takes direction
Runway attacked the same problem from the prompt side. Gen-4.5 caps out at five or ten seconds, shorter than either rival, but what it does inside that window is denser: detailed camera choreography, multiple subjects, sequential actions, and the timing of events all specified in one instruction. Scenes with several things happening at once land closer to what you asked for.
Runway is betting the sequence lives in the language, not the duration. You write the way a director writes a shot description, and the model executes the movement and the beats. It runs 12 credits per second in the app and $0.12 per generated second through the API, and the fixed 5- or 10-second durations are a deliberate choice for edits where predictable pacing beats raw length.
The skill that just depreciated
Here is the part that will annoy people who got good at the old way. The craft of AI video in 2025 was prompt-per-clip: write one strong sentence, roll it fifty times, keep the best generation. That skill is worth less every month. What matters now is whether you can specify a sequence: what the coverage is, where the cuts land, and what has to carry across them.
That's not a prompting skill. It's a shot list, which is the oldest tool in the room. The models converged on it because filmmaking converged on it a century ago, and the labs have finally built enough coherence to hold more than one shot in mind at a time.
It also moves the hard problem up a level. When you controlled every cut yourself in an editor, mediocre continuity was your problem to fix in post. When the model plans the cuts, continuity becomes a property of how well you briefed it. That is less fiddly and considerably less forgiving.
What to do about it
Stop writing prompts and start writing coverage. Before you generate anything, decide the shots, their order, and what must stay identical across them: face, wardrobe, lighting direction, the sound of the room. Hand that over as structure, not as one paragraph of adjectives.
Then match the model to the job. Kling if you need real cuts and dialogue. Seedance if a shot has to run long without a seam. Runway if the value sits in one dense, well-directed take. Nobody wins all three, and choosing per shot rather than per project is now ordinary practice.
If your pipeline already reasons about the whole film, from script and cast through shot list to final cut, this convergence is a tailwind rather than a retooling. That's roughly the bet Promvie makes.
The models learned to cut. The next scarce skill is knowing what to tell them to cut to.