ByteDance's Seedance 2.5, now rolling out after its June keynote, will take up to 50 reference images in a single generation. Google's Veo 3.1 lets you feed it three, one each for a character, a prop, and a setting, then composites them into one shot. A year ago the whole game was writing a sharper sentence. Now the sentence is losing ground to the pictures you hand the model alongside it.
Why this is the shift that actually helps
If you have ever spent an afternoon rewording a prompt to keep a character's jacket the same shade of gray from shot to shot, you already know the problem. Text is a lossy way to describe a face. "Weathered detective in a charcoal trench coat" returns a different detective every time you press generate, because the words leave a thousand small decisions to the model, and it makes them differently each run.
A reference image closes that gap. It pins the thing down. What changed this year is not that reference support exists, but that the models are finally built to accept enough references, from enough angles, to hold an identity together across a whole clip instead of a single frame.
What the top models now accept
Seedance 2.5 is the headline. It takes up to 50 multimodal references, images, video, audio, and even 3D assets, and uses them to keep characters and scenes aligned across a native 30-second clip. It also adds region-level editing: draw a box over a face or a product label and re-render only that patch, leaving the rest of the frame untouched.
Veo 3.1 takes the opposite approach with its "ingredients" feature. You supply three images, tag each one with a role, and the model reads all of them at once, one for the character's face and outfit, one for a prop, one for the location or style. Three is a smaller budget than fifty, but the role-mapping is explicit, which makes the results easier to predict.
Runway and Kling sit in the middle, leaning on reference-driven identity control that creators use as a casting step before rendering the actual scenes elsewhere. The common thread across all of them: the reference image is no longer a hint. It is the instruction.
The skill is moving from writing to casting
Here is the part nobody puts on the feature page. Fifty references is not fifty units of free control. It is fifty things you now have to art-direct. Feed the model muddy, inconsistent, or off-model reference frames and you get muddy, inconsistent output back, faster and at higher resolution than before. The failure mode has changed, not disappeared.
What this really does is move the craft. The hard part used to be prompt engineering, coaxing the right image out of a text box. Now the hard part is assembling a clean reference kit: a character sheet shot from several angles, a prop plate, a location plate, a frame that carries the color and grain you want. That is closer to a costume-and-set job than a writing job, and it rewards people who think like production designers rather than copywriters.
It also quietly raises the floor on continuity. Reuse the same reference set across ten generations and the tenth shot still looks like the first. For serialized content, brand work, or anything with a recurring cast, that consistency is the whole ballgame, and it is the thing text prompts never delivered.
What to do about it
Before you generate anything, build the reference kit. Lock your main character with three or four images from different angles and in consistent lighting. Add a plate for any prop that has to stay recognizable, and one for the location. Pull a single frame that defines your look, the palette and the texture, and keep it in every generation so the grade does not drift.
Then treat that kit as a reusable asset, not a one-off. The whole point is that shot two, shot twenty, and next week's scene all pull from the same source. This is the logic Promvie's pipeline is built around, where cast and settings persist as assets across the whole film rather than being re-described shot by shot. You can see how that fits together on the how it works page.
The prompt box is not going away
Text still owns intent, action, and timing, the things a picture cannot say. But identity, style, and continuity are migrating to the reference layer, and that is where the leverage now lives. The creators who adapt fastest will be the ones who stop treating references as a nice-to-have and start treating them as the shot list.