The short answer: don't ask one generation to carry a conversation. Give each speaker their own shot, write every line in straight double quotes, name who says it, and keep each line shorter than the clip. Veo 3.1, Kling 3.0 and Seedance 2.5 all reward the same rules, and the two-person scenes that fall apart usually fall apart because the prompt ignored them.
Monologues were solved last spring. A single face saying a single line with matching lips is now table stakes, which we covered when native audio arrived. Conversations are the harder problem, because a conversation is two faces, two voices and a rhythm of turns, and the model has to know at every frame whose mouth is moving. When it guesses wrong you get the AI-dialogue tell everyone recognises: both mouths flapping at once, or the wrong character mouthing the line.
Why two speakers is a different problem
Higgsfield ran this across tools and put it bluntly: syncing two or more faces in the same shot with overlapping dialogue produces visible errors on every tool tested. Single-speaker clips are the sweet spot. The reason is mechanical. The model attributes speech to a face, and when two faces share a frame and two lines share a timeline, the attribution has to be exact. Nothing in "two friends argue in a garage" tells it which one is talking.
The second problem is time. Veo 3.1 clips run 4, 6 or 8 seconds. Spoken English lands around two and a half words a second, so an 8-second clip holds one solid line or two short ones. A back-and-forth exchange doesn't fit, and when you cram it in, the model rushes the delivery or quietly drops a line.
Use a shot per speaker
The fix is old. Coverage. Shoot the conversation the way a crew would: a two-shot to establish, then singles that cut between speakers on each line.
Google's own Veo 3.1 guide demonstrates it with a detective scene built from reference images. The first prompt asks for a medium shot of the detective looking up and saying, in a weary voice, "Of all the offices in this town, you had to walk into mine." The second prompt, using the same references, is a shot on the woman, a slight smile, replying "You were highly recommended." Two generations, one line each, and the cut hides everything the model would have struggled to hold in one frame.
Seedance 2.5 does the same thing inside one generation, because it runs up to 30 seconds and accepts shot-by-shot structure. Label each shot, put one speaker in it, and reference the characters by tag so they stay on model across cuts. A working prompt looks like this:
@Image1 is A, grey sweater. @Image2 is B, green jacket. Kitchen, morning light.
Shot 1, two-shot: A asks, "Did you move the chair?"
Shot 2, single on B: B looks at the empty corner, pauses, answers, "No."
Shot 3, wider two-shot: the empty corner comes into view. Room tone, no music.
The common mistake, according to the guides that have stress-tested Seedance, is stacking both speakers in the same shot block. Split the speakers and the model splits the lips.
Name the speaker every time
Kling 3.0 Omni calls this speaker mapping: dialogue gets assigned to the right character in multi-speaker scenes, and its lip sync guide asks you to keep each line right next to the character's name and add a delivery note only when it matters. "The man with glasses, nervous: 'Are you sure about this?'" beats a paragraph of mood. Kling handles five languages natively (English, Chinese, Japanese, Korean and Spanish) and can switch mid-scene.
Two rules carry across every model. Straight double quotes mark speech, so the model treats the text as a line to say instead of a caption to print. And describe a voice with adjectives, never with an actor's name, because none of these models will produce a voice from a name.
What audio references actually do
If you have a recorded read, Seedance's @Audio reference will sync mouths to it. It will not preserve the voice. The uploaded track works as a timing guide for lip movement, and the voice you hear in the output is generated. That is fine for pacing a scene and useless if the whole point is a specific actor's voice, in which case you dub in post.
Long takes are the other trap. Reviewers who put Gemini Omni Flash through dozens of generations found lip sync drifts after six or seven seconds. Which is, again, an argument for cutting.
What to do about it
Write the scene as a shot list before you write a prompt. Two-shot, single, single, two-shot. One line per single, timed with a stopwatch, quoted and attributed. Lock the faces with reference images so the cut back to the two-shot doesn't introduce a third person. Generate the singles first. They are the shots that fail, and you will want to roll them more than once.
This is the workflow built into Promvie's pipeline: the script gets broken into coverage per line before any video model sees it, because that is the only shape of dialogue the models reliably deliver.
The models learned to talk. They still haven't learned to take turns, so you cut for them.