Every AI filmmaker knows the specific frustration of trying to write a performance. You can see the shot in your head: the actor tilts her chin, holds a beat, then breaks into a grin. You spend forty minutes typing that into a prompt box and get back a character who blinks like a mannequin and smiles a full second too late. Motion is the hardest thing to describe in words, so a growing set of tools has quietly stopped asking you to describe it. Instead of writing the performance, you perform it, and the model paints your character on top.
This is motion capture without the suit, the volume stage, or the six-figure rig. Point a webcam at yourself, act the line, and the system reads your head movement, your expressions, your hand gestures, and your voice, then retargets all of it onto a character image you supply. Runway calls its version Act-Two. Kling calls it Motion Control. Hedra builds its entire avatar product around the idea. The mechanics differ, but they all solve the thing prompts are worst at: producing a specific, intentional performance instead of a generic one.
What acting-driven capture actually does
The pattern is consistent across tools. You bring two things: a reference image of your character, and a driving performance video of a real person, usually you. The model extracts pose, facial expression, hand gesture, and audio from the driving take and maps them onto the character while holding its identity steady. Runway's Act-Two tracks head, face, body, and hands, and added a Voices feature so you can swap the vocal performance without re-recording the physical one. It is currently limited to enterprise customers and creative partners, which tells you where the fidelity bar now sits.
Kling took the same route from its own direction. Its Motion Control feature, upgraded through the 2.6 and 3.0 releases this year, lets you upload a character image and a reference video, then retargets the movement frame by frame while preserving expression and gesture. Hedra's Character-3 pushes the audio side, reasoning across image, text, and sound at once so lip-sync and head motion stay locked to a voice track. Three different companies, one shared bet: the fastest way to get a believable performance is to record one.
Where it beats prompting
The obvious win is dialogue. Lip-sync generated from a text prompt is a guess; lip-sync driven by your own voice track is a transcription. If a character has to speak, act-driven capture removes the single ugliest tell in AI video, the mouth that moves like it is dubbed from another language.
The subtler win is timing. Comedy, tension, and grief all live in the pause, and a pause is nearly impossible to prompt. When you perform the beat yourself, the hold lands exactly where you put it. You also get repeatability. Feed the same driving take to two shots and the character behaves the same way in both, which is the kind of continuity that stitching independent clips never gives you for free.
Where it still bites
Capture is not a cheat code. The first catch is that you have to be able to act, at least a little. Flat delivery in means flat delivery out, and the camera is honest about it. If you cannot sell the line, neither can your character.
The second is retargeting drift. When your face and the character's differ a lot in proportion, or when the performance involves extreme angles, fast turns, or a hand crossing the face, the mapping can smear or pop. You feel it most in the exact expressive moments you reached for capture to get.
The third is scope. These tools animate a character; they do not build the scene around it. Framing, lighting, environment, and the cut still come from elsewhere in your pipeline. Capture solves the performance and hands you back a moving figure, not a finished shot.
What to do about it
If a shot lives or dies on performance, a monologue, a reaction, a sung chorus, stop fighting the prompt box and record it. Shoot a clean driving take with even light on your face and audio you would be happy to keep, since the good ones keep your voice. Match your framing roughly to the character's proportions to cut down on retargeting artifacts. Then treat the animated performance as one input among several, dropping it back into the shot, the grade, and the cut. That handoff between a captured performance and everything else in the sequence is exactly the orchestration problem Promvie's pipeline is built to manage.
The prompt was never the right instrument for a performance. Directing has always meant showing an actor what you want, not describing it in a paragraph. The actor is synthetic now, and the direction is you, on camera, doing the thing.