All posts
guides

Movement Tracking for AI Video Control

You don't draw tracking points. You hand it a video whose motion you want and an image of who should perform it. How motion transfer actually works, and what doesn't exist.

By Flixly TeamJune 15, 2026
Movement Tracking for AI Video Control

TL;DR

There is no keypoint editor. Motion Control takes a driving video (MP4 or MOV, at least three seconds, up to 100MB) and a character image (JPG or PNG, up to 10MB), and transfers the performance onto the character. The character_orientation toggle decides whose framing wins, image or video, and is the setting most often left wrong. It runs on Kling 3.0 Motion Control at 720p or 1080p, with an optional prompt as a scene hint.

You do not draw tracking points. There is no point mode, no mask mode, no keypoint editor.

What actually exists is better, and simpler to explain: you hand the system a video whose motion you want, and an image of who should perform it. The motion transfers from one to the other.

That is Motion Control, and once you understand that it takes two inputs rather than a set of dials, most of the confusion around steering AI video disappears.

How motion transfer actually works

Two files, and one decision.

A driving video. MP4 or MOV, at least three seconds, up to 100MB. This is the performance: a walk, a gesture, a dance, a camera move. Its content does not matter beyond the movement itself.

A character image. JPG or PNG, up to 10MB. This is who ends up on screen.

The orientation toggle. This is the decision, and it is the one setting people miss. character_orientation chooses whose framing wins: the image or the video. Default is the image, meaning your character keeps their own facing and the motion is mapped onto them. Switch it to the video when you want the driving clip's staging to dominate.

Resolution is 720p or 1080p. A prompt is optional and acts as a scene hint, not a command. That is the entire control surface.

It runs on Kling 3.0 Motion Control, a PREMIUM-tier model.

Why this beats drawing points

Keypoint editing sounds like control. In practice it is a way of describing motion badly.

A video already contains the motion, perfectly, at every frame. Handing that over is a far higher-bandwidth instruction than a dozen anchor points a human placed by eye. You are not approximating the movement, you are supplying it.

The consequence for your workflow: the quality of your output is mostly the quality of your driving clip. Not a setting you failed to find.

Getting a good driving clip

This is where the real effort belongs.

Keep the subject in frame throughout. Motion that leaves the frame has nothing to transfer for those frames.

Steady framing beats dramatic framing. A locked-off shot of a person moving transfers cleanly. A handheld shot of a person moving transfers the handheld shake too.

One subject. The system maps a performance onto a character. A crowd is ambiguous about which performance you meant.

Match the body. A full-body driving clip onto a headshot has nowhere to put the legs. Frame the driving video roughly the way you want the output framed.

Three seconds minimum. Shorter clips are rejected before anything runs.

Shoot the driving clip yourself if you can. Your phone, good light, one clean take. It is usually faster than hunting for stock footage that happens to contain the exact movement you pictured.

When you want something else entirely

Motion Control is for transferring a performance. Three neighbouring problems have their own tools, and reaching for the wrong one is the most common mistake here.

You want a specific person or product to stay itself across shots. That is reference-to-video, supported by thirteen models. Start at reference to video.

You want a still to come alive without a driving clip. Image to video animates a picture from a text description of the motion.

You know exactly where the shot begins and ends. First-to-last frame generates the movement between two images you supply. Only Seedance 2.0 and Seedance 2.0 Fast support it, and it is the most deterministic option available.

You want continuity across a whole sequence. The Long Video Generator chains segments so each one continues from the last, which is a different problem from transferring a single performance.

What will not work

Worth stating plainly, because the internet is confident about several things that are false.

There is no camera-path editor. You cannot draw a trajectory and have the camera follow it. Camera movement is described in the prompt or inherited from a driving clip.

There is no per-model point budget, because there are no points. Comparison tables listing "max points" per model are describing software that does not exist.

Image models do not track motion. Nano Banana Pro, GPT-Image 2.0 and the rest generate stills. They have no video output path and no optical-flow fallback.

And there is no drift percentage to quote. Results vary with the driving clip, which is exactly why the clip is where your attention should go.

A workflow that holds up

  1. Get the character image right first. Everything downstream inherits it, so a mediocre reference propagates through every frame.
  2. Shoot or find the driving clip. Steady, one subject, framed like the output, three seconds or more.
  3. Set the orientation toggle deliberately. Image-driven keeps your character's facing. Video-driven adopts the clip's staging. This is the setting most worth an extra test.
  4. Generate at 720p while testing. Move to 1080p once the motion is landing.
  5. Add dialogue afterwards with Lip Sync, never before, since regenerating throws away sync work.

If the result is close but the framing is wrong, flip the orientation toggle before changing anything else. It is the single largest lever in the tool and the one most often left at its default.

Where it fits in a longer piece

Motion Control produces one clip of one performance. For a sequence, generate several and cut them together, or use the chaining approach for continuous action.

Motion Poster is the neighbouring tool when you want moving key art from the same character rather than a full performance.

Costs depend on model and settings, and the app quotes each generation before you run it. Pack prices are on the pricing page, and the full catalog is at Models.

Frequently Asked Questions

How do I draw motion tracking points on a subject?

You don't. There is no point mode, mask mode or keypoint editor. Motion Control works by transferring the movement from a driving video onto a character image, so the motion is supplied rather than approximated by hand-placed anchors.

What files does Motion Control need?

Exactly two: one driving video in MP4 or MOV, at least three seconds long and up to 100MB, and one character image in JPG or PNG up to 10MB. Resolution is 720p or 1080p, and a prompt is optional, acting as a scene hint rather than a command.

What does the character orientation toggle do?

It decides whose framing wins. Set to image, your character keeps their own facing and the motion maps onto them. Set to video, the driving clip's staging dominates. It defaults to image and is the largest lever in the tool, so it is worth testing both before changing anything else.

What makes a good driving clip?

Keep the subject in frame throughout, use steady framing since handheld shake transfers too, feature one subject rather than a crowd, and frame it roughly the way you want the output framed. A full-body clip onto a headshot has nowhere to put the legs. Three seconds is the minimum.

How do I keep a character consistent across several shots?

That is reference-to-video rather than motion transfer, supported by thirteen models. For a continuous sequence, the Long Video Generator chains segments so each continues from the last. Motion Control produces one clip of one performance.

Can I edit the camera path?

No. There is no camera-path editor. Camera movement is either described in the prompt or inherited from the driving clip. Guides describing a trajectory editor or per-model point budgets are describing software that does not exist.

Can image models like Nano Banana Pro track motion?

No. Image models generate stills and have no video output path. Motion transfer runs on Kling 3.0 Motion Control, and the other motion-related capabilities are reference-to-video, image-to-video and first-to-last frame, all of which are video models.

Tools mentioned in this post

guidesvideomotion-controlreference-to-video

Ready to create with guides?

Jump straight into Flixly's AI studio and try guides with 50+ models — free to start.