Motion Prompt

Describe how it moves—not merely that it moves

Jingyu separates character action, object motion, environmental change, physical feedback, and camera movement, then reconstructs them as a motion prompt with a start, trajectory, speed, rhythm, and resolved end state.

Short answer

Reverse motion prompting does not add words such as “slowly” or “cinematic” to a still description. It identifies who starts in what state, follows which path, changes at what speed, and ends where, while constraining the camera and environment so the principal event stays readable.

Five parts of an executable motion prompt

Motion language needs temporal relationships. Every part should answer a concrete question and remain consistent with visible keyframe evidence.

Starting state

Define pose, gaze, hand and foot placement, object state, and camera state before the action. A vague start makes the model invent preparation.

Primary trajectory

State direction, body-part order, and intermediate poses. “Lower head, hold, then raise it” preserves sequence better than a generic command to look up.

Speed and rhythm

Specify constant motion, acceleration, deceleration, pauses, or phases. Timing should support one principal event instead of several competing large actions.

Physical feedback

Explain how fabric, hair, dust, water, or props respond. Include only visible or physically justified effects rather than inventing wind, particles, or impacts.

Camera response

Choose fixed, tracking, panning, pushing, or restrained handheld behavior. Establish subject motion first so unnecessary camera movement does not damage readability.

Resolved end state

Say where the subject and camera stop, where the gaze lands, and how the frame settles. A defined endpoint reduces loops, repeated actions, and sudden rebounds.

From keyframe evidence to motion language

Jingyu first asks each keyframe to describe only its static moment. Cross-frame analysis then handles change. This keeps process words out of the start frame and unfinished action out of the end frame.

1

Lock the endpoints

The start frame freezes the scene before action; the end frame freezes the completed result. Neither uses words such as “gradually,” “currently,” or “about to.”

2

Extract intermediate phases

Middle frames reveal order, peak pose, pauses, occlusion, and speed change. Similar frames are removed unless they change the motion finding.

3

Compile in causal order

Write subject action first, then prop and environmental response, then camera behavior and final state. Both people and models can identify the principal event quickly.

Motion analysis from a real case

Why “lowering the gaze, then looking up again” needs a timeline

In the genuine 15-second portrait case, the result was not a vague line about an emotional woman. The head slowly lowers, gaze moves down, eyelids close, the pose holds briefly, then the head rises smoothly, eyes reopen, and the gaze stabilizes toward the foreground person. The camera stays essentially fixed and hair movement remains subtle.

Correct principal motion

Keep one continuous shot. The only major visual event is avoidance followed by renewed engagement. Attention travels from the moist eyes to the lowered head and back to the re-established gaze.

Explicit drift prevention

Do not reverse the order, repeat the motion, teleport the head, add an embrace, turn, dialogue, or unsupported camera push, and do not change clothing, hair, or subject count.

Structure: shot mode → director task → start state → subject trajectory → speed curve and pauses → physical/environmental response → camera response → end state → negative motion constraints.

Limits of reverse motion analysis

Sampled frames support many temporal findings, but they cannot prove every motion detail. Observation, inference, and generation control should remain separate.

Usually reliable

Visible position and pose changes, action order, strong pauses, gaze direction, object state changes, shot cuts, and most clear camera movements.

Needs caution

Tiny camera corrections, psychological motive, fully occluded trajectories, instantaneous motion between samples, and any rhythm or dialogue unsupported by audio.

Frequently asked questions

How is a motion prompt different from a video prompt?

A video prompt includes appearance, scene, composition, and style. A motion prompt focuses on state change, trajectory, speed, physical feedback, and camera response.

Should camera movement come before or after the action?

Usually define principal subject action first, environmental and physical response second, and camera response last. This makes causality and visual priority clearer.

Why does the end state need to be explicit?

Without an endpoint, models may repeat, rebound, or keep changing at the end. A resolved pose, gaze, and camera state improve stability.