Video to Prompt

Put keyframes, action, and camera on one timeline

A video is not merely a moving image. Jingyu builds keyframe evidence, separates static appearance from subject action, environmental response, and camera motion, then compiles frame and motion prompts.

Short answer

Video-to-prompt reverse engineering converts observable frames and temporal change into a generative structure: each keyframe says what the scene looks like at that moment, while cross-frame analysis explains how it moves from start to finish.

Why video analysis needs separate layers

Compressing a whole clip into one caption often loses order, pauses, speed curves, and camera stability. A layered process gives each type of evidence a clear job.

1

Select keyframes

Choose the beginning, end, and meaningful intermediate changes. Start and end frames act as generation anchors; intermediate frames prove the motion path.

2

Isolate frame facts

Record subjects, clothing, pose, composition, lighting, and background per frame. Motion analysis should not rewrite already verified appearance.

3

Compile temporal motion

Merge action, speed curve, physical response, camera behavior, and final state in causal order, then adapt the output for Kling, Dreamina, or another target.

Real local processing record · 2026-08-03

A 15-second single-shot gaze and expression change

This is one genuine reverse-engineering record from the local project. The source lasts about 15.045 seconds and produced seven selected frames, with the start and end used as hard anchors. The page shows actual analysis evidence, not an unverified generated-after comparison, and it does not present two runs of the same source as two different cases.

15.045 ssource duration
7 framesselected evidence
2 anchorshard start and end
1 shotcontinuous-shot finding
Case start frame: a young woman looking toward a person in the left foreground
Start: gaze held toward the person on the left.
Case middle frame: the young woman lowers her head and briefly closes her eyes
Middle: head lowered, eyes closed, brief hold.
Case end frame: the young woman raises her head and looks left again
End: head raised and gaze stabilized again.

What the system verified

The main subject has long black hair, blunt bangs, and a navy sailor-style top with white stripes. The person in the left foreground is only a blurred shoulder-and-back shape. The shot is an eye-level over-the-shoulder close-up with shallow depth of field and a mostly stable camera.

What remained explicitly unknown

The relationship, psychological motive, actual tears, and tiny camera corrections cannot be confirmed from sampled frames. They remain outside the fact layer, while unsupported dialogue, an embrace, or a reconstructed foreground face are prohibited.

Motion prompt excerpt: Continuous single shot, no cut, no new subject. The young woman slowly lowers her head and gaze, holds briefly, then raises her head and looks toward the person in the left foreground again. Her eyelids move from half-open to closed and open again; hair movement stays subtle. Keep the camera essentially fixed.

What belongs in the final video prompt

A useful prompt tells the model the principal event, starting state, path, and ending state without pasting every audit note or frame observation into the generation text.

Keep in the prompt

Shot continuity or cuts, principal event, action order, speed changes, key poses, gaze, environmental motion, physical feedback, camera motion, and a clear resolved endpoint.

Keep in the audit

Uncertainties, identity-alignment warnings, quality checks, source paths, internal schema fields, and evidence notes. They help human review but do not belong in the copyable prompt.

Frequently asked questions

Why are the first and last frames not enough?

They prove endpoints but not intermediate action, pauses, speed curves, gaze changes, or temporary occlusion. Selected middle frames reduce invented transitions.

Are more keyframes always better?

No. Similar frames waste context. Keep frames that show meaningful state changes and stable start/end anchors.

Does video reverse engineering reproduce audio?

Visual and audio analysis should be separate. Do not invent dialogue, music, or effects when no audio evidence is present.