Select keyframes
Choose the beginning, end, and meaningful intermediate changes. Start and end frames act as generation anchors; intermediate frames prove the motion path.
Video to Prompt
A video is not merely a moving image. Jingyu builds keyframe evidence, separates static appearance from subject action, environmental response, and camera motion, then compiles frame and motion prompts.
Video-to-prompt reverse engineering converts observable frames and temporal change into a generative structure: each keyframe says what the scene looks like at that moment, while cross-frame analysis explains how it moves from start to finish.
Compressing a whole clip into one caption often loses order, pauses, speed curves, and camera stability. A layered process gives each type of evidence a clear job.
Choose the beginning, end, and meaningful intermediate changes. Start and end frames act as generation anchors; intermediate frames prove the motion path.
Record subjects, clothing, pose, composition, lighting, and background per frame. Motion analysis should not rewrite already verified appearance.
Merge action, speed curve, physical response, camera behavior, and final state in causal order, then adapt the output for Kling, Dreamina, or another target.
Real local processing record · 2026-08-03
This is one genuine reverse-engineering record from the local project. The source lasts about 15.045 seconds and produced seven selected frames, with the start and end used as hard anchors. The page shows actual analysis evidence, not an unverified generated-after comparison, and it does not present two runs of the same source as two different cases.



The main subject has long black hair, blunt bangs, and a navy sailor-style top with white stripes. The person in the left foreground is only a blurred shoulder-and-back shape. The shot is an eye-level over-the-shoulder close-up with shallow depth of field and a mostly stable camera.
The relationship, psychological motive, actual tears, and tiny camera corrections cannot be confirmed from sampled frames. They remain outside the fact layer, while unsupported dialogue, an embrace, or a reconstructed foreground face are prohibited.
A useful prompt tells the model the principal event, starting state, path, and ending state without pasting every audit note or frame observation into the generation text.
Shot continuity or cuts, principal event, action order, speed changes, key poses, gaze, environmental motion, physical feedback, camera motion, and a clear resolved endpoint.
Uncertainties, identity-alignment warnings, quality checks, source paths, internal schema fields, and evidence notes. They help human review but do not belong in the copyable prompt.
They prove endpoints but not intermediate action, pauses, speed curves, gaze changes, or temporary occlusion. Selected middle frames reduce invented transitions.
No. Similar frames waste context. Keep frames that show meaningful state changes and stable start/end anchors.
Visual and audio analysis should be separate. Do not invent dialogue, music, or effects when no audio evidence is present.