First-frame and last-frame AI video: shot design guide

Runbo Li
Runbo Li
·
· 9 min read
Storyboard paper with a clearly distinct opening frame and closing frame, arrows showing a physical camera move between them

Quick answer

First-frame and last-frame AI video uses two images to guide where a shot begins and ends. Prepare compatible subject identity, framing and scene geometry, then describe the motion connecting them. Choose a model and settings that support an end frame. Review the middle of the clip as carefully as the endpoints: matching images do not guarantee a believable transition or exact product details.

This is a shot-planning guide, not a controlled model comparison. The example briefs and prompts below are illustrative; they are not reported generation results. Magic Hour’s linked product and API documentation were checked on October 2, 2026 for end-frame support and the handoff described here.

What first and last frames control

Treat the first image as your opening composition and the last image as the intended landing. The creative task is the journey between them: which object moves, how the camera behaves, and what stays consistent. An endpoint pair is useful when the ending matters, such as an open product box, a person facing the viewer, or a room with its lights on.

An end-frame reference is different from an ordinary reference image. A reference helps describe a subject or style; an end frame specifies the desired closing image. It is also different from adding a still image at the end in an editor. That edit can show the exact image, but it does not make the generated motion naturally arrive there. Choose the method that serves your actual deliverable.

Choose a route before preparing the pair

Open Magic Hour’s Image-to-Video creator. Its current interface offers “Start & End Frame” and “Keyframes.” For this two-image workflow, use Start & End Frame. Select a model and inspect its available duration and resolution before uploading the closing image. If an end-frame control is unavailable for the selected settings, change to a supported combination instead of assuming every image-to-video model accepts two endpoints.

The current API reference documents Kling 3.0 end frames at 720p, 1080p and 4k. A five-second, 720p Kling 3.0 shot is therefore one documented starting configuration, not a quality ranking. Veo 3.1 and Veo 3.1 Lite end-frame requests are limited to eight seconds or less. MiniMax H3 and Wan 2.2 do not currently support an end image in this API. Recheck the linked schema when changing models; a resolution available for ordinary image-to-video may still be unavailable with an end frame.

Prepare two images that describe one shot

1. Write the change in one sentence

Start with “The box opens,” “The person turns toward camera,” or “The room’s lamps switch on.” Then list what must stay fixed: box design, face and clothing, furniture layout, or camera position. If your sentence requires a new location, a new outfit and a moving camera, split the idea into shots. This makes the visual problem easier to inspect and gives you an edit point if one transition fails.

2. Compare identity and geometry

Place both images side by side at the same display size. Inspect the product silhouette, face shape, clothing, object count and relative positions. In a room, compare doorframes, windows and the edges of furniture. In a product shot, compare hinge placement, lid size and the tabletop. Correct contradictions in the images before generation; asking the prompt to preserve an object cannot resolve two different object designs.

3. Align the crop and camera logic

Use the same intended aspect ratio for both frames. Check whether the subject occupies a compatible area of the image and whether the horizon and perspective make sense together. A close-up ending after a wide opening needs a deliberate camera move; a fixed-camera brief should not end at a radically different angle. Crop thoughtfully instead of stretching an image to match dimensions, and leave space for the action inside the final delivery crop.

4. Separate lighting changes from scene changes

For a simple action, keep the lighting direction and important shadows compatible. For a planned time-of-day change, keep the room geometry fixed while defining what changes in brightness and color. Do not combine a lighting transformation with new furniture, an altered window view and a new camera angle unless those changes are part of the concept. Mark any readable labels or logos as details to review independently; endpoints are not a guarantee of accurate text between them.

Write the connecting motion

Use a compact prompt with three parts: the subject action, camera behavior and ending. Describe the missing movement rather than adding a long inventory of what both images already show. For a product opening, specify the hinge action and whether the camera is fixed. For a turn, specify who turns and where the pose should finish. When motion looks wrong, revise that movement first rather than appending more unrelated style adjectives.

A reusable structure is: “[Subject] [single action] at [pace]. The camera [stays fixed or makes one described move]. The shot ends at the supplied closing composition.” Treat this as a writing template, not special API syntax. Avoid mutually incompatible instructions, such as a fixed camera paired with a rotating viewpoint, or a gradual reveal paired with an immediate final pose.

Connect your opening and closing images

Prepare one compatible image pair, select a model with end-frame support, describe the connecting action, and review the whole clip before scaling.

Try Image-to-Video

Three illustrative shot briefs

Example 1: A product box opens

Opening image: one closed box on a table. Closing image: the same box, same viewpoint and table, with the lid raised on its visible hinge. Keep the logo, corners and surrounding objects in the same positions. If hands are unnecessary, leave them out of both frames; adding a hand introduces another identity and contact relationship to solve.

Prompt: “The box lid opens slowly on its rear hinge while the box stays on the table. The camera remains fixed. The shot settles at the supplied open-box image.” Review the hinge, the lid’s path and the box interior. Reject a result where the lid separates from the box, passes through it, or changes the packaging. If those defects recur, correct the endpoint geometry or simplify the opening angle before retrying.

Example 2: A character turns toward camera

Opening image: a three-quarter portrait. Closing image: the same person facing camera in the same clothes and location. Use authorized identity images. Compare facial proportions, earrings, glasses, hair and the crop; an unrelated front-facing portrait creates a second identity target even if both people look similar.

Prompt: “The person gently turns their head toward the camera, ending in the supplied frontal pose. The shoulders remain in place and the camera stays fixed.” Inspect the face during the turn, including the profile and the moment both eyes become visible. A convincing first and last frame can hide identity drift in the middle. Reduce the turn or prepare a closer endpoint pair if the intervening face repeatedly changes.

Example 3: A room changes from day to evening

Opening image: daylight in a furnished room. Closing image: the same room and viewpoint with evening lighting. Build the second frame from the same scene so windows, doorways and furniture remain aligned. Decide whether the concept needs a gradual light change or a conventional cut; do not describe an accelerated sunset when a simple lamp reveal is the actual goal.

Prompt: “Daylight fades gradually as the room’s lamps illuminate. The furniture and camera position remain fixed. The shot ends at the supplied evening composition.” Inspect straight architectural edges and the placement of chairs and lamps throughout the clip. If the room reconstructs itself, repair the pair or use separate shots with an edit. A dramatic morph can be useful creatively, but label and judge it as that intended effect.

Review the journey, then the landing

Watch the complete exported clip at normal speed, then scrub the busiest part of the motion and the approach to the final frame. Check subject count, identity, object contact, perspective, shadows and important text. Look for the moment the video abruptly snaps to the closing composition. The ending should make sense after the preceding movement, not merely resemble your final image in isolation.

Write acceptance criteria that match the brief. For the box, the lid stays attached and packaging stays recognizable. For the portrait, the same person remains identifiable through the turn. For the room, fixed architectural edges stay in place. Check the final crop at delivery size as well. If the exact closing still is a contractual requirement, inspect the exported final frame against it and use an editing workflow when necessary; do not promise pixel-perfect delivery from a reference alone.

Fix the cause before another attempt

The middle morphs into a different object

Compare endpoints for conflicting shape, wardrobe, object count or background layout. Repair the images first. If the pair is consistent, reduce how much changes in the shot. Adding another sentence about “consistency” is less actionable than removing the incompatible product angle or the unexplained second person.

The video rushes into the final frame

Check whether your chosen duration can accommodate the action and whether the prompt requests a pause or an overly complex sequence. Simplify the action or choose another supported duration, then review the new landing. If the ending must remain on screen for reading, reserve that hold in an editor rather than assuming the generator will allocate enough time.

The closing-image control disappears or the request fails

Verify the selected model, resolution and duration support an end frame. An unsupported input combination is a capability problem, not evidence that the images are poor. Read the returned validation message before retrying. Preserve the original pair and settings so changing the model does not silently change the intended shot.

Text, faces or straight edges change

Review those details across the entire clip, not only in the reference images. Simplify the motion or use a closer endpoint pair when feasible. For factual product labels, keep approved artwork in your finishing workflow and inspect the delivered composition. For identity-sensitive material, reject drift rather than treating a plausible stranger as an acceptable substitute.

API handoff: name both endpoints explicitly

For a programmatic workflow, submit to POST /v1/image-to-video with assets.image_file_path for the opening image and assets.end_image_file_path for the closing image. Set model, resolution, end_seconds and style.prompt explicitly. For the documented starting configuration above, these are kling-3.0, 720p and 5. Supply your actual uploaded asset paths; the illustrative file names in documentation are placeholders.

Store the returned project ID and retrieve the result after completion. Inspect the downloaded clip before accepting it. Two uploaded endpoint files, a queued request or a finished job establish different stages; none alone establishes that the motion passes your creative criteria. Keep the original images, exact request settings and review note together so the next operator knows which pair produced the accepted export.

Decide when to use a different workflow

Choose two endpoints when the shot’s beginning and ending are meaningful and a plausible connecting action exists. Use a single opening image when the end pose is flexible. Use separate generated shots and an edit when the scene changes location or needs a clear cut. Use a timeline or compositing workflow when exact labels, an exact still-frame hold or precise edit timing are the main requirements.

This guide owns endpoint-pair preparation and the movement between two compositions. For broader source selection, use How to use reference images for image-to-video; for more motion patterns, use Image-to-video prompts. Start with one prepared pair, inspect the complete result, and keep the simplest workflow that meets the shot’s acceptance criteria.

Official sources and product references

https://docs.magichour.ai/tools/video/image-to-video

https://docs.magichour.ai/api-reference/video-projects/image-to-video

Frequently asked questions

For a continuous shot, keep the same subject, proportions, aspect ratio and scene layout. Change only what the planned action needs, such as a lid opening or a person turning. Compare the images side by side before generation; a changed object or room layout is a source problem, not a motion instruction.

You can propose a transformation between different scenes, but two images alone do not establish a continuous physical path. Decide whether you want a visible morph, a stylized transition or a normal edit. For a documentary-style scene change, separately generated shots joined with a cut may be the clearer choice.

Name the connecting action, camera behavior and intended ending. For example: “The lid opens slowly on its hinge while the camera stays fixed, ending at the supplied open-box image.” Keep the prompt consistent with both images; do not ask for a movement that contradicts the final crop or pose.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Use Reference Images in Image-to-Video
How to use reference images for image-to-video
Keep Characters Consistent Without Manual Editing
Best reference image-to-video tools (2026): character and product consistency
Image-to-video prompt workflow shown as an editorial motion contact sheet
Image-to-video prompts: 30 examples, templates and fixes
Keep Characters Consistent in AI Video
How to keep characters consistent in AI video
Image-to-video generator comparison cover with a portrait animation interface
9 best image-to-video AI generators (2026): models & costs
Product-video shot concepts: an amber bottle, a white shoe, and a cream jar
AI product video prompts: 8 shots for product photos