AI Image Generation for Storyboards: First Frame, Last Frame and Character References
Most AI video creators hit the same wall. The script is written, the shots are planned, and then the first generation comes back looking nothing like what you imagined. The character's face changed. The scene drifted. The cut between two shots feels like two different films stitched together.
The fix usually isn't a better prompt. It's better storyboard images. In this walkthrough, I'll show how I use AI image generation inside 墨川导演台 (MiniMax H3 AI Director's Desk) to lock down three things before I ever hit "generate video": the first frame, the last frame, and the character reference. Think of these as the anchor points of a shot — if the anchors are solid, the video has somewhere to land.
Why Storyboard Images Come Before Video
A storyboard image is a contract. It tells the model: *this is where the shot starts, this is where it ends, and this is who's in it.* Without that contract, the video model improvises, and improvisation across multiple shots is how you get continuity chaos.
Here's the analogy I keep coming back to: generating video from text alone is like asking an actor to perform a scene they've never read. Generating video from a storyboard image is like handing them a marked-up script with blocking notes. Same actor, wildly different result.
Inside the director's desk, storyboard images feed directly into the 剧本库 (Script Library) as storyboard cards, and those cards connect to the 项目库 (Project Library) where generation tasks and versions live. So the images aren't decorative — they're the input layer for everything downstream.
Step 1: Build the Character Reference First
Before you touch a single scene, build your character. This is the step people skip, and it's the step that saves the most time.
Open the 资产库 (Asset Library). This is where characters, scenes, props, and reference videos live. Create an entry for each main character. Then head to AI 制图 (AI Image Generation) and use text-to-image to produce a clean, front-facing reference — neutral expression, even lighting, plain background. That image becomes your character reference.
A few practical tips from experience:
- One character per reference image. Don't try to sneak two people into a single reference. It muddies the signal.
- Describe the face, not the mood. "Mid-30s, narrow jaw, deep-set eyes, short black hair" works better than "a mysterious man." Mood belongs in the shot prompt, not the identity reference.
- Generate a few variants, then pick one. Upload the winner to the Asset Library and treat it as canon. If you later need a different angle, use image-to-image starting from that same reference rather than starting over from text.
If you're running this on a cloud GPU box, this is also the moment to confirm your 云部署 (Cloud Deployment) connection is live — SSH verified, cloud ComfyUI showing green. Reference images are cheap to make; regenerating a whole batch because your session dropped is not.
Step 2: Generate the First Frame
Now you're ready for the actual shot. In AI 制图, switch to text-to-image and write a prompt that describes the *opening moment* of the shot — not the whole scene, just the first beat.
A structure that works well:
[Shot type] + [Character + reference] + [Action at frame 1] + [Environment] + [Lighting] + [Lens feel]
Example: *"Medium shot, [character ref], standing at a rain-soaked window, hand resting on the frame, dim interior with cool blue window light, shallow depth of field, 35mm feel."*
Notice what's doing the heavy lifting: the character reference handles identity, and the prompt handles everything else. That division of labor is the whole trick. If you try to describe the character's face again in the prompt, you're competing with your own reference image.
Once you have a first frame you like, save it to the storyboard card in the 剧本库. This is your frame 1.
Step 3: Generate the Last Frame
Here's where most tutorials stop, and where continuity actually gets won.
The last frame is the destination. It's the pose, expression, and composition you want the shot to *arrive* at. Generate it the same way — text-to-image, same character reference, same environment language — but describe the end state of the action.
If the shot is "she turns from the window to face the room," the first frame is her at the window in profile, and the last frame is her facing camera, room visible behind her.
Two rules I follow:
- Keep environment and lighting language identical between first and last frame. Only the action and framing should change. Drift in the environment is what makes cuts feel wrong.
- Use image-to-image when the two frames are close. If the last frame is a small evolution of the first, feed the first frame in as the base and describe the change. It holds composition much more tightly than a fresh text-to-image pass.
Some creators also generate a mid-frame or a character reference sheet at a different angle for tricky shots. That's optional, but the Asset Library is there if you want to keep those variants organized.
Step 4: Wire the Frames Into a Shot and Generate
With first frame, last frame, and character reference all saved, go to the 项目库. Create or open the project, add the shot as a generation task, and attach the frames. You can run shots one at a time (单镜生成) or queue them in batch (批量生成).
A few workflow notes:
- Batch when the shots share assets. If five shots use the same character and location, batch them so the references stay consistent across the queue.
- Single-shot when you're still dialing in. Early in a project, generate one shot, watch it, adjust the frames, then move on. Batch is a reward for having your anchors right.
- Watch the AI 日志 (AI Logs). Real-time generation logs and system status live there. If a shot fails, the log usually tells you why before you waste a second attempt.
For dialogue-heavy shots, this is also where AI 配音 (IndexTTS 2.5 multilingual voiceover with emotion control) and 对口型 (lip sync) come in — but that's a separate pass. Get the visuals locked first.
Common Mistakes and How to Avoid Them
Mistake 1: Regenerating the character every shot. Fix: build the reference once, reuse it everywhere.
Mistake 2: Letting the prompt fight the reference. Fix: if you've attached a character reference, don't re-describe the face in text.
Mistake 3: First and last frames that don't share a world. Fix: copy the environment and lighting phrases verbatim between the two prompts.
Mistake 4: Skipping the storyboard card. Fix: save every approved frame to the 剧本库. If you don't, you'll be re-generating the same shot next week from memory.
Mistake 5: Running out of GPU mid-batch. Fix: verify your 云部署 SSH connection and cloud ComfyUI status before kicking off a long queue. The director's desk supports 本地部署, 云部署 (including AutoDL, Alibaba Cloud, Tencent Cloud, 仙宫云, and custom SSH servers), and API 调用 — pick the one that matches your hardware and budget. For specifics on setup, check the official site or ask.
Putting It Together
The pattern is simple: character reference → first frame → last frame → shot generation. Each step narrows the model's freedom a little more, until the only thing left to vary is the motion itself. That's exactly what you want.
If you're just starting, do this on a single shot today. One character, one location, two frames. Watch how much more stable the result is compared to a pure text-to-video attempt. Then scale it to a full scene, then a full episode.
For setup questions, cloud deployment help, or to see the current feature set, head to the official site at www.ymcdirector.com. You can also reach the team on WeChat at ymcandai, or find them on Taobao by searching 「杨墨川AI」. The blog at blog.ymcdirector.com is worth bookmarking too — that's where workflow updates tend to land first.
Storyboarding with AI isn't about replacing the director. It's about giving the director better anchors. Build the anchors, and the video takes care of itself. 🎬