Brand Logo

How to Use Nano Banana 2's Video Reference Feature to Generate Thumbnails From Your Own Footage (2026)

Aerin Kim

Written by

Aerin Kim

Nano Banana 2 can now use your own video clips as a reference for image generation. Here is the exact workflow, copy-paste prompts, and mistakes to avoid.

If you make thumbnails for a living, you already know the annoying part of the job is not the AI generation step. It is translating what actually happened in your video into a text prompt an image model can understand. You filmed the exact reaction, the exact product angle, the exact lighting you want on the thumbnail, and then you sit there typing "surprised expression, dramatic lighting, three-quarter angle" trying to describe a moment the model never actually saw.

Google just closed that gap. As of May 28, 2026, when Nano Banana 2 (officially Gemini 3.1 Flash Image) went generally available, it shipped a preview feature that lets you attach an actual video file as a reference for image generation, the same way you would attach a reference photo. The model uses what Google calls deep video understanding to analyze the visual context, the specific subjects, and the actions inside your clip, then generates an image guided by what it actually saw, not what you described from memory. Google's own announcement names thumbnails and rich infographics as the flagship use cases, which is about as direct a signal as a creator-facing feature gets.

This post is a full walkthrough of how the feature actually works, the exact workflow for using it, real prompt phrasing you can copy and adapt, and the mistakes that waste a generation. If you already have a library of general Nano Banana 2 prompts, this is not that post, see Best Nano Banana 2 Prompts 2026 or the Nano Banana 2 prompt templates for YouTube thumbnails for that. This post is specifically about the new capability to hand the model your own footage instead of describing it, and how to build that into your actual thumbnail and product-photo workflow using the AI image generator in Miraflow AI.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-hero-video-reference-tabletop-1.png

What the Video Reference Feature Actually Does

Before this update, an image model like Nano Banana 2 only understood a video the way you could describe it in words, or through a single static frame you manually extracted and uploaded as an image reference. Both of those approaches lose information. Describing a moment in text forces you to translate lighting, framing, and expression into adjectives, which is lossy by definition. Manually grabbing one frame with a screenshot tool gives the model a flat image with no sense of what happened just before or after it, no sense of motion, no sense of which moment in a longer clip actually matters.

Video-to-image generation removes that translation step. You attach the video file itself, either by uploading it directly or, on the underlying Gemini API, by pointing at a public YouTube URL, and the model processes the actual footage. Google's documentation describes the model analyzing video frames in context to extract visual themes and key events, which is a meaningfully different process than looking at one still image. It is watching a sequence, understanding what is happening across it, and then generating a new image informed by that understanding rather than by a single frozen instant you happened to pick.

Two things make this specifically useful for a thumbnail or product-photo workflow rather than just a technical curiosity:

It reads actions, not just objects. The model's stated capability includes recognizing specific subjects and actions within the clip. That means a prompt can reference something that happens in the video, "the moment they hold up the product," "the reaction right after the reveal," rather than only referencing what is visible in one frame.

It carries real, unstaged lighting and camera information. A video shot on your actual desk, in your actual studio, with your actual lighting setup, contains information a text prompt cannot recreate: the real color temperature of your lamp, the real shadow direction, the real lens compression from wherever your camera or phone was actually positioned. When the model pulls from that footage instead of guessing, the output tends to match your real setup instead of defaulting to a generic, stock-photo version of "studio lighting."

The feature currently sits in preview, which in practice on Google's own platforms means it is available to try but still being refined, the same status the standalone 1K/2K generally-available and 4K preview resolution options shipped alongside for Nano Banana 2 and Nano Banana Pro. Preview status is worth knowing about mainly because it means the exact mechanics (video length limits, supported formats) may still shift, so treat the workflow below as the current best practice rather than a permanently fixed spec.

Why This Actually Matters for Creators

It is easy to read "video as a reference input" as a minor technical footnote. For anyone who regularly produces thumbnails, product catalog shots, or a recurring video series, it changes three specific parts of the workflow that used to be genuinely tedious.

Consistent thumbnails pulled from your own best moment, not described from memory. Most creators already know which second of their video has the best reaction, the cleanest product reveal, or the most dramatic camera angle. Until now, turning that specific second into a polished thumbnail meant either screen-recording a frame and cleaning it up manually, or trying to describe it well enough in text that the model could reconstruct something close. Now you can point directly at that moment and ask the model to build the thumbnail from what is actually there.

Product shots that match your demo video's real lighting and angle instead of a generic studio look. If you already filmed an unboxing or demo video, that footage contains real information about how your product actually looks in the setting you shot it in. A generic "studio product photo" prompt throws that away and reconstructs a plausible-looking but disconnected scene. Referencing the actual video keeps your real lighting and angle intact while still letting you clean up the composition for a catalog-ready shot.

A consistent visual "look" extracted from a whole video series, not rebuilt from scratch each episode. Series creators run into a specific version of this problem every single upload: episode six's thumbnail needs to feel like it belongs next to episodes one through five, but episode six's raw footage is a completely different scene, different lighting, different framing than what you filmed months ago. Referencing each episode's own footage while holding the established style constant is a much more direct workflow than trying to redescribe your channel's whole aesthetic in a paragraph every single week.

There is a genuine difference here between a solo creator and a small team running multiple channels. A solo creator mainly benefits from speed, skipping the frame-extraction and manual editing step entirely. A small team running several series benefits more from consistency at scale, being able to hand any editor a video clip and a short style instruction and get something on-brand back, without that editor needing to have memorized the channel's exact color grade and framing rules.

The Actual Workflow, Step by Step

The mechanics are simple once you have done it once, but getting a good result depends on being specific at each step rather than treating the video attachment as a magic "make it good" button.

Step 1: Attach your video reference

Upload the actual clip you want the model to look at, either the full video or, if your tool supports it, a trimmed section around the moment you care about. A shorter, more focused clip generally gives the model less irrelevant footage to sort through than a full ten-minute upload, so trimming to the relevant 15 to 30 seconds before uploading, when you can, tends to produce more targeted results than handing over an entire raw video and hoping the model finds the right second.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-attach-video-reference-contact-sheet-1.png

Step 2: Attach an image reference too, if you need one

This is the step that is easy to miss. Video and image references are not either-or, they can be combined in the same request. If you need a specific, exact likeness locked in, your own face, a specific product's exact packaging, upload a clean photo reference alongside the video. The model can then take one element from the photo and a different element from the video rather than forcing you to choose a single source for everything.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-mixed-references-subject-setting-tags.png

Step 3: Tell the model exactly what to take from each source

This is the single most important part of the whole workflow, and it is also the step most people get vague about. Do not just attach a video and say "make a thumbnail from this." Name, explicitly, what should come from the video and what should come from any image reference: the subject, the pose, the expression, the setting, the lighting, or the camera angle. The table below is a quick reference for how to phrase that.

Reference typeWhat it supplies bestHow to phrase it in your prompt
Video clipA real subject's pose, expression, or action at a specific moment; real ambient lighting and camera angle; a genuine setting"Take the subject's pose from the video at [timestamp]" or "take the lighting and camera angle from this clip"
PhotoA locked, exact likeness, product shape, or outfit you need to stay identical"Take the subject/face/product only from the attached photo"
Video + photo mixedA consistent subject (from the photo) placed inside a real setting or lighting style (from the video)"Subject from the photo, setting and lighting from the video"
Text only, no referenceA scene that does not exist in your own footage yet, at the cost of losing your real lighting and framingFull written description with no attached media

Step 4: Generate, then review against the actual source

Once you generate, compare the result against the real moment in your video rather than just judging it as a standalone image. Does the expression actually match what happened at that timestamp? Does the lighting direction make sense given where your light source actually was? Because this feature is explicitly built to stay grounded in real footage, a result that drifts noticeably from what the video actually shows is usually a sign the prompt was not specific enough about what to pull, not a random model failure to just regenerate and hope fixes itself.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-side-by-side-frame-vs-generated-thumbnail.png

Real Creator Use Cases

The YouTuber turning raw footage into a matching thumbnail

The most direct use case is also the most common one: you have a finished video, you know the exact second that would make a great thumbnail, and you want a polished version of that second rather than the flat, uncolor-graded frame your camera actually captured. Reference that footage directly, name the timestamp or the moment in your prompt ("the moment right after they open the box," "the second their expression changes"), and ask for the lighting and composition boost a real thumbnail needs, more contrast, better framing, reserved space for text, while keeping the actual subject and action true to the source.

This pairs naturally with a clipping workflow if your channel repurposes long-form content into shorts. If you are already running your uploads through AI Clipping on Miraflow to pull viral moments out of a longer video, those same identified moments, the ones already scored highly for being the most engaging seconds of the footage, are exactly the kind of specific, already-proven moments worth referencing for a thumbnail too, instead of re-scrubbing the whole video by hand to find a good frame.

The product seller matching catalog images to their unboxing video's real setting

Anyone selling a physical product on a marketplace or storefront usually has to solve two separate problems: a demo or unboxing video that shows the product in a real, appealing context, and a set of clean catalog images that need to look professional and consistent. Those two assets used to require two entirely separate shoots, or a lot of manual frame-grabbing and Photoshop cleanup to turn video footage into something catalog-ready.

Referencing the unboxing video directly for the catalog shot keeps the real lighting and angle that made the video look good in the first place, while letting the prompt ask for the cleanup a catalog photo actually needs, a plain backdrop, no hands, no background clutter. If you are building a broader catalog and want more general prompt coverage beyond this specific video-reference workflow, our guide to AI prompts for Shopify product photography and the ecommerce product image prompt pack cover the general text-only version of this same job.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-product-catalog-from-unboxing-video.png

The series creator locking a consistent look across every episode's own footage

A recurring series has a specific version of the consistency problem that a one-off video does not: every episode is genuinely different footage, different guest, different location, different lighting, shot weeks or months apart, but the thumbnails need to look like they belong to the same show. Referencing each new episode's actual footage while explicitly instructing the model to match the established color grade, framing, and layout from previous thumbnails keeps the new thumbnail grounded in real, current footage instead of a generic restatement of "match my brand style," which tends to drift more the more times you repeat it.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-series-consistency-episode-thumbnails.png

If your series already has a defined visual identity, our guide on building a consistent YouTube thumbnail style with AI and how to make consistent AI characters across multiple images are both worth pairing with this workflow, since holding a subject's likeness consistent across episodes and holding your channel's visual style consistent are related but slightly different problems, and this feature helps with both at once when you reference both an image and a video source together.

Example Prompts You Can Copy and Adapt

These are written to show exactly how to phrase a video-reference instruction, not as a generic prompt list. Swap the timestamps, subjects, and details for your own footage.

Basic thumbnail from a specific video moment

Attach my uploaded video clip as a reference. From this video, take the subject's pose and the exact framing at the moment they hold up the product around the 0:42 mark. Generate a 16:9 YouTube thumbnail using that pose and framing, boost the lighting to a punchier golden-hour contrast than the original footage, add empty negative space in the upper right third for text, keep the subject's face and outfit exactly as shown in the video, no on-image text baked in.

Naming the exact moment that matters

Use this video as a reference and pull specifically from the moment around 1:15 where the creator's expression is most surprised. Generate a close-up 16:9 thumbnail built from that exact expression and head angle, keep the real background blurred softly behind them, increase color saturation and contrast for thumbnail impact, do not invent a different expression or outfit than what actually appears in the clip at that timestamp.

Mixing a photo reference with a video reference

I am attaching two references: a photo of me and a video clip of my workspace. Take the subject, my face and outfit, only from the attached photo. Take the setting, lighting, and camera angle only from the attached video, specifically the wide desk shot around the 0:20 mark. Combine them into a single 16:9 image of me sitting at that same desk in that same lighting, natural pose, no text, no elements invented outside what these two references actually show.

Product catalog shot matching your demo video's lighting

Attach my product unboxing video as a reference. Take the real lighting, camera angle, and desk surface from the close-up product shot around 0:35 in the video. Generate a clean 4:3 product catalog photo of the same product using that exact lighting direction and angle, remove any hands or background clutter, keep the product's true proportions and color as shown in the video, no invented product features, no text overlay.

Turning a video into a rich infographic

Infographics are the second named use case in Google's own announcement, and the workflow is the same principle applied to a summary instead of a single frame: the model needs to understand the sequence of what happened, not just one still.

Use this video as a reference and analyze the three main steps the creator demonstrates. Generate a single rich infographic-style image that visually summarizes those three steps in a left-to-right sequence, using the real objects and setting shown in the video for each step, clean modern infographic layout, short numbered stage markers, no fabricated steps that do not appear in the footage.

Locking a series style while pulling from this episode's own footage

Attach this episode's raw footage as a video reference. Take the subject and key moment from this clip's 1:05 mark, but match the exact color grade, framing style, and text-safe layout used in my previous thumbnails, warm high-contrast look, tight three-quarter framing, empty space reserved on the left for the episode title, keep this episode's real subject and setting from the video, do not change the established visual style.

Keeping a host consistent across weekly episodes

Attach this week's episode video as a reference. Pull the host's real pose and setting from the segment around 2:30. Keep the host's face, hairstyle, and outfit consistent with how they appear across my last three uploaded episode thumbnails, apply the same punchy contrast and color treatment used on those, 16:9 thumbnail composition, no invented facial features, no altered outfit colors beyond what this episode's footage actually shows.

Pulling only the camera angle from a demo clip

Attach my product demo video as a reference. Take only the camera angle and hand-held tilt from the moment around 0:50 where the product is turned to show its side profile. Generate a static hero product shot using that same angle and tilt on a clean studio backdrop, soft studio lighting instead of the video's ambient light, keep the product's true shape and detail exactly as shown, no text, no added props not present in the source.

Common Mistakes Creators Make With Video Reference Prompts

Being vague about what to take from which source. "Use this video and this photo to make a thumbnail" leaves the model guessing which element comes from where. Every prompt in this post explicitly names what to pull from the video and what to pull from any image reference, and that specificity is the difference between a result that looks like your actual footage and one that looks like a generic average of both inputs.

Expecting the model to invent facts the video never showed. If your clip never actually shows the product's back side, asking for "the product from a back angle" is asking the model to guess rather than reference. The feature is grounded in what deep video understanding can actually extract from the footage you gave it, not a general knowledge base about what your product probably looks like from every angle. For anything the video does not show, either film that angle too or switch to a text description for that specific detail, and cross-check the output against your source before publishing.

Using a low-quality or shaky source clip and expecting a sharp result. Motion blur, poor lighting, or a heavily compressed low-resolution export all limit how much real information the model has to work with. A clean, reasonably well-lit source clip, even a short one, generally outperforms a longer but shaky or dim clip.

nano-banana-2-video-reference-feature-thumbnails-guide-2026-common-mistakes-shaky-vs-sharp-source.png

Not specifying which frame or moment actually matters. Handing over ten minutes of raw footage with no timestamp guidance forces the model to guess which second you care about. Naming a rough timestamp, or describing the specific action ("the moment they turn to face the camera"), narrows that down dramatically and is worth the extra sentence in your prompt every time.

Forgetting that this is still a generation, not an extraction. Even when grounded in real footage, the output is a new generated image, not a literal enhanced screenshot. Review results the same way you would review any AI generation before publishing, checking that details like text, logos, or small product features rendered correctly rather than assuming accuracy because the source was real.

Where This Fits Into a Miraflow Workflow

If you already produce your own footage inside Miraflow, this feature is a natural extension of tools you may already be using rather than a separate step bolted on. Creators using AI Clipping to turn a long upload into several ranked, viral-moment shorts already have a set of specific, high-scoring moments identified in their own footage, exactly the kind of clip worth referencing directly when generating a matching thumbnail for one of those shorts. The same applies to short-form videos built with Text2Shorts: once you have your own script-to-visual short generated, that footage is a legitimate reference source for a thumbnail that actually matches what the short shows, rather than a generic thumbnail generated from a text description of the topic.

From there, the AI image generator in Miraflow AI is where the actual generation happens, including the image-to-image and inpainting tools worth combining with this workflow: generate a first version from your video reference, then use inpainting to mask and adjust just the one region that needs a touch-up rather than regenerating the whole image. If you want a deeper look at that masking workflow specifically, see how to use Nano Banana image inpainting on Miraflow AI. And once you have a thumbnail you're happy with, the YouTube Thumbnail Maker on Miraflow is built specifically for adding clean, readable text and finishing touches on top of it.

Frequently Asked Questions

Do I need a special plan or setting to use video references? The capability is tied to the underlying model itself rather than being a separate toggle, so as long as the tool you are using supports attaching a video file as a reference, alongside the usual text, image, and PDF reference options, you can use it the same way you would attach an image reference.

Can I use a YouTube link instead of uploading my own file? On Google's own platform, a public YouTube URL can be passed directly as a video reference in addition to uploading a local file. That is worth knowing if you want to reference a public video, though for your own unpublished footage, direct upload is the more reliable path.

Does the model actually watch the whole video, or just grab one frame? It is meaningfully different from a single-frame grab. Google describes the process as deep video understanding, analyzing visual context, subjects, and actions across the footage rather than freezing on one instant, which is why naming an action or a described moment in your prompt works better than only naming a timestamp.

Can I combine more than one video reference in the same request? The core capability described by Google is video plus image plus text references combined, with the model told what to take from each. Whether multiple separate video files can be combined in one request is more likely to depend on the specific tool's interface than on the underlying model, so check the upload limits of whatever tool you are using.

Is this better than just describing my video in a text prompt? For anything where the real lighting, a specific expression, or an exact camera angle matters, yes, because a text description is always a lossy translation of what actually happened on camera. For a completely new scene that does not exist in any of your footage yet, a text-only prompt is still the right and only tool for the job.

Will this replace manually picking a good thumbnail frame? Not entirely. You still need to know which moment in your video is worth building a thumbnail around, that editorial judgment does not go away. What changes is the step after that: instead of manually extracting and cleaning up that frame, you can reference the moment directly and let the generation do the cleanup and lighting work for you.

Conclusion

The gap between "what actually happened in my video" and "what I can describe well enough in a text prompt" has been a quiet, constant tax on creators using AI image tools since generative thumbnails became popular. Video-to-image generation on Nano Banana 2 is a direct fix for that specific problem, not a new gimmick feature, and Google naming thumbnails as one of its two flagship use cases is a strong signal of exactly who it was built for. The workflow itself is simple: attach the real footage, name what to pull from it, mix in an image reference when you need an exact likeness locked in, and generate. If you already have raw footage sitting in your library, whether from a full upload, a short built with Text2Shorts, or a clip pulled out with AI Clipping, that footage is now a direct input for your next thumbnail or product shot, not just raw material you have to describe from memory.