TikTok's New AI Create Tools: Image-to-Video, Text-to-Video, and AI Transitions Explained
Written by
Aerin Kim

TikTok is rolling out native Image-to-Video, Text-to-Video, and AI Transitions tools in its Create flow. Here is how each one works, real example prompts, and when to use a dedicated tool instead.
TL;DR
TikTok is rolling out three native AI video generation tools built directly into its Create flow: Image-to-Video, Text-to-Video, and AI Transitions. Instead of opening a separate app, you tap the "+" button to start a new post, choose Create, and an AI Create option now surfaces all three inside the same composer you already use to film and edit. Image-to-Video animates a photo you already have. Text-to-Video builds a clip from a written description with no starting image at all. AI Transitions generates a creative morph between two clips or frames instead of a plain cut or swipe.
This is a separate system from TikTok's 2025 AI Alive feature, which only lives inside TikTok Stories and only does one thing, animating a single still photo with ambient motion. The new Create-flow toolset sits in the main video composer, covers three distinct generation modes, and produces clips meant for full TikTok posts rather than a 24-hour Story. Below is a real walkthrough of how each tool works, example prompts you can actually copy, how to find the tools in your own app, an honest comparison against a dedicated platform like Miraflow, and the mistakes creators are already making with all three.

What Changed: TikTok's Native AI Create Tools
App researcher Jonah Manzano was the first to surface screenshots of this rollout, and Social Media Today's Andrew Hutchinson covered it on November 4, 2025, describing TikTok as "rolling out some new generative AI creation options, built into the TikTok composer, including image-to-video and text-to-video generation," alongside a third option for AI-generated transitions between clips. That reporting is worth taking at face value rather than treating as a rumor, because it lines up with something TikTok had already shipped elsewhere. The same three capabilities, image-to-video, text-to-video, and AI transitions, were built for TikTok Symphony, the platform's generative AI suite for advertisers, back in June 2025. What changed in late 2025 and through 2026 is that TikTok started extending those same generation tools into the ordinary creator composer, the one anybody uses to post a normal video, not just the ad-buying tools inside TikTok's business suite.
That distinction matters for how you should think about the rollout. Rather than building a brand new model from scratch for regular creators, TikTok took generation technology that was already proven out on the advertiser side and started making it available, gradually, to everyone else. Gradual is the operative word. Screens and reports through the first half of 2026 describe a staged rollout rather than a single global switch-flip, so it is entirely normal for one creator to have the AI Create option sitting in their composer while a friend on the same TikTok version does not see it yet. If you do not have it today, that is far more likely to mean your account has not been reached by the rollout than that TikTok removed the feature or reserved it for a narrow group.
It is worth being precise about what this toolset actually is, because the most common mix-up creators will make is assuming this is just a bigger version of AI Alive, TikTok's photo-animation feature from May 2025. AI Alive lives specifically inside the Story Camera, the tool behind the blue plus icon on your Inbox or Profile tab, a different part of the app entirely from the main video composer. It only takes a single still photo, applies TikTok's own ambient motion effects (drifting clouds, a subtle blink, a slow push-in), and posts the result as a 24-hour Story. Our full breakdown of that feature, TikTok's AI Alive and 20 living photo prompts, covers exactly how it works and where its limits sit.
The Create-flow tools covered in this guide are a genuinely different, broader system. They live in the composer you use for a normal TikTok video post, not the Story Camera. They cover three separate generation modes rather than one. Two of the three, Text-to-Video and AI Transitions, do not require a starting photo at all, something AI Alive cannot do under any circumstances. And the output is meant to become part of an actual TikTok video post, not an ephemeral Story that disappears after a day. If you have already tried AI Alive and found it charming but limited, treat this newer toolset as a separate system built for a different, more general purpose rather than an upgrade to that same feature, and understanding that distinction up front will save you time looking in the wrong part of the app.
Image-to-Video: How It Works
Image-to-Video takes a photo you already have, a product shot, a portrait, an old family photo, a screenshot of a graphic, and turns it into a short generated video clip. Inside the composer, you select the tool, choose a photo from your camera roll or from within TikTok, and add a short text description of the motion you want. TikTok's model then generates a clip built around that still image, adding camera movement, subject motion, or environmental effects based on what you asked for, rather than the fixed, non-customizable animation style AI Alive applies automatically. Early reporting on the equivalent Symphony version of this tool describes it producing short, TikTok-first clips built from an image and a brief prompt, of a length suited to being stitched together with other footage rather than standing alone as a full-length video.
The most practical use case is turning a photo you cannot easily reshoot into something that moves. A small business posting a product photo that was never filmed as video, a creator with a great single frame from an old shoot but no matching footage, or someone who wants to give a static quote card or announcement graphic a bit of life instead of posting it as a flat image, all of these are exactly what this tool is built for. You are not asking the AI to imagine a scene from nothing, you are asking it to extend a real photo you already trust into motion.
A useful example is a product photo you want to turn into something that feels like a short ad clip rather than a still listing photo. You would upload the photo, then describe the kind of subtle motion that reads as intentional rather than gimmicky.
A slow, subtle push-in on the product, soft studio light shifting gently across the surface, a faint reflection moving across the glass, gentle steam or light dust motion in the background, no other movement, camera stays level and centered, no warped edges or melting details.

Portraits and lifestyle photos work a little differently, and it is worth being deliberate about what you ask for, since a face is the fastest way to expose a shaky generation. Small, natural motion tends to hold up far better than anything ambitious.
Gentle head turn toward the camera, natural blink, hair moving slightly as if in a light breeze, soft background bokeh with a subtle warm light flare passing across the frame, keep facial features stable and consistent, no distortion around the eyes or mouth.
The genuine limitation to know about before you rely on this tool for something important: image-to-video generation across every consumer tool in this category, TikTok's included, tends to hold up best with a single clear subject against a relatively simple background. A photo with several people, cluttered detail near the edges of the frame, or a lot of fine texture close to the camera is more likely to produce warping around hands, fabric, or background elements once motion is introduced, the same failure mode that shows up across nearly every image-to-video model on the market right now, not something unique to TikTok's implementation. Treat a cluttered or complex source photo as a candidate for a redo rather than assuming a bad first result means the tool is broken. It is also worth remembering that TikTok's version is generating a short clip meant to live inside a normal post's timeline, not a long-form video, so plan for it to be one beat inside a larger edit rather than the entire video on its own.
Text-to-Video: How It Works
Text-to-Video skips the photo step entirely. You type a description of a scene, TikTok generates a video clip built from that description alone, with no reference image involved at all. This is the tool for a video idea you have in your head but no footage or photo to start from, a scene you cannot realistically film, an establishing shot to open a longer edit, or a visual metaphor that would otherwise require stock footage you do not have rights to use.
Where Image-to-Video is really an extension of a photo you already trust, Text-to-Video is closer to describing a shot to a cinematographer who cannot ask follow-up questions, so the more concrete and specific the prompt, the better the result tends to be. Vague prompts describing a mood ("something calming") tend to produce generic, forgettable footage. Prompts that specify the actual scene, the lighting, and what is or is not moving in frame tend to produce something usable.
A cozy rain-streaked coffee shop window at dusk, warm string lights reflecting on wet glass, steam rising from a cup on a nearby table, a slow static wide shot with soft ambient light, calm and quiet mood, no visible people, no text overlays.

Text-to-Video is also a genuinely good fit for abstract or conceptual visuals that would otherwise be difficult to source at all, not just realistic scenes. A creator explaining a concept, growth, transformation, decay, momentum, can describe a literal visual metaphor and get a usable establishing clip for a video essay or explainer without ever picking up a camera.
A single seed cracking open in rich dark soil and a small green sprout unfurling upward in time-lapse style motion, soft overhead natural light, shallow depth of field with the sprout in focus, no other objects entering frame, no flickering or warped stem shape.
The edge case worth planning around here is consistency, and it is the single biggest reason Text-to-Video is not a full replacement for scripted, planned production. Because there is no reference image or character anchor involved, generating the same described person or object twice, in two separate prompts, is not guaranteed to produce a visually consistent result each time. A face, an outfit, or a specific object can shift in small but noticeable ways between generations, which matters enormously if you are trying to build a multi-shot sequence that is supposed to feature the same subject throughout. TikTok's tool, like most consumer text-to-video generators right now, is built around a single clip at a time, not a continuous, identity-locked scene across several generations. If your video idea depends on the same character or product appearing consistently across multiple shots, that is a real constraint worth knowing about before you commit a whole concept to it.
AI Transitions: How It Works
AI Transitions is the third tool, and it behaves differently from the other two because its actual job is generating a bridge between two things you already have, rather than conjuring a scene from scratch. You provide two clips, or effectively a first frame and a last frame, and describe how the first should transform into the second. TikTok's model then generates a short, AI-driven morph connecting the two, rather than a standard hard cut, dissolve, or the built-in swipe and wipe transitions TikTok's editor already offers. Jonah Manzano's own framing of the feature, shared alongside his screenshots, captures the mechanic simply: describe how your first frame should transform into the last.
Social Media Today's coverage flagged this as the tool with the most obvious staying power of the three, and that assessment holds up once you think through the use case. A creative transition effect solves a real, recurring editing problem, connecting two shots in a way that feels intentional rather than abrupt, instead of offering novelty for its own sake. An outfit change reveal, a before-and-after transformation, a day-to-night scene shift, or a product unboxing moment are all classic transition-driven content formats that TikTok creators have been building manually for years using timed cuts and practiced hand movements. AI Transitions is TikTok's attempt to generate that same payoff automatically, from a description, instead of requiring the shot-matching and timing precision a manual transition normally demands.
First frame: person standing in casual everyday clothing in a plain room. Last frame: same framing and pose, now in a formal outfit in the same room. Transition should feel like a smooth stylistic wardrobe change, not a hard cut, keep the background and pose steady throughout, no warped limbs or facial distortion during the morph.

A second real use case worth naming directly is a product reveal, since it is one of the cleanest possible applications of a first-frame-to-last-frame morph.
First frame: closed product box sitting centered on a table. Last frame: same table and camera angle, product now unboxed and displayed upright. Transition should feel like the box dissolving away to reveal what's inside, smooth and continuous, no flickering, no warped product shape.
The edge case to plan for here is how much the two endpoints have in common. A transition works best when the first and last frame share real structural similarity, the same framing, the same subject position, a similar camera angle, because the model is essentially interpolating a path between two known points. Feed it two clips with wildly different framing, an unrelated subject, or a completely different camera angle, and the resulting morph is far more likely to look messy or visually confusing rather than smooth, since there is no coherent path for the model to invent between two genuinely unrelated images. Shooting or selecting your two source clips with the transition already in mind, matching the framing deliberately before you ever open the AI Transitions tool, is the single biggest thing you can do to get a clean result rather than a muddled one.
How to Access These Tools
The rollout is staged, so the exact menu wording and icon placement can vary slightly by app version and by how far the feature has reached your account, but the general path reported across multiple early testers is consistent enough to walk through step by step.
- Make sure your TikTok app is fully updated. Since this is an actively rolling-out feature, an outdated app build is one of the most common reasons someone does not see it yet.
- Tap the "+" icon at the bottom of the screen to start a new post, the same button you would tap to film or upload a normal video.
- Look for a "Create" option, followed by an "AI Create" or similarly labeled AI video option inside that flow. On some accounts this surfaces as a small dedicated AI icon at the top of the composer rather than a separate menu screen.
- Choose which of the three tools you want: Image-to-Video, Text-to-Video, or AI Transitions. These typically appear as distinct options within the same AI Create entry point rather than three separate buttons scattered around the app.
- Follow that tool's specific steps. For Image-to-Video, select a photo and add your motion prompt. For Text-to-Video, type your scene description directly with no photo required. For AI Transitions, select your two clips or frames and describe the transformation between them.
- Set your preferred aspect ratio before generating, since the composer lets you adjust this rather than locking you into one fixed shape.
- Generate the clip. Reports describe generation taking anywhere from under a minute to several minutes depending on load, and you can typically back out of the screen while it processes, with the finished clip landing in your drafts rather than forcing you to wait on the loading screen the whole time.
- Review the result inside your normal TikTok editing timeline. From here you can trim it, add more clips, apply captions, add a sound, and post it exactly like any other video.
If you have updated the app and still do not see any AI Create option after checking carefully, the most likely explanation is simply that the staged rollout has not reached your account yet, not that anything is misconfigured on your end. Checking back every week or two as the rollout continues through 2026 is a more useful strategy than assuming you are permanently excluded.

One practical detail worth flagging while you are testing these tools: TikTok has required visible labels on realistic AI-generated content since 2024, using Content Credentials metadata built on the C2PA standard to automatically detect and mark AI-made media, including content produced through TikTok's own in-house AI effects. TikTok's own explanation at the time of that rollout was direct about extending the same labeling to its own tools, not just third-party AI content uploaded from elsewhere. It is reasonable to expect clips made with Image-to-Video, Text-to-Video, and AI Transitions to carry that same automatic labeling once posted, particularly anything that reads as a realistic person, place, or event, so plan your captions accordingly rather than being caught off guard by a label appearing on a post you did not manually disclose.
When to Use TikTok's Native Tools vs a Dedicated Platform Like Miraflow
The honest answer here is not that one option is simply better than the other. It is that TikTok's native tools and a dedicated generation platform are built to solve different problems, and knowing which problem you actually have decides which one to reach for.
TikTok's Image-to-Video, Text-to-Video, and AI Transitions win on convenience and speed above everything else. There is no export step, no separate app to open, no account to set up somewhere else. You are already inside the app you are about to post to, the generated clip lands directly in your existing timeline, and you can go from an idea to a posted video in a single session without ever leaving TikTok. For quick trend participation, a one-off transition effect, or turning a single photo into a moving clip for a post that does not need a lot of planning, that immediacy is a genuine advantage a separate tool cannot match, because a separate tool always adds at least one extra step of exporting footage and bringing it back in.
Where the native tools run into real limits is anything that requires planning across more than one shot. Each of the three tools generates one clip, or one transition, at a time, with no built-in way to plan a full script, break it into scenes, and generate a consistent sequence of shots that are supposed to work together as a single narrative. There is also no meaningful choice of underlying model. You get whatever TikTok's in-house system produces, with no option to switch to a different engine suited to a different visual style or level of realism. And for a video that involves narration, TikTok's native tools do not offer any voice or voiceover generation at all, only the video clip itself.
This is exactly the gap a dedicated platform closes, and it is worth naming specifically rather than vaguely. Text2Shorts on Miraflow starts from the opposite direction of a single quick clip: you enter a topic, Miraflow generates a full script first, you can edit or regenerate that script before anything visual gets made, then Miraflow generates scene-by-scene visual prompts based on that finished script, lets you edit those prompts individually, and only then produces the full video with a chosen voice and speed. That is an actual pre-production step, script first, then scenes, then voice, that TikTok's in-app tools skip entirely by design, because they were built for a single quick generation rather than a planned multi-scene video. Our walkthrough on going from a prompt to a finished one-minute AI Short with Text2Shorts covers that full workflow in more depth.
Model choice is the second real gap. Miraflow's cinematic AI video generator runs on Veo3 and Veo3.1, models built for cinematic, higher-fidelity output, the kind of look creators reach for on product ads, storytelling sequences, and short cinematic scenes where a five-second in-app clip is not the goal, a genuinely polished shot is. TikTok's native tools have no equivalent option to pick a different underlying model suited to that kind of output. If your Image-to-Video attempt inside TikTok keeps producing a starting photo that is not clean enough to animate well, Miraflow's AI Image Generator also gives you text-to-image, image-to-image editing, and masked inpainting to fix or refine that source image first, replacing a cluttered background, adjusting an object, or cleaning up a detail, before you ever try to animate it, something the TikTok composer has no equivalent tool for at all.
There is a genuinely reasonable middle path here too, and it is worth naming instead of treating this as an either-or choice. Plenty of creators will use TikTok's native AI Create tools for fast, in-the-moment posts and trend participation, while turning to a dedicated platform specifically for anything meant to be a planned, recurring content pillar, a weekly explainer series, a product launch video, a scripted narrative short, where the extra planning step actually pays off. If you already have long-form content sitting around and want more short clips out of it rather than generating brand new footage from scratch, AI Clipping analyzes a full video, finds and scores the strongest moments, and automatically produces several captioned, vertical clips from a single upload, a different but related problem from generating a new clip out of nothing. Our guide to turning long videos into viral shorts automatically with AI Clipping walks through that workflow specifically. And once a video is finished, whether it came from TikTok's native tools or a full Text2Shorts production, adding an original soundtrack through Miraflow's AI Music Generator is a fast way to give it a finished, less generic-sounding feel than whatever trending audio everyone else on your For You Page is already using.

Common Mistakes Creators Will Make With These Tools
- Looking for AI Create inside TikTok Stories. Because AI Alive already lives in the Story Camera, some creators will assume the new Image-to-Video, Text-to-Video, and AI Transitions tools live there too. They do not. They sit inside the main video composer behind the "+" button and Create flow, an entirely different part of the app.
- Trying to build a multi-scene narrative with a single tool built for one clip at a time. Generating five separate Text-to-Video clips and expecting them to feel like a coherent, planned sequence usually produces something visibly disjointed, since nothing in the tool plans continuity between separate generations. A scripted, scene-by-scene workflow is the actual fix for that problem, not generating more individual clips and hoping they cut together.
- Assuming the feature is broken instead of checking for a staged rollout. Given how gradually this has expanded since late 2025, not seeing the option yet is far more often a rollout timing issue than a bug, and updating the app and checking again in a week or two is a better first move than concluding the feature does not exist for your account.
- Feeding Image-to-Video a cluttered or busy source photo and being surprised by warped results. Complex scenes with multiple subjects or dense background detail are the most common source of distorted hands, faces, or edges once motion gets added, a pattern shared across nearly every image-to-video tool right now, not a TikTok-specific flaw.
- Pairing two visually unrelated clips for AI Transitions and expecting a clean morph. The tool interpolates between two known points, so mismatched framing, angle, or subject position between the first and last frame tends to produce a messy result rather than a smooth one. Matching your two source clips deliberately before you generate is the actual fix.
- Posting AI-generated content without accounting for the automatic AI label. Since TikTok's Content Credentials system has labeled AI-made media, including content from TikTok's own tools, since 2024, assuming a clip you generated will post without any visible disclosure is a planning mistake that shows up as a surprise label on a finished post rather than something you controlled up front.

Frequently Asked Questions
Is TikTok's new AI Create toolset the same thing as AI Alive? No. AI Alive is a separate, older feature from May 2025 that lives specifically inside the Story Camera and only animates a single photo into an ambient 24-hour Story. The Image-to-Video, Text-to-Video, and AI Transitions tools covered here live inside the main video composer behind the "+" and Create flow, cover three distinct generation modes, and produce clips meant for regular TikTok posts.
How do I actually find these tools in my app? Update TikTok to the latest version, tap "+" to start a new post, and look for a "Create" or "AI Create" option inside that flow, sometimes shown as a small AI icon at the top of the composer. Since the rollout is staged, not every account sees it at the same time.
Does TikTok charge for these tools, or are there usage limits? TikTok has not published specific credit costs or generation caps for these tools as of this rollout, so treat any exact number you see elsewhere as unconfirmed. What is consistent across early reporting is that the tools sit inside the free, standard TikTok app rather than a separate paid product.
Will videos made with these tools show an AI-generated label? TikTok has automatically labeled realistic AI-generated content since 2024 using Content Credentials metadata, and has said this covers content made with its own AI effects, not just AI content uploaded from elsewhere. It is reasonable to expect clips made with these Create-flow tools to carry that same labeling, especially anything depicting a realistic person, place, or event.
Can I use any photo I want for Image-to-Video? You select from your own camera roll or existing TikTok content, so the tool works with whatever photo you provide. Results tend to hold up best with a clear single subject and a relatively simple background, and get less reliable with cluttered, busy source images.
Do these tools support a full multi-scene video, or just one clip? Each tool generates one clip, or one transition, at a time. There is no built-in scripting or scene-planning step across multiple generations, which is the main reason a dedicated platform is a better fit for anything longer than a single clip or transition.
What's the real difference between using TikTok's native tools and a platform like Miraflow? TikTok's tools are fastest for a single quick clip made without leaving the app. A platform like Miraflow adds an actual pre-production step, a full script you can edit, scene-by-scene visuals, a choice between models like Veo3 and Veo3.1, and voice generation, which matters once your idea needs more than one consistent shot.
Are these tools available worldwide yet? Reporting through 2026 describes a progressive, staged global rollout rather than a single simultaneous launch everywhere, so availability by region and by account varies as TikTok continues expanding access.
Conclusion
TikTok folding Image-to-Video, Text-to-Video, and AI Transitions directly into its Create flow is a meaningful expansion of what any creator can do without leaving the app, and it is worth learning properly rather than dismissing as a novelty or confusing with AI Alive. Each of the three tools solves a genuinely distinct problem, animating a photo you already trust, building a scene from a written idea alone, or generating a creative bridge between two clips, and each one has real, specific limits worth planning around rather than discovering by accident. For fast, in-the-moment posts and trend participation, these native tools are hard to beat on convenience. For anything that needs a planned script, a consistent multi-scene sequence, a specific model like Veo3.1, or an actual voiceover, a dedicated workflow through tools like Miraflow's Text2Shorts, cinematic AI video generator, and AI Image Generator is where the extra planning step actually pays off. Knowing which problem you have, a quick single clip or a planned piece of content, is the real skill here, and it is worth revisiting our guides on the TikTok algorithm in 2026, TikTok best practices for creators, what actually makes AI videos go viral on TikTok and YouTube, and TikTok RPM and monetization in 2026 as you build these new generation tools into a real strategy rather than a one-off experiment.

