Brand Logo

DreamX-Creator Explained: How a 7B Model Syncs Audio and Video

Aerin Kim

Written by

Aerin Kim

DreamX-Creator's 7B model jointly generates audio and video through Gated Cross-Modal Attention, then refines to 2K in one denoising step. Here is how the mechanism actually works.

Open any AI video tool right now, generate a ten-second clip of someone talking, and watch the mouth. On most tools, including the ones creators reach for every day, the lips move roughly in time with the words, but not quite. A door slams half a beat after it visually closes. A footstep lands slightly off the frame where the foot actually touches the ground. None of this is a bug in the traditional sense. It is the direct, structural consequence of how almost every AI video system alive today is built: a video model generates the picture, and then a completely separate audio model, or a completely separate lip-sync pass, gets bolted on afterward to guess where the sound should go [8][9].

A paper published on arXiv on August 31, 2026 tackles that exact structural problem head-on, and it is worth understanding in real technical depth rather than skimming the abstract [1]. DreamX-Creator 1.0, from a team led by Jiashu Zhu and nine co-authors, is titled plainly: "DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution" [1]. It describes a 7-billion-parameter generator that takes a first frame and a text prompt and jointly denoises audio and video latents together, in the same forward pass, coupled through a mechanism the authors call Gated Cross-Modal Attention, then post-trains that generator with a reinforcement learning scheme that routes feedback separately to video, audio, and cross-modal signals, and finally upsamples the result to 2K resolution using a distilled refiner that needs only one denoising step per chunk of video at inference time [2].

That is a lot of new terminology in one paragraph, and the rest of this post exists to actually unpack it, mechanism by mechanism, rather than repeat the abstract in different words. Why does joint denoising plus gating actually fix synchronization, instead of just being a fancier way to combine two signals that are still fundamentally separate? What does "autoregressive 1-step" distillation really mean, and why does it matter that the model is 7B when comparable systems run to 22B, 33B, or over 100B parameters [2]? Where does this system actually fall short of the larger models it gets compared against, according to the paper's own numbers? And for anyone actually building with AI video tools today, including Miraflow's own Cinematic AI Video Generator, which currently runs on Veo 3 and Veo 3.1 for the video itself with voiceover handled as its own separate step, what does this paper actually explain about why some tools nail audio timing and others don't. Miraflow does not use DreamX-Creator internally, but understanding what "solving" native audio-video sync actually requires is genuinely useful context for anyone generating video with any tool right now.

dreamx-creator-native-audio-video-generation-2k-explained-2026-hero.png

Step 1: Why Video and Audio Generation Split Apart in the First Place

To understand why DreamX-Creator's approach matters, it helps to be specific about what "generating audio and video separately" actually means under the hood, because it is not one single failure mode, it is at least three distinct ones stacked on top of each other.

The post-hoc pipeline problem. The dominant pattern in AI video generation today looks like this: a diffusion or autoregressive video model generates silent frames first, informed only by a text prompt and maybe a reference image. Then, in a completely separate stage, an audio model looks at those already-finished frames and tries to guess what sound should accompany them. MMAudio, a well-regarded open video-to-audio model published in late 2024 and still widely used as a component in production pipelines, is a clean example of this pattern done well: it takes a finished silent video plus an optional text prompt and generates synchronized audio for it using a conditional synchronization module that aligns video conditions with audio latents at the frame level [7]. That is genuinely good engineering, and MMAudio's own benchmarks show it beating earlier video-to-audio systems on semantic alignment and synchrony. But structurally, it is still working from a fixed, already-decided video. It cannot go back and adjust a mouth shape or a footstep's timing to better match the sound it is about to generate, because the video already exists by the time audio generation starts. The DreamX-Creator paper frames this directly as a limiting design choice across the field: "recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events" [1].

The lip-sync-as-a-patch problem. A second, closely related pattern shows up specifically for talking-head and dialogue content, where a separate lip-sync model (rather than a general audio model) reanimates a mouth region to match an already-recorded or already-generated audio track. Industry coverage of this space in 2026 describes lip sync itself as largely a solved surface-level problem on most leading tools, but notes that the harder issues sit just underneath it: voice tone that does not match facial expression, frozen or generic expressions outside the mouth region, and avatar inconsistency across cuts, because the audio and the face were never actually reasoned about together as one generative process [8]. A side-by-side test of leading AI video models specifically for lip sync on ad-style UGC content reached a similar conclusion: syncing two or more faces with overlapping dialogue in the same shot produces visible errors on every tool tested, and subtle emotional tone rarely transfers reliably from voice to face on any current system, precisely because the "sync" step is a correction applied after generation rather than something generated jointly [9].

The joint-generation attempt that still separates too cleanly. The most direct prior attempt at solving this the "right" way, generating audio and video together from the start, is Ovi, published in October 2025 under the title "Twin Backbone Cross-Modal Fusion for Audio-Video Generation" [6]. Ovi couples two architecturally matched diffusion transformers, one for video and one for audio, through bidirectional cross-modal attention inserted into every single transformer block, with a shared frozen T5 text encoder conditioning both branches and scaled rotary position embeddings reconciling their different natural time resolutions [6]. This is real, joint generation, not a post-hoc patch, and it was a genuine step forward when it shipped. But cross-modal attention applied uniformly, in full strength, at every single block, has its own cost: it risks the two modalities constantly overwriting each other's representations even in the many moments where a video frame and an audio frame have very little to actually say to each other, like a wide establishing shot with no dialogue and only ambient room tone. DreamX-Creator's own introduction names this exact tension as one of the field's four unresolved challenges: cross-modal interaction "must be introduced without overwhelming modality-specific representations" while its actual strength should vary across layers, heads, tokens, and samples, not stay fixed [2]. That single sentence is the entire justification for the gating mechanism covered in Step 3, and it is worth holding onto as you read the rest of this post: the problem was never whether to let audio and video talk to each other, prior work like Ovi already answered that question, the problem is how much, when, and for which specific tokens and attention heads that conversation should actually happen.

Step 2: Inside the Dual-Stream 7B Generator

DreamX-Creator's core generator is a 7-billion-parameter system, which the paper describes as "the smallest open-weight native joint model by disclosed total backbone size" among the systems it compares against, a group that ranges from NAVA's 6.3B active parameters up through MAGI-2 Preview's 114B total parameters [2]. Before getting to how the two modalities talk to each other, it is worth being precise about how they are kept separate, because that separation is just as deliberate a design choice as the coupling is.

Two Backbones, Not One Shared Sequence

Unlike an "everything is one token stream" omni-modal design, DreamX-Creator keeps video and audio as genuinely separate transformer backbones. "Each stream retains its own token rate, positional encoding, and transformer backbone," per the paper's own description of the architecture [2]. This matters because video and audio are not just different content, they are different signals at fundamentally different natural rates: video frames arrive at something like 24 to 30 samples per second of footage, while raw or lightly-compressed audio latents need to represent tens of thousands of samples per second to carry meaningful acoustic detail. Forcing both into one shared token rate and one shared positional scheme, the way a true single-stream omni-transformer would, means either wastefully padding the sparser modality or destructively downsampling the denser one. Keeping two specialized backbones lets each modality use a token rate and positional encoding that actually fits its own statistics, which is the same design logic that led Ovi to reach for scaled RoPE reconciliation between its own twin backbones rather than a single shared sequence [6].

Text conditioning is applied through modality-specific paths rather than one shared cross-attention block feeding both streams identically [2], which lets the video backbone attend to the parts of a prompt that describe visual content (a locked door, a rain-streaked window, a particular camera move) somewhat differently than the audio backbone attends to the parts that describe sound (footsteps, a slamming door, ambient rain), even when both are reading the same underlying prompt.

First Half Independent, Second Half Coupled

The architecture's real structural signature is a split down the depth of the network. "The first half processes streams independently; the latter half adds cross-modal paths" [2]. In plain terms, for roughly the first half of the network's transformer blocks, the video stream and the audio stream never see each other at all. Each backbone is free to build up its own modality-specific representation, working out what a coherent early sketch of "a hand knocking on a wooden door" looks like purely in visual terms, and what a coherent early sketch of "three sharp knocks in a quiet room" sounds like purely in acoustic terms, without either signal distorting the other before either representation is mature enough to be worth sharing.

Only in the latter half of the network does Gated Cross-Modal Attention get inserted, letting the two streams start exchanging information, by which point each stream already has a reasonably well-formed internal representation to offer the other. This is a meaningfully different bet than Ovi's design of inserting bidirectional cross-attention into every single block from the start [6]: DreamX-Creator is explicitly choosing to let each modality "think alone" first before it has to reconcile its output with the other modality's evolving representation, on the reasoning that forcing that reconciliation too early, before either stream has anything coherent to share, is exactly the kind of premature cross-modal interference the paper's introduction warns against [2].

dreamx-creator-native-audio-video-generation-2k-explained-2026-dual-stream-encoding.png

Step 3: Gated Cross-Modal Attention, the Mechanism That Actually Does the Sync Work

This is the paper's central architectural contribution, and it deserves a slower walkthrough than a headline term like "Gated Cross-Modal Attention" usually gets in coverage, because the actual mechanism is specific and mathematically described, not just a marketing label for "the two streams talk to each other now."

Two Directions, One Shared Clock

Once the network reaches its latter half, two cross-modal attention paths become available: Audio-to-Video (A2V) and Video-to-Audio (V2A). "A2V uses video queries with audio keys and values; V2A reverses these roles" [2]. Concretely, in the A2V path, a video token asks a question of the entire audio sequence ("what is happening acoustically around the moment I represent visually?") and pulls in a weighted combination of audio information to inform how it updates itself. The V2A path runs the same operation in reverse, letting audio tokens query the video stream for corresponding visual context. Because video and audio use different token rates, as covered in Step 2, the paper applies temporal rotary position encoding to align the two streams onto "a shared temporal coordinate system" [2], essentially giving both modalities a common clock so that a video token representing frame 47 and an audio token representing the corresponding fraction of a second can actually find each other in the attention computation, rather than the model having to implicitly learn an alignment from scratch with no shared notion of "now."

The Gate Itself, in Actual Formula Form

Here is where the mechanism earns the word "gated" rather than just "cross-modal." After the two streams compute their scaled dot-product cross-modal attention (the standard attention computation you would find in any transformer), the raw output does not simply get added straight into the target stream's representation. Instead, it passes through a learned gate first. The paper gives the gate's formula explicitly: g equals the sigmoid of a token-dependent term plus a context-dependent term plus a bias, computed per attention head, applied "after scaled dot-product attention and before the output projection" [2]. Each resulting gate value sits between 0 and 1, and "scales each attention-head output before head concatenation" [2].

Unpack what that buys the model. A gate value near 0 means that specific attention head, for that specific token, effectively contributes almost nothing from the other modality this pass, letting the target stream's own representation pass through largely untouched. A gate value near 1 means that head is fully trusting and incorporating the cross-modal signal it just computed. Crucially, because the gate depends on both the token itself and the specific attention output it just received (the paper's formula includes a term derived from "LN(C_i,h)," the layer-normalized cross-modal context for that head), the model can learn to open the gate wide for a token where the cross-modal signal is actually informative, like a mouth-region video token receiving strong signal from a corresponding speech-audio token, while keeping it nearly closed for a token where cross-modal information would just be noise, like a background wall texture token that has nothing meaningful to gain from the audio stream. This is precisely the "varying in strength across layers, heads, tokens, and samples" property the paper's introduction identifies as the actual unsolved problem [2], now solved with a concrete, learnable, per-token, per-head mechanism rather than a single fixed cross-attention strength applied everywhere, which is the design Ovi uses across its twin backbones [6].

Three Training Modes, and Why Corruption Order Matters

The gating mechanism gets trained under three distinct noise-ordering regimes, which matters because flow-matching and diffusion-style generation depends heavily on how much each modality's latent has been corrupted relative to the other at any given training step. In "A2V mode," video is corrupted more strongly than audio, and only the A2V path is active, forcing the model to learn to lean on relatively clean audio to help denoise a noisier video. In "V2A mode," the relationship reverses, with audio corrupted more heavily than video and only the V2A path active. In "Joint mode," both modalities are corrupted equally and both cross-modal paths run simultaneously [2]. Training across all three regimes means the model never gets to assume one modality is always the reliable "anchor" and the other is always the noisy one needing help, which is closer to how the model actually needs to behave at real inference time, when both audio and video start from pure noise together and have to denoise in lockstep.

One more detail worth calling out because it explains why this training scheme is stable rather than collapsing into one modality dominating the other: a stop-gradient is applied to the conditioning stream in each mode, explicitly designed "to block the target-stream loss from updating the conditioning backbone through cross-modal attention" [2]. In A2V mode, for example, the video stream's loss can flow back through the gate and adjust how the gate weighs audio information, but it cannot reach back and start rewriting the audio backbone's own weights just because doing so happened to make video denoising slightly easier on that particular batch. Without that stop-gradient, a joint model like this genuinely risks one modality's training signal quietly degrading the other modality's independent quality over the course of training, a known failure mode in multi-task and multimodal joint training generally.

dreamx-creator-native-audio-video-generation-2k-explained-2026-gated-cross-modal-attention.png

Progressive Joint Training: Getting the Two Streams to This Point at All

None of the mechanism above works from a cold start; the paper describes a three-stage progressive training curriculum to get there. Stage 1 is LoRA-based pretraining: the model initializes from separate modality-specific backbones, freezes the first half of each entirely, and applies rank-256 LoRA adapters to the latter half while training the new cross-modal attention modules directly, at learning rates of 1×10⁻⁴ for the LoRA adapters and 2×10⁻⁵ for the cross-modal modules [2]. This is a deliberately cheap, low-risk way to let a mostly-frozen, already-competent model start experimenting with cross-modal gating without risking catastrophic forgetting of whatever each backbone already knew how to do unimodally.

Stage 2 merges those LoRA weights back into the backbones and moves to full-parameter joint pretraining, with both diffusion-transformer backbones now jointly optimized at a backbone learning rate of 1×10⁻⁵ and the same 2×10⁻⁵ rate for cross-modal modules [2]. Stage 3 is a high-quality finetuning pass on a curated data subset specifically filtered by audio-video synchronization metrics and aesthetic scores, using fine-grained shot-boundary detection through a tool the authors call OmniShotCut to avoid training on clips that straddle a scene cut, with group-wise learning rates progressively reduced and the video-versus-audio loss balance shifted from an even 0.5/0.5 split during pretraining to a 0.5/0.1 split favoring video quality during this final stage [2]. All three stages optimize a token- and feature-normalized flow-matching objective, the same family of training objective used across most modern diffusion-transformer video models. The practical takeaway is that a three-stage curriculum like this, cheap adaptation first, full joint training second, curated high-quality polish third, is a genuinely reusable pattern for anyone trying to bolt a new modality onto an existing strong unimodal backbone without a full from-scratch training run, and it is a big part of why a 7B model can compete with systems several times its size on the metrics covered in Step 7.

Step 4: Audio-Video Reinforcement Learning and Modality-Aware Multimodal Feedback

Getting a joint model to denoise reasonably well is only the first half of DreamX-Creator's pipeline. The paper's introduction identifies a second unresolved problem directly: "likelihood-based training doesn't directly optimize perceptual quality, prompt adherence, semantic consistency, or fine-grained synchronization" [2]. In other words, a model trained purely to predict the right denoising direction can still produce output that is technically plausible but perceptually mediocre, or subtly out of sync in ways the training loss never explicitly penalized. The paper's answer is a reinforcement learning post-training stage the authors call Audio-Video Reinforcement Learning, built around what they term Modality-Aware Multimodal Feedback.

Why a Single Global Reward Actively Hurts a Joint Model

The naive approach to RL post-training for a joint audio-video model would be a single scalar reward per generated sample, some blend of "does this look good" and "does this sound good" averaged together. DreamX-Creator's authors explicitly reject that design, and the reasoning is worth internalizing because it generalizes well beyond this one paper: a single blended reward creates a real risk that "an improvement in one modality from masking degradation in the other" [2] goes completely unnoticed by the optimizer. Imagine a training step where the policy discovers that making the video slightly more visually striking, brighter colors, more dramatic camera motion, raises the blended reward more than the resulting slight degradation in audio quality lowers it. A single averaged reward has no way to see that trade happening; it just sees the net number go up and reinforces the behavior, even though one of the two modalities is quietly getting worse.

The paper's fix is to decompose feedback into three separate advantage signals rather than one: A_v for video-specific quality improvements, A_a for audio-specific quality improvements, and A_av specifically for cross-modal semantic consistency and synchronization [2]. Each of these gets computed from its own dedicated reward model rather than folded into one number early. Then, critically, the cross-modal signal gets routed to both streams rather than sitting off on its own: the paper's routing formula sets the effective video-stream advantage to A_v plus A_av, and the effective audio-stream advantage to A_a plus A_av [2]. That means a synchronization failure genuinely and separately penalizes both the video policy update and the audio policy update, rather than getting averaged away or attributed to only one side, which is exactly the design property that a single blended reward cannot provide.

Synchronization-Aware Token Weighting: Not Every Frame Needs Equal Attention

A further refinement addresses a genuinely subtle problem: not every moment in a clip is equally informative about synchronization quality. A wide shot of an empty room has essentially nothing to say about lip-sync accuracy; a close-up of a mouth mid-sentence has a great deal to say about it. The paper's synchronization-aware token weighting scheme identifies which video tokens actually matter for sync feedback by reusing "video-to-audio responses from selected interaction blocks to estimate the relevance of each video token" [2], essentially asking the model's own already-trained cross-modal attention which video regions the audio stream was actually paying attention to, and then weighting the RL feedback signal more heavily on exactly those regions using a frame-normalized sigmoid weighting formula [2]. This is an elegant reuse of a signal the model already computes for a different purpose (generation) to solve a second problem (which tokens deserve more RL attention) without introducing an entirely separate detection system.

The RL stage also applies depth-dependent gradient scaling specifically to the audio-to-video pathway, attenuating gradients entering the audio stream more strongly in shallow blocks [2], a stability measure that keeps early, foundational audio representations from being aggressively rewritten by late-stage RL feedback that is really about fine-grained synchronization rather than basic acoustic quality.

The Training Setup and What It Actually Bought

The RL stage initializes the policy from the pretrained DreamX-Creator 1.0 generator, using a training set of 1,000 first-frame-and-prompt pairs, generating multiple candidate outputs per condition through a rollout policy and normalizing rewards within candidates sharing the same condition, using separate reward models for video quality, audio quality, prompt consistency, and audio-visual synchronization [2]. Looking at the paper's own reported numbers, RL post-training moved the DeSync score, a lower-is-better desynchronization metric, from 0.1902 in the base model down to 0.1351 [2], a roughly 29% relative improvement, alongside a smaller gain in lip-sync confidence (LSE-C) from 7.8018 to 7.8361 and a small gain in ImageBind cross-modal similarity from 0.2608 to 0.2677 [2]. Interestingly, word-error-rate for generated speech got very slightly worse after RL, moving from 0.1112 to 0.1232, a real trade-off the paper's own tables show rather than hide, consistent with the modality-aware feedback design explicitly accepting that improving one axis (synchronization) is not guaranteed to be free on every other axis (raw speech intelligibility) [2]. Table 5 in Step 7 has the fuller picture of exactly which numbers moved and by how much.

dreamx-creator-native-audio-video-generation-2k-explained-2026-modality-aware-rl-feedback.png

Step 5: The Autoregressive 1-Step 2K Refinement Pipeline

The fourth unresolved challenge the paper names in its introduction is resolution: "joint latent generation and 2K refinement impose different computational requirements" [2]. This is a real, distinct problem from everything covered so far. Generating a coherent joint audio-video latent at a workable resolution is one computational regime; upsampling that result all the way to 2K, frame by frame, without breaking the timing and synchronization already achieved in the base generation, is an entirely different one, and naively running a large bidirectional video model at 2K resolution for every frame of every chunk would be prohibitively slow for anything resembling real product use. DreamX-Creator's answer is a three-stage distillation pipeline that is worth understanding in genuine technical depth, because "autoregressive 1-step" distillation is a real, specific technique, not a marketing phrase.

Stage 1: A Bidirectional Multi-Step Teacher

The pipeline starts by training what the paper calls a bidirectional multi-step teacher model on high-resolution video using standard flow matching. "Bidirectional" here means this teacher's attention mechanism can look both backward and forward in time when refining any given frame; "bidirectional spatiotemporal attention allows it to exploit both past and future frames" [2] when deciding how to sharpen or reconstruct a frame's fine detail. This is the natural, highest-quality way to build a refiner, because knowing what happens both before and after a given moment genuinely helps resolve ambiguous detail, similar to how a human colorist restoring old film footage benefits enormously from being able to scrub forward and backward through a scene rather than only ever seeing frames in strict forward order. The teacher is also trained against a deliberate degradation curriculum, progressively introducing blur, noise, compression artifacts, resampling artifacts, geometric distortion, and temporally correlated corruptions [2], so it learns to genuinely restore and enhance degraded low-resolution input rather than just naively upsampling clean input, which matters because the joint generator's raw output realistically does carry some of exactly these degradation types at its native resolution.

The catch with a bidirectional teacher, and the reason it cannot simply ship as the production refiner, is right there in the name: needing access to future frames to refine the current one means the entire clip, or at least a large forward-looking window of it, has to already exist before refinement of any given frame can even start. That is workable for offline batch processing, but it is exactly the wrong shape for a low-latency generation pipeline where a creator is waiting on a result.

Stage 2: Converting to an Autoregressive Multi-Step Refiner

Stage 2 converts that bidirectional teacher into what the paper calls an autoregressive multi-step refiner, using a technique called teacher forcing during training: "the current chunk is denoised conditioned on the LR video and ground-truth past HR chunks" [2]. Rather than needing the whole clip available at once, this refiner processes video in temporal chunks, moving strictly forward through time, refining chunk N using the already-finished high-resolution chunk N-1 as context rather than needing to peek ahead into chunk N+1 at all. This factorization over temporal chunks is what actually reduces the high-resolution computation burden, because the refiner never has to hold the entire bidirectional attention computation for the full clip in memory or compute at once, only a local window around the current chunk plus whatever HR context has already been produced. This stage's structure is explicitly built to "match the temporal structure of the final student" [2], meaning it is not just a standalone improvement, it is a deliberate intermediate step designed to make the final distillation in Stage 3 tractable.

Stage 3: Distribution Matching Distillation Into a 1-Step Student

The final stage distills that autoregressive multi-step refiner down into a student model that needs only one denoising evaluation per temporal chunk at actual inference time, using Distribution Matching Distillation (DMD), specifically following the DMD2 formulation [2]. DMD-style distillation works by training the student not to copy the teacher's exact intermediate denoising trajectory step by step, but to match the overall distribution of outputs the multi-step teacher would eventually produce, using an auxiliary "fake" score network trained alongside the student to estimate how far the student's current output distribution sits from the real target distribution, which lets the student learn to collapse an entire multi-step refinement trajectory into a single forward pass.

A real distillation-specific problem shows up here that the paper addresses directly: exposure bias. A student trained purely on teacher-provided (ground truth) past context learns to refine assuming its own previous outputs are perfect, but at actual inference time the student only has its own, imperfect, self-generated past chunks to condition on, and errors can compound chunk after chunk if the student was never trained to handle its own mistakes. DreamX-Creator's fix is self-rollout training: "complete videos are generated by the student itself, and the DMD objective is applied to these generated rollouts" [2], forcing the student to practice on its own, realistically imperfect chunk history during training rather than only ever seeing pristine ground-truth context it will never actually have access to at inference time. The final loss combines the distribution-matching signal with direct pixel-space supervision, specifically a DISTS perceptual loss and an L2 reconstruction loss, each weighted equally at 1.0 [2].

What This Actually Buys, in Inference-Cost Terms

The practical payoff is stated plainly: this refiner "performs efficient 1-step 2K refinement without changing the motion or audio-aligned timing produced by the joint generator" [2], requiring one denoising evaluation per temporal chunk rather than the many iterative steps a standard diffusion-based refiner would need. That distinction, one evaluation per chunk versus the typical 20 to 50 iterative denoising steps a bidirectional multi-step teacher would need per chunk, is the difference between a refinement stage that is a genuine, deployable part of a production pipeline and one that only ever exists as a research artifact too slow to actually ship. It is also the specific technical reason the paper's own comparison table (Step 7) shows the DreamX Refiner beating dedicated video super-resolution baselines like FlashVSR and SeedVR on quality metrics while running at a fraction of their computational cost per frame, since those baselines were not built around the same one-step-per-chunk autoregressive distillation design.

dreamx-creator-native-audio-video-generation-2k-explained-2026-autoregressive-2k-refinement.png

Step 6: How DreamX-Creator Actually Stacks Up on the Numbers

The paper reports two main quantitative comparisons, and the honest version of both is more interesting than a simple "our model wins" headline, because the authors are explicit that gaps remain against larger systems.

Against Similarly-Sized Research Baselines

Table 3 in the paper compares DreamX-Creator against four other native joint audio-video generation systems in a comparable parameter range: NAVA (6.3B), UniAVGen (7.1B), Ovi (10B), and two DaVinci-MagiHuman variants (15B each, at 256p and 512p) [2]. The metrics track video quality (VQ), two audio aesthetics dimensions called Content Enjoyment (CE) and Content Usefulness (CU), word error rate for generated speech (WER, lower is better), lip-sync confidence (LSE-C, higher is better), ImageBind cross-modal similarity (IB), and a DeSync score measuring overall audio-visual desynchronization (lower is better).

ModelParamsVQ (higher better)WER (lower better)LSE-C (higher better)DeSync (lower better)
NAVA6.3B0.61160.15747.72610.2342
UniAVGen7.1B0.66440.16614.95560.5371
Ovi10B0.65420.10537.30950.4730
DaVinci-MagiHuman-256p15B0.60670.13115.79720.4150
DaVinci-MagiHuman-512p15B0.60130.12286.89970.4010
DreamX-Creator (Base)7B0.65680.11127.80180.1902
DreamX-Creator (RL)7B0.65730.12327.83610.1351
DreamX-Creator (Refiner)7B0.69300.12327.69790.1731

A few specific numbers are worth pulling out rather than just skimming the table. DreamX-Creator's base model posts a DeSync score of 0.1902, already better (lower) than every other model in this comparison group, including the 15B DaVinci-MagiHuman variants at 0.4150 and 0.4010 [2]. Its LSE-C of 7.8018 also leads this group outright, ahead of Ovi's 7.3095 and comfortably ahead of UniAVGen's 4.9556 [2]. RL post-training then pushes DeSync down further to 0.1351, a genuinely large additional improvement on top of an already-leading number [2].

Against Larger Open-Weight Systems

Table 4 is the more interesting, and more honest, comparison, because it pits the 7B DreamX-Creator against two considerably larger open-weight systems: LTX-2.3 at 22B parameters and MiniMax-H3 at 33B parameters, a model covered in its own right on Miraflow's technical breakdown of MiniMax H3's omni-transformer architecture.

ModelParamsVQAudio CELSE-CImageBind IBDeSync
LTX-2.322B0.62855.18767.67680.30810.2412
MiniMax-H333B0.64295.27768.73540.31190.2708
DreamX-Creator (Base)7B0.65684.72107.80180.26080.1902
DreamX-Creator (RL)7B0.65734.74637.83610.26770.1351
DreamX-Creator (Refiner)7B0.69304.74637.69790.26750.1731

Here the picture is genuinely mixed, and the paper does not dress that up. DreamX-Creator's RL-tuned model still posts the best DeSync score of the three at 0.1351, versus 0.2412 for LTX-2.3 and 0.2708 for MiniMax-H3, and the best raw video-quality score at 0.6573 [2]. But on audio aesthetics specifically, both larger models pull ahead: LTX-2.3 scores 5.1876 on Content Enjoyment against DreamX-Creator's 4.7463, and MiniMax-H3 scores even higher at 5.2776 [2]. MiniMax-H3 also posts a meaningfully higher LSE-C at 8.7354 versus DreamX-Creator's 7.8361, and a higher ImageBind cross-modal similarity score at 0.3119 versus 0.2677 [2]. The paper's authors state this gap plainly in their own discussion: "relative to Ours (RL), LTX-2.3 is higher on CE... MiniMax-H3 further leads on CE... gaps show that the current 7B model has not reached across-the-board parity" [2].

The Refiner Against Dedicated Restoration Baselines

Table 5 compares the 2K refiner itself against three established restoration approaches: FlashVSR, a one-step streaming video super-resolution model built around its own distillation approach, SeedVR, a 2.48-billion-parameter diffusion transformer for video restoration, and LTX-2.5's own built-in refiner [2]. On the metrics that matter for this specific comparison, aesthetic score, MUSIQ perceptual quality, and MANIQA quality, the DreamX Refiner posts the best numbers in the group: 0.4911 aesthetic score, 0.7073 MUSIQ, and 0.4382 MANIQA, each ahead of FlashVSR, SeedVR, and the LTX-2.5 Refiner on the same metrics [2]. It is worth being precise about why this comparison matters: none of these three baselines were built to preserve audio-video timing in the first place, since they are general video restoration tools rather than components of a native joint audio-video pipeline, so beating them on raw visual quality metrics while also being the only one of the four explicitly engineered not to disturb the audio-aligned timing already established earlier in the pipeline is a genuinely meaningful result specific to this paper's actual design goal.

The Human Preference Study, Read Honestly

Beyond the automated metrics, the paper also reports a human preference study comparing DreamX-Creator against both the smaller research baselines and three industrial-scale closed systems: Wan 2.7, Kling v3, and MiniMax-H3. Against the smaller, comparably-sized models, DreamX-Creator wins decisively: a 73.7% win rate on video quality against UniAVGen, 68.2% against DaVinci, 61.7% against NAVA, and a 48.8% win rate against Ovi with a large 31.8% tie rate, meaning it rarely loses outright even in its closest comparison within this group [2]. Against the much larger industrial systems, the results are closer to competitive parity rather than a clear win: a 45.5% win rate against Wan 2.7 (with a 41.3% loss rate), 48.8% against Kling v3, and 39.5% against MiniMax-H3 (with a 41.8% loss rate) [2]. The authors are direct about where this study's losses concentrate: "losses exceed wins on AV-Align and audio quality against all three systems... audio fidelity and cross-modal alignment remain the main areas for improvement" [2]. They also flag a real limitation in their own test set: "highly dynamic and compositionally complex scenes are under-represented, so the study does not fully probe the capabilities of these industrial systems" [2], a genuinely important caveat that means the reported win rates against Wan 2.7, Kling v3, and MiniMax-H3 likely understate those industrial systems' real advantage on the harder, more dynamic content that test set does not cover well.

Step 7: Real Limitations and Failure Modes, Beyond What the Benchmark Tables Show

Reading a paper's own limitations honestly matters more for a technical explainer than repeating its strongest headline numbers, so it is worth listing out, specifically, where DreamX-Creator's own data says it still falls short, and where a careful reader should expect real-world failure modes the benchmark tables do not fully capture.

Audio aesthetics lag behind larger models. Across both Table 4 and the human preference study, audio quality specifically, not synchronization, is where DreamX-Creator trails LTX-2.3 and MiniMax-H3 most consistently [2]. This tracks with an intuitive explanation: a smaller model trained on a curated but still comparatively modest audio-video dataset likely has less raw acoustic modeling capacity than a 33B system trained with proportionally larger resources dedicated purely to audio quality, even when the smaller model's synchronization mechanism is genuinely more effective.

A real speech-intelligibility trade-off from RL. As covered in Step 4, word error rate for generated speech got measurably worse after RL post-training, moving from 0.1112 to 0.1232 [2]. This is a real, disclosed cost of optimizing hard for synchronization and cross-modal consistency rather than a hidden flaw the paper is glossing over, but it means a production deployment prioritizing crisp, intelligible dialogue above all else might reasonably want to weight that trade-off differently than the paper's own default RL configuration does.

Cross-modal semantic gap versus the largest models. ImageBind similarity, which measures how well the generated audio and video actually correspond semantically rather than just temporally, sits meaningfully lower for DreamX-Creator (0.2677 with RL) than for MiniMax-H3 (0.3119) [2]. This is a different failure mode than mistimed audio: it means the sounds and images can be well-synchronized in time while still occasionally being a slightly weaker semantic match for each other, a subtler problem than lip-sync drift and one likely to show up as generically plausible ambient sound rather than a sound specifically and precisely matched to the exact visual event happening on screen.

Test set bias likely understates the gap against the strongest industrial systems. The authors' own admission that "highly dynamic and compositionally complex scenes are under-represented" in their human preference test set [2] is worth taking at face value rather than treating the reported win rates against Wan 2.7, Kling v3, and MiniMax-H3 as a full picture. A creator generating genuinely complex, fast-moving, multi-subject scenes should expect the real-world gap against those larger, closed industrial systems to be somewhat wider than the reported 39 to 49 percent win rates suggest, since those numbers come disproportionately from simpler test cases.

This is a preprint, not a shipped consumer product. It is also worth being clear-eyed about what kind of release this actually is. DreamX-Creator ships as open research weights on Hugging Face and ModelScope alongside its evaluation code under an Apache 2.0 license [3][4], not as a polished consumer-facing generation tool with a hosted interface, rate limits handled for you, or a content moderation layer built in. Running it means setting up a GPU environment, downloading multi-gigabyte checkpoints, and working directly with the same research-grade tooling the authors used to produce their own benchmark numbers, which is a meaningfully different experience than opening a browser tab and typing a prompt into a finished product.

dreamx-creator-native-audio-video-generation-2k-explained-2026-benchmark-scale-comparison.png

Step 8: Running It Yourself, If You Have the GPU For It

For readers who want to go beyond reading about the architecture and actually try the released weights, the GitHub repository documents a genuinely straightforward two-part workflow: the 7B joint generator, and the separate 2K refiner, each with its own dependencies [3]. The repository credits the Wan Team, the OpenMOSS Team, and the VideoX-Fun Team for open-source foundations it builds on [3], and the checkpoints directory structure documented in the repo shares Wan2.2-TI2V-5B components (VAE, text encoder, tokenizer) between the two stages [3][4].

bash
/code # Clone the official DreamX-Creator repository and set up the joint # audio-video generator. Requires a CUDA GPU and the model checkpoints # downloaded separately from Hugging Face or ModelScope into checkpoints/. git clone https://github.com/AMAP-ML/DreamX-Creator.git cd DreamX-Creator/audio_video_generation python -m pip install -r requirements.txt # Run the default demo case, or point it at your own first frame and prompt CASE=case4 ./inference.sh # Or override the conditioning directly: INPUT_IMAGE=/path/to/first_frame.png \ PROMPT="A hand knocks three times on a wooden door in a quiet hallway" \ OUTPUT_PATH=./outputs/knock_scene \ ./inference.sh # On a memory-constrained GPU, offload the text encoder and VAE to CPU: ./inference.sh --text_encoder_cpu_offload --vae_cpu_offload

Once the base joint generation is done, the separate 2K refinement stage runs on the resulting clip:

bash
/code # Upscale a joint-generated clip to 2K using the distilled # autoregressive 1-step refiner (SR-DiT 5B). cd ../video_refiner python -m pip install -r requirements.txt INPUT=./outputs/knock_scene/output.mp4 bash run_inference.sh

The audio-video generation module also documents a set of environment variable overrides worth knowing about if you plan to actually script this rather than run the default demo case, including INPUT_IMAGE and PROMPT to swap in your own conditioning frame and text, CASE to select between the repository's six built-in evaluation scenarios, NUM_INFERENCE_STEPS and SEED for the usual generation controls, and flags for CPU-offloading the text encoder and VAE on more memory-constrained GPUs [3]. A multi-GPU sequence-parallel inference script is also documented for teams with more than one GPU available who want faster wall-clock generation on longer clips [3].

To see what a first-frame-and-prompt condition actually looks like written out for a joint audio-video generator like this, here is a prompt in the style you would hand to a Wan-based native audio-video model, since DreamX-Creator's own released checkpoints share Wan2.2-TI2V-5B components as their video backbone dependency [3].

Camera: static wide shot at eye level, framing a wooden front door dead center in the frame with a small entryway rug below it.
Subject: a plain wooden door in a quiet residential hallway, closed, warm light spilling faintly from underneath it.
Action: a hand enters from the right edge of frame and knocks three times against the upper door panel, then withdraws.
Setting: a narrow indoor hallway, late evening, a single warm ceiling light overhead, muted wallpaper in the background.
Style: naturalistic handheld-adjacent framing, soft indoor color grade, shallow depth of field on the background wallpaper, native 2K output.
Audio: generate jointly with the video in the same pass, not as a separate track — three distinct, sharply timed knock impacts landing exactly on each visual contact with the door, a faint room echo after each knock, quiet ambient hallway tone underneath, no dialogue, no music.

What This Actually Means If You Generate AI Video Today

None of this is about Miraflow using DreamX-Creator, and it does not. Miraflow's Cinematic AI Video Generator currently runs on Veo 3 and Veo 3.1 for the video generation itself, with voiceover and sound design handled as a distinct step in the creator's workflow rather than generated jointly with the picture in one pass. That is a genuinely different architecture from DreamX-Creator's native joint generation, and it is worth understanding why, now that you have seen what native joint generation actually requires under the hood.

What this paper actually explains, in concrete mechanistic terms, is why some AI video tools nail audio timing while others feel slightly off no matter how many times you regenerate. A tool that generates video first and adds audio, music, or a voiceover afterward is fundamentally doing post-hoc alignment, the same category of approach DreamX-Creator's own introduction identifies as structurally limited [2], and no amount of manual timeline nudging fully closes that gap for content with precise physical audio events, a door slam, a footstep, a clap, because the video was already finished by the time the audio decision got made. A tool built around genuinely joint generation, the way DreamX-Creator, Ovi [6], and MiniMax-H3 each approach it in their own way, has a structural shot at avoiding that specific failure mode, though as Step 6 covered, joint generation alone does not automatically guarantee every axis of quality, DreamX-Creator's own numbers show real trade-offs even within a joint-generation design.

For a creator using Miraflow's Cinematic AI Video Generator or building short-form content with Text2Shorts, the practical takeaway is not "wait for native audio-video models to replace your current tools." It is a genuinely useful piece of context for setting expectations correctly: content where audio timing has to be frame-precise, a physical impact, a specific spoken word landing on a specific gesture, is exactly the category where any post-hoc audio workflow, including a well-executed one, is working against a structural disadvantage that even the largest current joint-generation systems have not fully closed, per DreamX-Creator's own honest comparison against MiniMax-H3 and Wan 2.7 in Step 6. Content where audio timing has more natural slack, ambient music, general voiceover narration that does not need to land on a specific visual beat, is far less exposed to that gap, which is exactly the kind of content a tool like Text2Shorts's script-to-visual-to-voiceover pipeline handles well today. Knowing which category your specific shot falls into, before you generate it, saves a lot of regeneration cycles chasing a sync problem that a separate-audio-and-video workflow was never structurally going to fully solve in the first place.

dreamx-creator-native-audio-video-generation-2k-explained-2026-progressive-training-stages.png

Common Mistakes When Reading This Kind of Paper

  • Treating "native joint generation" as a single solved category. As Step 6's benchmark tables show, DreamX-Creator, Ovi, MiniMax-H3, and LTX-2.3 all generate audio and video jointly in some form, yet post meaningfully different numbers on synchronization, audio aesthetics, and cross-modal semantics. "Joint generation" describes an architectural category, not a single guaranteed quality level.
  • Assuming a smaller parameter count means worse quality across every axis. DreamX-Creator's own DeSync and LSE-C numbers beat every larger system it compares against in Tables 3 and 4, even while trailing on audio aesthetics and ImageBind similarity [2]. Parameter count correlates with capacity, not with which specific quality axis a team chose to optimize hardest during training and post-training.
  • Reading a single win-rate number without checking what test set produced it. The paper's own admission about under-represented dynamic scenes in its human preference study [2] is exactly the kind of caveat that changes how much weight a specific win-rate percentage deserves, and it is easy to miss if you only skim the headline chart.
  • Confusing the released open weights with a finished consumer product. As Step 7 covers, this is a research release with GPU setup, checkpoint downloads, and command-line tooling, not a hosted app, and evaluating it as if it should feel like a finished creator tool sets the wrong expectations for what an Apache 2.0 research weight release actually is.
  • Missing the real trade-off inside the RL stage itself. It is tempting to read "RL post-training improved the model" as a strictly positive step, but DreamX-Creator's own word-error-rate regression after RL [2] shows that optimizing hard for one axis, in this case synchronization, is not automatically free on every other axis, a genuinely common pattern in multi-objective RL post-training worth watching for in any similar paper going forward.

What to Watch Next in This Space

The gap the paper's own authors are most explicit about, audio aesthetics and cross-modal semantic alignment trailing larger models like MiniMax-H3 and LTX-2.3 [2], is the most likely area for a near-term follow-up release to target directly, since the synchronization mechanism itself already leads the field on the metrics that measure timing specifically. It is also worth watching whether other teams adopt the token- and head-wise gating pattern described in Step 3 for their own joint generation systems, given how directly it targets the specific "overwhelming modality-specific representations" problem the paper's introduction names, a problem that is not unique to audio-video generation and shows up in essentially any architecture trying to fuse two modalities without one drowning out the other's independent signal. And given how much of DreamX-Creator's compact size comes from reusing Wan2.2-TI2V-5B components rather than training every piece from scratch [3], it would not be surprising to see the same reuse-strong-open-components strategy show up again as more teams try to build compact, reproducible joint generation systems rather than committing to a full from-scratch training run on the scale MiniMax or ByteDance can afford.

Frequently Asked Questions

What is DreamX-Creator 1.0 in simple terms? It is a research system, described in a paper published on arXiv on August 31, 2026, that generates video and audio together in one process from a starting frame and a text prompt, using a 7-billion-parameter model, rather than generating silent video first and adding sound afterward [1].

How is this different from a video model with a lip-sync feature added on top? A lip-sync feature reanimates a mouth region to match audio that already exists, working on a video that is already finished. DreamX-Creator's video and audio streams are denoised together from the start, with a Gated Cross-Modal Attention mechanism letting the two modalities inform each other's generation throughout, not just at the mouth region after the fact [2].

Does Miraflow use DreamX-Creator? No. Miraflow's Cinematic AI Video Generator currently runs on Veo 3 and Veo 3.1, which generate video, with audio and voiceover produced as a separate step. This paper is useful background for understanding why that separation creates the specific timing gap it does, not a description of Miraflow's own technology stack.

What does "gated" actually mean in Gated Cross-Modal Attention? It means the cross-modal attention output, computed per attention head for every token, gets multiplied by a learned value between 0 and 1 before it is allowed to influence the target modality, so the model can let strongly relevant cross-modal signal through fully while nearly blocking it where it would not help, rather than always applying the same fixed cross-modal attention strength everywhere [2].

Why does 2K refinement need a whole separate distillation pipeline instead of just running the base model at higher resolution? Running a large bidirectional model at full 2K resolution for every frame is computationally expensive and, because a bidirectional model needs to see future frames to refine the current one, poorly suited to a low-latency generation pipeline. The paper's three-stage pipeline trains a high-quality bidirectional teacher first, converts it into an autoregressive refiner that only needs past context, then distills that into a student needing just one denoising pass per chunk, cutting the computational cost dramatically while preserving the timing already established by the joint generator [2].

Does DreamX-Creator beat every other model on every metric? No, and the paper is explicit about this. It leads on synchronization-specific metrics like DeSync and lip-sync confidence even against much larger models, but LTX-2.3 (22B) and MiniMax-H3 (33B) both score higher on audio aesthetics and cross-modal semantic similarity, gaps the authors themselves describe as showing the 7B model "has not reached across-the-board parity" [2].

Can I actually run DreamX-Creator myself? Yes, the generator and the 2K refiner are both released as open weights on Hugging Face and ModelScope under an Apache 2.0 license, with inference scripts documented in the GitHub repository, though running it requires a real GPU setup and command-line comfort rather than a hosted consumer interface [3][4].

What is the practical difference between "bidirectional" and "autoregressive" in this context? A bidirectional model can use information from both earlier and later frames when refining any given frame, generally giving higher quality but requiring more of the clip to already exist before processing can start. An autoregressive model only ever looks backward at already-finished earlier frames, which is less information to work with per frame but lets it process video strictly in forward order, which is what actually makes low-latency, chunk-by-chunk generation possible [2].

Conclusion

DreamX-Creator 1.0 is a genuinely useful case study in what "solving" native audio-video synchronization actually takes at the architecture level, not just a headline number to skim past. The Gated Cross-Modal Attention mechanism gives a concrete, learnable answer to a problem prior joint-generation systems like Ovi mostly left as a fixed design choice, letting each token and attention head decide for itself how much cross-modal signal is actually worth trusting rather than applying one fixed strength everywhere [2][6]. The Modality-Aware Multimodal Feedback design during RL post-training shows a genuinely careful approach to a subtle failure mode, one modality's improvement quietly masking another's degradation, that a single blended reward would never catch. And the autoregressive distillation pipeline behind the 2K refiner is a real, technically specific answer to why high-resolution generation and joint audio-video generation have historically pulled against each other computationally. None of that means the system is finished; the paper's own numbers show real gaps against larger models on audio aesthetics and cross-modal semantics, and the authors say so themselves rather than burying it. For anyone generating AI video today, on Miraflow's own Cinematic AI Video Generator, on Text2Shorts, or on any other tool, understanding exactly what native joint generation requires, and exactly where even the best current systems still fall short of it, is the difference between guessing why a clip's timing feels slightly off and actually knowing.

References

[1] arXiv. "DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution" (2608.31106).

[2] arXiv. "DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution, full HTML (v1).

[3] GitHub. "AMAP-ML/DreamX-Creator Repository."

[4] Hugging Face. "GD-ML/DreamX-Creator Model Card."

[5] Hugging Face. "Paper Page: DreamX-Creator (2608.31106)."

[6] arXiv. "Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation" (2510.01284).

[7] arXiv. "MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis" (2412.15322).

[8] Higgsfield. "How to Make Realistic AI Talking and LipSync Videos in 2026."

[9] Maxfusion. "I Tested Every Leading AI Video Model on Lip Sync for AI UGC Ads (2026)."