DMAD: A Faster Approach to AI Video Generation With Native Audio

DMAD: A Faster Approach to AI Video Generation With Native Audio

OverviewDMAD is a set of rank-128 LoRA adapters that distills the 33-billion-parameter MiniMax-H3 text-to-audio-video model from 50 denoising steps to 4. It generates 1344×768 video with native stereo audio; the repository provides two 1.4 GB adapter checkpoints, not a standalone base model. The maintainer is ZhengmingYu. The key practical point is that four-step sampling depends on downloading and running MiniMax-H3 as well as applying the adapter, so the small adapter files do not indicate low total hardware requirements. The project describes Diffusers-compatible weights and provides inference code and a Diffusers-pipeline example, but does not state a GPU-memory requirement or measured inference time for this release.Best use casesPrompt-to-video clips with synchronized sound. Choose DMAD when a prompt should produce both video and native stereo audio in one generation. The adapter targets MiniMax-H3, a joint text-to-audio-video model, and the supplied sampling configuration generates 124 frames at 24 fps. The README’s teaser notes that its displayed clips show video only, although each clip also has generated audio.Rapid iteration on short video concepts. Four denoising steps, rather than the teacher’s 50, make this checkpoint a fit for testing prompt variations, storyboards, or visual directions when generation latency matters. The README does not provide a wall-clock speed benchmark, so the step-count reduction should not be treated as a specific seconds-per-clip guarantee.Experiments on few-step distillation. Researchers can use the released adapters to examine adversarial distribution matching on a large joint audio-video generator. The paper describes two discriminator heads on a shared backbone, which distinguish real data and teacher samples from student samples; linear losses on their logits train the student without fitting an auxiliary diffusion score model.Comparing student checkpoints. The repository offers a paper checkpoint and a second checkpoint trained with a fully trained critic backbone. The latter is from iteration 1600, uses live weights, and is reported to score higher on AVGen-Bench; the paper checkpoint is an EMA of the student at iteration 800 of the main run. This makes the release useful for evaluating the effect of those training choices, though the README does not provide benchmark tables or scores for the two files.LimitationsDMAD is an adapter, not a self-contained generator. You need the MiniMax-H3 base model and the project’s inference code; the base download command excludes FL2VA/, Ref2VA/, and transformer_ref/*. The 1.4 GB adapter size therefore does not describe the full storage or memory footprint. The README gives no VRAM minimum, supported GPU list, throughput, latency, batch-size guidance, or hardware benchmark for this checkpoint.The documented output configuration is 1344×768, 124 frames, and 24 fps. The provided information does not establish support for other resolutions, clip lengths, frame rates, or audio controls. Four steps reduce sampling work relative to the 50-step teacher, but the README does not quantify the resulting quality tradeoff or identify specific visual or audio failure modes. The paper abstract reports aggregate results and human preference comparisons, not a detailed error analysis.The license is the MiniMax H3 Community License Agreement. The repository calls these weights model derivatives of MiniMax H3, modified through LoRA fine-tuning; it does not state that unrestricted commercial use is allowed. Review the agreement before deployment. The supplied material also does not describe safety filters, bias evaluations, or a maintenance schedule.How it comparesvdn-minimax-h3: Pick DMAD when you want a four-step adapter with a documented 1344×768, 124-frame, 24-fps joint audio-video configuration and the DMAD distillation method. Pick VDN-H3 when its hybrid-attention approach and published speed result better match your deployment: its description reports a 14.4-second clip in 11.23 seconds using 8 B200 GPUs and 8 denoising steps. That is a concrete speed benchmark for VDN-H3, not a directly comparable benchmark against DMAD; the available information does not establish relative output quality or cost.TaoMate-H3: Pick DMAD for four-step, fixed-clip generation and experiments with adversarial distillation. Pick TaoMate-H3 for low-latency streaming audio-video in small chunks, continuous long-form generation, and its stated 480p, 768p, and 1080p options. TaoMate-H3 is a runtime built on MiniMax H3 with a rank-128, alpha-128 EMA LoRA adapter; its description does not provide a directly comparable speed number, so the tradeoff is documented fixed-step generation versus a streaming-oriented runtime rather than a proven speed ranking.DMD2: Pick DMAD when you need the released MiniMax-H3 joint audio-video student or want to test adversarial distribution matching. Pick DMD2 when your task and pipeline match its image-synthesis model card and standard Diffusers usage. The DMAD paper reports human preference rates of 79.1% over DMD2 for joint audio-video generation, excluding ties; this is a reported comparison for that task, not evidence that DMAD is better for image generation or every workload.MiniMax-H3: Pick DMAD when four-step sampling is the priority and you can accept using a LoRA adapter with the base model. Pick MiniMax-H3 when you want the 50-step teacher rather than the distilled student. The README identifies H3 as a 33B text-to-audio-video model; it does not provide a direct speed, cost, or quality benchmark between the two checkpoints beyond the reduction from 50 to 4 steps.minimax-h3-4step-turbo-loras-comfyui-exp: The provided listing has no description, so there is not enough information to compare its quality, speed, cost, or intended use with DMAD. Choose DMAD when you need documented checkpoint files, LoRA layout, sampling settings, and an inference-code repository; evaluate the other adapter from its own documentation before choosing it.Technical specificationsDMAD applies low-rank adapters to the MiniMax-H3 transformer. The base model is identified as a 33B text-to-audio-video model. The adapters target attention projections and feed-forward layers across 50 transformer blocks and 2 token-refiner blocks, for 312 modules total. The stated rank and alpha are both 128.The paper describes DMAD as Distribution Matching as Adversarial Distillation. It replaces the auxiliary score fitting used by Distribution Matching Distillation with classification: two discriminator heads share a backbone and distinguish real data and teacher samples from student samples. Linear losses on discriminator logits train the student. The paper states that, at the discriminator optimum, these losses recover the distribution-matching gradient underlying DMD through the identity between discriminator logits and log-density ratios. It also introduces gap-based reweighting, which adapts teacher supervision across noise levels using the real-data head’s empirical logit gap between real and teacher samples.The paper abstract reports results across several tasks: FID 1.04 for one-step ImageNet-64×64 generation; 14.47 for four-step SDXL on COCO-10K; VBench total score 85.15 for four-step Wan2.1-T2V-14B, reported as the best values among the compared few-step methods and multi-step teachers; and human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation on MiniMax-H3-33B, excluding ties. These are paper-level results across models and tasks, not all measurements of the released MiniMax-H3 adapter.Base model: MiniMax-H3, 33B parameters; text-to-audio-video.Adapter files: dmad_minimax_h3_4step_lora_critic.safetensors and dmad_minimax_h3_4step_full_critic.safetensors, each 1.4 GB.Paper checkpoint: EMA of the student at iteration 800 of the main run.Full-critic checkpoint: iteration 1600, live weights; critic backbone fully trained rather than frozen under a LoRA; reported to score higher on AVGen-Bench.Adapter layout: Diffusers keys; .lora.down.weight is A with shape [128, in], and .lora.up.weight is B with shape [out, 128]. Alpha equals rank at 128.Target modules: attn.to_q, attn.to_k, attn.to_v, attn.to_out.0, ff.net.0.proj, and ff.net.2 across all 50 transformer blocks and 2 token-refiner blocks.Sampling: 4 steps; time shift 12 for video and 2 for audio; no classifier-free guidance; 124 frames at 24 fps.Output resolution: 1344×768 video with native stereo audio.Inference implementation: inference.py includes the paper’s re-noise step rule and a Diffusers-pipeline example. The repository README covers environment setup.Weight format: .safetensors; metadata repeats the LoRA layout.Training data, training compute, GPU memory, quantization options, and measured inference speed: not specified in the provided model information.Model inputs and outputsInputsText prompt: supplied through a prompt file in the example command (prompts/dmad_sweater.txt). The README does not specify a prompt schema or maximum prompt length.Seed: the example uses --seed 42.Base model directory: MiniMax-H3 weights, downloaded separately.LoRA checkpoint: one of the released .safetensors adapters.Sampling configuration: the documented setup uses 4 steps, video time shift 12, audio time shift 2, no classifier-free guidance, and 124 frames at 24 fps.OutputsVideo: 1344×768, 124 frames at 24 fps.Audio: native stereo audio generated jointly with the video.Output destination: set with --output-dir; the example writes to outputs/dmad_sweater.Post-processing: no required post-processing steps are specified.Getting startedThe repository’s example uses the project inference script rather than a standalone Diffusers call. Clone the code, download the base model and adapter, then run inference:git clone https://github.com/Yzmblog/DMAD.git cd DMAD hf download MiniMaxAI/MiniMax-H3 \ --local-dir models/MiniMax-H3 \ --exclude "FL2VA/*" \ --exclude "Ref2VA/*" \ --exclude "transformer_ref/*" hf download ZhengmingYu/DMAD \ --include "minimax_h3/*" \ --local-dir ckpt python inference.py \ --model-dir models/MiniMax-H3 \ --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \ --prompt-file prompts/dmad_sweater.txt \ --seed 42 \ --output-dir outputs/dmad_sweater Use the code repository’s README for environment setup. The provided information does not specify package versions or a minimal GPU configuration.Frequently asked questionsQ: Can I use DMAD commercially?A: The checkpoints are model derivatives of MiniMax H3 and are distributed under the MiniMax H3 Community License Agreement. The provided information does not establish unrestricted commercial rights; review that agreement before commercial use.Q: What GPU or VRAM do I need to run DMAD?A: The README does not state a GPU model or VRAM requirement. The 1.4 GB figure applies to each LoRA checkpoint, not the full MiniMax-H3 model or total inference memory.Q: How many steps does DMAD use, and how long does generation take?A: The documented sampler uses 4 steps, compared with the 50-step MiniMax-H3 teacher. No wall-clock inference time is provided for DMAD.Q: What input format does the inference script expect?A: The example passes a text prompt through a prompt file using --prompt-file. The README does not specify the file’s schema or a maximum prompt length.Q: Does DMAD generate audio as well as video?A: Yes. It is a student of MiniMax-H3’s joint text-to-audio-video model and is described as producing video with native stereo audio. The documented output is 124 frames at 24 fps.Q: Which checkpoint should I use?A: The dmad_minimax_h3_4step_lora_critic.safetensors file is the paper checkpoint: an EMA of the student at iteration 800. The dmad_minimax_h3_4step_full_critic.safetensors file is from iteration 1600 with live weights and a fully trained critic backbone; the README says it scores higher on AVGen-Bench.Q: Can I fine-tune or adapt the released weights?A: The release consists of rank-128 LoRA adapters in Diffusers key layout, so the weights are adapter-based. The provided README does not give fine-tuning commands, training data, or a supported fine-tuning recipe.Q: What quality failures or safety issues are documented?A: The provided model information does not list specific visual or audio failure modes, bias evaluations, or safety behavior. The paper reports aggregate benchmark and preference results, which do not establish performance on every prompt or use case.Q: How does DMAD compare with DMD2 for joint audio-video generation?A: The paper abstract reports that the four-step MiniMax-H3 student receives 79.1% human preference over DMD2, excluding ties, for joint audio-video generation. That result is specific to the reported comparison and does not establish superiority for other tasks. The related paper is DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation.This is a simplified guide to an AI model called DMAD maintained by ZhengmingYu. If you like these kinds of analysis, join AIModels.fyi or follow us on Twitter.

Original Source

Read the full article at Hackernoon →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.