Comfy-Org/MiniMax-Music-3

Hugging Face Models Trending Models

Summary

This article provides instructions for placing repackaged MiniMax-Music-3 model files into ComfyUI directories for music generation.

Tags: comfyui, license:apache-2.0, region:us
Original Article
View Cached Full Text

Cached at: 08/16/26, 09:32 AM

Comfy-Org/MiniMax-Music-3 Β· Hugging Face

Source: https://huggingface.co/Comfy-Org/MiniMax-Music-3 Repackaged model files for ComfyUI.

Original model repository:https://huggingface.co/MiniMaxAI/MiniMax-Music3

Place the files in the following folders:

πŸ“‚ ComfyUI/
β”œβ”€β”€ πŸ“‚ models/
β”‚   β”œβ”€β”€ πŸ“‚ diffusion_models/
β”‚   β”‚   β”œβ”€β”€ minimax_music3_dit_fp16.safetensors
β”‚   β”‚   β”œβ”€β”€ minimax_music3_dit_fp32.safetensors
β”‚   β”‚   └── minimax_music3_dit_int8_convrot.safetensors
β”‚   β”œβ”€β”€ πŸ“‚ text_encoders/
β”‚   β”‚   β”œβ”€β”€ minimax_music3_text_encoder_bf16.safetensors
β”‚   β”‚   β”œβ”€β”€ minimax_music3_text_encoder_pruned_bf16.safetensors
β”‚   β”‚   └── minimax_music3_text_encoder_pruned_int8_convrot.safetensors
β”‚   β”œβ”€β”€ πŸ“‚ vae/
β”‚   β”‚   └── minimax_music3_dav.safetensors

Similar Articles

Comfy-Org/MiniMax-H3

Hugging Face Models Trending

Comfy-Org repackaged MiniMax-H3 model files for ComfyUI, including diffusion models, text encoders, and VAEs, with workflow templates for text-to-video, image-to-video, and reference-to-video generation.

drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

Hugging Face Models Trending

Third-party ComfyUI-compatible LoRA conversions for MiniMax-H3 Turbo 4-step audio-video generation, including further-trained checkpoint-500 variants and an example workflow.

Kijai/MiniMax-H3_comfy

Hugging Face Models Trending

A Hugging Face repository hosting MiniMax-H3 models converted for ComfyUI usage, along with a Lightx2v distill LoRA for faster inference.

MiniMax-Music3 released!

Reddit r/LocalLLaMA

MiniMax releases Music 3, a high-performance music generation model that creates complete songs up to five minutes long using lyrics and detailed descriptions, with an 8B global LLM and 0.6B local LLM for long-range coherence and acoustic detail.