SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX

Reddit r/LocalLLaMA Models

Summary

NVIDIA Cosmos3, a 64B parameter image generation model, is released with INT4 quantization for local deployment on CUDA and MLX, with code and weights available and performance demonstrated on Apple Silicon.

Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda HF weights - https://huggingface.co/JuliaML/Cosmos3-Super-Text2Image-4Step-INT4-G64-BF16 Single clip took approximately 5m on M4 MAX 128 GB Mac Cosmos3 - a 64B params model
Original Article

Similar Articles

nvidia/Cosmos3-Super

Hugging Face Models Trending

NVIDIA released Cosmos3, a collection of omnimodal world foundation models for Physical AI, capable of generating video, image, audio, and action commands from various inputs, with versions for different tasks like policy learning and image-to-video generation.

Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang

Reddit r/LocalLLaMA

This is a quantized version of the Qwen3.8-27B AI model using IQ4_XS quantization, optimized for 16GB RAM systems, with instructions for local deployment using various tools like llama.cpp and Ollama.

@swyx: roundup of links:

X AI KOLs Following

NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.

@basecampbernie: https://x.com/basecampbernie/status/2074262192304832535

X AI KOLs Timeline

This post details the author's setup and benchmarks for running NVFP4-quantized image and video generation models on a GIGABYTE AI TOP ATOM (DGX Spark) workstation, achieving impressive performance with models like FLUX.2, Qwen-Image, and LTX-2.3 for video with synchronized audio.