ModelExpress: Distributing Model Artifacts at the Speed of Light - NVIDIA Technical Blog
Summary
NVIDIA introduces ModelExpress, a solution for rapidly distributing AI model artifacts, as described in their technical blog.
Similar Articles
ModelExpress: Distributing Model Artifacts at the Speed of Light (12 minute read)
NVIDIA introduces ModelExpress, a tool that accelerates distribution of model artifacts across GPU clusters by using GPU-to-GPU RDMA transfers and optimized streaming from storage, reducing model startup times from 8 minutes to under 2 minutes for large models like DeepSeek-V4 Pro.
NVIDIA/Model-Optimizer
NVIDIA Model Optimizer is a library offering advanced model optimization techniques like quantization, pruning, and distillation to accelerate AI models, with integration into NVIDIA's ecosystem for deployment.
NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
NVIDIA announced Nemotron 3.5 Lightning, a 30B mixture-of-experts open model optimized for high-volume agentic AI workloads, alongside NeMo Switchyard, an open-source library for intelligent model routing across heterogeneous model ecosystems.
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
NVIDIA NeMo AutoModel leverages HuggingFace Transformers v5 to deliver 3.4-3.7x higher training throughput and 29-32% less GPU memory for fine-tuning Mixture-of-Experts models, with no code changes beyond a single import.
@dhruvtwt_: Why is no one talking about this? @nvidia is offering around 80 AI models via hosted APIs absolutely for free. You get …
Nvidia quietly provides ~80 free hosted AI model APIs including MiniMax M2.7, GLM 5.1, Kimi 2.5, DeepSeek 3.2, GPT-OSS-120B, ready to integrate with popular dev tools like OpenClaude and Zed IDE.