@AdinaYakup: Hy-MT2 New translation model family from @TencentHunyuan 1.8B / 7B / 30B-A3B MoE Supports 33 languages 1.8B > 440MB wit…
Summary
Tencent Hunyuan released Hy-MT2, a family of translation models up to 30B parameters with MoE, supporting 33 languages and quantized for on-device use.
View Cached Full Text
Cached at: 05/22/26, 03:53 AM
Hy-MT2 🔥 New translation model family from @TencentHunyuan
✨ 1.8B / 7B / 30B-A3B MoE ✨ Supports 33 languages ✨ 1.8B > 440MB with 1.25-bit quantization ✨ Runs on device with faster inference ✨ 1.8B outperforms some commercial APIs https://t.co/Ep7gL8wedk
Similar Articles
SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX
NVIDIA Cosmos3, a 64B parameter image generation model, is released with INT4 quantization for local deployment on CUDA and MLX, with code and weights available and performance demonstrated on Apple Silicon.
Desert Ant Labs: local, fast models that run on device
Desert Ant Labs, a European AI lab, launches a suite of small, specialized on-device models for audio, vision, and text, offering fast inference and privacy benefits by running locally on devices.
GitHub - coder543/minnow: Fast LLaDA2.2 inference server
Minnow is a high-performance Rust inference server for LLaDA2.2 models, offering optimized inference with quantization, GPU acceleration, and significant speed improvements over standard Transformers implementations.
Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
The article announces the release of Qwen3.8-Flash-Next on MLX-serve, supporting 1 million token context with efficient performance on M5 Max hardware using quantized weights.
How many agents can 2×4090 actually run at once? Three weeks of llama.cpp concurrency data — soft cap 5 @ 64k, hard cap 9, and why.
A benchmarking report on running multiple AI agents concurrently using llama.cpp on 2× RTX 4090 GPUs, revealing performance limits and optimal configurations for Qwen models.