yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash

Reddit r/LocalLLaMA Models

Summary

Yandex has released AliceAI Foundation 80B A3B Base, a custom AI model competing with Qwen 35B and DeepSeek V4 Flash, featuring its own architecture without post-training.

https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base It's not a Qwen3 finetune, it's actually its own fully custom architecture. No Llama.cpp support yet sadly (Also note that this model is NOT post-trained like Qwen3.5/3.6)
Original Article

Similar Articles

yandex/AliceAI-Foundation-80B-A3B-Base

Hugging Face Models Trending

Yandex has released AliceAI-Foundation-80B-A3B-Base, a foundation language model with a hybrid architecture and MoE layers, featuring 80 billion parameters (3 billion active per token), trained from scratch and excelling in Russian factual knowledge tasks.

Qwen 3.8 27b vs Deepseek Flash

Reddit r/LocalLLaMA

The post compares the open-source AI models Qwen 3.8 (27B) and Deepseek Flash, discussing benchmarks and seeking user experiences to evaluate their performance.

Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash-GGUF

Hugging Face Models Trending

This entry describes Qwen3.5-9B-DeepSeek-V4-Flash, a distilled AI model that transfers reasoning capabilities from DeepSeek-V4 into a smaller 9B parameter space for efficient inference.

DeepSeek v4.1 Flash

Hacker News Top

DeepSeek has introduced DeepSeek-V4.1-Flash, a new AI model designed for enhanced capability, faster inference, native visual understanding, and scalability as part of their latest architecture family.

deepseek-ai/DeepSeek-V4-Flash-0731

Simon Willison's Blog

DeepSeek released DeepSeek-V4-Flash-0731, a 304B parameter model with enhanced agentic capabilities, priced at $0.14/M input and $0.27/M output, punching above its weight and ranking as the best value-per-intelligence model according to Artificial Analysis.