efficient

Tag

Cards List
#efficient

@sheriyuo: A 35B-parameter MoE agentic model with only 3B active that claims to match or surpass 100B-class models through post-tr…

X AI KOLs Timeline · 2026-07-13 Cached

A 35B-parameter MoE model with only 3B active parameters matches or surpasses 100B-class models using post-training RL, achieving significant efficiency gains.

0 favorites 0 likes
#efficient

@NFTCPS: Fraud call centers have a new weapon — voice cloning has been pushed to new heights again. LuxTTS, a lightweight TTS model, after seeing it I can only say: truly insane. Fast: 150x real-time on a single GPU, even runs faster than real speech on CPU. Clear: 48kHz directly, most models are still stuck at 24kHz…

X AI KOLs Timeline · 2026-07-05 Cached

LuxTTS is a lightweight voice cloning TTS model, supporting 48kHz high-fidelity output, achieving 150x real-time speed on a single GPU, requiring only 1GB VRAM for local operation, with performance comparable to models ten times its size.

0 favorites 0 likes
#efficient

gemma4 e2b is really good, what other small models work on crappy computers?

Reddit r/LocalLLaMA · 2026-07-03

A user praises the Gemma 4 e2b model for its speed and output quality on low-end hardware, comparing it favorably to ChatGPT 3.5 and 4, and asks for recommendations on other small models that work well on older computers.

0 favorites 0 likes
#efficient

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

Hugging Face Daily Papers · 2026-06-29 Cached

InnerZoom proposes a single-forward framework for cross-layer evidence bridging in GUI grounding, achieving state-of-the-art performance on multiple benchmarks while reducing latency by up to 31.8%.

0 favorites 0 likes
#efficient

@nickfrosst: now seems like a good day to remind people we have an apache 2.0 coding model you can run with 20 gigs of ram locally f…

X AI KOLs Following · 2026-06-26 Cached

Cohere Labs releases North Mini Code, a 30B parameter (3B active) open-source coding model under Apache 2.0, optimized for code generation and agentic tasks, capable of running locally with 20GB RAM via 4-bit quantization.

0 favorites 0 likes
#efficient

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

Hugging Face Daily Papers · 2026-06-23 Cached

Lite Any Stereo V2 presents an efficient stereo matching approach achieving state-of-the-art accuracy with significantly reduced latency through optimized architecture and training strategies, including a 2D-only cost aggregation framework and a three-stage training strategy.

0 favorites 0 likes
#efficient

@0x0SojalSec: Imagine fine-tuning a 31B parameter multimodal model for free,, on Kaggle. Now you can train this massive 31B dense mul…

X AI KOLs Timeline · 2026-06-20 Cached

Unsloth enables free fine-tuning of a 31B parameter multimodal model on Kaggle using 4-bit quantization, requiring only 22-24GB VRAM for local runs.

0 favorites 0 likes
#efficient

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

Hugging Face Daily Papers · 2026-06-18 Cached

Introduces SEVRA, a selective verification controller for budget-aware reasoning that decides when to accept a model's initial answer versus spending extra compute on verification, improving accuracy and reducing unnecessary tokens on benchmarks like MATH500 and GSM8K.

0 favorites 0 likes
#efficient

@cjzafir: A 3B parameter SLM: VibeThinker (fine-tuned on Qwen 2.5) matches Claude Opus 4.5 performance. Same performance as: > De…

X AI KOLs Timeline · 2026-06-17 Cached

VibeThinker, a 3B parameter model fine-tuned on Qwen 2.5, achieves performance comparable to Claude Opus 4.5 and much larger models like DeepSeek v3 through innovative post-training that includes multi-path thinking and staged training on math, coding, and science.

0 favorites 0 likes
#efficient

@rionaifantasy: Unbelievable! How Can a 34.5M Parameter OCR Beat a 235B Large Model? Let me tell you something ridiculous: I used to believe the future of OCR would inevitably be devoured by ever-larger multimodal large models. But after seeing PP-OCRv6 released by Baidu Wenxin, I've changed my mind. Because it doesn't follow the path of "continuing to pile on parameters..."

X AI KOLs Timeline · 2026-06-16 Cached

Baidu Wenxin releases PP-OCRv6, offering three model tiers: Tiny, Small, and Medium, supporting over 50 languages. The Tiny version is only 1.5MB and can run locally in a browser, with the fastest single-image inference at 97ms, proving that small specialized models can outperform large models on OCR tasks.

0 favorites 0 likes
#efficient

@nickfrosst: this model is the opposite of mythos. Its small, cost effective, apache 2.0, and locally deployable. This is the way LL…

X AI KOLs Following · 2026-06-09 Cached

Cohere released North Mini Code, its first open-source coding model under Apache 2.0, designed to be small, cost-effective, locally deployable, and focused on agentic performance.

0 favorites 0 likes
#efficient

@LeonEnglaender: We're just 8 people on our core code team and our 30B-A3B model lands on par with Claude Haiku 4.5 and ahead of NVIDIA'…

X AI KOLs Timeline · 2026-06-09 Cached

A team of 8 released a 30B-A3B coding model under Apache 2.0 that matches Claude Haiku 4.5 performance and beats NVIDIA's 120B-A12B Nemotron 3 Super on the Artificial Analysis Coding Index.

0 favorites 0 likes
#efficient

@cohere: Introducing Cohere's first open-source coding model: North Mini Code Small & efficient, designed for agentic performanc…

X AI KOLs Following · 2026-06-09 Cached

Cohere released its first open-source coding model, North Mini Code Small, designed for efficient agentic performance and community input.

0 favorites 0 likes
#efficient

You don't need a GPU to run gemma-4-26B-A4B

Reddit r/LocalLLaMA · 2026-06-07

The author demonstrates that the Gemma-4-26B-A4B model runs efficiently on a CPU-only system using Koboldcpp, achieving 7 tokens per second on an old desktop, suggesting that powerful GPUs may not be necessary for local LLM inference.

0 favorites 0 likes
#efficient

@_philschmid: We just launched a Gemma 4 12B! Our first mid-sized model with native audio inputs. Gemma 4 12 B is a unified, encoder-…

X AI KOLs Following · 2026-06-03 Cached

We just launched Gemma 4 12B, a mid-sized multimodal model with native audio inputs, requiring only 16GB memory and released under Apache 2.0.

0 favorites 0 likes
#efficient

WeCon: An Efficient Weight-Conditioned Neural Solver for Multi-Objective Combinatorial Optimization Problems

arXiv cs.LG · 2026-05-25 Cached

Presents WeCon, a weight-conditioned neural solver for multi-objective combinatorial optimization problems that achieves comparable hypervolume to the state-of-the-art while reducing inference time by 40%.

0 favorites 0 likes
#efficient

@FeitengLi: OpenBMB open-sources MiniCPM-V 4.6, 1.3B parameters (SigLIP2-400M + Qwen3.5-0.8B), 262k context, visual encoding FLOPs 50%+ less than previous generation. Token cost for the same task is lower than Qwen3.5-0…

X AI KOLs Timeline · 2026-05-16 Cached

OpenBMB releases MiniCPM-V 4.6, a 1.3B-parameter multimodal LLM with 262k context and significantly reduced visual encoding FLOPs, achieving strong benchmark performance and broad inference framework support.

0 favorites 0 likes
#efficient

Δ-Mem: Efficient Online Memory for Large Language Models

Hacker News Top · 2026-05-16 Cached

Proposes delta-Mem, a lightweight online memory mechanism that uses a compact state matrix updated by delta-rule learning to improve long-context performance of frozen LLMs without full fine-tuning or context extension.

0 favorites 0 likes
#efficient

@songhan_mit: Explore SANA World Model, using hybrid linear attention, efficient and fast!

X AI KOLs Following · 2026-05-15

SANA World Model is a new AI model that uses hybrid linear attention for efficiency and speed.

0 favorites 0 likes
#efficient

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

Hugging Face Daily Papers · 2026-05-14 Cached

SANA-WM is a 2.6B-parameter open-source world model that generates high-fidelity 720p minute-scale videos with precise camera control, achieving industrial-level quality while significantly reducing computational requirements.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback