model-optimization

Tag

Cards List
#model-optimization

Show HN: Optimize and serve models with Fable quality at half the cost

Hacker News Top · 2026-07-26 Cached

World Model Optimizer is an open-source CLI tool that optimizes AI models from agent traces and serves them with a router to maintain frontier model quality at reduced cost.

0 favorites 0 likes
#model-optimization

Our 1-bit quant of Hy3 295B runs 2.2x faster than the cloud API with no quality loss

Reddit r/LocalLLaMA · 2026-07-20

A 1-bit quantized version of the Hy3 295B model achieves 2.2x faster inference speed compared to the cloud API with no quality loss.

0 favorites 0 likes
#model-optimization

DFlash makes Qwen3.6 27B 2.2x faster with no quality loss

Reddit r/LocalLLaMA · 2026-07-16

DFlash is a method that accelerates Qwen3.6 27B model inference by 2.2x without quality degradation.

0 favorites 0 likes
#model-optimization

@sakurayukiai: Gemma 4 wakes ~4B of 26B weights per token, so the fix was beautifully boring: split one unsupported fused op into two …

X AI KOLs Timeline · 2026-07-15

A fix splitting a fused operation into two matmuls allows Gemma 4 to run on a 13-year-old dual-Xeon CPU at 5.2 tok/s, using only 4B of 26B weights per token.

0 favorites 0 likes
#model-optimization

@0xkeenz: Interesting, I was just studying the differences between Unsloth and NVIDIA's Qwen3.6 27B NVFP4 yesterday, and today Unsloth updated! The new Unsloth's quantization approach is very similar to NVIDIA's official solution: instead of choosing between BF16 and ...

X AI KOLs Timeline · 2026-07-10 Cached

Unsloth releases a new version of Qwen3.6 27B NVFP4 quantization scheme, introducing FP8_E4M3 intermediate precision layer and refined weight protection, achieving 2.5x speed improvement on 24GB VRAM, while improving accuracy and tool-calling capabilities.

0 favorites 0 likes
#model-optimization

Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization

arXiv cs.LG · 2026-07-09 Cached

This paper proposes a latency-oriented neural network learning method that uses zerorized batch normalization to optimize DNN architectures for edge systems under strict latency constraints. Experiments show significant latency reduction with minimal accuracy loss on NVIDIA Jetson devices.

0 favorites 0 likes
#model-optimization

@akshay_pachaar: You're in an ML Engineer interview at Anthropic. The interviewer asks: "Our model generates 100 tokens in 42 seconds. H…

X AI KOLs Timeline · 2026-07-07 Cached

Explains how KV caching speeds up LLM inference by eliminating redundant recomputation of attention keys and values, trading off speed for memory, and introduces production-scale cache management challenges.

0 favorites 0 likes
#model-optimization

@DeRonin_: Fable 5 leaves your plan tonight steal its brain before it's gone.. then a cheaper model keeps its output for you captu…

X AI KOLs Timeline · 2026-07-07 Cached

A workflow strategy for capturing and reusing AI model outputs, including setting quality bars, strategy roadmaps, knowledge notes, and reasoning records to reduce reliance on frontier models.

0 favorites 0 likes
#model-optimization

LivePortrait distilled model that can run at 25fps in the browser

Reddit r/LocalLLaMA · 2026-07-05 Cached

LivePortrait has been distilled into a model that can run at 25 frames per second directly in the browser, as shown in a Hugging Face Spaces demo by slperez.

0 favorites 0 likes
#model-optimization

Using "applications" to make a smaller model more effective at bigger tasks.

Reddit r/LocalLLaMA · 2026-07-05

A discussion on using applications to enhance the effectiveness of smaller AI models on larger tasks, balancing efficiency and performance.

0 favorites 0 likes
#model-optimization

Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train

Hacker News Top · 2026-07-02 Cached

This paper systematically studies layer-wise contribution in RL post-training for LLMs, finding that training a single middle transformer layer can recover or even surpass full-parameter RL gains, with consistent patterns across models and tasks.

0 favorites 0 likes
#model-optimization

@MiaAI_lab: Nvidia did it again! @NVIDIAAI's Qwen 3.6 27B NVFP4 is faster than Unsloth's Qwen 3.6 27B NVFP4 by a whopping ~41% on D…

X AI KOLs Timeline · 2026-06-30 Cached

Nvidia's optimized Qwen 3.6 27B NVFP4 model achieves 41% faster single-session inference and 23-25% faster concurrent inference on DGX Spark compared to Unsloth's version.

0 favorites 0 likes
#model-optimization

@MSFTResearch: AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Lea…

X AI KOLs Following · 2026-06-30 Cached

SkillOpt turns AI agent skill editing from manual modification into a training process, improving agent reliability without changing model weights, achieving consistent gains across benchmarks.

0 favorites 0 likes
#model-optimization

@karminski3: DeepSeek truly excels in both cost-effectiveness and technology... Some classmates don't understand what DSpark is, so here's a quick tutorial. Speculative decoding is a technique to improve the output speed of large models. The essence is to let a small model generate text for the large model to check. Because currently...

X AI KOLs Timeline · 2026-06-29 Cached

DeepSeek proposes the DSpark technique, which implements speculative decoding by inserting a mini Transformer after the Final RMSNorm, boosting large model output speed by 60%-85%.

0 favorites 0 likes
#model-optimization

Compressed Whisper large-v3-turbo to 368 MB with Q3_K-matched QAT — multilingual WER results

Reddit r/openclaw · 2026-06-28

Whisper large-v3-turbo has been compressed to 368 MB using Q3_K-matched quantization-aware training, with multilingual word error rate results reported.

0 favorites 0 likes
#model-optimization

I Figured Out What Causes 'Super Weights'

Reddit r/ArtificialInteligence · 2026-06-23

Explains that super weights in large language models arise from the SoftMax-Attention interaction creating a 'Nothing Dump' token that serves as a stable reference point; removing these weights cripples performance.

0 favorites 0 likes
#model-optimization

@Zephyr271828: You want a strong small LLM. Would you start small — or inherit from something bigger? New paper: Small LLMs: Pruning v…

X AI KOLs Timeline · 2026-06-23 Cached

A new paper investigates whether it's better to prune a larger LLM or train a small LLM from scratch, finding that pruning provides more than just a good initialization.

0 favorites 0 likes
#model-optimization

@eliebakouch: every infra piece you need to know to do RL on GLM-5 https://primeintellect.ai/blog/rl-at-1t-scale…

X AI KOLs Timeline · 2026-06-23 Cached

Prime Intellect releases prime-rl v0.6.0, enabling efficient reinforcement learning at trillion-parameter scale on large Mixture-of-Experts models, with sub-5-minute step times and optimizations for asynchronous RL.

0 favorites 0 likes
#model-optimization

@RayFernando1337: “The selected runtime uses NVFP4 weights for maximum performance. From the original FP8 weights, we performed an in-hou…

X AI KOLs Following · 2026-06-23

Discusses using NVFP4 4-bit floating point weights for maximum performance, achieved via in-house quantization from FP8 using NVIDIA ModelOpt, highlighting the data format's dual scale factors for high dynamic range.

0 favorites 0 likes
#model-optimization

@jakevin7: Recently I've been reading about GLM 5.2 and found some interesting things to share. GLM-5.2 uses MTP (Multi-Token Prediction) to accelerate inference: a lightweight "draft model" quickly predicts multiple tokens, then the main model verifies them all at once; if accepted, it skips the decoding steps.

X AI KOLs Following · 2026-06-19 Cached

GLM-5.2 adopts MTP (Multi-Token Prediction) technology to accelerate inference and fixes a training-inference discrepancy in GLM-5.1's MTP that caused KV cache mixing issues.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback