llm-optimization

Tag

Cards List
#llm-optimization

@TeksEdge: Someone just made ordinary Qwen behave a LOT more like Jev without training a new model. TOP Jev Clone (according to HF…

X AI KOLs Timeline ↗ · 2d ago Cached

A new technique called JEVfire enables existing LLMs like Qwen to behave more like Jev by modifying decision-making processes without retraining, resulting in significantly faster JSON generation and enabling local AI agents to run efficiently on consumer hardware.

0 favorites 0 likes
#llm-optimization

CoVeR: Coverage-Based Routing of Verifier Calls in Agentic Retrieval

arXiv cs.CL ↗ · 3d ago Cached

CoVeR is a coverage-based routing method that reduces LLM verifier calls by 62-68% in agentic retrieval systems while maintaining accuracy on multi-hop QA benchmarks.

0 favorites 0 likes
#llm-optimization

ClusterFewshot: Improving Few-shot Optimization for LLMs workflow

arXiv cs.CL ↗ · 3d ago Cached

ClusterFewshot is a novel method for improving few-shot demonstration selection in LLM workflows by integrating semantic clustering and utility scoring, which reduces optimization costs and enhances accuracy in DSPy-based pipelines.

0 favorites 0 likes
#llm-optimization

OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling

arXiv cs.AI ↗ · 3d ago Cached

OptiSkill is a framework that builds a hierarchical and evolving SkillBank for LLM-based operations research modeling, improving formulation accuracy by storing and validating reusable skills across problems.

0 favorites 0 likes
#llm-optimization

@0xCarnagee: https://x.com/0xCarnagee/status/2101019382009004249

X AI KOLs Timeline ↗ · 2026-09-18 Cached

This article provides 10 steps to optimize AI agent decision-making by using the langchain-typesafe tool, reducing LLM call costs to $0.042 per million tokens and latency to 70-500ms while preventing hallucinations.

0 favorites 0 likes
#llm-optimization

D-Quant: Driftable Entropy Coding for KV Cache Quantization

arXiv cs.CL ↗ · 2026-09-18 Cached

D-Quant is a flexible KV cache quantization framework using driftable entropy coding to convert variable-length representations into fixed-size bitstreams, enabling efficient parallel dequantization in attention kernels and achieving up to 7x memory reduction and 3.5x throughput improvement while preserving near BF16 performance.

0 favorites 0 likes
#llm-optimization

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

Hugging Face Daily Papers ↗ · 2026-09-10 Cached

COBRA-Skills introduces a method for optimizing agent skills using contextual bandits and evolutionary operators, achieving 55-58% lower optimization costs across multiple AI agent benchmarks.

0 favorites 0 likes
#llm-optimization

@somi_ai: If this holds up, you stop picking one model per app. Cheap model does the boring turns. Hard turn comes up, you hand t…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

NVIDIA's paper introduces a method to transfer KV cache between AI models, allowing target models to skip prefill entirely and achieve 2.7 to 25x faster conversions, which could improve efficiency in multi-model applications.

0 favorites 0 likes
#llm-optimization

Megathread for listing latest open source projects, research papers that are helping optimizations, efficiencies and accessibility to Open Source LLM and related hardware, software ?

Reddit r/LocalLLaMA ↗ · 2026-09-03

This megathread compiles a list of the latest open-source projects, research papers, and hardware innovations focused on optimizing inference, efficiency, and accessibility for open-source LLMs and related technologies.

0 favorites 0 likes
#llm-optimization

Optimizing On-Device Inference for Apple Silicon (20 minute read)

TLDR AI ↗ · 2026-09-02

Apple's Lily engine optimizes on-device LLM inference for Apple silicon by leveraging unified memory and hardware, outperforming MLX-LM, and is tuned for the Qwen3.6-35B-A3B model's architecture.

0 favorites 0 likes
#llm-optimization

Dynamic Thinking

Reddit r/ArtificialInteligence ↗ · 2026-08-27

OpenFreedom's Dynamic Thinking feature reduces task costs by 86% and execution time by 80% compared to OpenClaw, with 74% fewer LLM calls, while maintaining quality.

0 favorites 0 likes
#llm-optimization

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

arXiv cs.AI ↗ · 2026-08-25 Cached

LLM4LLM introduces a deployment-aware closed-loop optimization framework to bridge kernel benchmarks and real LLM inference, achieving up to 6.98x speedups on H100 GPUs.

0 favorites 0 likes
#llm-optimization

@DeRonin_: i built a system to run LLMs without limits... my Token Router it's why Claude Code runs 10+ hours a day and my API spe…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

The author describes a token routing system to optimize LLM usage, reducing API costs by up to 70% through techniques like cost-based routing, prompt caching, and output discipline.

0 favorites 0 likes
#llm-optimization

@LambLabs: Meet Woolly We post-trained Qwen3-8B to run up to 2–3× faster on math & coding prompts and, more importantly, to fit on…

X AI KOLs Timeline ↗ · 2026-08-21 Cached

LambLabs introduces Woolly, a post-training technique that accelerates Qwen3-8B by 2-3x for math and coding tasks, enabling it to run on their chip, with a live demo available.

0 favorites 0 likes
#llm-optimization

@inco_ai: Hope you enjoy our first release! (and more to come)

X AI KOLs Timeline ↗ · 2026-08-18 Cached

Inco AI announces DFlash 2, a next-generation decoding technique that achieves up to 4.6× speed increase for Qwen3.8-27B on M5 Max MacBook Pro, improving efficiency without altering output.

0 favorites 0 likes
#llm-optimization

@marfinxx: This Stanford and MIT paper is f*cking insane A new research paper proves that optimizing the Python harness around an …

X AI KOLs Timeline ↗ · 2026-08-18 Cached

A Stanford and MIT research paper shows that optimizing the Python harness around LLMs can yield up to a 6x performance gap without changing model weights, with systems like Meta-Harness automating context evolution.

0 favorites 0 likes
#llm-optimization

Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper proposes a reinforcement learning method for bootstrapping niche multilingual code translation with execution-based supervision, introducing a new benchmark HumanEval-X++ and demonstrating significant improvements over baselines using Qwen-3.5 models.

0 favorites 0 likes
#llm-optimization

Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation

Reddit r/LocalLLaMA ↗ · 2026-08-13

A developer created a task-aware GGUF quantization pipeline that uses tensor-level bit allocation to improve Gemma 4 12B Q3 coding performance by 8.55% over a hand-tuned imatrix while increasing model size by only 0.119%.

0 favorites 0 likes
#llm-optimization

Your agent isn't expensive. Your context window is. Here's the math

Reddit r/AI_Agents ↗ · 2026-08-10

Explains why agent API bills grow quadratically with context length because each turn re-reads the full history, and shares practical techniques like expiring tool results, shrinking tool schemas, and compacting context to cut costs.

0 favorites 0 likes
#llm-optimization

Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization

arXiv cs.AI ↗ · 2026-07-28 Cached

Proposes a two-stage structured pruning framework for LLMs that jointly optimizes latency and model size using multi-objective depth pruning and parallel Bayesian optimization, achieving favorable trade-offs for edge deployment.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback