model-inference

Tag

Cards List
#model-inference

The model fills the blank. Nobody gave it that authority.

Reddit r/AI_Agents · 4d ago

The article discusses the risk of AI models inferring missing data in agent systems, potentially leading to unauthorized executions, and proposes separating decision authority from models to ensure only explicitly declared instructions are followed.

0 favorites 0 likes
#model-inference

Dual AMD Radeon AI Pro R9700 or dual NVIDIA or RTX 3090.

Reddit r/LocalLLaMA · 2026-09-15

The author is building a dual GPU system for running large language models, evaluating AMD Radeon AI Pro R9700 versus NVIDIA RTX 3090 options while aiming to reduce AI subscription costs.

0 favorites 0 likes
#model-inference

Don't Sleep on EXL3 Quants

Reddit r/LocalLLaMA · 2026-08-30

The author shares their experience running a 30B parameter model with EXL3 quantization on a 12GB VRAM GPU, achieving efficient performance and speed for coding and agent tasks.

0 favorites 0 likes
#model-inference

It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.

Reddit r/LocalLLaMA · 2026-08-28

A user successfully used the mmap function in llama.cpp to fit the Qwen3.8-Flash-Next IQ3_XSS model into 16GB+64GB RAM, achieving a speed of 26 tokens per second, which outperforms a larger non-MOE 30B model.

0 favorites 0 likes
#model-inference

Spent a day seeing how far extreme MoE models can be pushed on a 4070 Ti + 32GB RAM. Kimi K3, DeepSeek V4 Flash, and Qwen3.5-122B results + research paper🔧

Reddit r/LocalLLaMA · 2026-08-25

This article details experiments with extreme Mixture-of-Experts models on consumer hardware using a custom runtime CRANE V2, and presents a research paper with results from Kimi K3, DeepSeek V4 Flash, and Qwen3.5-122B models.

0 favorites 0 likes
#model-inference

@no_stp_on_snek: Check out Buun's work, he cookin.

X AI KOLs Following · 2026-08-19 Cached

A user highlights Buun's work on optimizing AI models, achieving high-speed inference of Qwen 3.6 on a single 3090 GPU and developing DFlash2 for Qwen 3.8.

0 favorites 0 likes
#model-inference

Qwen 3.8 27b with DSH(DeepSeek Harness) is Amazing!! Experiences so far and perfomance.

Reddit r/LocalLLaMA · 2026-08-16

A user shares positive experiences using the Qwen 3.8 27b model with DeepSeek Harness, praising its stability and long-context handling, but mentions speed limitations and hopes for future model releases.

0 favorites 0 likes
#model-inference

@TheAhmadOsman: Prediction We’re gonna get Kimi K3 equivalent intelligence running on a single RTX PRO 6000 in less than 18 months How?…

X AI KOLs Following · 2026-08-15 Cached

A Twitter user predicts that AI intelligence comparable to Kimi K3 will run on a single RTX PRO 6000 GPU within 18 months, later noting that Opus 4.6 Max quality already fits on a single RTX 5090.

0 favorites 0 likes
#model-inference

@victormustar: I deployed FREE public endpoint for Qwen3.8-27B no token needed, OpenAI-compatible, light rate limiting. Powered by Hug…

X AI KOLs Following · 2026-08-15 Cached

A free public endpoint for the Qwen3.8-27B AI model has been deployed, offering an OpenAI-compatible API with vision support, tool calls, and a large context window, powered by Hugging Face Inference Endpoints for at least 72 hours.

0 favorites 0 likes
#model-inference

@TheAhmadOsman: I built this AI Server in 2023 with models like Qwen 3.8 27B and beyond in mind

X AI KOLs Timeline · 2026-08-15

A user built an AI server in 2023 designed to run models such as Qwen 3.8 27B.

0 favorites 0 likes
#model-inference

I implemented the YOLO26n model inference from scratch using ARM64 Assembly Language (No framework) [P]

Reddit r/MachineLearning · 2026-07-26

The author implemented YOLO26n model inference from scratch using ARM64 assembly language without any external frameworks, demonstrating low-level AI inference techniques.

0 favorites 0 likes
#model-inference

@DanKornas: Building video-generation workflows is easier when inference, model configs, and integration paths live in one place. L…

X AI KOLs Timeline · 2026-07-24 Cached

LTX-Video is an open-source Python repository by Lightricks for generating and conditioning videos locally using LTX-Video models, with support for text/image inputs, multi-condition workflows, and integration with ComfyUI and Diffusers.

0 favorites 0 likes
#model-inference

@Youssofal_: 72+ TPS on Qwen 3.6 27B on a Macbook pro M5 max. MTPLX V2 out now! The fastest way to run models on MLX.

X AI KOLs Following · 2026-07-07 Cached

MTPLX V2 is released, claiming 72+ tokens per second on Qwen 3.6 27B running on a Macbook Pro M5 Max via MLX.

0 favorites 0 likes
#model-inference

@yishan: Meta really fumbled this guy.

X AI KOLs Timeline · 2026-07-07 Cached

John Carmack comments on memory cost and capacity issues for AI accelerators, noting that model inference can have deterministic memory access patterns, contrasting with game rendering.

0 favorites 0 likes
#model-inference

DFlash support merged into llama.cpp

Reddit r/LocalLLaMA · 2026-06-28 Cached

DFlash support has been merged into llama.cpp, improving inference performance for compatible models.

0 favorites 0 likes
#model-inference

Best Settings for 48GB VRAM + Qwen 3.6 27B

Reddit r/LocalLLaMA · 2026-06-20

A user shares optimized settings for running Qwen3.6 27B (Q8_0) on a dual GPU setup (RTX 4090 + RTX 3090) with llama.cpp, achieving 75-100 t/s and 1500 pp with 250k context.

0 favorites 0 likes
#model-inference

Added an old 2070 Super to my rig and I can't go back...worse, now I need more

Reddit r/LocalLLaMA · 2026-05-31

A user shares their experience of adding an old NVIDIA 2070 Super GPU to their rig for extra VRAM, enabling them to run larger LLMs like Qwen3.6-27B at high quantization and context size with good performance, and now considering upgrading to a 3090 for even more VRAM.

0 favorites 0 likes
#model-inference

qwen3.6-35b-a3b-mtp running on GTX 1060 6GB

Reddit r/LocalLLaMA · 2026-05-24

A user successfully runs the Qwen3.6-35B-a3b-MTP model on a decade-old workstation with a GTX 1060 6GB using LMStudio under Windows, achieving acceptable chat speeds.

0 favorites 0 likes
#model-inference

Qwen3.6-35B-A3B Q4 262k context on 8GB 3070 Ti = +30tps

Reddit r/LocalLLaMA · 2026-05-22

The author shares detailed tuning tips for running the Qwen3.6-35B-A3B MoE model on an 8GB RTX 3070 Ti with up to 262k context using llama.cpp, achieving 30+ tps, and notes a 25% speed boost when switching from Windows to Ubuntu Server.

0 favorites 0 likes
#model-inference

Looking for early users to try our OpenClaw model plans and tell us what's broken (15–30 min)

Reddit r/openclaw · 2026-05-20

OpenClaw is seeking early users to test their open-source model inference plans, sold by concurrency slot with high throughput and no shared pool, in exchange for free access and feedback.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback