inference-benchmarks

Tag

Cards List
#inference-benchmarks

@LottoLabs: This should allow agents to purchase pro plans directly at http://localmaxxing.com

X AI KOLs Following · 2026-06-16 Cached

Nous Research announces Stripe integration for Hermes Agent, allowing AI agents to purchase pro plans and provision services directly, with configurable safety limits.

0 favorites 0 likes
#inference-benchmarks

Strix Halo ROCm + MTP Notes (May 2026)

Reddit r/LocalLLaMA · 2026-05-17

Technical benchmark comparing ROCm and Vulkan backends for LLM inference on Strix Halo hardware after MTP merged into llama.cpp, revealing ROCm suffers severe performance drops at full context while Vulkan remains stable.

0 favorites 0 likes
#inference-benchmarks

Qwen 35B-A3B is very usable with 12GB of VRAM

Reddit r/LocalLLaMA · 2026-05-08

A user benchmarks Qwen 35B-A3B (a 35B MoE model) on a 12GB RTX 3060, finding that 12GB VRAM is a practical sweet spot for running the model with 32k context, achieving ~47 t/s generation.

0 favorites 0 likes
← Back to home

Submit Feedback