ai-performance

Tag

Cards List
#ai-performance

Best case Voice AI Latency

Reddit r/AI_Agents ↗ · 14h ago

A tweet discussing a claim of 80ms latency in Voice AI systems, raising questions about its feasibility with custom models and on-prem inference.

0 favorites 0 likes
#ai-performance

OpenAI still leads in the benchmarks that matter

Reddit r/singularity ↗ · yesterday

The article discusses how OpenAI maintains its leading position in key AI benchmarks, highlighting its continued dominance in performance metrics.

0 favorites 0 likes
#ai-performance

@rohanpaul_ai: 19 Unitree humanoids joined 120 dancers in Shanghai before 10,000+ people in the world’s largest full-size live humanoi…

X AI KOLs Timeline ↗ · yesterday Cached

The world's largest live humanoid robot performance took place in Shanghai, featuring 19 Unitree humanoids joining 120 dancers in an AI-driven coordination showcase streamed globally.

0 favorites 0 likes
#ai-performance

@yoheinakajima: got object detection down to below 0.35 sec latency locally

X AI KOLs Timeline ↗ · 2d ago Cached

A developer shares their achievement of reducing object detection latency to below 0.35 seconds on local hardware, highlighting progress in AI performance optimization.

0 favorites 0 likes
#ai-performance

GPT-6 Astra makes a massive leap on ZeroBench (an extremely difficult vision benchmark), surpassing the human baseline across all three metrics

Reddit r/singularity ↗ · 2d ago

GPT-6 Astra achieves a massive leap on ZeroBench, an extremely difficult vision benchmark, surpassing the human baseline across all three metrics.

0 favorites 0 likes
#ai-performance

Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW

Reddit r/LocalLLaMA ↗ · 3d ago

The article compares GSQ and ByteShape quantizations of the Qwen 3.8 27B model on an RTX 3060, revealing that ByteShape's quant underperformed despite claims of high similarity to the original model.

0 favorites 0 likes
#ai-performance

Googlebooks launch October 4 starting at $899—here are the five models you can preorder today

Ars Technica ↗ · 4d ago Cached

Google is launching Googlebooks, a new line of Android-powered laptops with premium hardware and AI performance, starting at $899, with five models available for preorder.

0 favorites 0 likes
#ai-performance

@VraserX: Figure’s robot only completed 56% of its zero-shot household trials. That’s nowhere near a product I’d trust at home. I…

X AI KOLs Timeline ↗ · 6d ago Cached

Figure's robot achieved 56% success in zero-shot household trials, a major improvement from 9% without human-behavior pretraining, indicating progress but not yet ready for consumer use.

0 favorites 0 likes
#ai-performance

@10xmylife: Tried it out today, and it really works, so we'll go with this plan.

X AI KOLs Timeline ↗ · 6d ago Cached

A user tested the Jev model's performance in the game 'Slay the Spire 2', where its decision-making speed was extremely fast, taking only 0.7 seconds, far surpassing humans and GPT-6 Astra, which is astonishing.

0 favorites 0 likes
#ai-performance

600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"?

Reddit r/LocalLLaMA ↗ · 2026-09-17

Achieving 600 tokens per second on the Qwen3.6 model using Ninfer on an RTX Pro 6000, noted as useful for brute-force tasks despite not being the most advanced model.

0 favorites 0 likes
#ai-performance

Finally understood why my coding agent types fast on boilerplate and slow on new logic

Reddit r/AI_Agents ↗ · 2026-09-09

The author explains how speculative decoding affects coding agent speed, with higher acceptance rates on boilerplate code leading to faster typing, and discusses other factors like cache misses that impact performance.

0 favorites 0 likes
#ai-performance

@TechByMarkandey: What if progress is no longer about adding parameters?

X AI KOLs Timeline ↗ · 2026-09-07 Cached

MiniCPM5-2B is a 2B-parameter open-source language model that achieves high intelligence density for edge deployment and ranks highly in several AI performance benchmarks.

0 favorites 0 likes
#ai-performance

@Modular: People ask what our secret is when they see our performance numbers. Brendan Hansknecht, AI Performance Engineering Man…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

Brendan Hansknecht from Modular explains that their high performance numbers stem from treating performance as a full-stack problem, rather than relying on piecemeal components in production.

0 favorites 0 likes
#ai-performance

Any predictions on what GPT-6 Astra will score?

Reddit r/singularity ↗ · 2026-09-03

The content asks for predictions on the expected performance scores of the GPT-6 Astra model.

0 favorites 0 likes
#ai-performance

$60k in Macs for Local LLM vs $10 Subscription

Reddit r/AI_Agents ↗ · 2026-08-31

The article discusses a YouTuber's experiment demonstrating that expensive local hardware for running LLMs is not cost-effective compared to affordable cloud subscriptions, emphasizing the current practical limitations of local AI for everyday use.

0 favorites 0 likes
#ai-performance

@HannesVonEssen: How fast can we make it run?? My current best is 1.6 m/s

X AI KOLs Timeline ↗ · 2026-08-30 Cached

A tweet asking about the speed of an unspecified system, with the author sharing their personal best performance of 1.6 meters per second.

0 favorites 0 likes
#ai-performance

@0xSero: Qwen3.8-27B on 1500$ oh hardware. This is the realistic local AI experience for most people, for tons of reasons includ…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

The tweet discusses the feasibility of running the Qwen3.8-27B AI model on $1500 hardware, highlighting its usability and cost-effectiveness for most people.

0 favorites 0 likes
#ai-performance

I wonder when people are going to realize we need to bring this back...

Reddit r/AI_Agents ↗ · 2026-08-27

The author argues for reviving 'Needle in a haystack' benchmarks to evaluate AI capabilities, sharing private test results that show many models performing poorly in remembering instructions, questioning their trustworthiness for real-world tasks.

0 favorites 0 likes
#ai-performance

@0xSero: Opus at home, at 200+ tok/s I love ZAI

X AI KOLs Timeline ↗ · 2026-08-26 Cached

The user @0xSero shares running Anthropic's Opus model at home with over 200 tokens per second using ZAI, expressing enthusiasm for ZAI.

0 favorites 0 likes
#ai-performance

Ran our Apache 2.0 Gepard TTS through Coval's public benchmark. 68.7 ms to first audio on one RTX 4090.

Reddit r/LocalLLaMA ↗ · 2026-08-26 Cached

The article reports on benchmarking the Gepard 1.0 text-to-speech model, achieving 68.7 ms to first audio on a single RTX 4090 GPU, outperforming several commercial APIs in latency while being open-source under Apache 2.0.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback