coding-models

Tag

Cards List
#coding-models

Ngram and world knowledge - why are we just building a coding model?

Reddit r/LocalLLaMA ↗ · 2026-09-22

The author discusses the need for AI models with better world knowledge, leveraging N-gram technology to fit more knowledge into smaller models, and questions why development focuses more on coding capabilities than broader world knowledge.

0 favorites 0 likes
#coding-models

Open-source coding models are getting really good — and dev can be almost free now

Reddit r/ArtificialInteligence ↗ · 2026-09-12

A developer outlines a workflow using expensive AI models for planning and cheap or open-source models for coding tasks, showing that with clear specifications, the quality gap between models narrows, making development nearly cost-free.

0 favorites 0 likes
#coding-models

@rohanpaul_ai: A strong coding model is not enough if the model managing its work does not know when to redirect, verify, or stop. Loo…

X AI KOLs Following ↗ · 2026-09-03 Cached

LoopArena benchmarks models as runtime controllers for coding tasks, revealing that even GPT-5.5 only achieves a 24.69% success rate, emphasizing the need for better control mechanisms in agent systems.

0 favorites 0 likes
#coding-models

@danshipper: Tea leaves: In order to be competitive today Google needs to catch up on frontier coding. Demis believes different fund…

X AI KOLs Timeline ↗ · 2026-08-05 Cached

A tweet commenting that Google needs to catch up on frontier coding to stay competitive, while Demis Hassabis focuses on fundamental research like world models for long-term goals.

0 favorites 0 likes
#coding-models

Using multiple coding models to develop an open-source static AI Agents Capability & Risk Analyzer

Reddit r/AI_Agents ↗ · 2026-07-26

Describes the development of an open-source static analyzer that leverages multiple coding models to evaluate the capabilities and risks of AI agents.

0 favorites 0 likes
#coding-models

I run 35B–480B coding models on my 36 GB MacBook by streaming MoE experts from SSD — self-contained app, and I publish the benchmarks that *failed* too

Reddit r/LocalLLaMA ↗ · 2026-07-25

Slipstream streams MoE expert weights from SSD instead of RAM, enabling large coding models (35B–480B) on 36 GB MacBooks. Benchmarks show ~13–19 tok/s for 35B models and ~2.8 tok/s for 118B, with honest reporting of failed approaches.

0 favorites 0 likes
#coding-models

@swyx: one thing i think people dont appreciate enough about @poolsideai is their unusual degree of openness — not only have t…

X AI KOLs Timeline ↗ · 2026-07-23 Cached

A tweet highlights PoolsideAI's unusual openness, praising their release of a small coding model, publication of papers, and full evaluation datasets, setting a standard for transparency in AI.

0 favorites 0 likes
#coding-models

@heyshrutimishra: Good models are everywhere now. Running them is the hard part. ClinePass is $10/month for a curated set of the best ope…

X AI KOLs Following ↗ · 2026-06-29 Cached

ClinePass is a $10/month subscription offering curated open-weights coding models (GLM 5.2, Kimi k2.7-code, DeepSeek V4 Pro, etc.) for use within Cline CLI and IDE, providing discounted access.

0 favorites 0 likes
#coding-models

@GergelyOrosz: What happens when the most capable coding model (Fable / GPT-5.6) is banned by the US government, and the sending most …

X AI KOLs Following ↗ · 2026-06-28 Cached

A speculative tweet by Gergely Orosz ponders the impact if the US bans the most capable coding model (Fable/GPT-5.6), suggesting businesses would shift to the next best open model (GLM-5.2) via inference providers for a cheaper and better alternative.

0 favorites 0 likes
#coding-models

@GergelyOrosz: This is from a popular inference provider GLM-5.2 plus the US banning the most capable new models means open source cau…

X AI KOLs Following ↗ · 2026-06-27 Cached

GLM-5.2 is a new open-source coding model that has caught up to closed-source SOTA models, potentially disrupting revenues of OpenAI and Anthropic.

0 favorites 0 likes
#coding-models

DeepReinforce releases Ornith-1.0 open-source coding models (2 minute read)

TLDR AI ↗ · 2026-06-26 Cached

DeepReinforce open-sources Ornith-1.0, a family of self-improving coding models from 9B to 397B parameters, trained on Gemma 4 and Qwen 3.5 foundations, featuring a novel RL approach that learns to generate its own scaffolds.

0 favorites 0 likes
#coding-models

Gemma 4 beats Qwen 3.5 (UPDATE), and Qwen 3.6 27B + MiniMax M2.7 is the best OpenCode setup

Reddit r/LocalLLaMA ↗ · 2026-04-23

Personal benchmark shows Gemma-4E4B tops for routing, Qwen-3.6 27/30B beats Gemma-4 for coding, and MiniMax M2.7 MXFP4 replaces giant Qwen-3.5 quants in an OpenCode llama-swap workflow.

0 favorites 0 likes
#coding-models

Google ramps up agentic AI efforts amid pressure from Anthropic

Reddit r/singularity ↗ · 2026-04-20

Google has formed a dedicated strike team to improve its coding AI models, ramping up agentic AI efforts amid competitive pressure from Anthropic. This signals an intensifying race in AI coding capabilities between major AI labs.

0 favorites 0 likes
#coding-models

Why we no longer evaluate SWE-bench Verified

OpenAI Blog ↗ · 2026-02-23 Cached

OpenAI announces it will no longer report SWE-bench Verified scores, citing two critical issues: 59.4% of failed problems have flawed test cases that reject correct solutions, and frontier models have seen benchmark problems during training, making improvements reflect training data exposure rather than genuine capability gains.

0 favorites 0 likes
← Back to home

Submit Feedback