Tag
A user shares an observation that Qwen and Gemma tokenize code very differently, with Qwen using far fewer tokens for the same HTML/JS input, which may explain differences in coding and language performance. They also note a potential retraining project by LiquidAI using a more efficient tokenizer.
A blog post argues that the common saying "code was never the hard part" insults programmers, defending the difficulty of coding and questioning the overemphasis on requirements gathering.
OpenAI's new model Astra is rumored to have made a breakthrough in coding ability, but because its capabilities are too strong and may pose cyberattack risks, its release has been postponed.
In his talk, Carson Gross discussed the impact of AI on university computer science education, arguing that in the AI era, students still need to be taught to write and read code, while also noting that AI brings an assessment crisis and opportunities for pedagogical change.
BigBang-v1 is a self-evolving 36B LLM from Endless Frontier Lab in Shanghai, trained with AI-generated frontier tasks and achieving strong performance with only 10K high-quality examples across science, coding, tool use, and long context.
A discussion questioning whether older AI models like GLM 5.2 and Kimi 2.7 remain relevant for coding now that newer models such as Kimi K3, Qwen 3.8 Max, and DeepSeek V4 Pro are arriving.
Yohei Nakajima describes upgrading his AI-agent stack to work hands-free from anywhere, with a chief of staff assigning tasks to project managers and builders, all coordinated via repos.
HAR is an open-source harness for orchestrating multi-agent coding workflows.
Meta launched Muse Code, a terminal-based AI coding agent in beta that handles complex software engineering tasks across large codebases, competing with OpenAI's Codex and Anthropic's Claude Code.
Prime Intellect introduced Prime Agent, a self-improving RLM harness for coding and long-running autonomous tasks, featuring programmatic tool calling, context as a variable, multi-agent messaging, and self-modifiable harness state.
Ling-3.0-flash MXFP4, a quantized model, has been released and runs locally on a single DGX Spark, achieving ~80 tok/s decoding and 2,500-3,500 tok/s long-input prefilling, enabling private on-device inference for coding, agents, and offline batch jobs.
AI has made learning to code significantly easier by acting as a patient, 24/7 personal tutor that explains concepts, debugs code, and guides learners step-by-step without judgment.
User benchmarks DeepSeek v4 Flash against Qwen3.6-27B, Qwen3.5-122B, and Gemma 4 31B on a local coding benchmark, finding Flash wins overall but Qwen 122B performs surprisingly well with better first-try success and lower token usage.
Epoch AI and METR introduce MirrorCode, a benchmark that tests AI models on reimplementing entire programs end-to-end over long horizons. Early results show Claude Opus 4.7 solving a bioinformatics toolkit in 14 hours at $251, though memorization caveats remain.
Qwen3.8-Max, a 2.4T parameter open-weight model, matches Kimi K3 and DeepSeek V4 Flash on benchmarks, excelling in coding and software tasks. Weights release next week, with pricing of $2/M input and $6/M output tokens.
AI9Stars released G9v3-39A5B, an open-weights 39B MoE language model with 5 active experts, targeting reasoning, coding, and assistant tasks under Apache 2.0.
Alibaba announces Qwen3.8-Max, a 2.4T-parameter MoE frontier model with open weights coming next week, claiming autonomous operation for 16 days and significantly lower cost than GPT-5.6 Sol and Claude Fable 5.
An analysis of the 'AI productivity gap' in software engineering, arguing that AI mainly speeds up the coding portion of developers' jobs while leaving other crucial tasks like design, reviews, and meetings largely unchanged, leading to only modest overall gains. It also notes juniors benefit more than seniors, contrary to some leaders' assumptions.
A tweet jokes that while Anthropic is cautious about releasing powerful AI, Qwen has released Qwen3.8-Max, a new model for coding and cowork, asking users to star their repo.
Qwen 3.8-Max official version released. The video demonstrates the Qwen Office and Qwen Code tools, focusing on programming and collaborative office capabilities.