Tag
An agentic CUDA kernel optimizer that automates GPU implementation generation through iterative code generation, correctness checks, benchmarking, and refinement, powered by LangGraph and OpenAI models.
ELF-REG scales continuous diffusion language models to reasoning tasks by using representation alignment and entanglement, achieving improved performance on benchmarks like GSM8K and HumanEval compared to other diffusion models.
The author reflects on the decline of plan modes in AI-assisted coding, arguing that as models improve, these modes are becoming obsolete, while the need for human comprehension of software systems becomes more critical.
GPT-6 Astra generated a 4K video about time using p5.js code, spanning from the Big Bang to writing its own code, showcasing editable AI-created content.
A user created a 50-shot painted music video using Claude Opus 5.5 by providing an MP3 and art direction, with the model writing all code for audio analysis, storytelling, and rendering without any image or video generation models.
This article compares the game generation capabilities of AI models Opus 4.6 and Opus 5.5 by having them create a complete Wolfenstein 3D-style game in a single HTML file, highlighting advancements in AI code generation over recent months.
Opus 5.5 demonstrated advanced video creation capabilities by autonomously generating a SNES-style video game video with code, music, and characters, featuring a scenario involving Sydney, Altman, and Claude.
OpenWiki 0.6.0 is released with new features including shared workspaces, faster code wiki generation through concurrency controls, and integration with @antigravity and oh my pi.
GPT-6 Luna (Max) is praised as a cost-effective AI model with strong performance in automations and creative tasks, highlighted in Rails agent evals where it competes well with other models at a low price.
ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.
Jev State is a free and open-source tool that transforms AI conversations into tests and runnable code, enabling developers to build, test, and export conversational workflows to TypeScript and workflow JSON.
Anthropic engineer Dan Fein showcased Opus 5.5's capabilities in generating complex code-based animations with 13,000 tiles, highlighting significant advancements in visual code generation.
HyperFrames announces the Code2Video Bench, developed in collaboration with Google DeepMind and Kaggle, to evaluate and improve code-to-video generation in AI.
AutoRecLab is an autonomous Python-based system that automates recommender systems experiments from natural-language prompts, using retrieval-augmented generation (RAG), static verification, and tree search to generate and validate executable code.
This paper introduces SWE-Proof, a benchmark of formally verified code patches for real-world software issues, demonstrating that formal verification improves error detection in LLM-generated code and identifies specification synthesis as a key open problem.
This paper introduces CoVer, a co-training framework for code generation that addresses self-play RL failures by using information-gain rewards and diversity-pruned tests, achieving significant pass rate improvements on benchmarks.
The article details a test of the Qwen3.8-Flash-Next AI model running locally on Intel V620 GPUs, where it generated a 3D game from a sloppy prompt in about 3 hours using the OMP harness.
The tweet promotes QuixiAI's tool that allows users to create a running game in just 30 seconds with 30 lines of code, likely leveraging AI-assisted development.
This article tests various quantized versions of the Qwen 3.8 27B AI model on limited VRAM setups, comparing their performance on tasks like animation generation, app development, and word generation.
Kody is a tool that converts skills into runnable code for any agent, device, or custom application.