Tag
This article discusses a high-priced course aimed at preparing individuals for emerging AI jobs that are not yet fully defined, highlighting the growing demand and uncertainty in the AI job market.
Training of the Marin 535B-A23B AI model has begun with an open process, involving pretraining and midtraining on 18.75T tokens using GB200 NVL72 hardware over about 3 months.
Runtime by Modal is a conference event for AI and ML engineers to discuss running AI in production, featuring talks on inference, training, and agents.
Matt Pocock shares a template letter for developers to convince their bosses to invest in his AI Hero course, which teaches a structured engineering process for using AI coding assistants effectively.
A tweet announces the live training of a 535B parameter (23B active) large language model, with links to follow the process on Weights & Biases and GitHub.
An experiment training three from-scratch LLMs with the same GRPO recipe yielded inconsistent results, with GRPO degrading performance in some models, particularly the middle-sized one, and no clear relationship to scale.
LEGO-RL presents a framework to bridge native coding-agent harnesses with scalable policy-gradient reinforcement learning, improving performance on benchmarks like SWE-bench.
The article discusses challenges in training browser-agent models for sequential decision-making, such as error recovery and memory, and seeks input on optimizing training objectives. The author mentions working on the 'mako' model at tinyfish and invites community feedback.
Registrations are open for an 8-week AI Engineering Cohort focusing on RAG & Agents, aimed at software engineers upskilling into AI roles.
Introduces 'Training Under Challenge', an executable-certificate framework that uses architecture-valid procedures to construct alternative models and estimate the empirical global-optimality gap of neural network checkpoints, with theoretical guarantees and experiments on ResNet-18 distillation and quantized denoising.
Andrej Karpathy built an LLM from scratch and open-sourced all code, covering the full pipeline of dataset, training loop, and inference loop, continuing his work on simplifying AI education with micrograd and makemore.
The author discusses a paper that demystifies the warm-up process for OPD (likely on-policy distillation), explaining how warm-up enables well-defined student-sampled sequences and educational token-level dense rewards from the teacher.
This paper proposes an adaptive supervised anchoring framework for on-policy self-distillation, addressing the problem of rollout-conditioned signal degradation in language model training. The method separates rollout-conditioned distribution matching from canonical-context supervision, improving task acquisition while preserving general capabilities.
Multiverse Computing announces a paper on making LLM knowledge distillation cheaper via offline top-K logits and a fused chunked KL loss, cutting VRAM usage for distillation at scale.
Pathway's BDH, a post-transformer architecture, reportedly matches GPT-2 scaling from 10M to 1B parameters while training from scratch on standard GPUs.
Prime RL now supports expressing and training multi-agent systems, enabling use cases like agentic judge, self-play, user simulation, and complex agent collaboration.
Tesla Europe announces FSD Supervised has been trained and tested on 2.2 million km across 19 EU countries, handling rare edge cases.
ByteDance is in the early stages of training a large language model with up to 10 trillion parameters, signaling a massive scale-up in AI development.
Gregory Kurtzer praises Lucas Atkins' talk on lessons learned from training a large sparse Mixture-of-Experts model and life at an AI lab startup.
据金融时报报道,字节跳动正在训练一个估算参数量达10万亿级别的AI大模型,规模接近Anthropic的先进系统,旨在缩小与美国顶级AI实验室的差距。