Tag
This paper analyzes how soft skills are expressed in CVs of ML engineers, data scientists, and software engineers, using an LLM-based pipeline to distinguish explicit keywords from narrative descriptions, and tests demand-side hypotheses against candidate-side data.
Announcing Ray Summit 2026, co-presented with vLLM, taking place August 24-26 in San Francisco. The full agenda is live, featuring tracks on foundation model training, multimodal pipelines, and RL at scale.
HASTE introduces a hierarchical multi-agent system for ML engineering that organizes cross-competition knowledge into three tiers, achieving 77.3% medal rate on MLE-Bench Lite while reducing compute by 52% and demonstrating that structured knowledge transfer outperforms flat memory approaches.
A practical guide on setting up iterative loops for AI coding agents with defined stop conditions, cloud execution, and notification channels to offload work without constant babysitting.
Researchers benchmarked 7 frontier models on autoresearch tasks. Fable-5 won overall, but the open model Kimi-K2.7-Code surpassed others on ML engineering tasks.
This article discusses best practices for LLM application development using Arize Phoenix, specifically highlighting the importance of using train/validation/test splits for honest evaluation and tracking regressions.
An ML team documents practical challenges encountered while fine-tuning and deploying Gemma-4, including incompatibilities with PEFT, SFTTrainer, DeepSpeed ZeRO-3, and lack of runtime LoRA serving support, along with workarounds for each issue.