Tag
The authors have created a lightweight, performant asynchronous reinforcement learning stack in pure JAX, sharing a work log with insights on inference, RDMA weight transfer, memory optimizations, and sharding for scaling RL systems.
A case study of an LLM coding agent implementing a multi-component data system, analyzing defects and evaluating retrieval strategies on HotpotQA, highlighting gaps in automated versus empirical testing.
The article lists multiple ways to earn income in robotics through selling systems, services, artifacts, and owning verticals, highlighting the industry's reliance on practical engineering skills over pure AI development.
The article argues that the biggest AI shift is not larger models but better systems around them, such as context, model routing, caching, agent workflows, and evaluation, making the model the engine and the system the product.
The author shares their experience building a prototype to verify AI-generated financial claims, focusing on systems and engineering challenges like evidence reconciliation and deterministic verification, and invites conversations with like-minded engineers.
This paper presents a black-box evaluation framework to assess LLMs' ability to generate Design Structure Matrices (DSMs) from structured technical documentation. It introduces reproducible metrics and a composite quality score, showing that while LLMs can produce plausible DSMs, they remain sensitive to ambiguity and prompt formulation.
Harvard has open-sourced a comprehensive two-volume Machine Learning Systems textbook that covers engineering AI systems for real-world constraints, including distributed training, production inference, edge deployment, and governance, with hands-on components like TinyTorch, hardware kits, and interactive tools.
DeepSeek released DSpark, a system where the main model rapidly generates a sentence while a tiny editor fixes coherence before verification, pushing LLM systems engineering beyond new architecture.
This paper reviews the progress in AI for Systems Engineering (AI4SE) and Systems Engineering for AI (SE4AI) over the past decade, identifies five critical research gaps, and provides a human-AI agreement dataset and web explorer for relevance judgments.