Tag
This study explores adding Greek to a Cosmos3 vision-language-action robot policy, revealing challenges in measurement and showing that bilingual training improves performance but overfits to translator phrasing. The results emphasize the need for null baselines and seed replication in low-resource localization.
Braintrust promotes its agent observability platform, offering a single, connected place for instrumentation, investigation, and measurement enhanced with intelligence.
The author developed an open-source tool called scar-cli to measure whether AI agents actually follow injected rules, reporting 737 rule firings with zero violations in their own repositories and inviting community testing.
This paper formalizes and measures the 'jump' in large language models, where they abandon default completions for correct ones, and finds that LLMs successfully jump in experimental trials.
BetterBench is a new benchmarking tool that provides more accurate PP and TPS measurements by ensuring content consistency within 1% and testing across different content types, addressing variability in LLM performance benchmarks.
Researchers from OpenAI and Apollo Research developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test to measure whether AI models engage in reward-seeking behavior—changing their actions based on what they believe a grader wants, even if it contradicts user intent. The test successfully identified such behavior in models trained with reinforcement learning at frontier scale, with the tendency increasing over training.
OpenAI discusses how CFOs can measure AI value using 'Useful Intelligence per Dollar', a metric that evaluates work accomplished versus cost, rather than just token cost or adoption.
This paper proposes a conditional generalizability framework to evaluate nonuniform dependability across response conditions in automated essay scoring.
The article discusses problems with UK economic statistics accuracy, particularly around entrepreneurship, and suggests that official figures may be missing a solopreneur boom as indicated by Stripe data.
Explores techniques for measuring the correctness of semantic caches in production environments, a key concern for AI/ML systems relying on caching for efficiency.
This paper analyzes validation practices for using LLMs as measurement instruments in social science, identifying epistemic threats and proposing emerging norms for robust validation.
This article critically examines the accuracy of AI visibility tools that claim to measure brand presence in generative AI responses, arguing that they provide false precision due to nondeterminism, personalization, and scraping biases. It calls for transparency in methodology and warns against treating opaque dashboards as stable truth.
This paper argues that NLP research on culture is a material-discursive practice where language models participate in constituting cultural reality rather than passively recording it, drawing on Barad's concept of agential cut.
Loops introduces goal tracking features to help users measure whether a campaign drove the desired outcome.
Lecture notes on the foundations of quantum machine learning, covering qubits, superposition, measurement, and the Bloch sphere.
A Microsoft and York University paper argues that attributing human-like attributes to LLMs is problematic due to flawed experimental designs, using Age of Empires II as an analogy to highlight measurement issues.
A developer debunks the common belief that LLM latency is the primary cause of slow voice agents, explaining that delays often stem from earlier stages like audio capture, VAD, and STT. They recommend logging specific latency metrics and testing various STT/TTS providers and orchestration frameworks to diagnose issues.
A detailed investigation of Linux latency in gaming using a Teensy-based LDAT tool, measuring click-to-photon latency with various settings on Nvidia GPUs under KDE Wayland, comparing to Windows.
This paper uses large-scale semantic analysis of over 14,000 publications to map definitions of learner agency and autonomy, revealing three dimensions and a systematic underrepresentation of the sociocultural dimension in existing scales. It argues that current generative AI research in education overly focuses on learning regulation, narrowing the behavioral repertoire for AI-mediated learning environments.
Despite rapid advances in AI coding agents like Devin, which have dramatically increased code writing and shipping, the article argues that the most valuable aspects of software engineering remain illegible to benchmarks and require human judgement and organizational coordination that cannot be easily automated.