Tag
A person announces joining @amilabs as a Member of Technical Staff in Paris, expressing belief that this team will help AI better understand the real world.
The inaugural 'Mathematics in the Age of AI' seminar at NYU Courant was packed, featuring Tristan Buckmaster's talk on breakthroughs in Euler equations with smooth forcing.
The article explores Dynamic Abliteration, a non-destructive method to suppress refusal behavior in open-weight LLMs like Qwen3-4B at runtime without permanently altering model weights, using multi-layer engram steering via PyTorch forward hooks.
Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.
Research finds that GPT-4 can produce empty responses to null prompts while GPT-3.5 cannot, with cross-vendor studies confirming similar behavior in other models and an open-source tool introduced for controlling EOS token behavior.
Jeff Dean highlights research on computing with rat neurons and a partnership between BioComputingCo and Amazon to bring this technology to customers.
The article likely presents research on contrastive language models, exploring the use of contrastive learning techniques in language model development.
The paper proposes LWCal, a CPU-only post-hoc calibration method for tabular classifiers that handles noisy calibration labels without requiring clean data, showing improved calibration error and scoring metrics.
The paper investigates how editorial framing in prompts affects LLMs' factual and tonal responses in data analysis, finding that factual errors occur in specific scenarios, while tonal shifts are more common.
This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.
This survey paper reviews developments in brain-to-language decoding, translating neural activity into linguistic outputs for communication restoration and scientific study, covering tasks, methods, evaluation, and future directions.
The paper introduces Emergi-PersonaOS, a psychology-grounded operating system for persona agents that enables situational adaptation and controllable evolution through a three-layer persona representation and mechanisms for belief updates and experience development.
PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.
PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.
The paper investigates fine-tuning strategies for customer support LLMs, comparing multi-task training, sequential updates, and model merging across multiple model families. It concludes that multi-task full fine-tuning is the most robust default, while specialist models degrade off-task and require reliable routing.
The paper proposes principled methods for context representation in large-scale AI reasoning, introducing R3Con which outperforms baselines and enables smaller models to achieve performance comparable to larger ones at lower cost.
This paper investigates whether large language models internally organize mathematical reasoning by topic or approach, finding evidence that approach is the key organizing principle, challenging traditional benchmarking methods.
The paper quantifies how improvements in fact-verification scores are partitioned between answer accuracy and evidence quality, using trained DeBERTa checkpoints and LLMs across multiple benchmarks.
The paper proposes AV-GRPO, a modality-anchored reinforcement learning framework for joint audio-video generation that improves generation quality, semantic alignment, and cross-modal synchronization over existing methods.
The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.