This paper presents accurate software models of AMD GPU matrix cores for CDNA 1/2/3 architectures, validated for bit-level reproducibility against hardware, and demonstrates their use in numerical applications to compare accuracy with NVIDIA tensor cores.
GoBench is a benchmark for evaluating large language models on 9x9 Go games, demonstrating strong correlation with ARC-AGI and featuring a leaderboard with KataGo opponents from random to superhuman levels.
A study found that emergency department visits for gambling disorders nearly doubled in Ontario after the expansion of online gambling, particularly among young men, indicating increased harm from legalized sports betting.
Linum AI introduces JiT-DDT, a novel encoder-decoder architecture that trains text-to-image models 3.6× faster than previous methods while generating images with higher resolution.
This paper compares context trimming strategies for AI agents, finding that protocol-aware trimming with adaptive budget guardrails maintains high task success while reducing tokens, though it relies on gold annotations for implementation.
A paper in Physical Review D suggests that neutrinos changing flavor inside supernovae could carry energy away, potentially explaining discrepancies in core-collapse supernova models and observed rates.
Researchers led by neuroscientist Sergiu Pașca have created mice with human brain cells transplanted into genetically modified cortices, showing improved cognition in maze tests and raising ethical concerns about cross-species brain research.
LARA is a research project and PyTorch library that enables modular, composable behaviors for frozen large language models using low-rank residual adapters, allowing efficient training and inference-time blending of multiple behaviors.
The article critiques AI detectors like Pangram and GPTZero for their low accuracy against fine-tuned AI models and references a study indicating readers prefer AI outputs trained on copyrighted books.
Researchers developed smart nanoparticles to deliver mRNA directly to tumor-associated macrophages, reprogramming them to enhance immune response and slow tumor growth in mouse models of breast cancer.
A 3-year-old boy with metastatic hepatoblastoma achieved complete cancer regression after two doses of experimental CAR T cell therapy, administered outpatient without systemic toxicity.
This study compares memory systems to full conversation history in AI agents over simulated days, showing 23-62x fewer context tokens with similar or better recall on personal facts, but memory loses on numerical data and specific details like identifiers.
Chimpanzees learn tool use through social teaching, where adults model behavior and transfer tools to offspring, as per a new study in Frontiers in Psychology.
This paper introduces a neuro-symbolic Hierarchical Planning Decoder (HPD) for anticipating human intentions from multimodal episodes. It demonstrates improved performance in goal inference and logic constraint satisfaction on a benchmark dataset.
The paper introduces MASA, a method that uses frozen multimodal large language models to break the self-referential loop in wild test-time adaptation by providing structured semantic descriptions for more reliable adaptation.
SKIP is a self-knowledge-guided step-wise preference learning framework that improves reasoning compression in large language models, mitigating performance degradation and reducing overthinking by using DPO to guide efficient reasoning.
ORDER introduces a framework that dynamically adapts indexing and retrieval strategies in RAG systems based on the query, improving performance in expert domains.
ThinkFlow is a novel end-to-end latent memory framework for lifelong conversational agents that uses probabilistic vectors to overcome textual memory bottlenecks. It enables autonomous personalization through self-evolution and test-time learning, outperforming existing memory systems.
FlexEE introduces a self-speculative and KV-cache-compatible early exiting framework for efficient LLM inference in offloading deployments, achieving significant speedups on Llama models with minimal accuracy degradation.
This paper proposes an Affect-Prototype-Conditioned Fusion (APCF) framework for open-vocabulary multimodal emotion recognition with incomplete modalities, using an affect-prototype library to guide feature fusion and an LLM decoder for generating natural language emotion labels.