ai-research

Tag

Cards List
#ai-research

@tommiekerssies: I’ve joined @amilabs as a Member of Technical Staff in Paris! AI has yet to truly understand the real world. I’m convin…

X AI KOLs Following ↗ · yesterday Cached

A person announces joining @amilabs as a Member of Technical Staff in Paris, expressing belief that this team will help AI better understand the real world.

0 favorites 0 likes
#ai-research

@NYU_Courant: An absolutely packed house at yesterday's inaugural 'Mathematics in the Age of AI' seminar with Tristan Buckmaster! Pro…

X AI KOLs Following ↗ · yesterday Cached

The inaugural 'Mathematics in the Age of AI' seminar at NYU Courant was packed, featuring Tristan Buckmaster's talk on breakthroughs in Euler equations with smooth forcing.

0 favorites 0 likes
#ai-research

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

Hacker News Top ↗ · yesterday Cached

The article explores Dynamic Abliteration, a non-destructive method to suppress refusal behavior in open-weight LLMs like Qwen3-4B at runtime without permanently altering model weights, using multi-layer engram steering via PyTorch forward hooks.

0 favorites 0 likes
#ai-research

Accelerating vision-language models with LFM2.5-VL-DSpark

Hugging Face Blog ↗ · yesterday Cached

Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.

0 favorites 0 likes
#ai-research

What Changed from GPT-3.5 to GPT-4? Same Prompts. GPT-3.5: 0/30 Empty Nulls. GPT-4: 30/30. Run It Yourself.

Reddit r/ArtificialInteligence ↗ · yesterday

Research finds that GPT-4 can produce empty responses to null prompts while GPT-3.5 cannot, with cross-vendor studies confirming similar behavior in other models and an open-source tool introduced for controlling EOS token behavior.

0 favorites 0 likes
#ai-research

@JeffDean: Did you know you can learn interesting things from computations in real rat neurons? I love what @AlexKsendzovsky and t…

X AI KOLs Timeline ↗ · 2d ago Cached

Jeff Dean highlights research on computing with rat neurons and a partnership between BioComputingCo and Amazon to bring this technology to customers.

0 favorites 0 likes
#ai-research

Contrastive Language Models

Hacker News Top ↗ · 2d ago

The article likely presents research on contrastive language models, exploring the use of contrastive learning techniques in language model development.

0 favorites 0 likes
#ai-research

LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels

arXiv cs.LG ↗ · 2d ago Cached

The paper proposes LWCal, a CPU-only post-hoc calibration method for tabular classifiers that handles noisy calibration labels without requiring clean data, showing improved calibration error and scoring metrics.

0 favorites 0 likes
#ai-research

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

arXiv cs.AI ↗ · 2d ago Cached

The paper investigates how editorial framing in prompts affects LLMs' factual and tonal responses in data analysis, finding that factual errors occur in specific scenarios, while tonal shifts are more common.

0 favorites 0 likes
#ai-research

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

arXiv cs.AI ↗ · 2d ago Cached

This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.

0 favorites 0 likes
#ai-research

Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond

arXiv cs.CL ↗ · 2d ago Cached

This survey paper reviews developments in brain-to-language decoding, translating neural activity into linguistic outputs for communication restoration and scientific study, covering tasks, methods, evaluation, and future directions.

0 favorites 0 likes
#ai-research

Emergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution

arXiv cs.AI ↗ · 2d ago Cached

The paper introduces Emergi-PersonaOS, a psychology-grounded operating system for persona agents that enables situational adaptation and controllable evolution through a three-layer persona representation and mechanisms for belief updates and experience development.

0 favorites 0 likes
#ai-research

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

arXiv cs.CL ↗ · 2d ago Cached

PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.

0 favorites 0 likes
#ai-research

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

arXiv cs.AI ↗ · 2d ago Cached

PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.

0 favorites 0 likes
#ai-research

Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs

arXiv cs.CL ↗ · 2d ago Cached

The paper investigates fine-tuning strategies for customer support LLMs, comparing multi-task training, sequential updates, and model merging across multiple model families. It concludes that multi-task full fine-tuning is the most robust default, while specialist models degrade off-task and require reliable routing.

0 favorites 0 likes
#ai-research

Realize What Matters: Principled Context Representation for Large-Scale Reasoning

arXiv cs.CL ↗ · 2d ago Cached

The paper proposes principled methods for context representation in large-scale AI reasoning, introducing R3Con which outperforms baselines and enables smaller models to achieve performance comparable to larger ones at lower cost.

0 favorites 0 likes
#ai-research

Math Reasoning in LLMs is Organized by Approach, Not Topic

arXiv cs.AI ↗ · 2d ago Cached

This paper investigates whether large language models internally organize mathematical reasoning by topic or approach, finding evidence that approach is the key organizing principle, challenging traditional benchmarking methods.

0 favorites 0 likes
#ai-research

What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs

arXiv cs.CL ↗ · 2d ago Cached

The paper quantifies how improvements in fact-verification scores are partitioned between answer accuracy and evidence quality, using trained DeBERTa checkpoints and LLMs across multiple benchmarks.

0 favorites 0 likes
#ai-research

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Hugging Face Daily Papers ↗ · 2d ago Cached

The paper proposes AV-GRPO, a modality-anchored reinforcement learning framework for joint audio-video generation that improves generation quality, semantic alignment, and cross-modal synchronization over existing methods.

0 favorites 0 likes
#ai-research

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Hugging Face Daily Papers ↗ · 2d ago Cached

The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback