Newest

All articles, most recently crawled first.

Cards List

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

arXiv cs.LG · 10h ago Cached

The paper proposes mechanistic tomography as a unified framework for designing measurements to recover internal mechanisms in AI models, improving control-oriented interpretability through interventions and calibration.

0 favorites 0 likes

Improved Confidence Estimates for Black-Box Large Language Models

arXiv cs.LG · 10h ago Cached

This paper presents a method to improve confidence estimates for black-box large language models by building classifiers that predict response correctness, outperforming existing zero-shot methods with minimal computational overhead.

0 favorites 0 likes

Quantum Kernel Estimation for the Discovery of Early Lung Cancer Detection

arXiv cs.LG · 10h ago Cached

This paper evaluates quantum-classical hybrid machine learning for lung cancer detection using cfDNA fragmentomics and methylation data, showing competitive performance of quantum kernel models against classical baselines.

0 favorites 0 likes

Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis

arXiv cs.LG · 10h ago Cached

Holtercare-Bench is a multimodal benchmark introduced to evaluate long-term dynamic ECG analysis using the Holtercare-23K dataset, revealing performance gaps in current MLLMs and providing improvements through fine-tuning for clinical applications.

0 favorites 0 likes

Triangular Fuzzy Rescaling Distance

arXiv cs.LG · 10h ago Cached

This paper proposes the Triangular Fuzzy Rescaling Distance (d_TR), a metric that integrates normalization directly into distance calculations for comparing triangular fuzzy numbers, suitable for heterogeneous data applications.

0 favorites 0 likes

Towards On-Board Implementation of ML-Based Helicopter Weight Estimator

arXiv cs.LG · 10h ago Cached

The paper proposes a supervised machine learning model for estimating helicopter weight during takeoff, using data from Airbus's fleet, and details its implementation for on-board use with regulatory compliance.

0 favorites 0 likes

The Single English County Saying No to Palantir

Wired · 8h ago Cached

Greater Manchester's NHS board refuses to adopt Palantir's federated data platform, citing superior local systems and public trust concerns, amid broader UK and European debates on Palantir's contracts.

0 favorites 0 likes

Best Early Tech Labor Day Sales I’d Shop Myself (2026): AirTags, Dyson, and More

Wired · 4h ago Cached

The article lists early Labor Day tech deals on gadgets such as Apple AirTags, Dyson vacuums, and Sony headphones, featuring discounts on tested and recommended products.

0 favorites 0 likes

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

arXiv cs.CL · 10h ago Cached

The paper introduces ConceptGuard, a benchmark for evaluating context-sensitive unlearning in large language models using dual-use concepts, revealing that current unlearning techniques perform poorly under this practical evaluation framework.

0 favorites 0 likes

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

arXiv cs.CL · 10h ago Cached

This paper introduces G-CARL, a grounded checklist-aligned reinforcement learning framework for patient-oriented medical report interpretation, along with the MMedReport benchmark, demonstrating improved factuality and alignment with patient needs.

0 favorites 0 likes

ContractScrub: A benchmark for final review of legal contracts

arXiv cs.AI · 10h ago Cached

Introduces ContractScrub, a benchmark for evaluating LLMs on legal contract scrubbing tasks, revealing that current frontier models perform poorly on this domain-specific challenge.

0 favorites 0 likes

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

arXiv cs.CL · 10h ago Cached

Task-CoEvolve is a novel approach for efficient LLM agent harness optimization that adaptively selects validation tasks to reduce evaluation costs while maintaining performance, achieving an 80% reduction in evaluations on benchmarks.

0 favorites 0 likes

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

arXiv cs.CL · 10h ago Cached

FormalTCS is a benchmark for evaluating large language models on end-to-end theoretical computer science research, revealing significant limitations, especially in autoformalization.

0 favorites 0 likes

The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents

arXiv cs.AI · 10h ago Cached

The paper argues that software is undergoing a third paradigm shift to Software 3.0, where context and reasoning determine behavior, converging to three core elements: generalized storage, large models, and agents. It formalizes this thesis and analyzes its conditions and boundaries.

0 favorites 0 likes

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

arXiv cs.CL · 10h ago Cached

This paper introduces a benchmark to study how large language models arbitrate conflicting evidence from text and numerical sources, finding that models use heuristic strategies with biases towards recency and external tools.

0 favorites 0 likes

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

arXiv cs.AI · 10h ago Cached

DARS is a reinforcement learning framework for dual-level credit assignment in instruction-based image editing, improving performance by routing updates between planner and renderer modules through structured reasoning and adaptive curriculum.

0 favorites 0 likes

OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

arXiv cs.CL · 10h ago Cached

OenoBench introduces a wine-domain benchmark for evaluating large language models on knowledge-grounded multiple-choice questions, built from verified sources to address limitations in existing benchmarks.

0 favorites 0 likes

SABET-QA: Temporal Knowledge Graph Question Answering

arXiv cs.CL · 10h ago Cached

SABET-QA introduces an iterative framework for temporal knowledge graph question answering that enhances multi-hop reasoning through bidirectional entity-temporal scoring and contextualization, showing consistent improvements over baselines on benchmarks like CronQuestions and TimeQuestions.

0 favorites 0 likes

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

arXiv cs.AI · 10h ago Cached

The paper introduces DECOWAM, a decoupled whole-body world-action model for legged mobile manipulation that improves video and action prediction performance over existing models like FastWAM through dedicated conditional interfaces and a new dataset.

0 favorites 0 likes

Auditing Cross-Lingual Fairness in Language Model Watermarking

arXiv cs.CL · 10h ago Cached

The paper proposes an evaluation framework for cross-lingual fairness in language model watermarking, revealing that disparities are structural to language typology rather than idiosyncratic to specific languages.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback