Articles from arXiv
Introduces the Epistemic Probabilistic Language Agents (EPLA) model, a neuro-symbolic architecture for enhancing multi-agent LLM coordination by integrating epistemic logic, probabilistic belief engines, and symbolic guards to ensure protocol-governed actions under uncertainty.
The paper proposes XACT, a framework for learning sparse attribution masks over time-frequency transforms to explain time-series classifiers, demonstrating improved precision and interpretability over baselines.
The paper proposes framing machine learning problems as input-output tasks to leverage automated programming assessment tools for AI education, enhancing the integration of theory and practice without innovating AI itself.
BridgeMem proposes a new method for temporal knowledge graph forecasting by modeling dyadic transition residuals, demonstrating consistent improvements over baselines on standard benchmarks.
This paper presents a task-substitution framework for automating enterprise governance reviews using AI agents and software, with benchmarks like DGF-Bench showing that models such as Gemini 3.8 Flash can achieve high success rates in replacing human execution for specified review tasks.
FB-GDM introduces a fully-Bayesian guided diffusion method for high-dimensional linear inverse problems, eliminating per-task hyperparameter tuning and demonstrating robust performance improvements over existing techniques.
This paper presents SkinAgent AI, a multimodal agentic framework for non-diagnostic skincare support that combines visual analysis with safe, auditable orchestration using large language models.
This paper presents an Edge AI system for binary sleep-wake classification on resource-constrained devices, combining inertial sensing and visual pose classification to achieve high accuracy in dynamic environments.
The paper addresses accountability challenges in human-AI collaboration when contributions become indistinguishable, complicating the assignment of responsibility for judgments.
This paper proposes a particle-swarm-assisted gradient meta-learning algorithm for joint optimization of transmit precoding and STAR-RIS coefficients in multi-user wireless systems, achieving improved weighted sum rates over conventional methods.
PixelJev is introduced as a native-image decision interface using small open multimodal models to map images, instructions, and candidate sets to structured choices. The study demonstrates adaptation improves accuracy on benchmarks like Pets and highlights challenges in calibration and generalization.
This paper introduces Selective Supervision for Direct-OPD (S2D-OPD), a method that improves knowledge distillation by masking low-divergence states, enhancing accuracy on math reasoning benchmarks without extra computation.
ALOE introduces a semantically addressed low-rank operator for knowledge editing in language models, improving edit scope and precision with high efficacy and locality on standard benchmarks.
The paper derives a uniform concentration bound for two-timescale actor-critic algorithms with function approximation in reinforcement learning, analyzing the actor parameter's behavior with high probability.
This paper presents a data-centric fine-tuning approach for a Polish multi-label content-safety classifier, comparing it with existing systems and analyzing calibration and robustness issues.
This paper benchmarks transformer models for Bangla medical Named Entity Recognition, demonstrating that fine-tuned XLM-RoBERTa achieves state-of-the-art performance and outperforms language-specific models, emphasizing the importance of domain diversity in low-resource settings.
This paper presents a coroutine-bridge harness that reduces model calls for LLM agents on CAR-bench, enabling policy enforcement as code and achieving top performance with minimal latency.
This paper proposes a method for improving frozen forecasters by combining static and online correctors to control downside risk, demonstrating gains up to 11.5% in benchmarks and reduced errors in electricity load forecasting.
This paper proposes an LLM-driven unified conversion framework that enables automatic and semantically consistent transformation between behavior trees and finite state machines in autonomous intelligent systems, improving scalability and maintainability.
This paper investigates where exactly-once behavior should be enforced in LLM agents to prevent duplicate side effects, using a deterministic sandbox named Limbo to test across models, harnesses, and tool contracts under various fault conditions.