arXiv

Articles from arXiv

Cards List

Epistemic-Probabilistic Model for Guarded Multi-Agent LLM Coordination

arXiv cs.AI ↗ · yesterday Cached

Introduces the Epistemic Probabilistic Language Agents (EPLA) model, a neuro-symbolic architecture for enhancing multi-agent LLM coordination by integrating epistemic logic, probabilistic belief engines, and symbolic guards to ensure protocol-governed actions under uncertainty.

0 favorites 0 likes

Learnable Time-Frequency Masks for Explaining Time-Series Classifiers

arXiv cs.LG ↗ · yesterday Cached

The paper proposes XACT, a framework for learning sparse attribution masks over time-frequency transforms to explain time-series classifiers, demonstrating improved precision and interpretability over baselines.

0 favorites 0 likes

Beyond Simple Input-Output Assessment Tasks: Leveraging Automated Programming Assessment for Non-Trivial Courses

arXiv cs.AI ↗ · yesterday Cached

The paper proposes framing machine learning problems as input-output tasks to leverage automated programming assessment tools for AI education, enhancing the integration of theory and practice without innovating AI itself.

0 favorites 0 likes

BridgeMem: Causal Dyadic Transition Residuals for Temporal Knowledge Graph Forecasting

arXiv cs.LG ↗ · yesterday Cached

BridgeMem proposes a new method for temporal knowledge graph forecasting by modeling dyadic transition residuals, demonstrating consistent improvements over baselines on standard benchmarks.

0 favorites 0 likes

The Last Human Gate: Forward Deployed Engineering for Governance Automation

arXiv cs.AI ↗ · yesterday Cached

This paper presents a task-substitution framework for automating enterprise governance reviews using AI agents and software, with benchmarks like DGF-Bench showing that models such as Gemini 3.8 Flash can achieve high success rates in replacing human execution for specified review tasks.

0 favorites 0 likes

FB-GDM: Fully-Bayesian Guided Diffusion Models for High-Dimensional Linear Inverse Problems via Unsupervised Variational Inference

arXiv cs.LG ↗ · yesterday Cached

FB-GDM introduces a fully-Bayesian guided diffusion method for high-dimensional linear inverse problems, eliminating per-task hyperparameter tuning and demonstrating robust performance improvements over existing techniques.

0 favorites 0 likes

SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support

arXiv cs.AI ↗ · yesterday Cached

This paper presents SkinAgent AI, a multimodal agentic framework for non-diagnostic skincare support that combines visual analysis with safe, auditable orchestration using large language models.

0 favorites 0 likes

Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments

arXiv cs.LG ↗ · yesterday Cached

This paper presents an Edge AI system for binary sleep-wake classification on resource-constrained devices, combining inertial sensing and visual pose classification to achieve high accuracy in dynamic environments.

0 favorites 0 likes

When No One Owns the Judgment: Accountability Under Contribution Dissolution in Human-AI Collaboration

arXiv cs.AI ↗ · yesterday Cached

The paper addresses accountability challenges in human-AI collaboration when contributions become indistinguishable, complicating the assignment of responsibility for judgments.

0 favorites 0 likes

A Particle-Swarm-Assisted Gradient Meta-Learning Algorithm for Joint Transmit Precoding and STAR-RIS Coefficient Optimization

arXiv cs.LG ↗ · yesterday Cached

This paper proposes a particle-swarm-assisted gradient meta-learning algorithm for joint optimization of transmit precoding and STAR-RIS coefficients in multi-user wireless systems, achieving improved weighted sum rates over conventional methods.

0 favorites 0 likes

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

arXiv cs.AI ↗ · yesterday Cached

PixelJev is introduced as a native-image decision interface using small open multimodal models to map images, instructions, and candidate sets to structured choices. The study demonstrates adaptation improves accuracy on benchmarks like Pets and highlights challenges in calibration and generalization.

0 favorites 0 likes

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

arXiv cs.LG ↗ · yesterday Cached

This paper introduces Selective Supervision for Direct-OPD (S2D-OPD), a method that improves knowledge distillation by masking low-divergence states, enhancing accuracy on math reasoning benchmarks without extra computation.

0 favorites 0 likes

ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing

arXiv cs.AI ↗ · yesterday Cached

ALOE introduces a semantically addressed low-rank operator for knowledge editing in language models, improving edit scope and precision with high efficacy and locality on standard benchmarks.

0 favorites 0 likes

A Concentration Bound for Two-Timescale Actor-Critic Algorithm

arXiv cs.LG ↗ · yesterday Cached

The paper derives a uniform concentration bound for two-timescale actor-critic algorithms with function approximation in reinforcement learning, analyzing the actor parameter's behavior with high probability.

0 favorites 0 likes

Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier

arXiv cs.AI ↗ · yesterday Cached

This paper presents a data-centric fine-tuning approach for a Polish multi-label content-safety classifier, comparing it with existing systems and analyzing calibration and robustness issues.

0 favorites 0 likes

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

arXiv cs.LG ↗ · yesterday Cached

This paper benchmarks transformer models for Bangla medical Named Entity Recognition, demonstrating that fine-tuned XLM-RoBERTa achieves state-of-the-art performance and outperforms language-specific models, emphasizing the importance of domain diversity in low-resource settings.

0 favorites 0 likes

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

arXiv cs.AI ↗ · yesterday Cached

This paper presents a coroutine-bridge harness that reduces model calls for LLM agents on CAR-bench, enabling policy enforcement as code and achieving top performance with minimal latency.

0 favorites 0 likes

Downside-Controlled Online Forecast Combination under Delayed and Revised Outcomes

arXiv cs.LG ↗ · yesterday Cached

This paper proposes a method for improving frozen forecasters by combining static and online correctors to control downside risk, demonstrating gains up to 11.5% in benchmarks and reduced errors in electricity load forecasting.

0 favorites 0 likes

Towards An LLM-Driven Unified Conversion Framework for BT and FSM in Autonomous Intelligent Systems

arXiv cs.AI ↗ · yesterday Cached

This paper proposes an LLM-driven unified conversion framework that enables automatic and semantically consistent transformation between behavior trees and finite state machines in autonomous intelligent systems, improving scalability and maintainability.

0 favorites 0 likes

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

arXiv cs.LG ↗ · yesterday Cached

This paper investigates where exactly-once behavior should be enforced in LLM agents to prevent duplicate side effects, using a deterministic sandbox named Limbo to test across models, harnesses, and tool contracts under various fault conditions.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback