最新

全部文章,按抓取时间从新到旧排列。

Cards List

CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence

arXiv cs.AI · 6小时前 缓存

CTIFoundry introduces an agent-native corpus scaffold for cyber threat intelligence that improves LLM agent performance through structured ontology graphs and procedural skills, achieving higher accuracy and efficiency in investigations.

0 人收藏 0 人点赞

Aslema at NADI 2026: Augmentation through Fewshot for SLU

arXiv cs.CL · 6小时前 缓存

This paper presents Aslema, a system for the NADI 2026 shared task on spoken language understanding, using fine-tuned audio LLMs and synthetic data augmentation to improve intent recognition and slot filling for Tunisian Derja.

0 人收藏 0 人点赞

Learning What to Fail On: Failure-Mode Contextual Bandits for Adversarial Data Curation

arXiv cs.CL · 6小时前 缓存

The paper introduces a failure-aware adversarial retrieval-augmented framework using contextual bandits to improve robustness in natural language understanding, with significant improvements on benchmarks like SNLI, ANLI, and MultiNLI.

0 人收藏 0 人点赞

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

arXiv cs.AI · 6小时前 缓存

This study introduces BudgetDoc, the first multimodal benchmark for evaluating model-budget-performance trade-offs in document tasks, and develops DRB, a lightweight estimator that predicts reasoning performance to optimize compute allocation and reduce costs in LLMs.

0 人收藏 0 人点赞

X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance

arXiv cs.CL · 6小时前 缓存

X2Streaming-TTS presents a causal token-level text-to-speech framework for true streaming synthesis, using causal commitment and speech-state inheritance to handle uncertain text prefixes and maintain acoustic continuity in low-latency spoken dialogue systems.

0 人收藏 0 人点赞

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

arXiv cs.CL · 6小时前 缓存

The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.

0 人收藏 0 人点赞

Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement

arXiv cs.AI · 6小时前 缓存

The paper presents a scalable framework that bridges search and CRM workflows using AI-powered Product Research Agents for proactive customer re-engagement in e-commerce, evaluated in a production deployment with improved CTR and sales.

0 人收藏 0 人点赞

From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

arXiv cs.CL · 6小时前 缓存

VAKE is a two-stage reinforcement-learning framework that externalizes latent parametric knowledge in LLMs through explicit priming and implicit reasoning, achieving superior performance across multiple benchmarks.

0 人收藏 0 人点赞

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

arXiv cs.AI · 6小时前 缓存

FinRCA-Bench is a benchmark designed to evaluate evidence retrieval and reasoning capabilities in financial AI systems, providing a standardized approach for assessment and improvement.

0 人收藏 0 人点赞

Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs

arXiv cs.CL · 6小时前 缓存

The paper finds that bitsandbytes INT4 quantization significantly amplifies proactive interference in LLMs, degrading accuracy in contexts with repeated overwrites and highlighting deployment risks for semantically dense applications.

0 人收藏 0 人点赞

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

arXiv cs.AI · 6小时前 缓存

This paper proposes separating generation from selection in explainable-recommendation systems to reduce serving costs, using a frozen candidate pool of explanations and a small CPU-resident selector. It benchmarks offline-pool selectors and finds that pairwise learning-to-rank outperforms single-action RL formulations like PPO, GRPO, and DPO in terms of F1 scores.

0 人收藏 0 人点赞

Beyond LLM-Based Reasoning: Lightweight GNNs for Agent Failure Attribution

arXiv cs.CL · 6小时前 缓存

This paper introduces AFANet, a lightweight graph-based framework for agent failure attribution in multi-agent systems, which matches or outperforms LLM-based methods with significantly lower computational cost.

0 人收藏 0 人点赞

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

arXiv cs.AI · 6小时前 缓存

The paper introduces HN-CLIP, a method that uses the text encoder's text-text geometry to create adaptive similarity margins for dense-caption retrieval, addressing saturation issues in contrastive learning and improving performance over existing methods.

0 人收藏 0 人点赞

Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

arXiv cs.CL · 6小时前 缓存

This research investigates how multilingual large language models internally handle subject-verb agreement across languages, finding that models reuse shared computational structure for languages with overt inflection, indicating cross-lingual overlap in morphosyntactic processing.

0 人收藏 0 人点赞

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

arXiv cs.AI · 6小时前 缓存

UMER introduces a unified framework for multimodal retrieval that combines embedding and ranking via pair-aware discriminative reasoning, achieving state-of-the-art performance on the MMEB-V2 benchmark.

0 人收藏 0 人点赞

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

arXiv cs.CL · 6小时前 缓存

DART-SD proposes a topology-aware retrieval and tuning framework for self-distillation of LLM-based tool-calling agents, improving policy diversity by correcting only critical topological breakpoints while preserving valid reasoning.

0 人收藏 0 人点赞

MissDiag: Diagnostic Evaluation of Incomplete-Knowledge Robustness in KGQA and KG-RAG

arXiv cs.CL · 6小时前 缓存

MissDiag introduces a diagnostic evaluation framework that decomposes robustness in KGQA and KG-RAG systems under incomplete knowledge into typed evidence interventions for more interpretable comparisons.

0 人收藏 0 人点赞

Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions

arXiv cs.AI · 6小时前 缓存

The paper introduces SDDL, a neuro-symbolic framework that improves combinatorial optimization accuracy in resource-constrained language models by translating natural-language problems into formal representations, resulting in higher feasibility rates compared to direct-generation and solver-code baselines.

0 人收藏 0 人点赞

WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing

arXiv cs.CL · 6小时前 缓存

WhiteMatter introduces all-to-all cross-layer connections in Transformers via KV mixing, reducing memory footprint and improving performance over standard architectures in pretraining experiments.

0 人收藏 0 人点赞

OmniAlign: A Unified Multilingual Aligner for Word and Sentence Alignment

arXiv cs.CL · 6小时前 缓存

OmniAlign is a unified multilingual aligner that supports both word-level and sentence-level alignment using a single lightweight model, achieving competitive performance on benchmarks and generalizing to unseen language pairs.

0 人收藏 0 人点赞
← 上一页
下一页 →
← 返回首页

提交意见反馈