最新

全部文章,按抓取时间从新到旧排列。

Cards List

When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification

arXiv cs.AI · 6小时前 缓存

This paper introduces the Individual Conformal Coupling Monitor (ICCM), a pre-inference method to detect structural ambiguity in wearable stress classification, improving safety by routing uncertain signals for abstention or deferral.

0 人收藏 0 人点赞

Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage

arXiv cs.CL · 6小时前 缓存

This paper introduces a tri-stream fine-tuned LLM framework for automated clinical supervision in mental health, achieving high technique identification accuracy and reducing supervisory triage latency from 72 hours to real-time.

0 人收藏 0 人点赞

A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations

arXiv cs.AI · 6小时前 缓存

This paper evaluates the robustness of AI code agents when codebases are perturbed with semantics-preserving transformations, revealing a jagged frontier where model performance varies unpredictably across different scaffolds and benchmarks.

0 人收藏 0 人点赞

Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text

arXiv cs.CL · 6小时前 缓存

This paper presents the first systematic study of word segmentation for the extinct Tangut language, combining traditional lexicons, unlabeled text, and a pretrained character encoder to achieve high performance despite extreme resource scarcity.

0 人收藏 0 人点赞

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

arXiv cs.AI · 6小时前 缓存

The paper introduces THPT-Ladder, a benchmark that evaluates AI models using Vietnam's 2025 convex marking scheme, revealing how partial credit gaps lead to inaccurate assessments compared to standard accuracy metrics.

0 人收藏 0 人点赞

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

arXiv cs.CL · 6小时前 缓存

This research investigates whether fine-tuning large language models on cultural data improves figurative language understanding and vice versa, finding that while poetry fine-tuning enhances idiom comprehension, cultural fine-tuning can reduce proverb accuracy, highlighting a non-straightforward relationship.

0 人收藏 0 人点赞

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

arXiv cs.AI · 6小时前 缓存

This paper investigates using governance records from machine-verifiable workflows as supervision for training language models to perform structured workflow repair, demonstrating that verifier-selected self-training enhances execution efficiency and validity.

0 人收藏 0 人点赞

面向自主科学代理的以工件为中心的主张感知可观测性

arXiv cs.CL · 6小时前 缓存

本文提出了一种针对自主科学代理的主张感知可观测性框架,强调工件沿袭和审计关系,以提高科学工作流程的透明度和错误检测能力。

0 人收藏 0 人点赞

利用 PyTorch 原生栈在服务器 CPU 上高效进行小型 NLP 模型的 INT8 推理

arXiv cs.CL · 6小时前 缓存

本文将 SmoothQuant 集成到 PyTorch 的原生栈中,以在 Intel Xeon CPU 上高效进行小型 NLP 模型的 INT8 推理,实现了高达 5.8 倍的加速,且精度损失可忽略不计。

0 人收藏 0 人点赞

ComponentBench: 诊断计算机使用代理中的组件级故障

arXiv cs.AI · 6小时前 缓存

ComponentBench 引入了一个基准和诊断管道,用于评估计算机使用代理在现代网页用户界面中的组件级交互,通过关注现实、短暂的交互来诊断跨模型的故障,从而填补当前评估方法的空白。

0 人收藏 0 人点赞

大语言模型在文本之外安全吗?表情符号是否暴露了安全评估的漏洞?

arXiv cs.CL · 6小时前 缓存

本文通过研究表情符号增强的提示,探讨了大语言模型在文本输入之外的安全性,揭示了当前安全评估的漏洞和模型依赖性漏洞。

0 人收藏 0 人点赞

SESSE: 草图、扩展、排序、总结、评估 -- 通过结构化分解的LLM-as-Judge评估

arXiv cs.AI · 6小时前 缓存

SESSE是一个无需训练的框架,将整体性的LLM-as-Judge评估分解为结构化子问题,从而提高可解释性并诊断标签模糊性,同时与微调模型达到竞争性的性能。

0 人收藏 0 人点赞

LLMs 何时真正有效?评估 LLMs 作为数据质量标注器

arXiv cs.CL · 6小时前 缓存

本研究评估了 LLMs 在电子商务任务中作为数据质量标注器的表现,发现当需要背景知识时,它们优于基线方法,但对于具有强烈词汇信号的任务优势有限,同时展示了跨运行的高度一致性。

0 人收藏 0 人点赞

大型推荐解释中LLM-as-a-Judge的生命周期

arXiv cs.AI · 6小时前 缓存

本文提出了一个在Netflix使用大语言模型作为评判者来评估推荐解释的生命周期框架,涵盖从开发到部署和监控的各个阶段,并通过A/B测试取得了积极结果,显示用户参与度得到提升。

0 人收藏 0 人点赞

评估开源模型在高风险公共部门应用中的结构化信息提取

arXiv cs.AI · 6小时前 缓存

本文对开源OCR、LLM和VLM系统在高风险公共部门应用中的结构化信息提取进行了基准测试,发现VLM通常优于OCR+LLM流水线,但大多数配置在零样本设置中表现不佳,强调了输入质量的关键作用。

0 人收藏 0 人点赞

道义缺口:大型语言模型与义务的模态语言

arXiv cs.CL · 6小时前 缓存

研究考察了大型语言模型对道义情态词的使用情况,发现与人类相比,模型较少使用'must'和'should'等术语,这反映了其训练基于正式书面语料库。

0 人收藏 0 人点赞

葡萄牙语语言模型:一项系统映射研究

arXiv cs.CL · 6小时前 缓存

本调查系统映射了46个葡萄牙语语言模型,分析其发展、特征、演进,并识别该领域的研究空白和未来方向。

0 人收藏 0 人点赞

可通过设计实现缓存?针对边缘内存带宽墙的Mixture-of-Experts路由器局部性训练:一项预注册的负面结果与系统测量研究

arXiv cs.AI · 6小时前 缓存

本文提出了一项关于训练Mixture-of-Experts路由器以实现缓存局部性对抗内存带宽墙的预注册负面结果,表明尽管采用了训练机制,未命中率降低与语言建模质量之间存在权衡。

0 人收藏 0 人点赞

对齐就是一切:面向通用音频-语言模型的无指令训练

arXiv cs.CL · 6小时前 缓存

本文介绍了一种无指令的纯对齐方法,用于构建大型音频-语言模型,该方法通过冻结LLM和音频编码器,仅在自生成数据上训练一个轻量级投影器,实现了与传统多阶段流程相比更少数据下的竞争性性能。

0 人收藏 0 人点赞

Redakto - LLMs 的隐身标签

arXiv cs.AI · 6小时前 缓存

Redakto 是一个开源工具,旨在在将文本输入大型语言模型之前对其进行匿名化,通过 PII 编辑和假名化确保隐私,并进行了关于隐私和实用性的实证评估。

0 人收藏 0 人点赞
← 上一页
下一页 →
← 返回首页

提交意见反馈