标签
This paper presents MMLongBench-Doc-V2, a corrected and semantics-aware revision of the MMLongBench-Doc long-document QA benchmark, fixing annotation errors and replacing string matching with an LLM judge, along with a decision procedure for empty-set keys.
MARDoc是一种用于多模态长文档问答的记忆感知精炼代理框架,在MMLongBench-Doc和DocBench基准上使用Qwen3-VL模型进行评估,相比基于MLLM、RAG和代理的基线表现出持续改进。