arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-21 至 2026-04-21 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 文档图表理解 5 篇

2603.23885 2026-04-21 cs.CV 84%

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

通过现实场景合成和文档感知训练实现真实世界文档解析

Gengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen, Liang Wu, Xingyu Wan, Gangyan Zeng, Han Hu, Can Ma, Yu Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Tencent(腾讯) Nankai University(南开大学) University of Chinese Academy of Sciences(中国科学院大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 文档图表理解 :MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出数据训练协同框架,通过现实场景合成和文档感知训练提升文档解析的准确性和鲁棒性,构建了Wild-OmniDocBench基准,并在1B参数MLLM中实现了跨扫描/数字和真实世界场景的优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10154 2026-04-21 cs.CR cs.AI cs.MM 83%

PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models

PRISM-XR: 通过多模态大语言模型赋能隐私感知的扩展现实协作

Jiangong Chen, Mingyu Zhu, Bin Li

机构 * Department of Electrical Engineering, The Pennsylvania State University, University Park, PA 16802, USA(电气工程系,宾夕法尼亚州立大学,University Park,PA 16802,USA)

专题命中 文档图表理解 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出PRISM-XR框架,通过边缘服务器智能预处理过滤敏感数据,实现隐私保护的多用户XR协作,实验表明其在准确性和隐私保护方面表现优异。

Comments Accepted to the 2026 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR)

Journal ref 2026 IEEE Conference on Virtual Reality and 3D User Interfaces (VR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28554 2026-04-21 cs.CV cs.AI cs.IR 81%

Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model

Hydra:在单一视觉-语言模型中统一文档检索与生成

Athos Georgiou

机构 * Independent Researcher(独立研究者)

专题命中 文档图表理解 :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 Hydra通过单一视觉-语言模型实现文档检索与生成的统一,减少内存和系统复杂性,通过LoRA适配器在推理时切换,提升效率并保持生成质量。

Comments 21 pages, 4 figures, 10 tables, 1 algorithm. v3: two-scale release (4B, 0.8B); bitwise generation-equivalence (426/426 LM tensors at 4B); peak VRAM -62.7% at 4B, -59.1% at 0.8B; GritLM joint-training ablation; Qwen2.5-Omni-3B omni extension. Models: huggingface.co/collections/athrael-soju/hydra-dual-head-retrieval-and-generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17325 2026-04-21 cs.CL 50%

Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation

对齐文档与问题:面向问题的文档重写以增强检索增强生成

Jiaang Li, Zhendong Mao, Quan Wang, Yuning Wan, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 文档图表理解 :grounding(abstract)

AI总结 本文提出QREAM框架,通过风格控制重写使检索文档更符合问题需求,提升检索增强生成的准确性与效率。

Comments ACL'26 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13468 2026-04-21 cs.IR cs.CL 50%

From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines

从相关性到权威性:面向网页搜索引擎的权威感知生成检索

Sunkyung Lee, Jihye Back, Donghyeon Jeon, Soonhwan Kwon, Moonkwon Kim, Inho Kang, Jongwuk Lee

机构 * Sungkyunkwan University(成均馆大学) Naver Corporation(Naver公司)

专题命中 文档图表理解 :vision-language model(abstract)

AI总结 本文提出权威感知生成检索框架AuthGR,通过多模态权威评分、三阶段训练和混合集成流程提升检索的权威性和准确性,在实际应用中显著提高用户参与度和可靠性。

Comments ACL 2026 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏