Aiding Intra-Text Representations with Visual Context for Multimodal Named Entity Recognition
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments IEEE Intelligence Systems. arXiv admin note: substantial text overlap with arXiv:1707.09538
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:1811.07216
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Journal ref ICCV17: second workshop on Closing the Loop Between Vision and Language. Venice, Italy. 2017
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments EMNLP 2018
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments Published as a conference paper at ICLR 2018. 12 pages
专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI
Comments To appear in the 7th IEEE Symposium Series on Computational Intelligence (IEEE SSCI 2016), 8 pages, 6 figures. Minor revisions, in response to reviewers' comments
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Clarified main contributions, minor correction to Equation 8, additional comparisons in Table 2, added more related work
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 10 pages, ISERC 2013, IIE Annual Conference. Proceedings. Institute of Industrial Engineers
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Appears in NIPS 2016. The datasets introduced in this work will be gradually released on the project page
对话状态的多模态信号有多可靠?来自远程二元协作任务的证据
机构 * Colby College(科尔比学院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
AI总结 研究探讨从多模态行为测量对话状态的特征可靠性,提出三维评估框架用于视频会议二元对话特征评估,发现语言特征预测佳但跨任务通用性差,声学可靠性受说话者身份影响,交互特征是唯一可靠信号,强调相关评估对对话系统特征选择的重要性。
Comments Accepted, to appear in Proceedings of ACM International Conference on Multimodal Interaction 2026, 13 pages, 6 figures, 4 tables
GDP.pdf:针对专业PDF文档的基础多模态推理基准测试
机构 * Surge AI
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
AI总结 该研究针对专业PDF文档构建多模态推理基准测试GDP.pdf,由专业人员编写问题-文档对,通过严格筛选保留问题,有详细评分标准和能力分类。评估七个前沿模型,发现多数错误源于特定模式,公开了完整基准测试。
Comments 9 pages. v2: results updated to July 2026 leaderboard (17 models). Accepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 (non-archival), under the former title "PDFParse: A Benchmark for Grounded Multimodal Reasoning over Professional PDF Documents". Dataset: https://huggingface.co/datasets/surgeai/GDP.pdf ; Code: https://github.com/surge-ai/gdp-pdf
MM-tau-p$^2$: 人格自适应提示用于双控制设置中多模态代理的鲁棒性评估
机构 * Sprinklr AI
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI;multimodal(comments)
AI总结 本文提出MM-tau-p$^2$基准,通过12个新指标评估多模态代理在双控制环境中的鲁棒性,结合用户输入解决查询,并展示即使使用前沿LLM,多模态鲁棒性等指标仍需考虑。
Comments A benchmark for evaluating multimodal both voice and text LLM agents in dualcontrol settings. We introduce persona adaptive prompting and 12 new metrics to assess robustness safety efficiency and recovery in customer support scenarios
BanglaMM-Disaster: 一种基于Transformer的多模态深度学习框架,用于孟加拉语多类灾害分类
机构 * Department of Computer Science and Engineering(计算机科学与工程系) ; Chittagong University of Engineering and Technology(奇特格隆工程与技术大学) ; Department of Electronics and Telecommunication Engineering(电子与电信工程系) ; Wilmington University(维明顿大学) ; College of Graduate and Professional Studies(研究生与专业研究学院) ; Trine University(特林大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
AI总结 BanglaMM-Disaster通过结合文本和视觉数据,提出了一种多模态深度学习框架,用于孟加拉语多类灾害分类,提升了灾害响应效率。
Comments Presented at the 2025 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), November 21-22, 2025, University of Rajshahi, Bangladesh. 6 pages, 9 disaster classes, multimodal dataset with 5,037 samples
机构 * Harvard University(哈佛大学) ; MIT Media Lab(麻省理工学院媒体实验室)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
Comments Accepted at the Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2025
机构 * Purdue University(普渡大学) ; Indiana University(印第安纳大学) ; Curtin University(Curtin大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments The extended full version of the accepted paper in 2025 IEEE BHI conference with title: Evaluating Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata. Dataset is available at: https://skynet.ecn.purdue.edu/~coburn6/ACETADA/
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted by ICPR Multi-Modal Visual Pattern Recognition Workshop
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
Comments 29 pages, submited to conference, code available at: https://github.com/rxn4chemistry/multimodal-spectroscopic-dataset
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI
Journal ref S. González, A. K.-C. Yi, W.-T. Hsieh, W.-C. Chen, C.-L. Wang, V. C.-C. Wu, S.-H. Chang, Multi-modal heart failure risk estimation based on short ECG and sampled long-term HRV, Information Fusion 107 (2024) 102337
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
Comments Accepted by 25th ACM International Conference on Multimodal Interaction (ICMI '23), Late-Breaking Results
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments Published in Computer Vision and Pattern Recognition (CVPR) Workshops 2023 - 6th Multimodal Learning and Applications Workshop
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
Comments Accepted to The Third Workshop on Multimodal Artificial Intelligence (MAI-Workshop)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
Comments 8 pages of content, 11 pages total, 2 figures. Published as a workshop paper at ACL 2018, Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML). 2018
MELLA:弥合低资源语言多模态大语言模型的语言能力与文化根基
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; East China Normal University(东华大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Institute of High Performance Computing, A*STAR(高性能计算研究所,A*STAR)
专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 研究针对低资源语言MLLMs文化内涵不足问题,提出MELLA数据集,采用双源策略,结合母语网络图像-替代文本对与生成翻译的图像描述进行监督,经实验表明能减轻文化幻觉,强调数据对齐对低资源语言文化基础多模态理解的重要性。
基于历史的迭代视觉推理与自我校正
机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM
AI总结 H-GIVR框架通过迭代视觉推理与自我校正,显著提升多模态推理准确性并保持低计算成本。
专题命中 多模态评测 :multimodal(title,abstract)
Comments Version accepted by Computers in Biology and Medicine: 21 pages, 2 figures, code available under https://github.com/AI4HealthUOL/MDS-ED, dataset available under https://physionet.org/content/multimodal-emergency-benchmark/
Journal ref J.M. Lopez Alcaraz, H. Bouma, N. Strodthoff, Enhancing clinical decision support with physiological waveforms -- A multimodal benchmark in emergency care, Computers in Biology and Medicine, Vol. 192, Part A, 2025, 110196
专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI、cs.MM
Comments 19 pages, 6 figures, 12 tables
专题命中 多模态评测 :multimodal(title,abstract)
Comments Accepted in 1st International Workshop on Multiscale Multimodal Medical Imaging (MMMI 2019) - MICCAI 2019
Journal ref Multiscale Multimodal Medical Imaging. MMMI 2019. Lecture Notes in Computer Science, vol 11977
CFB-GBM v2.0:用于多模态胶质母细胞瘤分割、放射组学及RANO进展追踪的增强纵向数据集
机构 * Centre François Baclesse(弗朗索瓦·巴克莱斯中心) ; Université de Caen Normandie(卡昂诺曼底大学) ; ENSICAEN(卡昂高等工程师学院) ; GREYC(格雷计算机科学研究中心)
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV
AI总结 本文发布增强版CFB-GBM v2.0纵向数据集,含264名GBM患者数据,完成所有时间点GTV勾画(完成率达97%),提供相关标注、特征及WHO分类信息,可用于多模态GBM相关研究。
Comments 9 pages, 2 figures,
面向证据基础计算病理学的多模态智能体协同助手
机构 * Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) ; Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科) ; Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科) ; Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系) ; Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室) ; Jinfeng Laboratory(锦风实验室) ; Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) ; Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系) ; State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室) ; HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
AI总结 提出PathPocket,一种多模态AI协同助手,通过构建包含11万文档的病理证据语料库和455万实体的超图,实现基于证据的病理诊断,在20万真实案例上超越现有方法。