Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Beijing University of Chemical Technology(北京化工大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Beijing Institute of Technology (Zhuhai)(北京理工大学(珠海))
;
Tencent Hy(腾讯(深圳))
;
Peng Cheng Laboratory(鹏城实验室)
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
基于触觉的多模态融合在具身智能中的应用:视觉、语言和接触驱动范式的综述
Zhixiang Cao, Di Tian, Runwei Guan, Yanzhou Mu, Xiaolou Sun, Shaofeng Liang, Daizong Liu, Tao Huang, Yutao Yue, Henghui Ding, Bin Fang, Alex Zhou, Qing-Long Han, Hui Xiong
机构
*
School of Electronic Science and Engineering, Xi’an Jiaotong University, China(西安交通大学电子科学与技术学院)
;
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)人工智能研究所)
;
State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室)
;
Purple Mountain Laboratory, China(紫金山实验室)
;
Institute for Math & AI, Wuhan University, China(武汉大学数学与人工智能学院)
;
Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University, Australia(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院)
;
School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院)
;
Institute of Big Data, Fudan University, China(复旦大学大数据研究院)
;
Linkerbot (Beijing) Technology Co., Ltd, China(北京链动科技有限公司)
;
School of Engineering, Swinburne University of Technology, Melbourne(斯威本技术大学工程学院)
Modality-Native Routing in Agent-to-Agent Networks: A Multimodal A2A Protocol Extension
代理到代理网络中的模态本原路由:一种多模态A2A协议扩展
Vasundra Srinivasan
机构
*
AI Architect, Author—Data Engineering for Multimodal AI (O’Reilly)(人工智能架构师,作者—多模态AI的数据工程(O’Reilly))
;
Stanford School of Engineering (April 2026)(斯坦福大学工程学院(2026年4月))
MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems
MONETA:通过地理信息和多智能体系统进行多模态行业分类
Arda Yüksel, Gabriel Thiem, Susanne Walter, Patrick Felka, Gabriela Alves Werb, Ivan Habernal
机构
*
Trustworthy Human Language Technologies(可信人类语言技术实验室)
;
Technical University of Darmstadt, Germany(德国达姆施塔特工业大学)
;
Deutsche Bundesbank(德国联邦银行)
;
Frankfurt University of Applied Sciences, Germany(德国法兰克福应用技术大学)
;
Research Center for Trustworthy Data Science and Security, Ruhr University Bochum, Germany(德国波鸿鲁尔大学可信数据科学与安全研究中心)
MultiPress: A Multi-Agent Framework for Interpretable Multimodal News Classification
MultiPress:一种用于可解释多模态新闻分类的多智能体框架
Tailong Luo, Hao Li, Rong Fu, Xinyue Jiang, Huaxuan Ding, Yiduo Zhang, Zilin Zhao, Simon Fong, Guangyin Jin, Jianyuan Ni
机构
*
New York Institute of Technology(纽约理工学院)
;
University of Arizona(亚利桑那大学)
;
University of Macau(澳门大学)
;
Peking University(北京大学)
;
Juniata College(朱尼亚塔学院)
机构
*
Fudan University(复旦大学)
;
IFLYTEK CO.LTD(若lytek有限公司)
;
Zhejiang University(浙江大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
National University of Defense Technology(国防科技大学)
;
Hainan University(海南大学)
Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng
机构
*
Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院)
;
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)
CommentsAccepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version
Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection
事实还是假象?评估深度伪造检测器在多模态虚假信息检测中的作用
A S M Sharifuzzaman Sagar, Mohammed Bennamoun, Farid Boussaid, Naeha Sharif, Lian Xu, Shaaban Sahmoud, Ali Kishk
机构
*
The University of Western Australia(西澳大学)
;
Fatih Sultan Mehmet Vakif University(法提赫·苏丹·梅赫梅特·瓦基夫大学)
;
Aljazeera Media Network Investigative Department(半岛电视台调查部门)
CommentsThis manuscript (arXiv:2503.03215) is being withdrawn at the supervisor's request. The content is preliminary and needs further internal revision and approval before public release. We will resubmit a revised version after completion. Apologies for the inconvenience
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
DriveMLM: 通过行为规划状态对齐多模态大语言模型以实现自动驾驶
Erfei Cui, Wenhai Wang, Zhiqi Li, Jiangwei Xie, Haoming Zou, Hanming Deng, Gen Luo, Lewei Lu, Xizhou Zhu, Jifeng Dai
机构
*
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
;
Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
DrivePI: 基于空间感知的4D MLLM用于统一自动驾驶理解、感知、预测与规划
Zhe Liu, Runhui Huang, Rui Yang, Siming Yan, Zining Wang, Lu Hou, Di Lin, Xiang Bai, Hengshuang Zhao
机构
*
The University of Hong Kong(香港大学)
;
Yinwang Intelligent Technology Co. Ltd.(英维智能科技有限公司)
;
Tianjin University(天津大学)
;
Huazhong University of Science and Technology(华中科技大学)
机构
*
University of Science and Technology of China(科学技术大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
;
Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
University of Washington(华盛顿大学)