arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-09 至 2026-03-09 共收录 79 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 5 篇

2603.02795 2026-03-09 cs.CV 79%

VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning

VSearcher:通过强化学习实现长周期多模态搜索代理

Ruiyang Zhang, Qianguo Sun, Chao Song, Yiyan Qi, Zhedong Zheng

机构 * University of Macau(澳门大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 VSearcher通过强化学习实现多模态搜索代理,提升长周期多轮工具使用能力,在多模态网络搜索任务中表现优异。

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20230 2026-03-09 cs.AI cs.CV cs.MA 73%

A Multi-Agent System Enables Versatile Information Extraction from the Chemical Literature

多智能体系统实现从化学文献中灵活的信息提取

Yufan Chen, Ching Ting Leung, Bowen Yu, Jianwei Sun, Yong Huang, Linyan Li, Hao Chen, Hanyu Gao

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种基于多模态大语言模型的多智能体系统,实现了从化学文献中高效提取化学信息,F1分数达76.27%,显著提升信息提取效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05535 2026-03-09 eess.IV cs.CV cs.LG 70%

Clinical-Injection Transformer with Domain-Adapted MAE for Lupus Nephritis Prognosis Prediction

具有领域适应MAE的临床注射变换器用于系统性红斑狼疮肾炎的预后预测

Yuewen Huang, Zhitao Ye, Guangnan Feng, Fudan Zheng, Xia Gao, Yutong Lu

机构 * Sun Yat-sen University Department of Nephrology, Guangzhou Women Children's Medical Center, Guangzhou Medical University, Guangzhou, China

专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究提出了一种多模态计算病理学框架,利用临床数据和常规染色活检,通过临床注射变换器和领域适应MAE实现儿童系统性红斑狼疮肾炎预后预测的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06061 2026-03-09 cs.CV cs.RO 57%

Transforming Omnidirectional RGB-LiDAR data into 3D Gaussian Splatting

将全方位RGB-LiDAR数据转换为3D高斯散点

Semin Bae, Hansol Lim, Jongseong Brad Choi

机构 * Department of Computer Science, State University of New York(计算机科学系,纽约州立大学) Department of Mechanical Engineering, State University of New York(机械工程系,纽约州立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种将全方位RGB-LiDAR数据转换为3D高斯散点的重用管道,解决数据处理中的非线性失真和计算开销问题,提升复杂场景的渲染保真度。

Comments This work has been submitted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05270 2026-03-09 cs.RO cs.AI cs.HC cs.MA cs.SY eess.SY 57%

XR-DT: Extended Reality-Enhanced Digital Twin for Safe Motion Planning via Human-Aware Model Predictive Path Integral Control

XR-DT:增强现实增强型数字孪生用于通过人感知模型预测路径积分控制的安全运动规划

Tianyi Wang, Jiseop Byeon, Ahmad Yehia, Yiming Xu, Jihyung Park, Tianyi Zeng, Sikai Chen, Ziran Wang, Junfeng Jiao, Christian Claudel

机构 * Department of Civil, Architectural, and Environmental Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校土木、建筑与环境工程系) School of Architecture, The University of Texas at Austin(德克萨斯大学奥斯汀分校建筑学院) School of Civil and Construction Engineering, Purdue University(普渡大学土木与建设工程学院) Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 XR-DT通过结合增强现实与数字孪生技术,提出HA-MPPI控制模型,实现基于人类行为预测的安全高效人机交互。

Comments 8 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 11 篇

2603.05566 2026-03-09 cs.LG cs.CL 85%

Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment

对齐真实语义:基于约束解耦和分布采样的跨模态对齐

Xiang Ma, Lexin Fang, Litian Xu, Caiming Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CL

AI总结 本文提出CDDS算法,通过约束解耦和分布采样方法,解决跨模态对齐中的语义分离和模态差距问题,实验表明其优于现有方法。

Comments AAAI 2026 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06403 2026-03-09 cs.LG 85%

Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling

适配器增强的带状机用于在线多约束多模态推断调度

Xianzhi Zhang, Yue Xu, Yinlin Zhu, Di Wu, Yipeng Zhou, Miao Hu, Guocong Quan

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, 510006, China.(中山大学计算机科学与工程学院) School of Computing, Macquarie University, NSW 2109, Australia(麦考瑞大学计算机学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);MLLM(abstract)

AI总结 M-CMAB通过多适配器增强框架,实现多模态任务调度中的多约束优化,提升推断效率和预算利用效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06986 2026-03-09 cs.CV cs.CL 81%

Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs

重新思考多模态大语言模型中视觉编码器混合范式以提升视觉理解

Mozhgan Nasr Azadani, James Riddell, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 LEO通过轻量级融合设计提升多模态大语言模型的视觉理解能力,并在自动驾驶领域展现良好泛化性能。

Comments Accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10328 2026-03-09 cs.CV 79%

Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization

Fuse4Seg: 多模态医学分割的图像融合 via 两级优化

Yuchen Guo, Junli Gong, Hongmin Cai, Yiu-ming Cheung, Weifeng Su

机构 * Northwestern University(西北大学) Northeastern University(东北大学) South China University of Technology(华南理工大学) Hong Kong Baptist University(香港 Baptist大学) Beijing Normal - Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 Fuse4Seg通过两级优化实现多模态医学图像融合,解决视觉与语义间的差距问题,提升分割任务的准确性和临床可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05623 2026-03-09 cs.CV cs.AI 76%

Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection

融合后鸟瞰图特征稳定化用于鲁棒多模态3D检测

Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin

机构 * Department of Electrical Engineering, University of South Florida(佛罗里达州立大学电气工程系) Department of Mechanical Engineering, University of South Florida(佛罗里达州立大学机械工程系)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出了一种轻量级模块PFS,通过稳定特征统计、抑制传感器退化区域和残差校正,提升多模态3D检测在域转移和传感器故障下的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06186 2026-03-09 cs.CV 74%

SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region Detection

SpaCRD:多模态深度融合组织学与空间转录组学用于癌症区域检测

Shuailin Xue, Jun Wan, Lihua Zhang, Wenwen Min

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 SpaCRD通过多模态深度融合组织学与空间转录组学数据,实现跨样本、平台和批次的癌症区域检测,优于现有八种方法。

Comments Accepted by AAAI-2026-Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05528 2026-03-09 cs.MM cs.AI cs.CL cs.CV cs.SD eess.AS 73%

Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder

Omni-C:将异构模态压缩到单一密集编码器

Kin Wai Lau, Yasar Abbas Ur Rehman, Lai-Man Po, Pedro Porto Buarque de Gusmão

机构 * City University of Hong Kong(香港城市大学) TCL AI Lab(TCL人工智能实验室) University of Surrey, United Kingdom(英国萨里大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Omni-C通过单一密集Transformer编码器学习跨异构模态的共享表示,有效缓解跨模态冲突,提升多模态学习的效率与可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06250 2026-03-09 cs.CV 70%

Hierarchical Collaborative Fusion for 3D Instance-aware Referring Expression Segmentation

层次化协作融合用于3D实例感知指代表达分割

Keshen Zhou, Runnan Chen, Mingming Gong, Tongliang Liu

机构 * The University of Sydney(悉尼大学) The University of Melbourne(墨尔本大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 HCF-RES通过层次化视觉语义分解和渐进多级融合,实现了3D实例感知指代表达分割的高精度与细粒度定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17938 2026-03-09 cs.CL cs.LG 70%

SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization

SPINE:基于熵带正则化的令牌选择性测试时间强化学习

Jianghao Wu, Yasmeen George, Jin Ye, Yicheng Wu, Daniel F. Schmidt, Jianfei Cai

机构 * Monash University(墨尔本大学) Imperial College London(伦敦帝国理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 SPINE通过令牌选择性和熵带正则化提升测试时间推理稳定性与效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06531 2026-03-09 cs.CV cs.RO 57%

Spatial Calibration of Diffuse LiDARs

扩散式时间飞行激光雷达的空间校准

Nikhil Behari, Ramesh Raskar

机构 * MIT(麻省理工学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种扩散式时间飞行激光雷达的空间校准方法,通过恢复每像素响应图实现激光雷达与 RGB 的跨模态对齐与融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06081 2026-03-09 cs.CV 57%

Lyapunov Probes for Hallucination Detection in Large Foundation Models

Lyapunov探针用于大型基础模型中的幻觉检测

Bozhi Luan, Gen Li, Yalan Qin, Jifeng Guo, Yun Zhou, Faguo Wu, Hongwei Zheng, Wenjun Wu, Zhaoxin Fan

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University(北京未来区块链与隐私计算先进创新中心,人工智能学院,北航) School of Electronic and Information Engineering, State Key Laboratory of CNS/ATM, Beihang University(电子与信息工程学院, CNS/ATM 国家重点实验室,北航) National Key Laboratory of Information Systems Engineering, National University of Defense Technology(信息系统工程国家重点实验室,国防科技大学) Beijing Academy of Blockchain and Edge Computing(北京区块链与边缘计算研究院) Shanghai University(上海大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 通过Lyapunov探针检测大型基础模型中的幻觉,利用动力系统稳定性理论分析知识过渡区域的边界特征。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 3 篇

2511.17355 2026-03-09 cs.CV 79%

UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification

UAM: 多模态框架中用于肿瘤细胞分类的统一注意力-马amba骨干

Taixi Chen, Jingyun Chen, Nancy Guo

机构 * State University of New York at Binghamton(纽约州立大学布林莫尔分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 UAM提出了一种统一的注意力-马amba架构,用于多模态框架中的肿瘤细胞分类和图像分割,通过灵活结合两种模块提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05278 2026-03-09 cs.LG cs.CL 79%

Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs

解码偏微分方程:解码器-only模型在偏微分方程上的跨模态适应

Paloma García-de-Herreros, Philipp Slusallek, Dietrich Klakow, Vagrant Gautam

机构 * Saarland University(萨尔兰大学) DFKI Heidelberg Institute for Theoretical Studies(海德堡理论研究所)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CL

AI总结 本文研究了解码器-only模型在偏微分方程时间依赖模拟任务中的跨模态适应,提出并行翻转和序列加倍两种方法,提升模型性能,缩小与编码器-only模型的差距。

Comments ICLR 2026 Workshop on AI and Partial Differential Equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06525 2026-03-09 cs.RO 78%

Underactuated multimodal jumping robot for extraterrestrial exploration

欠驱动多模态跳跃机器人用于外星探索

Neil R. Wagner, Justin K. Yim

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 该研究提出了一种欠驱动单足机器人,通过两个控制器实现滚动、跳跃和着陆,适用于低重力环境下的多模态探索。

Comments 8 pages, 14 figures, Accepted for ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏