arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-29 至 2026-04-29 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 4 篇

2604.25563 2026-04-29 cs.RO 78%

Improving Sensing Coverage and Compliance of 3D-Printed Artificial Skins Through Multi-Modal Sensing and Soft Materials

通过多模感知和柔性材料提升3D打印人工皮肤的感知覆盖与合规性

Carson Kohlbrenner, Caleb Escobedo, Sayak Ray, Alexander Dickhans, Anna Soukhovei, Nickolaus Jackoski, Lyle Antieau, Alessandro Roncone

机构 * University of Colorado Boulder(科罗拉多大学波尔得分校)

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出结合时间飞行和自电容传感的柔性皮肤,实现多模感知、压力传感与电接口优化,提升3D打印人工皮肤的实用性和覆盖能力。

Comments This work was accepted at the "Towards Large-Area Tactile Sensing Skins: From Scalable Materials to Embodied Robotic Perception" workshop at the International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24793 2026-04-29 eess.IV cs.CV 74%

CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images

CRC-SAM:基于SAM的多模态结直肠癌在CT、结肠镜和组织学图像中的分割与量化

Daniel Lao

机构 * Independent researcher(独立研究者)

专题命中 多模态Agent :multi-modal(title);分类 cs.CV

AI总结 本文提出CRC-SAM框架,实现跨结肠镜、CT和病理图像的结直肠癌分割,通过LoRA层实现高效领域迁移,实验显示在多个数据集上优于现有方法。

Comments 4 pages, 3 figures, ISBI 2026 oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16918 2026-04-29 cs.CV 70%

AdaTooler-V: Adaptive Tool-Use for Images and Videos

AdaTooler-V:面向图像和视频的自适应工具使用

Chaoyang Wang, Kaituo Feng, Dongyang Chen, Zhongyu Wang, Zhixun Li, Sicheng Gao, Meng Meng, Xu Zhou, Manyuan Zhang, Yuzhang Shang, Xiangyu Yue

机构 * MMLab, CUHK(香港中文大学MML实验室) THU(清华大学) SJTU(上海交通大学) DB Group, CUHK(香港中文大学DB小组) UCF(佛罗里达大学) Sangfor(Sangfor公司) JMU(约翰·霍普金斯大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出AdaTooler-V,通过自适应工具使用提升多模态大语言模型的视觉推理能力,通过强化学习算法和定制数据集优化工具调用策略,实验证明其在多种视觉任务中表现优异。

Comments ACL 2026 Findings, Project page: https://github.com/CYWang735/AdaTooler-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25376 2026-04-29 cs.CV cs.AI 62%

CoRE: Concept-Reasoning Expansion for Continual Brain Lesion Segmentation

CoRE:基于概念推理的连续脑病变分割

Qianqian Chen, Anglin Liu, Jingyang Zhang, Yudong Zhang

机构 * Southeast University, Nanjing, China(东南大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CoRE框架,通过整合视觉特征与结构化概念,解决连续学习中容量限制和冗余参数增长问题,实现基于临床先验的模型进化。

详情

展开后加载摘要…

URL PDF HTML 收藏