arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2766 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2766 篇

2511.05169 2025-11-10 cs.LG 78%

Multimodal Deep Learning for Prediction of Progression-Free Survival in Patients with Neuroendocrine Tumors Undergoing 177Lu-based Peptide Receptor Radionuclide Therapy

Simon Baur, Tristan Ruhwedel, Ekin Böke, Zuzanna Kobus, Gergana Lishkova, Christoph Wetz, Holger Amthauer, Christoph Roderburg, Frank Tacke, Julian M. Rogasch, Wojciech Samek, Henning Jann, Jackie Ma, Johannes Eschrich

机构 * Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute(人工智能部门,弗劳恩霍夫海因里希·赫兹研究所) Department of Nuclear Medicine, Charité—Universitätsmedizin Berlin(核医学部门,柏林查理医院) Department of Hepatology and Gastroenterology, Charité—Universitätsmedizin Berlin(肝病与胃肠病学部门,柏林查理医院) Division of Interventional Radiology, Department of Radiology, Memorial Sloan Kettering Cancer Center(介入放射学部门,纪念斯隆-凯特琳癌症中心) Department of Endocrinology and Metabolism, Charité—Universitätsmedizin Berlin(内分泌与代谢学部门,柏林查理医院) Clinic for Gastroenterology, Hepatology and Infectious Diseases, University Hospital Düsseldorf, Medical Faculty of Heinrich Heine University Düsseldorf(胃肠病、肝病和传染病诊所,杜塞尔多夫大学医院,海因里希·海涅大学医学部) Department of Electrical Engineering and Computer Science, Technische Universität Berlin(电气工程与计算机科学部门,柏林技术大学) Berlin Institute of Health at Charité – Universitätsmedizin Berlin(柏林查理医院健康研究所)

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27443 2025-11-03 cs.LG 78%

MVeLMA: Multimodal Vegetation Loss Modeling Architecture for Predicting Post-fire Vegetation Loss

Meenu Ravi, Shailik Sarkar, Yanshen Sun, Vaishnavi Singh, Chang-Tien Lu

机构 * Georgetown University(乔治城大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments Accepted for 2025 ACM SIGSPATIAL conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00056 2025-10-31 math.OC 78%

Sustainable Multi-Modal Transportation and Routing focusing on Costs and Carbon Emissions Reduction

Saba Javanpour, A. Radman, Sarow Saeedi, Sina Feizi Karimabadi, Daniel A. Larson, Eric C. Jones

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Abstracted submitted in the Proceedings of the IISE Annual Conference & Expo 2025

Journal ref Proceedings of the IISE Annual Conference & Expo 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23032 2025-10-28 cs.CE 78%

P1GPT: a multi-agent LLM workflow module for multi-modal financial information analysis

Chen-Che Lu, Yun-Cheng Chou, Teng-Ruei Chen

专题命中 多模态Agent :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06278 2025-10-28 cs.RO cs.HC 78%

Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation

Tongfei Bian, Mathieu Chollet, Tanaya Guha

机构 * University of Glasgow(格拉斯哥大学) University of Glasgow School of Computer Science(格拉斯哥大学计算机科学学院)

专题命中 多模态Agent :multimodal(title,abstract)

Comments Accepted by ACM Multimedia 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15767 2025-10-20 cs.SE 78%

EASELAN: An Open-Source Framework for Multimodal Biosignal Annotation and Data Management

Rathi Adarshi Rammohan, Moritz Meier, Dennis Küster, Tanja Schultz

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19168 2025-09-24 cs.RO 78%

A Multimodal Stochastic Planning Approach for Navigation and Multi-Robot Coordination

Mark Gonzales, Ethan Oh, Joseph Moore

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 8 Pages, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11782 2025-09-16 cs.LG q-bio.BM 78%

Multimodal Regression for Enzyme Turnover Rates Prediction

Bozhen Hu, Cheng Tan, Siyuan Li, Jiangbin Zheng, Sizhe Qiu, Jun Xia, Stan Z. Li

机构 * AI Division, School of Engineering, Westlake University(西拉丘学院人工智能系,西湖大学) Zhejiang University(浙江大学) Oxford University(牛津大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 9 pages, 5 figures. This paper was withdrawn from the IJCAI 2025 proceedings due to the lack of participation in the conference and presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15917 2025-09-12 cs.SE 78%

Towards Test Generation from Task Description for Mobile Testing with Multi-modal Reasoning

Hieu Huynh, Hai Phung, Hao Pham, Tien N. Nguyen, Vu Nguyen

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Change the method and experimentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01161 2025-09-03 cs.LG 78%

Multi-Modal Machine Learning Framework for Predicting Early Recurrence of Brain Tumors Using MRI and Clinical Biomarkers

Cheng Cheng, Zeping Chen, Rui Xie, Peiyao Zheng, Xavier Wang

机构 * Department of Encephalopathy, Chengdu Pidu District Hospital of Traditional Chinese Medicine(脑病科,成都_pidu区中医医院) Department of Tuina, Chengdu Pidu District Hospital of Traditional Chinese Medicine(推拿科,成都_pidu区中医医院) Wuhan Hospital of Integrated Traditional Chinese and Western Medicine, Affiliated to Hubei University of Chinese Medicine(中西医结合医院,湖北中医药大学附属医院) Department of Electrical and Computer Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学)

专题命中 多模态Agent :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21364 2025-09-01 cs.RO cs.SY eess.SY 78%

Multi-Modal Model Predictive Path Integral Control for Collision Avoidance

Alberto Bertipaglia, Dariu M. Gavrila, Barys Shyrokau

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Accepted as an oral presentation at the 29th IAVSD. August 18-22, 2025. Shanghai, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21037 2025-08-29 cond-mat.mtrl-sci physics.app-ph 78%

Predicting Trends in $V_{OC}$ Through Rapid, Multimodal Characterization of State-of-the-Art p-i-n Perovskite Devices

Amy E. Louks, Brandon T. Motes, Anthony T. Troupe, Axel F. Palmstrom, Joseph J. Berry, Dane W. deQuilettes

专题命中 多模态Agent :multimodal(title,abstract)

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18292 2025-08-27 cs.RO 78%

Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration

Zhixuan Shen, Haonan Luo, Kexun Chen, Fengmao Lv, Tianrui Li

机构 * Zhixuan Shen, Haonan Luo, Kexun Chen, Fengmao Lv, Tianrui Li(作者)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 16 pages, 10 figures, Extended Version of accepted AAAI 2025 Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15043 2025-08-22 cs.HC 78%

LitForager: Exploring Multimodal Literature Foraging Strategies in Immersive Sensemaking

Haoyang Yang, Elliott H. Faa, Weijian Liu, Shunan Guo, Duen Horng Chau, Yalong Yang

专题命中 多模态Agent :multimodal(title,abstract)

Comments 11 pages, 10 figures, Accepted to IEEE ISMAR 2025 (TVCG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11723 2025-08-19 cs.LG 78%

From Heuristics to Data: Quantifying Site Planning Layout Indicators with Deep Learning and Multi-Modal Data

Qian Cao, Jielin Chen, Junchao Zhao, Rudi Stouffs

机构 * National University of Singapore (NUS)(新加坡国立大学) The Cambridge Centre for Advanced Research and Education in Singapore (CARES)(新加坡剑桥高级研究与教育中心) China Southwest Architectural Design and Research Institute Co., Ltd. (CSWADI)(中国西南建筑规划设计研究院有限公司)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract)

Comments 42 pages, 32 figures, submitted to Environment and Planning B: Urban Analytics and City Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17734 2025-07-24 cs.HC 78%

DataWink: Reusing and Adapting SVG-based Visualization Examples with Large Multimodal Models

Liwenhan Xie, Yanna Lin, Can Liu, Huamin Qu, Xinhuan Shu

专题命中 多模态Agent :multimodal(title,abstract)

Comments Accepted to the IEEE Visualization Conference (VIS'25). 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15293 2025-06-19 cs.HC cs.RO 78%

Designing Intent: A Multimodal Framework for Human-Robot Cooperation in Industrial Workspaces

Francesco Chiossi, Julian Rasch, Robin Welsch, Albrecht Schmidt, Florian Michahelles

机构 * Aalto University(阿alto大学) TU Wien(维也纳技术大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 9 pages

Journal ref The Future of Human-Robot Synergy in Interactive Environments: The Role of Robots at the Workplace @ CHIWork 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12710 2025-06-17 cs.RO 78%

Multimodal Large Language Models-Enabled UAV Swarm: Towards Efficient and Intelligent Autonomous Aerial Systems

Yuqi Ping, Tianhao Liang, Huahao Ding, Guangyu Lei, Junwei Wu, Xuan Zou, Kuan Shi, Rui Shao, Chiya Zhang, Weizheng Zhang, Weijie Yuan, Tingting Zhang

机构 * College of Informatics, Harbin Institute of Technology(信息学院,哈尔滨工业大学) Key Laboratory of Forest and Grassland Fire Risk Prevention, Ministry of Emergency Management, China Fire and Rescue Institute(森林和草原火灾风险预防重点实验室,应急管理部,中国消防救援学院) Southern University of Science and Technology(南方科技大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 8 pages, 5 figures,submitted to IEEE wcm

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00237 2025-06-05 cs.RO cs.LG cs.SY eess.SY 78%

Future-Oriented Navigation: Dynamic Obstacle Avoidance with One-Shot Energy-Based Multimodal Motion Prediction

Ze Zhang, Georg Hess, Junjie Hu, Emmanuel Dean, Lennart Svensson, Knut Åkesson

机构 * Chalmers University of Technology(查尔姆斯理工大学)

专题命中 多模态Agent :multimodal(title,abstract)

Comments Published in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03144 2025-05-27 cs.MA cs.DB cs.DS 78%

Group Trip Planning Query Problem with Multimodal Journey

Dildar Ali, Suman Banerjee, Yamuna Prasad

专题命中 多模态Agent :multimodal(title,abstract)

Comments 11 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02123 2025-05-06 cs.RO cs.DB 78%

DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving

Xinmeng Hou, Wuqi Wang, Long Yang, Hao Lin, Jinglun Feng, Haigen Min, Xiangmo Zhao

机构 * Chang’an University(长安大学) Agency for Science, Technology and Research (A*STAR)(科技研究局) University of California, Davis(加州大学戴维斯分校) CCNY Robotics Lab, The City College of New York(纽约城市学院机器人实验室)

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07837 2025-05-02 cs.RO cs.LG 78%

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model

Sicheng Wang, Sheng Liu, Weiheng Wang, Jianhua Shan, Bin Fang

机构 * Casbot Robotic Corporation(Casbot机器人公司) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Anhui University of Technology(安徽理工大学)

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14904 2025-03-25 cond-mat.mtrl-sci cs.LG 78%

Towards an automated workflow in materials science for combining multi-modal simulative and experimental information using data mining and large language models

Balduin Katzer, Steffen Klinder, Katrin Schulz

专题命中 多模态Agent :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16824 2025-03-24 cs.HC 78%

Toward AI-driven Multimodal Interfaces for Industrial CAD Modeling

Jiin Choi, Yugyeong Jang, Kyung Hoon Hyun

专题命中 多模态Agent :multimodal(title,abstract)

Comments 4 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13964 2025-03-19 cs.LG 78%

MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding

Siwei Han, Peng Xia, Ruiyi Zhang, Tong Sun, Yun Li, Hongtu Zhu, Huaxiu Yao

专题命中 多模态Agent :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06779 2025-03-11 cs.RO cs.SY eess.SY 78%

Chance-Constrained Trajectory Planning with Multimodal Environmental Uncertainty

Kai Ren, Heejin Ahn, Maryam Kamgarpour

专题命中 多模态Agent :multimodal(title,abstract)

Comments Published in IEEE Control Systems Letters

Journal ref in IEEE Control Systems Letters, vol. 7, pp. 13-18, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06733 2025-03-11 cs.RO 78%

Embodied multi-modal sensing with a soft modular arm powered by physical reservoir computing

Jun Wang, Suyi Li

专题命中 多模态Agent :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07434 2025-03-06 cs.LG cs.SY eess.SY 78%

Multi-Modal Conformal Prediction Regions with Simple Structures by Optimizing Convex Shape Templates

Renukanandan Tumu, Matthew Cleaveland, Rahul Mangharam, George J. Pappas, Lars Lindemann

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Accepted to L4DC 2024. 14 pages, 3 figures. The source code and toolbox are available at https://github.com/nandantumu/conformal_region_designer

Journal ref PMLR 242:1343-1356, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00401 2025-03-05 cs.CL cs.AI cs.CV cs.HC 78%

Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks

Zongru Wu, Pengzhou Cheng, Zheng Wu, Tianjie Ju, Zhuosheng Zhang, Gongshen Liu

专题命中 多模态Agent :MLLM(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06026 2025-02-11 cs.LG cs.NA math.NA 78%

A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions

Elisa Negrini, Yuxuan Liu, Liu Yang, Stanley J. Osher, Hayden Schaeffer

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏