Multi-Modal Conditioned High-Resolution Transformer for Urban Electromagnetic Field Map Prediction Download PDF
面向城市电磁场地图预测的多模态条件高分辨率Transformer
Do-Eon Kim, Dongryul Park, Seungyoung Ahn, Namwoo Kang, Seong-heum Kim, Seongsin Kim
机构
*
Soongsil University(崇实大学)
;
Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology(韩国科学技术院赵春植移动研究生院)
;
Department of Intelligent Semiconductors, Soongsil University(崇实大学智能半导体系)
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
SciOrch: 学习编排专家大语言模型以解决前沿多模态科学推理任务
Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin
机构
*
Imperial College London(伦敦帝国学院)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
University of Oxford(牛津大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
一种多模态深度感知方法用于具身参照理解
Fevziye Irem Eyiokur, Dogucan Yaman, Hazım Kemal Ekenel, Alexander Waibel
机构
*
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
KIT Campus Transfer GmbH (KCT)(KIT校园转移有限公司)
;
Istanbul Technical University(伊斯坦布尔技术大学)
;
Carnegie Mellon University(卡内基梅隆大学)
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
New Laboratory of Pattern Recognition(模式识别新实验室)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
Hong Kong Institute of Science & Innovation(香港科学创新研究院)
;
PolyU
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
Eco-Bee: A Personalised Multi-Modal Agent for Advancing Student Climate Awareness and Sustainable Behaviour in Campus Ecosystems
Eco-Bee:一种面向校园生态系统的个性化多模态代理,用于提升学生气候意识和可持续行为
Caleb Adu, Neil Kapadia, Binhe Liu, Jonathan Randall, Sruthi Viswanathan
机构
*
University of Hull(赫尔大学)
;
The Spaceship Academy(太空学院)
;
City St George’s, University of London(伦敦大学城市学院)
;
King’s College London(伦敦国王学院)
;
Lancaster University(兰卡斯特大学)
;
University of Oxford(牛津大学)
PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation via Multi-Agent LLMs
PosterGen:基于多智能体LLM的美观化论文到海报生成
Zhilin Zhang, Xiang Zhang, Jiaqi Wei, Yiwei Xu, Chenyu You
机构
*
Stony Brook University(石溪大学)
;
New York University(纽约大学)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Zhejiang University(浙江大学)
;
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
Mind over Space: Can Multimodal Large Language Models Mentally Navigate?
心灵超越空间:多模态大语言模型能否进行心理导航?
Qihui Zhu, Shouwei Ruan, Xiao Yang, Hao Jiang, Yao Huang, Shiji Zhao, Hanwei Fan, Hang Su, Xingxing Wei
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学计算机科学与技术系)
;
School of Automation Science and Electrical Engineering , Beihang University(北京航空航天大学自动化科学与电气工程学院)
;
college of AI, Tsinghua University(清华大学人工智能学院)
;
Department of Computer Science and Technology , Tsinghua University(清华大学计算机科学与技术系)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
在简单视觉规划任务中多模态大语言模型推理的分布外泛化
Yannic Neuhaus, Nicolas Flammarion, Matthias Hein, Francesco Croce
机构
*
Tübingen AI Center -- University of Tübingen(图宾根人工智能中心 -- 图宾根大学)
;
EPFL(瑞士联邦理工学院)
;
ELLIS Institute Finland -- Aalto University(芬兰ELLIS研究所 -- 阿尔托大学)
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
Xianzheng Ma, Brandon Smart, Yash Bhalgat, Shuai Chen, Xinghui Li, Jian Ding, Jindong Gu, Dave Zhenyu Chen, Songyou Peng, Jia-Wang Bian, Philip H Torr, Marc Pollefeys, Matthias Nießner, Ian D Reid, Angel X. Chang, Iro Laina, Victor Adrian Prisacariu
机构
*
University of Oxford(牛津大学)
;
King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)
;
Technical University of Munich(慕尼黑技术大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Simon Fraser University(西蒙·弗雷泽大学)
;
ETH Zurich(苏黎世联邦理工学院)
The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
Yuan Xun, Xiaojun Jia, Xinwei Liu, Hua Zhang
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
Nanyang Technological University(南洋理工大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation
Nima Fathi, Amar Kumar, Tal Arbel
机构
*
Center for Intelligent Machines, McGill University, Montreal, Canada(麦吉尔大学智能机器中心,加拿大蒙特利尔)
;
Mila - Quebec AI institute, Montreal, Canada(魁北克AI研究所)
专题命中
多模态Agent
:multi-modal(title);分类 cs.CV
Comments9 pages, 3 figures, International Conference on Medical Image Computing and Computer-Assisted Intervention
Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning
Ziqi Jia, Anmin Wang, Xiaoyang Qu, Xiaowen Yang, Jianzong Wang
机构
*
Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
Tsinghua University(清华大学)
;
Huazhong University of Science and Technology(华中科技大学)
专题命中
多模态Agent
:multi-modal(title);分类 cs.CV
CommentsAccepted by the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)