arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4884 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4884 篇

2404.13274 2025-05-19 cs.HC cs.AI 57%

Augmented Object Intelligence with XR-Objects

Mustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du, Andrea Colaço, Johnny Lee, Mar Gonzalez-Franco, David Kim

机构 * Google(谷歌)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 15 pages, 15 figures, 2024 ACM Symposium on User Interface Software and Technology (UIST)

Journal ref ACM UIST 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09952 2025-05-16 cs.LG cs.AI 57%

Task-Core Memory Management and Consolidation for Long-term Continual Learning

Tianyu Huai, Jie Zhou, Yuxuan Cai, Qin Chen, Wen Wu, Xingjiao Wu, Xipeng Qiu, Liang He

机构 * School of Computer Science and Technology, East China Normal University(东华师范大学计算机科学与技术学院) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) Computation and Artificial Intelligence Innovative College, Fudan University(复旦大学计算与人工智能创新学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Submitted to Neurips2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07250 2025-05-13 physics.ins-det cs.AI 57%

Lightweight Deep Learning Framework for Accurate Particle Flow Energy Reconstruction

Yu Wang, Yangguang Zhang, Shengxiang Lin, Xingyi Zhang, Han Zhang

机构 * School of Automation and Electrical Engineering, University of Science and Technology Beijing(北京科技大学自动化与电气工程学院) School of Control and Computer Engineering, North China Electric Power University(华北电力大学控制与计算机工程学院) Faculty of Electronic and Information Engineering, Xi’an Jiaotong University(西安交通大学电子与信息工程学院) School of Mechanical Engineering, Shanghai Jiao Tong University(上海交通大学机械工程学院) College of Aritificial Intelligence and Automation, Hohai University(河海大学人工智能与自动化学院)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.AI

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18988 2025-05-13 cs.HC cs.CL 57%

LINC: Supporting Language Independent Communication and Comprehension to Enhance Contribution in Multilingual Collaborative Meetings

Saramsh Gautam, Mahmood Jasim

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments The manuscript has been withdrawn by the authors due to ongoing revisions and substantial updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11895 2025-05-09 cs.CV 57%

Search is All You Need for Few-shot Anomaly Detection

Qishan Wang, Jia Guo, Shuyong Gao, Haofen Wang, Li Xiong, Junjie Hu, Hanqi Guo, Wenqiang Zhang

机构 * Academy for Engineering and Technology Fudan University(复旦大学工程与技术学院) School of Biomedical Engineering Tsinghua University(清华大学生物医学工程学院) School of Computer Science Fudan university(复旦大学计算机学院) College of Design & Innovation Tongji University(同济大学设计与创新学院) School of Physics and Mechanical & Electrical Engineering Hexi University(河西大学物理与机械电子工程学院) Engineering Research Center of AI&Robotics, Ministry of Education(教育部人工智能与机器人工程研究中心) Academy for Engineering&Technology Fudan university(复旦大学工程与技术学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03315 2025-05-07 cs.AI 57%

Artificial Behavior Intelligence: Technology, Challenges, and Future Directions

Kanghyun Jo, Jehwan Choi, Kwanho Kim, Seongmin Kim, Duy-Linh Nguyen, Xuan-Thuy Vo, Adri Priadana, Tien-Dat Tran

机构 * Department of Electrical, Electronic and Computer Engineering University of Ulsan(电气电子与计算机工程系韩国釜山大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 6 figures, Pre-print for IWIS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02121 2025-05-06 cs.AI 57%

Overview of AI Grading of Physics Olympiad Exams

Lachlan McGinness

机构 * Australian National University(澳大利亚国立大学) CSIRO, Australia(澳大利亚联邦科学与工业研究组织)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments International Conference on Artificial Intelligence in Education, Doctoral Consortium

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10276 2025-05-06 cs.LG cs.AI 57%

FEDKIM: Adaptive Federated Knowledge Injection into Medical Foundation Models

Xiaochen Wang, Jiaqi Wang, Houping Xiao, Jinghui Chen, Fenglong Ma

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Georgia State University(佐治亚州立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted by EMNLP'24 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01396 2025-05-05 cs.RO cs.AI cs.LG 57%

SIME: Enhancing Policy Self-Improvement with Modal-level Exploration

Yang Jin, Jun Lv, Wenye Yu, Hongjie Fang, Yong-Lu Li, Cewu Lu

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institution(上海创新研究所) Noematrix Ltd.(诺玛特有限公司)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21184 2025-05-01 cs.AI 57%

AffectEval: A Modular and Customizable Framework for Affective Computing

Emily Zhou, Khushboo Khatri, Yixue Zhao, Bhaskar Krishnamachari

机构 * Computer Science University of Southern California(计算机科学 华南大学) Electrical and Computer Engineering University of Southern California(电气与计算机工程 华南大学) USC Information Sciences Institute University of Southern California(USC信息科学研究所 华南大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments The short version is published in ACM/IEEE CHASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03118 2025-05-01 cs.HC cs.CV 57%

ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Ruiping Liu, Jiaming Zhang, Angela Schön, Karin Müller, Junwei Zheng, Kailun Yang, Anhong Guo, Kathrin Gerling, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15388 2025-05-01 cs.CV cs.LG eess.IV 57%

A Contrast-Agnostic Method for Ultra-High Resolution Claustrum Segmentation

Chiara Mauri, Ryan Fritz, Jocelyn Mora, Benjamin Billot, Juan Eugenio Iglesias, Koen Van Leemput, Jean Augustinack, Douglas N Greve

机构 * Department of Radiology, Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital, Charlestown, MA, USA(放射科,Athinoula A. Martinos生物医学成像中心,麻省总医院,查尔斯顿,马萨诸塞州,美国) Department of Radiology, Harvard Medical School, Boston, MA USA(放射科,哈佛医学院,波士顿,马萨诸塞州,美国) MIT Computer Science & Artificial Intelligence Laboratory, Cambridge, MA USA(麻省理工学院计算机科学与人工智能实验室,剑桥,马萨诸塞州,美国) Epione team, Inria, Sophia Antipolis, France(Epione团队,Inria,索菲亚安蒂波利斯,法国) UCL Centre for Medical Image Computing, London, United Kingdom(伦敦大学学院医学图像计算中心,伦敦,英国) Department of Neuroscience and Biomedical Engineering, Aalto University, Espoo, Finland(神经科学与生物医学工程系,阿alto大学,埃斯波,芬兰) Department of Computer Science, Aalto University, Espoo, Finland(计算机科学系,阿alto大学,埃斯波,芬兰)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 20 pages, 15 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12150 2025-04-29 cs.CV 57%

Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis

Hongyu Sun, Qiuhong Ke, Ming Cheng, Yongcai Wang, Deying Li, Chenhui Gou, Jianfei Cai

机构 * Department of Computer Science, Renmin University of China(中国人民大学计算机学院) Department of Data Science & AI, Monash University(墨尔本大学数据科学与人工智能系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025; 24 pages, 14 figures, 18 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16021 2025-04-23 cs.HC cs.AI 57%

Navigating the State of Cognitive Flow: Context-Aware AI Interventions for Effective Reasoning Support

Dinithi Dissanayake, Suranga Nanayakkara

机构 * Augmented Human Lab, National University of Singapore(增强人类实验室,新加坡国立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Presented at the 2025 ACM Workshop on Human-AI Interaction for Augmented Reasoning, Report Number: CHI25-WS-AUGMENTED-REASONING

Journal ref Proceedings of the 2025 ACM CHI Workshop on Human-AI Interaction for Augmented Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15970 2025-04-23 cs.HC cs.CV cs.MA 57%

Recent Advances and Future Directions in Extended Reality (XR): Exploring AI-Powered Spatial Intelligence

Baichuan Zeng

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments 7 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14920 2025-04-22 cs.CV 57%

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Geng Li, Jinglin Xu, Yunzhen Zhao, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) School of Intelligence Science and Technology, University of Science and Technology Beijing(北京科技大学智能科学与技术学院) Tencent Beijing Research(腾讯北京研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025 (Hightlight). Project page with code: https://github.com/PKU-ICST-MIPL/DyFo_CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19653 2025-04-21 cs.LG cs.CL cs.SY eess.SY 57%

SysCaps: Language Interfaces for Simulation Surrogates of Complex Systems

Patrick Emami, Zhaonan Li, Saumya Sinha, Truc Nguyen

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at ICLR 2025. 23 pages. Updated with final camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12869 2025-04-18 cs.CV 57%

SC3EF: A Joint Self-Correlation and Cross-Correspondence Estimation Framework for Visible and Thermal Image Registration

Xi Tong, Xing Luo, Jiangxin Yang, Yanpeng Cao

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Intelligent Transportation Systems, Early Access, 10.1109/TITS.2025.3542159

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13791 2025-04-18 cs.LG cs.AI 57%

Multi-omics data integration for early diagnosis of hepatocellular carcinoma (HCC) using machine learning

Annette Spooner, Mohammad Karimi Moridani, Azadeh Safarchi, Salim Maher, Fatemeh Vafaee, Amany Zekry, Arcot Sowmya

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05667 2025-04-18 cs.CR cs.AI cs.HC cs.IR cs.LG 57%

PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT

Sayak Saha Roy, Shirin Nilizadeh

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13878 2025-04-15 cs.HC cs.CV 57%

Eye Gaze as a Signal for Conveying User Attention in Contextual AI Systems

Ethan Wilson, Naveen Sendhilnathan, Charlie S. Burlingham, Yusuf Mansour, Robert Cavin, Sai Deep Tetali, Ajoy Savio Fernandes, Michael J. Proulx

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments To appear in ETRA '25: Proceedings of the 2025 Symposium on Eye Tracking Research and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08002 2025-04-14 cs.CL 57%

More diverse more adaptive: Comprehensive Multi-task Learning for Improved LLM Domain Adaptation in E-commerce

Tong Piao, Pei Tang, Zhipeng Zhang, Jiaqi Li, Qiao Liu, Zufeng Wu

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL

Comments Accepted by KDD workshop 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06604 2025-04-14 physics.geo-ph cs.CV 57%

Image registration of 2D optical thin sections in a 3D porous medium: Application to a Berea sandstone digital rock image

Jaehong Chung, Wei Cai, Tapan Mukerji

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00004 2025-04-14 cs.IR cs.AI cs.LG 57%

Navigating the Future of Federated Recommendation Systems with Foundation Models

Zhiwei Li, Guodong Long, Chunxu Zhang, Honglei Zhang, Jing Jiang, Chengqi Zhang

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 11 pages, position paper, survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05318 2025-04-09 cs.IR cs.AI 57%

Efficient Multi-Task Learning via Generalist Recommender

Luyang Wang, Cangcheng Tang, Chongyang Zhang, Jun Ruan, Kai Huang, Jason Dai

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Journal ref The 32nd ACM International Conference on Information and Knowledge Management (2023) pp. 4335-4339

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01459 2025-04-03 cs.LG cs.AI 57%

Probabilistic Curriculum Learning for Goal-Based Reinforcement Learning

Llewyn Salt, Marcus Gallagher

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07523 2025-04-02 cs.CV 57%

VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Zhangquan Chen, Xufang Luo, Dongsheng Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 18pages,11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24065 2025-04-01 cs.CV cs.RO 57%

COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation

Siqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan, Qi Wu, Zhihua Wei, Jing Liu

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18390 2025-04-01 cs.CL 57%

Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

Xinyu Wang, Wenbo Zhang, Sarah Rajtmajer

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16668 2025-04-01 cs.HC cs.AI 57%

Satori: Towards Proactive AR Assistant with Belief-Desire-Intention User Modeling

Chenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G Turakhia, Sonia Castelo Quispe, Dong Li, Leslie Welch, Claudio Silva, Jing Qian

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏