arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

2508.05068 2025-08-20 cs.CV cs.AI cs.LG eess.IV 62%

Automatic Image Colorization with Convolutional Neural Networks and Generative Adversarial Networks

Changyuan Qiu, Hangrui Cao, Qihan Ren, Ruiyu Li, Yuqing Qiu

机构 * University of Michigan(密歇根大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments All authors have equal authorship and equal contribution, ranked in alphabetic order. First version of this paper was completed and published in 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11747 2025-08-13 cs.CV cs.AI 62%

OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation

Phuc D. A. Nguyen, Minh Luu, Anh Tran, Cuong Pham, Khoi Nguyen

机构 * Movian AI Posts & Telecommunications Inst. of Tech(电信技术研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCVW'25 - OpenSUN3D: 5th Workshop on Open-World 3D Scene Understanding with Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02622 2025-08-11 cs.AI cs.CL cs.CY 62%

Noosemia: toward a Cognitive and Phenomenological Account of Intentionality Attribution in Human-Generative AI Interaction

Enrico De Santis, Antonello Rizzi

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This version has been extensively revised and revisited in light of feedback and further research. Several sections have been expanded or improved for greater clarity and completeness. Specifically, new clarification on complex system foundation related to Noosemia has been added (Secs. "2.4 and "2.5")

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05009 2025-08-08 cs.AI cs.CL 62%

Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses

Bin Han, Robert Wolfe, Anat Caspi, Bill Howe

机构 * University of Washington(华盛顿大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03562 2025-08-06 cs.CV cs.CL 62%

Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching

Muzhaffar Hazman, Susan McKeever, Josephine Griffith

机构 * School of Computer Science University of Galway(计算机科学学院 Galway大学) School of Computer Science Technological University Dublin(计算机科学学院 技术大学都柏林)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted for publication at IEEE International Conference on Image Processing Theory, Tools and Applications (IPTA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04400 2025-08-05 cs.LG cs.AI cs.CR cs.MM 62%

Adaptive Prototype Knowledge Transfer for Federated Learning with Mixed Modalities and Heterogeneous Tasks

Keke Gai, Mohan Wang, Jing Yu, Dongjue Wang, Qi Wu

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18678 2025-07-28 cs.CV cs.AI 62%

Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting

Xingyu Miao, Haoran Duan, Quanhao Qian, Jiuniu Wang, Yang Long, Ling Shao, Deli Zhao, Ran Xu, Gongjie Zhang

机构 * Durham University(杜ham大学) DAMO Academy, Alibaba Group(达摩院,阿里集团) Tsinghua University(清华大学) UCAS-Terminus AI Lab(北航-terminate人工智能实验室)

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18645 2025-07-28 cs.CV cs.AI 62%

Quantum-Cognitive Tunnelling Neural Networks for Military-Civilian Vehicle Classification and Sentiment Analysis

Milan Maksimovic, Anna Bohdanets, Immaculate Motsi-Omoijiade, Guido Governatori, Ivan S. Maksymov

机构 * Artificial Intelligence and Cyber Futures Institute, Charles Sturt University(人工智能与未来网络研究所,查尔斯·斯特劳特大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06270 2025-07-28 cs.LG cs.AI cs.CL 62%

XAI4LLM. Let Machine Learning Models and LLMs Collaborate for Enhanced In-Context Learning in Healthcare

Fatemeh Nazary, Yashar Deldjoo, Tommaso Di Noia, Eugenio di Sciascio

机构 * Polytechnic University of Bari(巴里理工大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15878 2025-07-23 cs.CV cs.AI 62%

Salience Adjustment for Context-Based Emotion Recognition

Bin Han, Jonathan Gratch

机构 * Computer Science, University of Southern California(计算机科学,南加州大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09567 2025-07-21 cs.AI cs.CL 62%

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, Wanxiang Che

机构 * LARG Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究中心) Harbin Institute of Technology(哈尔滨工业大学) School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) The University of Hong Kong(香港大学) Fudan University(复旦大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Paper list and Github tutorial are available at https://github.com/LightChen233/Awesome-Long-Chain-of-Thought-Reasoning. Update 250+ New Reference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09036 2025-07-15 cs.CV cs.AI cs.LG cs.SE 62%

BrainLesion Suite: A Flexible and User-Friendly Framework for Modular Brain Lesion Image Analysis

Florian Kofler, Marcel Rosier, Mehdi Astaraki, Hendrik Möller, Ilhem Isra Mekki, Josef A. Buchner, Anton Schmick, Arianna Pfiffer, Eva Oswald, Lucas Zimmer, Ezequiel de la Rosa, Sarthak Pati, Julian Canisius, Arianna Piffer, Ujjwal Baid, Mahyar Valizadeh, Akis Linardos, Jan C. Peeken, Surprosanna Shit, Felix Steinbauer, Daniel Rueckert, Rolf Heckemann, Spyridon Bakas, Jan Kirschke, Constantin von See, Ivan Ezhov, Marie Piraud, Benedikt Wiestler, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland(苏黎世大学定量生物医学系) Helmholtz AI, Helmholtz Zentrum München, Germany(海德堡人工智能,慕尼黑研究中心) Department of Diagnostic and Interventional Neuroradiology, School of Medicine, Klinikum rechts der Isar, Technical University of Munich, Germany(医学诊断与介入神经放射学系,慕尼黑技术大学) TranslaTUM - Central Institute for Translational Cancer Research, Technical University of Munich, Germany(TranslaTUM - 转化癌症研究中央研究所,慕尼黑技术大学) Chair for AI in Medicine, Klinikum Rechts der Isar, Munich, Germany(医学人工智能系,慕尼黑Klinikum Rechts der Isar) Department of Radiation Oncology, Klinikum rechts der Isar, Technical University of Munich, Germany(放射肿瘤学系,慕尼黑Klinikum Rechts der Isar) Institute of Clinical Neuroimmunology, Faculty of Medicine, Ludwig Maximilian University, Munich, Germany(临床神经免疫学研究所,路德维希-马克西米利安大学) School of Computation, Information and Technology, Technical University of Munich, Germany(计算、信息与技术学院,慕尼黑技术大学) Division of Computational Pathology, Indiana School of Medicine, Indianapolis, IN, USA(计算病理学系,印第安纳医学院) Research Center for Digital Technologies in Dentistry and CAD/CAM, Department of Dentistry, Faculty of Medicine and Dentistry, Danube Private University, Krems, Austria(牙科数字技术研究中心,医学与牙科学院,多瑙私人大学) Medical Research Group, MLCommons, San Francisco, CA, USA(医学研究组,MLCommons,旧金山,美国) Department of Neurology, Clinical Neuroscience Center and Brain Tumor Center, University and University Hospital Zurich, 8091 Zurich, Switzerland(神经学系,临床神经科学中心和脑肿瘤中心,苏黎世大学及苏黎世大学医院) Department of Medical Radiation Physics, Stockholm University, Solna, Sweden(医学辐射物理学系,斯德哥尔摩大学) Department of Oncology-Pathology, Karolinska Institutet, Solna, Sweden, Sweden(肿瘤病理学系,卡罗林斯卡研究所) Division of Oncology and Children's Research Center, University Children's Hospital Zurich, Switzerland(肿瘤与儿童研究中心,苏黎世儿童医院) Wallace H. Coulter Department of Biomedical Engineering at Georgia Tech and Emory University, Atlanta, USA(Wallace H. Coulter生物医学工程系,佐治亚理工学院和埃默里大学,美国) Institute of Radiation Medicine (IRM), Helmholtz Zentrum München (HMGU), Germany(放射医学研究所(IRM),海德堡研究中心(HMGU)) German Consortium for Translational Cancer Research (DKTK), Partner Site Munich, Munich, Germany(德国转化癌症研究联盟(DKTK),慕尼黑合作伙伴站点)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 16p, 3f

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24030 2025-07-11 cs.LG cs.AI cs.CV 62%

From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?

Ziming Zhao, ChengAo Shen, Hanghang Tong, Dongjin Song, Zhigang Deng, Qingsong Wen, Jingchao Ni

机构 * University of Houston(休斯顿大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Connecticut(康涅狄格大学) Squirrel Ai Learning(squirrel Ai 学习)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08740 2025-07-10 cs.CV cs.AI cs.IR 62%

Hespi: A pipeline for automatically detecting information from hebarium specimen sheets

Robert Turnbull, Emily Fitzgerald, Karen Thompson, Joanne L. Birch

机构 * The University of Melbourne(墨尔本大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21364 2025-06-27 cs.CV cs.AI 62%

CA-I2P: Channel-Adaptive Registration Network with Global Optimal Selection

Zhixin Cheng, Jiacheng Deng, Xinjun Li, Xiaotian Yin, Bohao Liao, Baoqun Yin, Wenfei Yang, Tianzhu Zhang

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Institute of Advanced Technology, University of Science and Technology of China(先进技术研究院,中国科学技术大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13476 2025-06-17 cs.CV cs.AI eess.IV 62%

ESRPCB: an Edge guided Super-Resolution model and Ensemble learning for tiny Printed Circuit Board Defect detection

Xiem HoangVan, Dang Bui Dinh, Thanh Nguyen Canh, Van-Truong Nguyen

机构 * Faculty of Electronics and Telecommunications, University of Engineering and Technology, Vietnam National University(电子与电信学院,工程与技术大学,越南国家大学) School of Information Science, Japan Advanced Institute of Science and Technology(信息科学学院,日本先进科学和技术研究院) Faculty of Mechatronics, SMAE, Hanoi University of Industry(机械学院,SMAE,河内工业大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Published in Engineering Applications of Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12156 2025-06-17 cs.LG cs.AI cs.CV 62%

Explaining Recovery Trajectories of Older Adults Post Lower-Limb Fracture Using Modality-wise Multiview Clustering and Large Language Models

Shehroz S. Khan, Ali Abedi, Charlene H. Chu

机构 * KITE Research Institute(KITE研究机构) Toronto Rehabilitation Institute(多伦多康复研究所) University Health Network(大学健康网络) College of Engineering and Technology(工程与技术学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21333 2025-06-17 cs.LG cs.AI cs.CL cs.CY 62%

Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse

Ryan Liu, Jiayi Geng, Addison J. Wu, Ilia Sucholutsky, Tania Lombrozo, Thomas L. Griffiths

机构 * Department of Computer Science, Princeton University, Princeton, NJ, USA(计算机科学系,普林斯顿大学) NYU Center for Data Science, New York University, New York, NY, USA(纽约大学数据科学中心) Department of Psychology, Princeton University, Princeton, NJ, USA(心理学系,普林斯顿大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10004 2025-06-13 cs.MM cs.AI cs.ET cs.NI 62%

Immersive Multimedia Communication: State-of-the-Art on eXtended Reality Streaming

Haopeng Wang, Haiwei Dong, Abdulmotaleb El Saddik

机构 * University of Ottawa(渥太华大学) University of Ottawa and Huawei Canada(渥太华大学和华为加拿大) Huawei Canada(华为加拿大)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments accepted by ACM Transactions on Multimedia Computing, Communications, and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06729 2025-06-10 cs.CV cs.CL 62%

Mitigating Object Hallucination via Robust Local Perception Search

Zixian Gao, Chao Yang, Zhanhui Zhou, Xing Xu, Chaochao Lu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Center for Future Media & School of Computer Science and Engineering, University of Electronic Science and Technology of China(未来媒体中心及电子科技大学计算机科学与工程学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11062 2025-05-20 cs.LG cs.AI cs.CL 62%

EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Ping Luo

机构 * The University of Hong Kong(香港大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main, camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03689 2025-05-16 cs.CV cs.CL 62%

Pose Priors from Language Models

Sanjay Subramanian, Evonne Ng, Lea Müller, Dan Klein, Shiry Ginosar, Trevor Darrell

机构 * University of California, Berkeley(加州大学伯克利分校) Google DeepMind(谷歌DeepMind) Toyota Technological Institute at Chicago(丰田技术研究所(芝加哥))

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00057 2025-04-28 cs.LG cs.AI cs.CV 62%

VisTabNet: Adapting Vision Transformers for Tabular Data

Witold Wydmański, Ulvi Movsum-zada, Jacek Tabor, Marek Śmieja

机构 * Faculty of Mathematics and Computer Science, Jagiellonian University(数学与计算机科学学院,杰兹维日大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11378 2025-04-22 cs.CV cs.AI 62%

Inspecting Explainability of Transformer Models with Additional Statistical Information

Hoang C. Nguyen, Haeil Lee, Junmo Kim

机构 * KAIST(韩国科学技术院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at Responsible Computer Vision workshop at ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10878 2025-04-16 cs.CV cs.AI cs.LG 62%

Large Language Model-Informed Feature Discovery Improves Prediction and Interpretation of Credibility Perceptions of Visual Content

Yilang Peng, Sijia Qian, Yingdan Lu, Cuihua Shen

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09456 2025-04-15 cs.AI cs.CV 62%

Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs

Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07954 2025-04-11 cs.CV cs.CL 62%

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

En Yu, Kangheng Lin, Liang Zhao, Jisheng Yin, Yana Wei, Yuang Peng, Haoran Wei, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang, Wenbing Tao

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV、cs.CL

Comments Github page: https://github.com/linkangheng/PR1

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05878 2025-04-09 cs.MM cs.CV 62%

KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection

Xingyuan Li, Ruichao Hou, Tongwei Ren, Gangshan Wu

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments This paper is accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05231 2025-04-08 cs.AI cs.CV cs.LG 62%

Mapping biodiversity at very-high resolution in Europe

César Leblanc, Lukas Picek, Benjamin Deneu, Pierre Bonnet, Maximilien Servajean, Rémi Palard, Alexis Joly

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04645 2025-04-08 eess.IV cs.AI cs.CV 62%

Here Comes the Explanation: A Shapley Perspective on Multi-contrast Medical Image Segmentation

Tianyi Ren, Juampablo Heras Rivera, Hitender Oswal, Yutong Pan, Agamdeep Chopra, Jacob Ruzevick, Mehmet Kurt

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏