arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4884 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4884 篇

2511.09944 2025-11-14 cs.CV 57%

TSPE-GS: Probabilistic Depth Extraction for Semi-Transparent Surface Reconstruction via 3D Gaussian Splatting

Zhiyuan Xu, Nan Min, Yuhang Guo, Tong Wei

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI26 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05162 2025-11-14 cs.CY cs.AI 57%

Artificial-Intelligence Grading Assistance for Handwritten Components of a Calculus Exam

Gerd Kortemeyer, Alexander Caspar, Daria Horica

机构 * Michigan State University(密歇根州立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07912 2025-11-12 cs.AI 57%

Neurophysiological Characteristics of Adaptive Reasoning for Creative Problem-Solving Strategy

Jun-Young Kim, Young-Seok Kweon, Gi-Hwan Shin, Seong-Whan Lee

机构 * Dept. of Artificial Intelligence(人工智能系) Korea University(韩国大学) Dept. of Brain and Cognitive Engineering(脑科学与认知工程系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 4 pages, 4 figures, 1 table,

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18108 2025-11-12 cs.CV 57%

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen, Zhang Liu, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Corning Inc(康宁公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Journal ref CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07004 2025-11-11 cs.CV cs.HC 57%

Exploring the "Great Unseen" in Medieval Manuscripts: Instance-Level Labeling of Legacy Image Collections with Zero-Shot Models

Christofer Meinecke, Estelle Guéville, David Joseph Wrisley

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06744 2025-11-11 cs.CV 57%

PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks

Da-Yeong Kim, Yeong-Jun Cho

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06297 2025-11-11 cs.HC cs.AI 57%

Decomate: Leveraging Generative Models for Co-Creative SVG Animation

Jihyeon Park, Jiyoon Myung, Seone Shin, Jungki Son, Joohyung Han

机构 * MODULABS

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at the 1st Workshop on Generative and Protective AI for Content Creation (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06256 2025-11-11 cs.CV 57%

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang, Xiang Wan, Xiaonan Luo, Guanbin Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Sun Yat-sen University(中山大学) Baidu Inc.(百度公司) Guilin University of Electronic Technology(桂林电子科技大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13765 2025-11-06 cs.HC cs.CL 57%

SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model

Xingbo Wang, Samantha L. Huey, Rui Sheng, Saurabh Mehta, Fei Wang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint version of the paper accepted to Campbell Systematic Reviews. Code is available at https://github.com/xingbow/SciDaEx

Journal ref Campbell Systematic Reviews 21 (2025): 1-16

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09163 2025-11-05 cs.CV 57%

CWSSNet: Hyperspectral Image Classification Enhanced by Wavelet Domain Convolution

Yulin Tong, Fengzong Zhang, Haiqin Cheng

机构 * School of Transportation Engineering(交通运输工程学院) East China Jiaotong University(东华交通大学) School of Coumputing and Information(计算与信息学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14269 2025-10-31 cs.LG cs.CV cs.IR 57%

Deep Learning for Technical Document Classification

Shuo Jiang, Jie Hu, Christopher L. Magee, Jianxi Luo

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages, 8 figures, 9 tables

Journal ref IEEE Transactions on Engineering Management 71 (2024): 1163-1179

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10940 2025-10-30 cs.IR cs.AI 57%

Who You Are Matters: Bridging Topics and Social Roles via LLM-Enhanced Logical Recommendation

Qing Yu, Xiaobei Wang, Shuchang Liu, Yandong Bai, Xiaoyu Yang, Xueliang Wang, Chang Meng, Shanshan Wu, Hailan Yang, Huihui Xiao, Xiang Li, Fan Yang, Xiaoqiang Feng, Lantao Hu, Han Li, Kun Gai, Lixin Zou

机构 * Wuhan University(武汉大学) Kuaishou Technology(快手科技)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments to be published in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24023 2025-10-29 cs.CL 57%

Success and Cost Elicit Convention Formation for Efficient Communication

Saujas Vaduguru, Yilun Hua, Yoav Artzi, Daniel Fried

机构 * Carnegie Mellon University(卡内基梅隆大学) Department of Computer Science and Cornell Tech, Cornell University(计算机科学系和康奈尔科技,康奈尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23648 2025-10-29 cs.SI cs.AI 57%

RoGBot: Relationship-Oblivious Graph-based Neural Network with Contextual Knowledge for Bot Detection

Ashutosh Anshul, Mohammad Zia Ur Rehman, Sri Akash Kadali, Nagendra Kumar

机构 * Indian Institute of Technology Indore(印度理工学院印多尔分校) University of Maryland, College Park, USA(美国马里兰大学学院公园分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17047 2025-10-28 cs.CL 57%

Modeling Bottom-up Information Quality during Language Processing

Cui Ding, Yanning Yin, Lena A. Jäger, Ethan Gotlieb Wilcox

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02366 2025-10-28 cs.LG cs.CL q-fin.TR 57%

Language Model Guided Reinforcement Learning in Quantitative Trading

Adam Darmanin, Vince Vella

机构 * University of Malta(马耳他大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL

Comments 12 pages (4 pages appendix and references) and 6 figures. Accepted for presentation at FLLM 2025, Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14677 2025-10-28 cs.CV 57%

Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning

Jiaer Xia, Yuhang Zang, Peng Gao, Sharon Li, Kaiyang Zhou

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanghai AI Lab(上海人工智能实验室) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17611 2025-10-27 cs.CV 57%

One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection

Jia Guo, Shuai Lu, Lei Fan, Zelin Li, Donglin Di, Yang Song, Weihang Zhang, Wenbing Zhu, Hong Yan, Fang Chen, Huiqi Li, Hongen Liao

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) University of New South Wales(新南威尔士大学) DZ Matrix(DZ矩阵) Fudan University(复旦大学) Rongcheer Co., Ltd.(荣彻科技有限公司)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Extended version of CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16034 2025-10-21 cs.CV 57%

VisualLens: Personalization through Task-Agnostic Visual History

Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Luna Dong

机构 * Meta University of Southern California(南加州大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13660 2025-10-17 cs.CV 57%

OmniGaze: Reward-inspired Generalizable Gaze Estimation In The Wild

Hongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao, Wenguan Wang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) Zhejiang University(浙江大学) Nanjing Forestry University(南京林业大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025; Project page: https://github.com/quhongyu/OmniGaze

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00527 2025-10-17 cs.RO cs.AI 57%

Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions

Guanwen Xie, Jingzehua Xu, Yimian Ding, Zhi Zhang, Shuai Zhang, Yi Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, 518055, China(清华大学深圳国际研究生院,清华大学,深圳,518055,中国) Department of Data Science, New Jersey Institute of Technology(数据科学系,新泽西理工学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted by IEEE/RSJ IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13029 2025-10-16 cs.AI 57%

Toward Reasoning-Centric Time-Series Analysis

Xinlei Wang, Mingtian Tan, Jing Qiu, Junhua Zhao, Jinjin Gu

机构 * University of Sydney(悉尼大学) University of Virginia(弗吉尼亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Institute of Artificial Intelligence and Robotics for Society(深圳人工智能与机器人社会研究院) INSAIT

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23179 2025-10-16 cs.CV 57%

DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes

Sungjune Park, Hyunjun Kim, Junho Kim, Seongho Kim, Yong Man Ro

机构 * Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11106 2025-10-14 cs.CV 57%

Compositional Zero-Shot Learning: A Survey

Ans Munir, Faisal Z. Qureshi, Mohsen Ali, Muhammad Haris Khan

机构 * Information Technology University(信息科技大学) University of Ontario Institute of Technology(安大略理工大学) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

Comments Survey paper with 36 pages, 8 plots and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21988 2025-10-09 cs.AI 57%

Functional Matching of Logic Subgraphs: Beyond Structural Isomorphism

Ziyang Zheng, Kezhi Li, Zhengyuan Shi, Qiang Xu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18331 2025-10-06 cs.CL 57%

BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle

EunJeong Hwang, Peter West, Vered Shwartz

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能矢量研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25603 2025-10-01 cs.CV 57%

GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification

Yijia Weng, Zhicheng Wang, Songyou Peng, Saining Xie, Howard Zhou, Leonidas J. Guibas

机构 * Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24873 2025-09-30 cs.LG cs.AI 57%

Uncertainty-Guided Expert-AI Collaboration for Efficient Soil Horizon Annotation

Teodor Chiaburu, Vipin Singh, Frank Haußer, Felix Bießmann

机构 * Einstein Center Digital Future Berlin Germany Einstein Center Digital Future

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 11 pages, 7 figures, presented at ECAI 2025, CLEAR-AI Workshop, Bologna

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24124 2025-09-30 cs.RO cs.AI cs.LG 57%

Ancestry Tree Clustering for Particle Filter Diversity Maintenance

Ilari Vallivaara, Bingnan Duan, Yinhuan Dong, Tughrul Arslan

机构 * Visiting Research Fellow University of Edinburgh Edinburgh, UK(访问研究员 爱丁堡大学 英国爱丁堡) School of Engineering University of Edinburgh Edinburgh, UK(工程学院 爱丁堡大学 英国爱丁堡)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 15th International Conference on Indoor Positioning and Indoor Navigation, 15-18 September 2025, Tampere, Finland Originally 8 pages. The online version with appendices is 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13920 2025-09-29 cs.AI 57%

Integrating Knowledge Graphs and Bayesian Networks: A Hybrid Approach for Explainable Disease Risk Prediction

Mbithe Nzomo, Deshendran Moodley

机构 * Centre for Artificial Intelligence Research (CAIR)(人工智能研究中心) Department of Computer Science(计算机科学系) University of Cape Town(开普敦大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments This work has been accepted for presentation at the 49th IEEE International Conference on Computers, Software, and Applications (COMPSAC 2025). The final published version will be available via IEEE Xplore

Journal ref Proceedings of the 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC), Toronto, ON, Canada, 2025, pp. 834-844

详情

展开后加载摘要…

URL PDF HTML 收藏