arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4884 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4884 篇

2512.14082 2025-12-17 cs.CL 57%

A Unified Sparse Attention via Multi-Granularity Compression

一种通过多粒度压缩的统一稀疏注意力机制

Siran Liu, Zane Cao, Yongchao He

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CL

AI总结 UniSparse通过多粒度压缩和块级选择,实现高效稀疏注意力机制,提升LLM在长上下文理解和推理中的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18577 2025-12-16 q-fin.CP cs.AI cs.LG 57%

Advancing Financial Engineering with Foundation Models: Progress, Applications, and Challenges

用基础模型推进金融工程:进展、应用与挑战

Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Xiaoyu Wang, Henglin Liu, Chuang Li, Kecheng Jiao, Jixuan Ying, Yang Veronica Liu, Qiang Yang, Xiu Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了金融基础模型(FFMs)的进展、应用与挑战,涵盖三种关键模态,并探讨了数据可用性、算法可扩展性和基础设施限制等关键问题。

Comments Accepted by [J]. Engineering, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12821 2025-12-16 cs.LG cs.AI physics.data-an 57%

On the continuity of flows

关于流的连续性

Congzhou M Sha

机构 * Penn State College of Medicine(宾夕法尼亚州立大学医学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文研究了流匹配中因分布拓扑不匹配导致的连续性问题,揭示了最优速度场的跳跃不连续性现象,并探讨了其对流匹配方法的影响。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12773 2025-12-16 cs.HC cs.AI 57%

Designing The Drive: Enhancing User Experience through Adaptive Interfaces in Autonomous Vehicles

设计驱动:通过自适应界面提升自动驾驶车辆中的用户体验

Reeteesha Roy

机构 * VIT Bhopal University(维特理工学院博帕尔分校)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文探讨了通过自适应界面提升自动驾驶车辆用户体验的方法,强调透明性和用户控制在增强信任和满意度中的作用。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12793 2025-12-16 cs.CV 57%

ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution

ViCO: 一种面向语义感知动态高分辨率的训练策略

Long Cui, Weiyun Wang, Jie Shao, Zichen Wen, Gen Luo, Linfeng Zhang, Yanting Zhang, Yu Qiao, Wenhai Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Nanjing University(南京大学) Donghua University(东华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 ViCO通过动态调整视觉标记数量以适应图像语义复杂度,有效降低推理成本的同时保持模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09393 2025-12-11 cs.CV cs.LG 57%

Detection and Localization of Subdural Hematoma Using Deep Learning on Computed Tomography

利用深度学习进行CT扫描中硬脑膜下出血的检测与定位

Vasiliki Stoumpou, Rohan Kumar, Bernard Burman, Diego Ojeda, Tapan Mehta, Dimitris Bertsimas

机构 * Operations Research Center, Massachusetts Institute of Technology(麻省理工学院运营研究中心) Boston University(波士顿大学) Massachusetts Institute of Technology(麻省理工学院) University of Connecticut School of Medicine(康奈尔大学医学学院) Hartford HealthCare(哈特福德医疗集团) Sloan School of Management, Massachusetts Institute of Technology(麻省理工学院斯隆管理学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种多模态深度学习框架,结合临床数据与CT影像,实现硬脑膜下出血的快速准确检测与定位,提升神经外科急诊处理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13151 2025-12-11 cs.CV cs.GR 57%

Foveation Improves Payload Capacity in Steganography

注视点增强隐写术的载荷能力

Lifeng Qiu Lin, Henry Kam, Qi Sun, Kaan Akşit

机构 * University College London(伦敦大学学院) New York University(纽约大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 通过高效潜在表示和注视点渲染,提升隐写术的载荷能力至500位,并在高准确率和视觉质量上取得显著成果。

Comments SIGGRAPH Asia 2025 Posters Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08567 2025-12-10 cs.LG cs.AI 57%

A Hybrid Model for Stock Market Forecasting: Integrating News Sentiment and Time Series Data with Graph Neural Networks

一种股票市场预测的混合模型:整合新闻情感与时间序列数据与图神经网络

Nader Sadek, Mirette Moawad, Christina Naguib, Mariam Elzahaby

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种结合新闻情感与时间序列数据的混合模型,利用图神经网络提升股票市场预测性能,实验显示GNN在准确率和精度上均优于LSTM基线。

Comments 11 pages, 6 figures. Published in the Proceedings of the 5th International Conference on Artificial Intelligence Research (ICAIR 2025). Published version available at: https://papers.academic-conferences.org/index.php/icair/article/view/4294

Journal ref Proceedings of the 5th International Conference on AI Research (ICAIR 2025), Vol. 5, No. 1, pp. 452-462 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08262 2025-12-10 cs.CV cs.RO 57%

RLCNet: An end-to-end deep learning framework for simultaneous online calibration of LiDAR, RADAR, and Camera

RLCNet: 一种用于同时在线校准激光雷达、雷达和摄像头的端到端深度学习框架

Hafeez Husain Cholakkal, Stefano Arrigoni, Francesco Braghin

机构 * Department of Mechanical Engineering, Politecnico di Milano(机械工程系,米兰理工学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 RLCNet通过端到端深度学习框架实现激光雷达、雷达和摄像头的同时在线校准,提升自动驾驶系统在动态环境中的感知可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08221 2025-12-10 cs.CV 57%

VisKnow: Constructing Visual Knowledge Base for Object Understanding

VisKnow: 构建视觉知识库以实现物体理解

Ziwei Yao, Qiyang Wan, Ruiping Wang, Xilin Chen

机构 * Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室、计算技术研究所)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 VisKnow通过构建视觉知识库,利用多模态数据提升物体理解能力,包含AnimalKB案例研究,用于零样本识别和知识图谱完成等任务。

Comments 16 pages, 12 figures, 7 tables. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07733 2025-12-09 cs.CV 57%

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery

SpatialDreamer: 通过主动心理意象激励空间推理

Meng Cao, Xingyu Li, Xue Liu, Ian Reid, Xiaodan Liang

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·宾·扎耶德人工智能大学) Sun Yat-sen University(孙中山大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 SpatialDreamer通过主动心理意象和几何策略优化,提升多模态大语言模型在复杂空间推理任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07040 2025-12-09 cs.LG cs.CV 57%

Transformation of Biological Networks into Images via Semantic Cartography for Visual Interpretation and Scalable Deep Analysis

通过语义制图将生物网络转换为图像以实现可视化解释和可扩展深度分析

Sakib Mostafa, Lei Xing, Md. Tauhidul Islam

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 Graph2Image通过将生物网络转换为图像,实现可扩展的深度分析和可视化解释,提升分类准确率并揭示生物一致性模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01268 2025-12-09 cs.CL cs.LG 57%

Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Unveiling AI's Potential Through Tools, Techniques, and Applications

深度学习与机器学习:推动大数据分析与管理,通过工具、技术和应用揭示人工智能的潜力

Pohsun Feng, Ziqian Bi, Yizhu Wen, Xuanhe Pan, Benji Peng, Ming Liu, Jiawei Xu, Keyu Chen, Junyu Liu, Caitlyn Heqi Yin, Sen Zhang, Jinlang Wang, Qian Niu, Ming Li, Tianyang Wang, Xinyuan Song, Zekun Jiang

机构 * National Taiwan Normal University(国立台湾师范大学) Indiana University(印第安纳大学) University of Hawaii(夏威夷大学) Kyoto University(京都大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) Georgia Institute of Technology(佐治亚理工学院) Rutgers University(罗格斯大学) Purdue University(普渡大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Emory University(埃默里大学) Sichuan University(四川大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文探讨深度学习与机器学习在大数据分析中的应用,通过工具和技术揭示AI潜力,强调伦理与公平性,推动各领域创新。

Comments This book contains 155 pages and 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03494 2025-12-04 cs.CL cs.LG 57%

A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention

关于原生Top-k稀疏注意力的潜力与挑战的初步研究

Di Xiu, Hongyin Tang, Bolin Rong, Lizhi Yan, Jingang Wang, Yifan Lu, Xunliang Cai

机构 * Meituan, Beijing, China(美团,北京,中国)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究探讨了原生Top-k稀疏注意力机制在解码和训练中的有效性,通过实验验证其在提升模型性能方面的潜力,并从熵的角度解释了其在下游任务中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03233 2025-12-04 cs.CV 57%

Object Counting with GPT-4o and GPT-5: A Comparative Study

基于GPT-4o和GPT-5的对象计数:比较研究

Richard Füzesséry, Kaziwa Saleh, Sándor Szénási, Zoltán Vámossy

机构 * Software Engineering Institute, Obuda University, Budapest, Hungary(奥布达大学软件工程研究所) Doctoral School of Applied Informatics(应用信息科学博士学院) Applied Mathematics, Obuda University, Budapest, Hungary(应用数学,奥布达大学) John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary(约翰·冯·诺伊曼信息学院,奥布达大学) Faculty of Economics(经济学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文比较了GPT-4o和GPT-5在零样本对象计数任务中的性能,展示了其在FSC-147和CARPK数据集上的表现。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01262 2025-12-02 cs.SI cs.AI cs.ET cs.LG 57%

Social Media Data Mining of Human Behaviour during Bushfire Evacuation

社交媒体中人类在森林火灾疏散中的行为挖掘

Junfeng Wu, Xiangmin Zhou, Erica Kuligowski, Dhirendra Singh, Enrico Ronchi, Max Kinateder

机构 * RMIT University(皇家墨尔本理工大学) CSIRO(澳大利亚国家科学研究院) Lund University(吕勒奥大学) National Research Council Canada(加拿大国家研究理事会)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨了利用社交媒体数据挖掘森林火灾疏散行为的挑战与方法,提出未来应用及开放问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07820 2025-12-02 cs.AI cs.LG 57%

AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift

AI应更善于感知,而非仅仅更大:适应性感知作为范式转变

Eunsu Baek, Keondo Park, Jeonggil Ko, Min-hwan Oh, Taesik Gong, Hyung-Sin Kim

机构 * SNU(首尔国立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出适应性感知作为AI范式转变,通过动态调节传感器参数以提升效率和公平性,推动可持续且稳健的人工智能发展。

Comments Published in NeurIPS 2025 (Position Paper Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00366 2025-12-02 cs.LG cs.AI 57%

S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting

S^2-KD:语义-频谱知识蒸馏时空预测

Wenshuo Wang, Yaomin Shen, Yingjie Tan, Yihao Chen

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 S^2-KD通过结合语义与频谱知识蒸馏,提升时空预测模型的性能,使其在复杂场景中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15953 2025-11-27 cs.CV 57%

Activator: GLU Activation Function as the Core Component of a Vision Transformer

Activator: GLU激活函数作为视觉变换器的核心组件

Abdullah Nazhat Abdullah, Tarkan Aydin

机构 * Bahcesehir University(巴切塞希尔大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出使用GLU激活函数替代传统MLP和注意力机制,以提高视觉变换器的计算效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20627 2025-11-26 cs.AI 57%

Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems

用AI对抗AI:利用基础模型确保AI赋能的安全关键系统

Anastasia Mavridou, Divya Gopinath, Corina S. Păsăreanu

机构 * KBR Inc.(KBR公司) NASA Ames(美国国家航空航天局阿姆斯研究中心)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出利用AI技术解决安全关键系统中AI保证问题,通过REACT和SemaLens两个组件实现需求工程与感知系统验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17158 2025-11-24 physics.med-ph cs.CV 57%

Exploring the added value of pretherapeutic MR descriptors in predicting breast cancer pathologic complete response to neoadjuvant chemotherapy

探讨术前MRI描述符在预测乳腺癌新辅助化疗病理完全缓解中的附加价值

Caroline Malhaire, Fatine Selhane, Marie-Judith Saint-Martin, Vincent Cockenpot, Pia Akl, Enora Laas, Audrey Bellesoeur, Catherine Ala Eddine, Melodie Bereby-Kahane, Julie Manceau, Delphine Sebbag-Sfez, Jean-Yves Pierga, Fabien Reyal, Anne Vincent-Salomon, Herve Brisse, Frederique Frouin

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本研究探讨术前MRI特征在预测乳腺癌新辅助化疗病理完全缓解中的作用,发现非分叶边缘和单发性是独立预测因素,可提高预测模型性能。

Journal ref European Radiology, 2023, 33 (11), pp.8142-8154

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15375 2025-11-20 cs.LG cs.AI 57%

Parameter Importance-Driven Continual Learning for Foundation Models

Lingxiang Wang, Hainan Zhang, Zhiming Zheng

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15194 2025-11-20 cs.RO cs.AI 57%

Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization

Jian Deng, Yuandong Wang, Yangfu Zhu, Tao Feng, Tianyu Wo, Zhenzhou Shao

机构 * Beijing Key Laboratory of Light Industrial Robot and Safety Verification, Capital Normal University(北京工业机器人与安全验证重点实验室,首都师范大学) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Software, Beihang University(北航软件学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 12 pages, 4 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12001 2025-11-20 cs.CL cs.HC 57%

Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations

Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, Maarten Sap, Motahhare Eslami

机构 * Seoul National University(首尔国立大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

Comments Under review; 16 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13565 2025-11-18 cs.AI 57%

Artificial Intelligence-driven Intelligent Wearable Systems: A full-stack Integration from Material Design to Personalized Interaction

Jingyi Zhao, Daqian Shi, Zhengda Wang, Xiongfeng Tang, Yanguo Qin

机构 * The second hospital of Jilin University(吉林大学第二医院) QMUL Digital Environment Research Institute(QMUL数字环境研究机构) UCL Institute Of Health Informatics(UCL健康信息学研究所)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

Comments 5 pages, l figure, l table. Accepted at AI4RWC@WI-IAT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06266 2025-11-18 cs.CV 57%

Spatially-Aware Mixture of Experts with Log-Logistic Survival Modeling for Whole-Slide Images

Ardhendu Sekhar, Vasu Soni, Keshav Aske, Shivam Madnoorkar, Pranav Jeevan, Amit Sethi

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05955 2025-11-17 cs.CV cs.LG 57%

CSGaze: Context-aware Social Gaze Prediction

Surbhi Madan, Shreya Ghosh, Ramanathan Subramanian, Abhinav Dhall, Tom Gedeon

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) The University of Queensland(昆士兰大学) University of Canberra(堪培拉大学) Curtin University(Curtin大学) Monash University(莫纳什大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10914 2025-11-17 cs.CV 57%

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

Zihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu, Hua Zhang, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Mathematical Sciences, Fudan University(复旦大学数学学院) School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络科学与技术学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10890 2025-11-17 cs.AI stat.ML 57%

LLM enhanced graph inference for long-term disease progression modelling

Tiantian He, An Zhao, Elinor Thompson, Anna Schroder, Ahmed Abdulaal, Frederik Barkhof, Daniel C. Alexander

机构 * Department of Computer Science(计算机科学系)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10648 2025-11-14 cs.CV 57%

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

Jiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou, Lewei Lu, Houqiang Li, Xiaohua Wang, Jinguo Zhu

机构 * Xi’an Jiaotong University(西安交通大学) University of Science and Technology of China(中国科学技术大学) SenseTime Research(商汤科技研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏