arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

2602.14788 2026-02-17 cs.CV cs.AI 62%

VIPA: Visual Informative Part Attention for Referring Image Segmentation

VIPA: 视觉信息部分注意力用于指认图像分割

Yubin Cho, Hyunwoo Yu, Kyeongbo Kong, Kyomin Sohn, Bongjoon Hyun, Suk-Ju Kang

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 VIPA通过视觉信息部分注意力和视觉表达生成器提升指认图像分割的细粒度分割性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13264 2026-02-17 cs.LG cs.AI cs.CL 62%

Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models

方向性集中不确定性:一种代表方法用于生成模型的不确定性量化

Souradeep Chattopadhyay, Brendan Kennedy, Sai Munikoti, Soumik Sarkar, Karl Pazdernik

机构 * Department of Mechanical Engineering, Iowa State University, Ames, IA, USA(机械工程系,爱荷华州立大学) Pacific Northwest National Laboratory, Richland, WA, USA(太平洋西北国家实验室) Department of Statistics, North Carolina State University, Raleigh, NC, USA(统计系,北卡罗来纳州立大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出方向性集中不确定性(DCU)方法,通过基于vMF分布的嵌入集中度量化,提升生成模型的不确定性量化性能,并在多模态任务中展现良好泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01034 2026-02-03 cs.AI cs.CL 62%

Discovering Process-Outcome Credit in Multi-Step LLM Reasoning

在多步骤LLM推理中发现过程-结果信用

Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Saman Halgamuge

机构 * The University of Melbourne(墨尔本大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的框架,通过分步边际信息增益机制和解耦掩码策略,提升多步骤LLM推理的样本效率和准确性,并增强模型的分布外鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02820 2026-01-30 cs.CV cs.AI cs.GR 62%

Mesh Neural Cellular Automata

网格神经元细胞自动机

Ehsan Pajouheshgar, Yitao Xu, Alexander Mordvintsev, Eyvind Niklasson, Tong Zhang, Sabine Süsstrunk

机构 * EPFL(瑞士联邦理工学院) Google Research(谷歌研究)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 MeshNCA是一种无需UV映射即可实时生成高质量3D动态纹理的神经元细胞自动机方法,通过多模态监督和用户交互实现了纹理合成的泛化能力。

Comments ACM Transactions on Graphics (TOG) - SIGGRAPH 2024

Journal ref ACM Transactions on Graphics (TOG), Volume 43, Issue 4 Article No.: 122, Pages 1 - 16; 19 July 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15347 2026-01-23 cs.AI cs.CL cs.LG 62%

Logic Programming on Knowledge Graph Networks And its Application in Medical Domain

知识图谱网络上的逻辑编程及其在医疗领域的应用

Chuanqing Wang, Zhenmin Zhao, Shanshan Du, Chaoqun Fei, Songmao Zhang, Ruqian Lu

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出知识图谱网络的系统理论与技术,探讨其在医疗领域的应用,通过多条件下的实验验证创新方法。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08226 2026-01-14 cs.CV cs.AI 62%

Knowledge-based learning in Text-RAG and Image-RAG

基于知识的学习在Text-RAG和Image-RAG中的应用

Alexander Shim, Khalil Saieh, Samuel Clarke

机构 * Florida International University(佛罗里达国际大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本研究通过比较基于文本和图像的RAG方法,探讨了如何利用外部知识减少幻觉问题并提升胸部X光图像疾病检测的准确性。

Comments 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17607 2026-01-06 cs.CV cs.CL cs.LG 62%

Robustness of Structured Data Extraction from Perspectively Distorted Documents

从透视变形文档中提取结构化数据的鲁棒性

Hyakka Nakada, Yoshiyasu Tanaka

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

AI总结 本研究探讨了透视变形对多模态LLMs提取文档数据准确性的影响,发现结构识别准确性显著下降,但可通过旋转校正提升。

Comments 8 pages, 12 figures

Journal ref 2025 10th International Conference on Intelligent Informatics and Biomedical Sciences (ICIIBMS), Okinawa, Japan, 2025, pp. 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14576 2025-12-17 cs.CL cs.AI 62%

Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies

低资源、高影响:为包容性语言技术构建语料库

Ekaterina Artemova, Laurie Burchell, Daryna Dementieva, Shu Okabe, Mariya Shmatova, Pedro Ortiz Suarez

机构 * Toloka AI

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本教程提供构建包容性语言技术的实用工具和方法,涵盖多语言和低资源语言的端到端NLP流水线建设。

Comments Tutorial is accepted to LREC2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04187 2025-12-05 cs.CV cs.AI 62%

OnSight Pathology: A real-time platform-agnostic computational pathology companion for histopathology

OnSight病理科:一种实时、平台无关的计算病理科辅助工具,用于组织病理学

Jinzhen Hu, Kevin Faust, Parsa Babaei Zadeh, Adrienn Bourkas, Shane Eaton, Andrew Young, Anzar Alvi, Dimitrios George Oreopoulos, Ameesha Paliwal, Assem Saleh Alrumeh, Evelyn Rose Kamski-Hennekam, Phedias Diamandis

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 OnSight病理科是一种实时、平台无关的计算病理科工具,通过本地运行的AI模型提供实时诊断支持,适用于多种病理学流程和场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00019 2025-12-02 cs.RO cs.AI cs.CV 62%

A Comprehensive Survey on Surgical Digital Twin

外科数字孪生的全面综述

Afsah Sharaf Khan, Falong Fan, Doohwan DH Kim, Abdurrahman Alshareef, Dong Chen, Justin Kim, Ernest Carter, Bo Liu, Jerzy W. Rozenblit, Bernard Zeigler

机构 * IEEE

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了外科数字孪生的技术现状与挑战,提出分类方法并识别了验证、安全性和数据治理等开放问题,旨在推动其在临床中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14341 2025-11-19 cs.RO cs.AI cs.CV 62%

Going Places: Place Recognition in Artificial and Natural Systems

Michael Milford, Tobias Fischer

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref Annual Review of Control, Robotics, and Autonomous Systems 2026, vol. 9

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12691 2025-11-18 cs.CV cs.AI 62%

R$^{2}$Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection

Shuaike Shen, Ke Liu, Jiaqing Xie, Shangde Gao, Chunhua Shen, Ge Liu, Mireia Crispin-Ortuzar, Shangqi Gao

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Cambridge(剑桥大学) Zhejiang University(浙江大学) ETH Zurich(苏黎世联邦理工学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02720 2025-11-05 cs.CV cs.AI 62%

LLEXICORP: End-user Explainability of Convolutional Neural Networks

Vojtěch Kůr, Adam Bajger, Adam Kukučka, Marek Hradil, Vít Musil, Tomáš Brázdil

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00411 2025-11-04 cs.LG cs.AI cs.CV 62%

Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling

Zenghao Niu, Weicheng Xie, Siyang Song, Zitong Yu, Feng Liu, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University, China(深圳大学计算机科学与软件工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China(广东省人工智能与数字经济发展实验室(深圳)) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University, China(广东省智能信息处理省级重点实验室) School of Computer Science, University of Exeter, U.K.(埃克塞特大学计算机科学学院) Department of Computing and Information Technology, Great Bay University, China(大鹏大学计算与信息技术系) Computer Vision Institute, School of Artificial Intelligence, Shenzhen University, China(人工智能学院计算机视觉研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments accepted by iccv 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22729 2025-10-28 cs.AI cs.CL 62%

Critical Insights into Leading Conversational AI Models

Urja Kohli, Aditi Singh, Arun Sharma

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 7 tables, 3 figures. Open-access preprint intended for journal or conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13563 2025-10-14 cs.LG cs.AI cs.CV 62%

Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression

Xiaohui Wang, Peng Ye, Chenyu Huang, Shenghe Zheng, Bo Zhang, Lei Bai, Wanli Ouyang, Tao Chen

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10293 2025-10-14 cs.CL cs.AI 62%

MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning

Hongwei Chen, Yishu Lei, Dan Zhang, Bo Ke, Danxiang Zhu, Xuyi Chen, Yuxiang Lu, Zhengjie Huang, Shikun Feng, Jingzhou He, Yu Sun, Hua Wu, Haifeng Wang

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10138 2025-10-14 cs.CL cs.AI 62%

Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task

Zilong Wang, Xiaoyu Shen

机构 * Ningbo Institute of Digital Twin(宁波数字孪生研究所) Eastern Institute of Technology(东部技术研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03349 2025-10-07 cs.LG cs.AI cs.CL physics.ao-ph 62%

AgentCaster: Reasoning-Guided Tornado Forecasting

Michael Chen

机构 * Department of Computing + Mathematical Sciences(计算与数学科学系) California Institute of Technology(加州理工学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25941 2025-10-01 cs.AI cs.CL 62%

Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA

Raphael Schumann, Stefan Riezler

机构 * Heidelberg University(海德堡大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19657 2025-09-25 cs.CL cs.AI cs.SI 62%

Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections

Yicheng Yang, Zixian Li, Jean Paul Bizimana, Niaz Zafri, Yongfeng Dong, Tianyi Li

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18141 2025-09-24 cs.LG cs.AI cs.CV stat.AP stat.ML 62%

KM-GPT: An Automated Pipeline for Reconstructing Individual Patient Data from Kaplan-Meier Plots

Yao Zhao, Haoyue Sun, Yantian Ding, Yanxun Xu

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16769 2025-09-23 cs.LG cs.AI cs.CL 62%

Geometric Mixture Classifier (GMC): A Discriminative Per-Class Mixture of Hyperplanes

Prasanth K K, Shubham Sharma

机构 * The Nilgiris, Tamil Nadu, India - 643005(印度泰米尔纳德邦尼尔吉里斯) SunitechAI

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10831 2025-09-23 cs.HC cs.AI cs.CL 62%

Creating General User Models from Computer Use

Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park, Diyi Yang, Michael S. Bernstein

机构 * Stanford University(斯坦福大学) Microsoft Research(微软研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 23 pages, 6 figures, 2 tables; see https://generalusermodels.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09869 2025-09-15 cs.CV cs.AI 62%

Surrogate Supervision for Robust and Generalizable Deformable Image Registration

Yihao Liu, Junyu Chen, Lianrui Zuo, Shuwen Wei, Brian D. Boyd, Carmen Andreescu, Olusola Ajilore, Warren D. Taylor, Aaron Carass, Bennett A. Landman

机构 * Department of Electrical and Computer Engineering, Vanderbilt University(维斯尼尔大学电气与计算机工程系) Department of Radiology and Radiological Science, Johns Hopkins Medical School(约翰霍普金斯医学学校放射学与放射科学系) Image Analysis and Communications Laboratory in the Department of Electrical and Computer Engineering, Johns Hopkins University(约翰霍普金斯大学电气与计算机工程系图像分析与通信实验室) Center for Cognitive Medicine, Department of Psychiatry and Behavioral Science, Vanderbilt University Medical Center(维斯尼尔大学医学中心认知医学中心) University of Pittsburgh, School of Medicine(匹兹堡大学医学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05034 2025-09-08 cs.CV cs.AI 62%

Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization

Jingqi Wu, Hanxi Li, Lin Yuanbo Wu, Hao Chen, Deyin Liu, Peng Wang

机构 * Jiangxi Normal University(江西师范大学) Southern University of Science and Technology(南方科技大学) Swansea University(斯旺西大学) Zhejiang University(浙江大学) Anhui University(安徽大学) Northwestern Polytechnical University(西北工业大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20758 2025-08-29 cs.CV cs.AI 62%

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan, Yachao Zhang, Yuan Xie, Yanyun Qu

机构 * School of Informatics, Xiamen University(厦门大学信息学院) School of Computer Science, Nanjing University(南京大学计算机科学学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02549 2025-08-29 cs.CV cs.AI 62%

Federated nnU-Net for Privacy-Preserving Medical Image Segmentation

Grzegorz Skorupko, Fotios Avgoustidis, Carlos Martín-Isla, Lidia Garrucho, Dimitri A. Kessler, Esmeralda Ruiz Pujadas, Oliver Díaz, Maciej Bobowicz, Katarzyna Gwoździewicz, Xavier Bargalló, Paulius Jaruševičius, Richard Osuala, Kaisar Kushibar, Karim Lekadir

机构 * Artificial Intelligence in Medicine Laboratory (BCN-AIM)(人工智能医学实验室(BCN-AIM)) Departament de Matemàtiques i Informàtica(数学与计算机科学系) Universitat de Barcelona(巴塞罗那大学) Medical University of Gdańsk (GUMed)(格但斯克医学大学(GUMed)) Hospital Clínic de Barcelona (HCB)(巴塞罗那医院(HCB)) Lithuanian University of Health Sciences(立陶宛卫生科学大学) Institució Catala(加泰罗尼亚机构)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments In review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00451 2025-08-26 cs.CL cs.AI 62%

Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities

Aishik Mandal, Tanmoy Chakraborty, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab) Department of Computer Science and Hessian Center for AI (hessian.AI) Technische Universität Darmstadt National Research Center for Applied Cybersecurity ATHENE, Germany(技术大学达姆施塔特应用网络安全国家研究中心、海森国家人工智能中心(hessian.AI)、计算机科学系、无处不在知识处理实验室(UKP Lab)) Department of Electrical Engineering, Indian Institute of Technology Delhi, India(印度德里印度理工学院电气工程系) Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi, India(印度德里印度理工学院人工智能学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 18 pages, 2 figures, Accepted in Nature Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15752 2025-08-22 cs.HC cs.AI cs.CV 62%

"Does the cafe entrance look accessible? Where is the door?" Towards Geospatial AI Agents for Visual Inquiries

Jon E. Froehlich, Jared Hwang, Zeyu Wang, John S. O'Meara, Xia Su, William Huang, Yang Zhang, Alex Fiannaca, Philip Nelson, Shaun Kane

机构 * University of Washington(华盛顿大学) Google Research(谷歌研究) UCLA(加州大学洛杉矶分校) Google DeepMind(谷歌DeepMind)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to the ICCV'25 Workshop "Vision Foundation Models and Generative AI for Accessibility: Challenges and Opportunities"

详情

展开后加载摘要…

URL PDF HTML 收藏