arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2606.15284 2026-06-16 eess.SP cs.AI cs.LG 新提交 70%

CAP: Towards PPG Universal Representation Learning with Patient-level Supervision

CAP:面向患者级监督的PPG通用表示学习

Chenyang He, Xinyi Shao, Shun Huang, Bosong Huang, Daoqiang Zhang, Ming Jing, Cheng Ding

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Peking University(北京大学) Independent Researcher(独立研究者) Jinling Clinical Medical College College of Artificial Intelligence Nanjing University of Aeronautics and Astronautics(金陵临床医学院人工智能学院南京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出CAP方法,通过构建大规模PPG-EHR多模态数据集和跨模态对比对齐,学习患者级临床语义的PPG表示,在四项下游任务中平均提升26.7%,呼吸率预测提升87.6%。

Comments Accepted as an Oral presentation at KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15104 2026-06-16 cs.CV 新提交 70%

Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

红外与可见光图像的文本驱动融合:在双曲空间实现图像场景自适应

Huan Kang, Hui Li, Tianyang Xu, Tao Zhou, Xiao-Jun Wu, Josef Kittler

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出一种文本驱动的红外与可见光图像融合框架,利用双曲流形学习嵌入层次语义,通过BLIP文本提示引导视觉-属性对齐,实现无文本输入的自适应融合,性能优于现有方法。

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10005 2026-06-16 cs.CV 版本更新 70%

TUNI: Unifying Pre-training and Fine-tuning with Modality-Aware Mutual Learning and Rectification for RGB-T Semantic Segmentation

TUNI:基于模态感知互学习和矫正的RGB-T语义分割统一预训练与微调框架

Xiaodong Guo, Xianda Guo, Tong Liu, Zhihong Deng, Yanlun Peng, Xiang Li, Wujie Zhou

机构 * School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学) School of Computer Science, Wuhan University(计算机学院,武汉大学) Great Wall Motor(长城汽车) School of Information and Electronic Engineering, Zhejiang University of Science and Technology(信息电子工程学院,浙江理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出TUNI框架,通过模态感知互学习与矫正统一预训练和微调,解决RGB-T语义分割中多模态特征提取融合、模态依赖不平衡及热信息利用不足问题,在五个数据集上优于15种SOTA模型。

Comments This paper is an extended version of the authors' work previously presented at the ICRA conference. To appear in IEEE Transactions on Circuits and Systems for Video Technology. DOl: 10.1109/TCSVT.2026.3701706

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11269 2026-06-11 cs.CV cs.HC 新提交 70%

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment

特质更深:面向人格评估的特质特异性非对称融合

Jia Li, Qian Chen, Wei Wang, Xinyu Li, Zhenzhen Hu, Dongsheng Shao, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Intelligent Interconnected Systems Laboratory of Anhui Province(安徽省智能互联系统实验室) Jianghuai Advanced Technology Center(江淮前沿技术中心) Anhui Provincial Industry Innovation Center of Humanoid Robots(安徽省人形机器人产业创新中心) Anhui Provincial Key Laboratory of Humanoid Robots(安徽省人形机器人重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出Traits Run Deeper框架,通过多模态基础表示、特质特异性非对称融合和分布校准回归模块,解决人格评估中模态偏好差异和标签偏差问题,在AVI Challenge 2026上MSE降低约25%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09347 2026-06-11 cs.CV 版本更新 70%

IB-HFN: Information Bottleneck-Driven SAR-Optical Fusion Network for High-Fidelity Cloud Removal

IB-HFN: 信息瓶颈驱动的SAR-光学融合网络用于高保真云去除

Haojun Guo, Fan Feng, Ziquan Wang, Yongsheng Zhang, Ying Yu

机构 * Institute of Geospatial Information, Information Engineering University(测绘信息研究院,信息工程大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出IB-HFN网络,通过双流骨干、空间信息瓶颈融合模块和联合优化策略,抑制SAR散斑噪声并保留光学细节,实现高保真云去除。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26615 2026-06-08 cs.CL 版本更新 70%

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

SlideAgent:用于多页视觉文档理解的分层代理框架

Yiqiao Jin, Rachneet Kaur, Zhen Zeng, Sumitra Ganesh, Srijan Kumar

机构 * Georgia Institute of Technology(佐治亚理工学院) J.P. Morgan AI Research(摩根大通AI研究)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CL

AI总结 提出SlideAgent,一种用于多模态多页文档(如幻灯片)理解的分层代理框架,通过全局、页面和元素三级推理构建结构化表示,在专有和开源模型上分别提升7.9%和9.8%的准确率。

Comments ACL 2026 Main Conference. https://slideagent.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05455 2026-06-05 cs.CV 70%

Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

面向不完整图像-表格分类的解缠细粒度原型学习

Feixiang Zhou, Jianyang Xie, Zhuangzhi Gao, Qinkai Yu, Fu Wang, Yuheng Fan, Jing Li, Zheheng Jiang, Yitian Zhao, Yanda Meng, He Zhao, Gregory Y. H. Lip, Yalin Zheng

机构 * School of Eye and Vision Sciences, University of Liverpool, U.K.(利物浦大学眼科与视觉科学学院) Department of Cardiovascular and Metabolic Medicine, University of Liverpool, U.K.(利物浦大学心血管与代谢医学系) School of Computer Science, University of Exeter, U.K.(埃克塞特大学计算机科学学院) School of Computer Science and Engineering, South China University of Technology, China(华南理工大学计算机科学与工程学院) School of Computing and Mathematical Sciences, University of Leicester, U.K.(莱斯特大学计算科学与数学科学学院) Ningbo Institute of Industrial Technology, Chinese Academy of Sciences, China(中国科学院宁波工业技术研究所) Bioengineering Program, Biological and Environmental Science and Engineering Division (BESE), King Abdullah University of Science and Technology (KAUST), Saudi Arabia(卡尔斯塔德大学科学与技术学院(KAUST)生物工程项目,沙特阿拉伯)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 针对图像-表格多模态学习中缺失模态问题,提出DFPL框架,通过共享-特定原型建模、原型级解缠和细粒度对齐,实现鲁棒分类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05232 2026-06-05 cs.LG cs.AI 70%

Differentiable Efficient Operator Search

可微分高效算子搜索

Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo, Weiguo Feng, Tao Huang, Cho-Jui Hsieh, Chang Xu

机构 * The University of Sydney(悉尼大学) ByteDance(字节跳动) Shanghai Jiao Tong University(上海交通大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 提出可微分高效算子搜索框架,统一解释多种token缩减算子,通过联合搜索缩减位置、保留数量和算子行为,在预算约束下优化多模态模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03581 2026-06-03 cs.CV cs.RO 70%

UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion

UnsOcc:非结构化场景下基于渲染融合的3D语义占用预测

Ye Wu, Ruiqi Song, Baiyong Ding, Nanxin Zeng, Junjie Cheng, Yunfeng Ai

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Waytous Inc.(Waytous公司)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出UnsOcc多模态框架,通过渲染融合模块和基于高斯溅射的细节感知辅助监督,解决非结构化场景中跨模态融合困难与长尾分布问题,在露天矿和nuScenes数据集上超越现有方法。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00471 2026-06-02 cs.CV 70%

MUSCLE-NET: Predicted-Multiscale-Aware Network for Pedestrian Trajectory Forecasting

MUSCLE-NET:面向行人轨迹预测的预测多尺度感知网络

Yu Liu, Ming Huang, Xiao Ren, Zhijie Liu, Youfu Li, He Kong

机构 * Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology, School of Automation and Intelligent Manufacturing, Southern University of Science and Technology (SUSTech), Shenzhen(广东省全主动系统控制理论与技术重点实验室,自动化与智能制造学院,南方科技大学(SUSTech),深圳) Department of Mechanical Engineering, City University of Hong Kong, Hong Kong SAR, China(香港城市大学机械工程系,香港特别行政区,中国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出MUSCLE-NET,通过多尺度多模态特征提取和尺度自适应预测机制,解决现有方法对观测信息利用不足及忽视未来运动尺度依赖的问题,在JAAD和PIE数据集上取得竞争性能。

Comments This manuscript has been accepted to the IEEE Transactions on Intelligent Transportation Systems as a regular paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21007 2026-06-01 cs.CV cs.RO 70%

LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation

LiteViLNet: 轻量级视觉-激光雷达融合网络用于高效道路分割

Daojie Peng, Bingtao Wang, Fulong Ma, Liang Zhang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Shandong University(山东大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出轻量级多模态网络LiteViLNet,通过双流编码器、深度可分离卷积和多尺度特征融合模块,在KITTI数据集上以14.04M参数达到96.36% MaxF,实现精度与效率的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26273 2026-05-27 cs.CV 70%

Frequency-Guided Fusion For RGB-Thermal Semantic Segmentation

频率引导的RGB-热红外语义分割融合

İsmail Emre Canıtez, Özgür Erkent

机构 * Hacettepe University(哈切特佩大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出一种基于双ConvNeXt V2骨干网络的多模态融合架构,通过频率分解和置信门控残差机制融合RGB与热红外特征,在MFNet和PST900上以较低参数量实现先进性能。

Comments 9 pages, 7 figures, To be Presented at Perception Beyond the Visible Spectrum workshop series (IEEE PBVS) at CVPR, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24532 2026-05-26 cs.CV 70%

Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation

图像条件实例提示网络用于遥感图像指代分割

Biaoyu Ren, Qingsheng Wang, Cun Xu, Dingkang Yang, Wenxuan Wang

机构 * School of Computer Science, Northwestern Polytechnical University, Xi'an, China(西北工业大学计算机科学学院,西安,中国) College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China(复旦大学智能机器人与先进制造学院,上海,中国) Shenzhen Research Institute of Northwestern Polytechnical University, Shenzhen, China(西北工业大学深圳研究院,深圳,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 提出图像条件实例提示网络(ICIPNet),通过自适应视觉语义表示和双边信息融合模块,缓解跨模态特征融合瓶颈,提升遥感图像指代分割性能。

Comments 6 pages, 3 figures. Equal contribution: Biaoyu Ren and Qingsheng Wang. Corresponding authors: Dingkang Yang and Wenxuan Wang

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01328 2026-05-26 cs.CV 70%

CSFNet: A Cosine Similarity Fusion Network for Real-Time RGB-X Semantic Segmentation of Driving Scenes

CSFNet: 用于驾驶场景实时RGB-X语义分割的余弦相似度融合网络

Danial Qashqai, Emad Mousavian, Shahriar Baradaran Shokouhi, Sattar Mirzakuchaki

机构 * Department of Electrical Engineering, Iran University of Science and Technology(伊朗科学技术大学电气工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出CSFNet,通过余弦相似度注意力融合模块(CS-AFM)高效融合双模态特征,实现实时且高精度的RGB-X语义分割。

Journal ref Engineering Applications of Artificial Intelligence, 174, 114362 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16923 2026-05-22 cs.CV 70%

Neuroscience-inspired Staged Representation Learning with Disentangled Coarse- and Fine-Grained Semantics for EEG Visual Decoding

受神经科学启发的分阶段表征学习:解纠缠的粗粒度和细粒度语义用于EEG视觉解码

Xiang Gao, Hui Tian, Yanming Zhu, Xuefei Yin, Alan Wee-Chung Liew

机构 * School of Information and Communication Technology, Griffith University(信息与通信技术学院,格里菲斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种受神经科学启发的分阶段表征学习框架,通过解纠缠的粗粒度和细粒度语义来改进EEG视觉解码,解决了现有方法在人类视觉处理分阶段和层次特性方面的不足。

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19660 2026-05-20 cs.LG cs.CL 70%

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond

OScaR:LLMs及更广泛场景中的极压缩KV缓存量化之奥卡姆之刀

Zunhai Su, Rui Yang, Chao Zhang, Yaxiu Liu, Yifan Zhang, Wei Wu, Jing Xiong, Dayou Du, Xialie Zhuang, Yulei Qian, Yuchen Xie, Yik-Chung Wu, Hongxia Yang, Ngai Wong

机构 * Tsinghua University(清华大学) Meituan LongCat Team(美团LongCat团队) The University of Hong Kong(香港大学) The University of Edinburgh(爱丁堡大学) UCAS(中国科学技术大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);omni-modal(abstract);分类 cs.CL

AI总结 本文针对LLMs中KV缓存极压缩时的量化保真问题,提出OScaR框架,通过Canalized Rotation和Omni-Token Scaling有效缓解Token Norm Imbalance,实现近无损的INT2量化性能,同时提升解码速度和吞吐量。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17894 2026-05-19 cs.AI 70%

Evaluating Cognitive Age Alignment in Interactive AI Agents

评估交互式AI代理的认知年龄对齐

Yifan Shen, Jiawen Zhang, Jian Xu, Junho Kim, Ismini Lourentzou, Xu Cao, Meihuan Huang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Shenzhen Children's Hospital(深圳儿童医院) Peking University(北京大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出ChildAgentEval,首个基于心理测量的交互式基准,用于评估基于多模态大语言模型的代理的认知年龄对齐,通过与年龄特定的人类发展阶段进行系统比较,揭示当前代理在模拟年龄特定认知行为方面的优劣。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17566 2026-05-19 cs.CV 70%

Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework

重新思考点云作为序列:一种因果性下一标记预测学习框架

Yumeng Yao, Jingzhi Dong, Haowen Gu, Tao Chen, Zonghan Wu, Xiaoshui Huang, Yazhou Yao

机构 * Nanjing University of Science and Technology(南京理工大学) Shanghai Jiao Tong University(上海交通大学) Hangzhou City University(杭州城市学院) East China Normal University(华东师范大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 本文提出PointNTP,将点云预训练重新定义为全因果、无解码器的潜在下一标记预测问题,通过局部补丁分割和结构化3D标记序列生成,实现对点云结构依赖的直接建模,无需重建解码器或显式几何恢复,实验表明其在多个下游任务中表现优异。

Comments 10 pages, 2 figures. Code will be released upon acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12056 2026-05-13 cs.AI 70%

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

OmniRefine: 一种面向对齐的协作压缩方法以提高多模态大语言模型的效率

Yuchen Deng, Zidang Cai, Hai-Tao Zheng, Jie Wang, Feidiao Yang, Yuxing Han

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);audio-visual(abstract);分类 cs.AI

AI总结 本文提出OmniRefine,一种无需训练的两阶段框架,通过跨模态对齐和协作压缩提升多模态大语言模型的推理效率与性能稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10676 2026-05-12 cs.CV cs.LG 70%

Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium

非盲目而是被沉默:通过对抗性反常识平衡视觉与语言

Qingxin Xiao, Peilin Zhao, Yangyang Zhao, Lingwei Dang, Qingyao Wu

机构 * South China University of Technology(华南理工大学) Institute for Super Robotics (Huangpu)(机器人研究所(黄埔)) Shanghai Jiao Tong University(上海交通大学) Changsha University of Science and Technology(长沙理工大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出ACE框架,通过对抗性反常识补丁扰动视觉上下文,动态平衡语言先验与视觉信息,提升模型可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05709 2026-05-08 cs.AI 70%

Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs

隐藏、重建、突破:在大规模语言模型中利用重建-隐藏权衡

Md Farhamdur Reza, Richeng Jin, Tianfu Wu, Huaiyu Dai

机构 * NC State University(北卡罗来纳州立大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract_cn);分类 cs.AI

AI总结 本文探讨了在多模态大语言模型中利用重建与隐藏的权衡进行意图混淆攻击,提出了一种基于字符移除的变体构造方法,并引入关键词相关的干扰图像以提高攻击效果。

Comments 39 pages, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00930 2026-05-05 q-bio.GN cs.AI 70%

CellxPert: Inference-Time MCMC Steering of a Multi-Omics Single-Cell Foundation Model for In-Silico Perturbation

CellxPert:一种多组学单细胞基础模型的推理时间MCMC引导方法用于计算机模拟扰动

Andac Demir, Erik W. Anderson, Jeremy L. Jenkins, Srayanta Mukherjee

机构 * Novartis Biomedical Research(诺华生物医学研究)

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 CellxPert通过整合多组学数据,实现了对单细胞和空间多组学的统一表示,支持细胞类型注释、扰动响应预测和多组学整合,采用MCMC方法提升生物可解释性。

Journal ref ICLR Machine Learning for Genomics Explorations Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08156 2026-05-04 cs.CV 70%

LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping

LandSegmenter:面向土地利用与覆盖制图的灵活基础模型

Chenying Liu, Wei Huang, Xiao Xiang Zhu

机构 * Chair of Data Science in Earth Observation, Technical University of Munich(地球观测数据科学教授职位,慕尼黑技术大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML))

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出LandSegmenter框架,解决土地利用与覆盖制图中数据需求高、模型泛化性差的问题,通过多模态数据集、跨模态特征提取和置信度引导融合策略提升模型性能。

Comments Accepted by ISPRS for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27507 2026-04-28 cs.CV 70%

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

Chat-Scene++: 利用语境丰富的物体识别用于3D大语言模型

Haifeng Huang, Yilun Chen, Zehan Wang, Jiangmiao Pang, Zhou Zhao

机构 * School of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出Chat-Scene++框架,通过将3D场景表示为语境丰富的物体序列,提升细粒度物体定位和上下文推理能力,在五个3D视觉语言基准测试中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20769 2026-04-24 cs.CV 70%

Information Bottleneck-Guided Heterogeneous Graph Learning for Interpretable Neurodevelopmental Disorder Diagnosis

信息瓶颈引导的异构图学习用于可解释性神经发育障碍诊断

Yueyang Li, Lei Chen, Wenhao Dong, Shengyu Gong, Zijian Kang, Boyang Wei, Weiming Zeng, Hongjie Yan, Lingbin Bian, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang

机构 * Department of Language Science and Technology, The Hong Kong Polytechnic University(香港理工大学语言科学与技术系) Laboratory of Digital Image and Intelligent Computation, Shanghai Maritime University(上海 Maritime 大学数字图像与智能计算实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出I2B-HGNN框架,结合信息瓶颈原理指导脑连接建模与跨模态特征融合,实现高分类准确率和可解释性生物标志物识别。

Journal ref Neurocomputing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18713 2026-04-22 cs.CV 70%

Align then Refine: Text-Guided 3D Prostate Lesion Segmentation

对齐后再细化:基于文本的3D前列腺病变分割

Cuiling Sun, Linkai Peng, Adam Murphy, Elif Keles, Hiten D. Patel, Ashley Ross, Frank Miller, Baris Turkbey, Andrea Mia Bejar, Halil Ertugrul Aktas, Gorkem Durak, Ulas Bagci

机构 * Department of Radiology, Northwestern University, Chicago, USA(放射科,西北大学,芝加哥,美国) Department of Urology, Northwestern University, Chicago, USA(泌尿科,西北大学,芝加哥,美国) Center for Cancer Research, National Cancer Institute, Bethesda, USA(癌症研究中心,国家癌症研究所,贝塞斯达,美国)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种多编码器U-Net架构,通过引入对齐损失、热图损失和置信度门控多头交叉注意力细化模块,提升前列腺病变分割的精度和多模态融合能力。

Comments Accepted to EMBC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16264 2026-04-20 cs.CV cs.LG 70%

Information Router for Mitigating Modality Dominance in Vision-Language Models

信息路由器用于缓解视觉-语言模型中的模态主导问题

Seulgi Kim, Mohit Prabhushankar, Ghassan AlRegib

机构 * OLIVES at the Center for Signal Information Processing CSIP, School of Electrical

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MoIR信息路由器,通过在融合前减少信息差异,缓解视觉-语言模型中模态主导问题,提升鲁棒性和下游性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14520 2026-04-17 cs.CV 70%

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

模态链:从静态融合到动态编排在多模态大语言模型中

Ziyang Luo, Nian Liu, Junwei Han

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);omni-modal(abstract);分类 cs.CV

AI总结 本文提出CoM框架,通过动态编排解决多模态融合的静态融合问题,提升模型在不同任务中的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16333 2026-04-14 cs.CV cs.LG 70%

RL makes MLLMs see better than SFT

强化学习使多模态大语言模型比监督微调看得更清楚

Junha Song, Sangdoo Yun, Dongyoon Han, Jaegul Choo, Byeongho Heo

机构 * NAVER AI Lab(NAVER AI实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文研究了强化学习在多模态大语言模型中的视觉编码器影响,发现RL能产生更精确的视觉表示,提出PIVOT方法提升视觉编码性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00748 2026-04-14 cs.CV 70%

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

通过强化学习提升多图像接地在MLLMs中的推理能力

Bob Zhang, Haoran Li, Tao Zhang, Jianan Li, Cilin Yan, Xikai Liu, Jiayin Cai, Yanbin Hao

机构 * Xiaohongshu Inc.(小红书公司) University of Science and Technology of China(中国科学技术大学) Wuhan University(武汉大学) Technical University of Munich(慕尼黑工业大学) Hefei University of Technology(合肥工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文通过强化学习策略提升多图像接地任务中MLLMs的推理能力,采用合成CoT数据、LoRA微调和规则引导的RL方法,实验表明在MIG-Bench和多个领域外基准上效果显著。

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏