arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-07-23 至 2026-07-23 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2607.20357 2026-07-23 cs.CV 新提交 83%

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

少看,快思考:多模态大语言模型的联合令牌-计算适配

Pengcheng Wang, Zhiquan Wang, Jayoung Lee, Zhuoyan Xu, Ran Xu, Saurabh Bagchi, Yin Li, Somali Chaterji

机构 * Purdue University(普渡大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 针对多模态大语言模型推理成本高的问题,提出SmartVL框架,通过视觉侧令牌控制器和LLM侧计算控制器联合控制视觉令牌数量和模型计算能力,实验证明该框架优于先前方法,实现更好的精度-效率平衡。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19600 2026-07-23 stat.ME cs.CV cs.LG q-bio.QM stat.ML 新提交 79%

Deep Shape Regression for Planar Curves with Multimodal Covariates

具有多模态协变量的平面曲线深度形状回归

Manuel Pfeuffer, Roshan Prakash Rane, Hadya Yassin, Kerstin Ritter, Sonja Greven

机构 * Humboldt-Universität zu Berlin, Berlin, Germany(柏林洪堡大学) Hertie Institute for AI in Brain Health, Universität Tübingen, Tübingen, Germany(人工智能与脑健康赫特研究所) Universität Tübingen, Tübingen, Germany(图宾根大学) Hasso Plattner Institut, Universität Potsdam, Potsdam, Germany(哈索·普拉特纳研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究针对平面曲线提出深度形状回归模型,允许多模态高维协变量。用复值函数表示曲线,提出含模态特定编码器的协方差平滑器,模型具多种不变性,还提供弹性均值估计算法,经模拟和实际应用验证了方法有效性。

Comments 17 pages, 4 figures, 1 algorithm. Submitted to the ShapeMI Workshop, MICCAI 2026. Code is available at https://github.com/mpff/dnn-shapes

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02533 2026-07-23 cs.RO cs.LG 78%

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

HMVLA:超几何多模态融合用于视觉-语言-动作模型

Kun Wang, Xiao Feng, Mingcheng Qu, Tonghua Su

机构 * Harbin Institute of Technology(哈尔滨工业大学) Chongqing Research Institute of HIT(重庆HIT研究 institute) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 HMVLA通过在双曲空间中融合视觉、语言和动作信息,提升多模态语义对齐的效率和准确性。

Comments 5 pages,5 figures,ICASSP

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21863 2026-07-23 cs.CV 版本更新 70%

Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation

Prompt-Calibrated SAM 3 用于开放词汇遥感语义分割

Yanghui Song, Nanqing Liu, Haonan Yin, Yingjie Gao, Chengfu Yang, Qi Ming

机构 * School of Information Science and Technology, Yunnan Normal University(云南师范大学信息科学与技术学院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) College of Computer Science, Beijing University of Technology(北京工业大学计算机学院)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出ProC-SAM3方法,通过离线提示池、缓存文本嵌入和存在引导残差融合,解决开放词汇遥感语义分割中提示语义覆盖不足、冗余编码和噪声传播问题,在8个基准上平均mIoU达56.1%。

Comments 5 pages, 5 figures. Accepted for publication in IEEE Geoscience and Remote Sensing Letters (GRSL)

Journal ref IEEE Geoscience and Remote Sensing Letters, vol. 23, 2026, Art. no. 3713378

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19889 2026-07-23 cs.CV 新提交 57%

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

LAVIFT:用于手术交互识别的潜在动作引导视觉微调

Jiajun Cheng, Subarna Tripathi, Sainan Liu, Xiaofan Yu, Shan Lin

机构 * Arizona State University(亚利桑那州立大学) University of California, Merced(加州大学默塞德分校) Intel Corporation(英特尔公司)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

AI总结 研究针对手术交互识别中预训练模型适应细粒度交互的挑战,提出LAViFiT框架,通过逆动力学模型、前向世界模型及补丁级SIG正则化器,实现视觉语言微调,提升了识别率和图像-文本对齐,增强了特征基础和空间连贯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19879 2026-07-23 cs.CV 新提交 57%

Current Injection Spiking Neural Network for Infrared and Visible Image Fusion

用于红外与可见光图像融合的电流注入脉冲神经网络

Rui Zhao, Zhuoyuan Li, Wenrui Li, Yanchen Dong, Yajing Zheng, Giuseppe Valenzise, Weisi Lin

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) The Hong Kong Polytechnic University(香港理工大学) Harbin Institute of Technology(哈尔滨工业大学) State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 研究红外与可见光图像融合问题,提出CIS - Fuse脉冲网络,通过电流注入脉冲算子在膜电位水平实现跨模态融合,构建双向跨模态融合模块并部署在双分支架构上,实验表明其融合质量与基于ANN的方法相当且更节能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17330 2026-07-23 cs.LG cs.AI 版本更新 57%

SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework

SubQuad:在自适应受体框架中利用分布平衡目标的近二次自由结构推断

Rong Fu, Zijian Zhang, Kun Liu, Jiekai Wu, Xianda Li, Simon Fong

机构 * University of Macau(澳门大学) University of Pennsylvania(宾夕法尼亚大学) University of Southampton(南安普顿大学) Juntendo University(顺天堂大学) University of Bologna(博洛尼亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 SubQuad通过结合抗原感知的近子二次检索、GPU加速的亲和力核、学习多模态融合和公平性约束聚类,解决大规模群体适应性免疫谱分析中的二次成本和数据不平衡问题,提升效率与公平性。

Comments 27 pages, 9 figures. In the previous version, Juntendo University was erroneously listed as the affiliation; we must clarify that this paper has absolutely no relation to Juntendo University. Therefore, we have replaced this affiliation in the new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00972 2026-07-23 cs.CV 版本更新 57%

SemICP: Semantic Non-Rigid Point Cloud Registration with Elastic Energy Regularization

SemICP:基于弹性能量正则化的语义非刚性点云配准

Wanwen Chen, Qi Zeng, Carson Studders, Zongze Li, Jamie J. Y. Kwon, Tara Kemper, Emily H. T. Pang, Eitan Prisman, Septimiu E. Salcudean

机构 * Department of Electrical and Computer Engineering, University of British Columbia(电气与计算机工程系,不列颠哥伦比亚大学) Faculty of Medicine, University of British Columbia(医学院,不列颠哥伦比亚大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 研究针对计算机辅助干预中点云配准问题,提出SemICP框架,结合语义信息点匹配与变形正则化,经多数据集测试,相比其他方法降低多种距离误差,结合AI分割后为多模态配准提供有效管道。

详情

展开后加载摘要…

URL PDF HTML 收藏