arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2512.09311 2025-12-11 cs.CV cs.CR 79%

Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance

基于Transformer的多模态融合用于视觉监控中的可解释性可疑性估计

Kuldeep Singh Yadav, Lalan Kumar

机构 * Big Data Research and Supercomputing Division, CSIR Fourth Paradigm Institute(CSIR第四范式研究所大数据研究与超级计算部门) Department of Electrical Engineering, Bharti School of Telecommunication, Yardi School of Artificial Intelligence, IIT Delhi(电信学院电子工程系、Yardi人工智能学院、德里理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于Transformer的多模态融合框架DeepUSEvision,结合YOLOv12、双深度卷积网络和Transformer判别器,实现高准确率和可解释性的可疑性估计。

Comments 12 pages, 10 figures, IEEE Transaction on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08135 2025-12-10 cs.CV 79%

CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning

CVP:基于中央-外围视觉的多模态模型用于空间推理

Zeyuan Chen, Xiang Zhang, Haiyang Xu, Jianwen Xie, Zhuowen Tu

机构 * UC San Diego(圣迭戈大学) Lambda, Inc(Lambda公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 CVP通过结合中央视觉和外围视觉的启发,提出了一种多模态模型,以提升复杂3D环境的空间推理能力。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05996 2025-12-09 cs.CV cs.CY cs.RO eess.IV 79%

FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting

FishDetector-R1: 基于统一MLLM框架的弱监督鱼类检测、分割与计数强化微调方法

Yi Liu, Jingyu Song, Vedanth Kallakuri, Katherine A. Skinner

机构 * University of Michigan(密歇根大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);分类 cs.CV

AI总结 FishDetector-R1通过统一MLLM框架和强化学习微调,实现了弱监督下的鱼类检测、分割与计数的高效准确提升。

Comments 18 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10573 2025-12-09 cs.CV 79%

Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration

通过病理级跨模态对齐和相关性探索提升医学视觉表示学习

Jun Wang, Lixing Zhu, Xiaohan Yu, Abhir Bhalerao, Yulan He

机构 * Department of Computer Science, University of Warwick(沃里克大学计算机科学系) Department of Informatics, King’s College London(伦敦国王学院信息学系) School of Computing, Macquarie University(麦考瑞大学计算机科学学院) Alan Turing Institute, UK(英国艾伦·图灵研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PLACE框架,通过病理级跨模态对齐和相关性探索提升医学视觉表示学习,实现多下游任务的性能提升。

Comments Accepted to IEEE Journal of Biomedical and Health Informatics (JBHI).Code: https://github.com/Markin-Wang/PLACE

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04943 2025-12-05 cs.CV 79%

Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition

面向人类动作识别的多模态深度网络自适应融合

Novanto Yudistira

机构 * Departemen Teknik Informatika, Fakultas Ilmu Komputer, Universitas Brawijaya(计算机科学系,信息学院,布拉格亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于多模态深度网络的自适应融合方法,通过门控机制提升人类动作识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16848 2025-12-05 cs.CV 79%

A re-calibration method for object detection with multi-modal alignment bias in autonomous driving

面向自动驾驶的多模态对齐偏差校准方法

Zhihang Song, Dingyi Yao, Ruibo Ming, Lihui Peng, Danya Yao, Yi Zhang

机构 * Department of Automation Tsinghua University Beijing(自动化系清华大学北京)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出一种重新校准模型,通过语义分割和定制损失函数提升自动驾驶中多模态检测的鲁棒性和性能,应对校准偏差带来的影响。

Comments Accepted for publication in IST 2025. Official IEEE Xplore entry will be available once published

Journal ref 2025 IEEE International Conference on Imaging Systems and Techniques (IST)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00530 2025-12-04 cs.LG cs.AI cs.SI 79%

Generic Multimodal Spatially Graph Network for Spatially Embedded Network Representation Learning

通用多模态空间图网络用于空间嵌入网络表示学习

Xudong Fan, Jürgen Hackl

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出通用多模态空间图卷积网络,通过多模态特征提升空间嵌入网络的表示准确性,实验显示在电力网络中边存在预测任务的准确率提高37.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03404 2025-12-04 cs.CV 79%

MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification

MOS:缓解光学-合成孔径雷达模态差距以实现跨模态舰船重识别

Yujian Zhao, Hankun Liu, Guanglin Niu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 MOS通过模态一致表示学习和跨模态数据生成与融合,有效缓解光学与SAR图像间的模态差距,提升舰船跨模态重识别的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01513 2025-12-04 cs.CR cs.CV 79%

SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism

SafePTR: 通过剪枝-恢复机制实现多模态大语言模型的令牌级 Jailbreak 防御

Beitao Chen, Xinyu Lyu, Lianli Gao, Jingkuan Song, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳研究院) Southwestern University of Finance and Economics(西南财经大学) Engineering Research Center of Intelligent Finance, Ministry of Education(教育部智能金融工程研究中心) Tongji University(同济大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SafePTR 提出一种无需训练的多模态大语言模型防御机制,通过剪枝有害令牌并恢复良性特征,有效提升安全性并保持效率。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01949 2025-12-02 cs.CV 79%

Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models

脚本:图结构和查询条件的语义令牌修剪用于多模态大语言模型

Zhongyu Yang, Dannong Xu, Wei Pang, Yingfang Yuan

机构 * BCML, Heriot-Watt University(赫瑞瓦德大学BCML中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 Script通过图结构和查询条件的语义令牌修剪,提升多模态大语言模型的效率和准确性,实现显著的性能提升。

Comments Published in Transactions on Machine Learning Research, Project in https://01yzzyu.github.io/script.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20469 2025-12-02 q-bio.QM cs.CV 79%

Prediction of Distant Metastasis in Head and Neck Cancer Patients Using Tumor and Peritumoral Multi-Modal Deep Learning

利用肿瘤及周围多模态深度学习预测头颈癌患者远端转移

Nuo Tong, Changhao Liu, Zizhao Tang, Feifan Sun, Yingping Li, Shuiping Gou, Mei Shi

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本研究提出多模态深度学习模型,结合CT影像、放射组学和临床数据,用于预测头颈癌患者远端转移风险,通过多模态融合显著提高预测性能。

Comments 23 pages, 6 figures, 7 tables. Nuo Tong and Changhao Liu contributed equally. Corresponding Authors: Shuiping Gou and Mei Shi

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01410 2025-12-02 cs.CL 79%

DyFuLM: An Advanced Multimodal Framework for Sentiment Analysis

DyFuLM:一种用于情感分析的先进多模态框架

Ruohan Zhou, Jiachen Yuan, Churui Yang, Wenzheng Huang, Guoyan Zhang, Shiyao Wei, Jiazhen Hu, Ning Xin, Md Maruf Hasan

机构 * Department of Applied Mathematics, Xi'an Jiaotong-Liverpool University(应用数学系,西安交通大学-利物浦大学) School of AI and Advanced Computing, XJTLU Entrepreneur College (Taicang)(人工智能与先进计算学院,XJTLU创业学院(太仓)) Department of Intelligent Science, Xi'an Jiaotong-Liverpool University(智能科学系,西安交通大学-利物浦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 DyFuLM通过动态融合和门控聚合模块提升多模态情感分析的准确率与稳定性

Comments 8 pages, 6 figures, preprint. Under review for a suitable AI conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00596 2025-12-02 cs.IR cs.AI 79%

DLRREC: Denoising Latent Representations via Multi-Modal Knowledge Fusion in Deep Recommender Systems

DLRREC: 通过深度融合多模态知识在深度推荐系统中进行潜在表示去噪

Jiahao Tian, Zhenkai Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 DLRREC通过深度融合多模态和协同知识,提升深度推荐系统中潜在表示的去噪能力,从而实现更精确的推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22961 2025-12-01 cs.CV 79%

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

HMR3D:用于大视觉-语言模型的层次多模态表示以实现3D场景理解

Chen Li, Eric Peh, Basura Fernando

机构 * Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore(前沿人工智能研究中心,科技研究局,新加坡) College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 HMR3D通过层次化多模态表示,结合多视图图像和文本描述,提升3D场景理解的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21937 2025-12-01 cs.CV 79%

Interpretable Multimodal Cancer Prototyping with Whole Slide Images and Incompletely Paired Genomics

可解释的多模态癌症原型生成:结合整张滑片图像与不完全配对的基因组学

Yupei Zhang, Yating Huang, Wanming Hu, Lequan Yu, Hujun Yin, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Electrical & Electronic Engineering, The University of Manchester, UK(曼彻斯特大学电气与电子工程系) Department of Pathology, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Sun Yat-sen University Cancer Center, China(南方医科大学肿瘤学国家重点实验室、广东省癌症临床研究中心、中山大学肿瘤中心病理学部) Department of Statistics and Actuarial Science, The University of Hong Kong, Hong Kong SAR, China(香港大学统计与精算科学系) Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学学院) School of Science and Engineering and School of Medicine, University of Dundee, UK

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种可解释的多模态原型生成框架,通过整合整张滑片图像和不完整的基因组学数据,提升精准肿瘤学中的多模态整合效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12099 2025-12-01 cs.CV 79%

TinyRS-R1: Compact Multimodal Language Model for Remote Sensing

TinyRS-R1:用于遥感的紧凑多模态语言模型

Aybora Koksal, A. Aydin Alatan

机构 * Center for the Image Analysis (OGAM) and Department of Electrical and Electronics Engineering of Middle East Technical University (METU)(图像分析中心(OGAM)和中欧技术大学(METU)电子与电气工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TinyRS-R1是一种专为遥感设计的紧凑多模态语言模型,通过四阶段训练实现高效性能,兼具推理增强与低资源消耗。

Comments Accepted to IEEE Geoscience and Remote Sensing Letters (GRSL). Code, models, and the captions for datasets are available at https://github.com/aybora/TinyRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21106 2025-11-27 cs.CV 79%

EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens

EM-KD: 通过不平衡视觉标记的知识蒸馏提升高效多模态大语言模型

Ze Feng, Sen Yang, Boqiang Duan, Wankou Yang, Jingdong Wang

机构 * Baidu VIS(百度视觉)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 EM-KD通过改进的知识蒸馏方法,有效提升高效多模态大语言模型的视觉理解和效率。

Comments accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20997 2025-11-27 cs.LG cs.AI 79%

FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning

FANoise:奇异值自适应噪声调制用于鲁棒多模态表示学习

Jiaoyang Li, Jun Fang, Tianhao Gao, Xiaohui Zhang, Zhiyuan Liu, Chao Liu, Pengzhang Liu, Qixia Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 FANoise通过自适应噪声注入策略提升多模态表示学习的鲁棒性和性能

Comments 13 pages, 5 figures, accept to AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10133 2025-11-27 cs.CV 79%

MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning

MANGO: 多模态基于注意力的归一化流方法用于融合学习

Thanh-Dat Truong, Christophe Bobda, Nitin Agarwal, Khoa Luu

机构 * CVIU Lab, University of Arkansas, USA(大学实验室,亚利桑那大学,美国) University of Florida, USA(佛罗里达大学,美国) COSMOS Research Center, University of Arkansas, Little Rock, USA(研究机构,亚利桑那大学,小石城,美国) ICSI, University of California, Berkeley, USA(研究机构,加州大学伯克利分校,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MANGO提出了一种基于归一化流和可逆交叉注意力机制的多模态融合学习方法,通过三种新型交叉注意力机制提升多模态数据的建模能力。

Comments Accepted to NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12121 2025-11-26 cs.LG cs.MM 79%

To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance

对齐还是不对齐:为最佳性能的战略多模态表示对齐

Wanlong Fang, Tianle Zhang, Alvin Chan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文研究了显式对齐对模型性能的影响,提出可控对比学习模块,发现最佳对齐强度取决于数据冗余量,为优化单模态编码器提供指导。

Comments Accepted by AAAI 2026. This arXiv version includes additional details and extended appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19023 2025-11-25 cs.LG cs.AI 79%

OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs

OrdMoE:通过多模态混合专家模型中的层次专家小组排名实现偏好对齐

Yuting Gao, Weihao Chen, Lan Wang, Ruihan Xu, Qingpei Guo

机构 * AntGroup(蚂蚁集团) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 OrdMoE通过利用混合专家架构中的内在信号,实现多模态混合专家语言模型的零成本偏好对齐,无需人工标注数据即可提升模型对齐和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18677 2025-11-25 cs.CV 79%

A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification

基于理论的少样本跨模态素描人重识别框架

Yunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep, Min Jiang

机构 * School of Informatics, Xiamen University(厦门大学信息学院) School of Electronic Science and Engineering, Xiamen University(厦门大学电子科学与技术学院) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院) Department of Artificial Intelligence, Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, School of Informatics, Key Laboratory of Digital Protection and Intelligent Processing of Intangible CulturalHeritage of Fujian and Taiwan, Ministry of Culture and Tourism, Xiamen University(人工智能系,教育部多媒体可信感知与高效计算重点实验室,厦门大学信息学院,福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室,文化和旅游部,厦门大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出KTCAA框架,通过理论指导的对齐增强和知识转移催化剂,解决少样本跨模态素描人重识别中的模态差距和数据稀缺问题。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23206 2025-11-25 cs.CV 79%

HyperPointFormer: Multimodal Fusion in 3D Space with Dual-Branch Cross-Attention Transformers

HyperPointFormer: 在三维空间中通过双分支交叉注意力变换器实现多模态融合

Aldino Rizaldy, Richard Gloaguen, Fabian Ewald Fassnacht, Pedram Ghamisi

机构 * Helmholtz-Zentrum Dresden-Rossendorf (HZDR)(德累斯顿-罗斯托克亥姆霍尔茨研究中心) Freie Universität Berlin(柏林自由大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 HyperPointFormer通过双分支交叉注意力Transformer在三维空间中融合多模态数据,提升城市场景土地利用分类的精度与灵活性。

Journal ref IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 21254-21274, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00728 2025-11-25 eess.IV cs.CV cs.LG 79%

MultiFusionNet: Multilayer Multimodal Fusion of Deep Neural Networks for Chest X-Ray Image Classification

MultiFusionNet:多层多模态融合的深度神经网络用于胸部X光图像分类

Saurabh Agarwal, K. V. Arya, Yogesh Kumar Meena

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MultiFusionNet通过多层多模态融合和FDSFM模块提升胸部X光图像分类的准确率,达到97.21%和99.60%

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18424 2025-11-25 cs.CV 79%

CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images

跨模态联合嵌入预测架构CrossJEPA:用于从2D图像高效学习3D表示的架构

Avishka Perera, Kumal Hewagamage, Saeedha Nazar, Kavishka Abeywardana, Hasitha Gallella, Ranga Rodrigo, Mohamed Afham

机构 * University of Moratuwa(摩图瓦大学) Technische Universität Darmstadt(达姆施塔特技术大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 CrossJEPA通过跨模态联合嵌入预测架构,利用图像基础模型知识,实现高效3D表示学习,达到SOTA性能,且训练高效、内存占用低。

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17904 2025-11-25 cs.CV cs.RO 79%

CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation

CUS-GS: 一种紧凑的统一结构高斯点扩散框架用于多模态场景表示

Yuhang Ming, Chenxin Fang, Xingyuan Yu, Fan Zhang, Weichen Dai, Wanzeng Kong, Guofeng Zhang

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院) CAD & CG, Zhejiang University(浙江大学计算机辅助设计与图形学研究所) School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 CUS-GS通过统一结构化高斯点扩散框架,实现多模态场景表示的高效建模与语义一致性。

Comments 15 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17587 2025-11-25 cs.LG cs.AI 79%

Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection

基于情感与意图引导的多模态学习用于贴纸响应选择

Yuxuan Hu, Jian Chen, Yuhao Wang, Zixuan Li, Jing Xiong, Pengyue Jia, Wei Wang, Chengming Li, Xiangyu Zhao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出EIGML框架,通过联合建模情感与意图,提升贴纸响应选择的准确性和理解深度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17308 2025-11-24 cs.CV 79%

SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion

SpatialGeo:通过几何-语义融合提升多模态大语言模型的空间推理能力

Jiajie Guo, Qingpeng Zhu, Jin Zeng, Xiaolong Wu, Changyong He, Weida Wang

机构 * School of Computer Science and Technology, Tongji University, Shanghai, China(计算机科学与技术学院,同济大学,上海,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 SpatialGeo通过几何-语义融合提升多模态大语言模型的空间推理能力,实验表明在空间推理任务中准确率提升8.0%且内存消耗减少50%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15433 2025-11-20 cs.CV 79%

Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection

基于模态解耦的表示空间约束学习用于多模态目标检测

YiKang Shao, Tao Shi

机构 * school of reliability and systems engineering, Beihang University(可靠性与系统工程学院,北京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出RSC-MD方法,通过模态解耦和表示空间约束学习解决多模态目标检测中的融合退化问题,提升各模态的优化效果。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12079 2025-11-20 cs.CV 79%

Point Cloud Quantization through Multimodal Prompting for 3D Understanding

Hongxuan Li, Wencheng Zhu, Huiying Xu, Xinzhong Zhu, Pengfei Zhu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2026. 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏