arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2602.01610 2026-02-03 cs.AI cs.LG 70%

ToPT: Task-Oriented Prompt Tuning for Urban Region Representation Learning

为城市区域表示学习设计的任务导向提示微调:ToPT

Zitao Guo, Changyang Jiang, Tianhong Zhao, Jinzhou Cao, Genan Dai, Bowen Zhang

机构 * College of Applied Science, Shenzhen University, Shenzhen, China(深圳大学应用科学学院) School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China(深圳科技大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 ToPT通过空间一致融合和任务对齐提升城市区域表示学习,实现任务导向的提示微调,提升多个城市任务的性能。

Comments The paper has been accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11512 2026-02-03 cs.RO cs.CV 70%

Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities

协同表征学习用于触觉、语言和视觉模态的对齐

Yiyun Zhou, Mingjing Xu, Jingwei Shi, Quanjiang Li, Jingyuan Chen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 TLV-CoRe通过协同表征学习提升触觉、语言和视觉模态的跨模态对齐与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21675 2026-01-30 cs.MM 70%

Rethinking Fusion: Disentangled Learning of Shared and Modality-Specific Information for Stance Detection

重新思考融合:为立场检测学习解耦的共享信息和模态特定信息

Zhiyu Xie, Fuqiang Niu, Genan Dai, Qianlong Wang, Li Dong, Bowen Zhang, Hu Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.MM

AI总结 DiME通过解耦共享和模态特定信息提升多模态立场检测性能,采用对比学习和余弦对齐的专家模块实现更准确的立场预测。

Comments ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20742 2026-01-29 cs.CV 70%

Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification

压缩揭示智能:视觉编码、视觉令牌技术与统一

Xin Jin, Jinming Liu, Yuntao Wei, Junyan Lin, Zhicheng Wang, Jianguo Huang, Xudong Yang, Yanxiao Liu, Wenjun Zeng

机构 * Eastern Institute of Technology, Ningbo(宁波东部技术研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文探讨了视觉编码与视觉令牌技术的统一,揭示压缩效率与模型性能的权衡,并预测了下一代视觉编解码器和令牌技术的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20168 2026-01-29 cs.CV 70%

Efficient Token Pruning for LLaDA-V

LLaDA-V的高效令牌修剪

Zhewen Wan, Tianchen Song, Chen Lin, Zhiyong Zhao, Xianpeng Lang

机构 * Li Auto Inc.(利自动公司) SJTU(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种针对LLaDA-V的结构化令牌修剪方法,通过在中层到后期层移除部分视觉令牌,减少计算开销并保持95%的任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06832 2026-01-28 cs.AI 70%

Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges

遥感图像智能解读的以语言为中心的视角:原理、方法与挑战

Haifeng Li, Wang Guo, Haiyang Wu, Mengwei Wu, Jipeng Zhang, Qing Zhu, Yu Liu, Xin Huang, Chao Tao

机构 * School of Geosciences and Info-Physics, Central South University(中南大学地球科学与信息物理学院) Faculty of Geosciences and Environmental Engineering, Southwest Jiaotong University(西南交通大学地质科学与环境工程学院) School of Earth and Space Sciences, Peking University(北京大学地球与空间科学学院) Institute of Remote Sensing Information Processing (IRSIP), Wuhan University(武汉大学遥感信息处理研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出以语言为中心的遥感图像解读框架,探讨LLMs作为认知核心的作用,并总结多模态表示、知识关联等核心挑战,为认知驱动的智能地理空间分析提供理论基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12869 2026-01-27 cs.LG cs.AI cs.DC cs.IT cs.MA math.IT 70%

On the Fundamental Limits of LLMs at Scale

在大规模下的大语言模型根本限制

Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal, Zeeshan Memon, Muhammad Ibtsaam Qadir, Sagnik Bhattacharya, Hassan Rizwan, Abhiram R. Gorle, Maahe Zehra Kazmi, Nukhba Amir, Ali Subhan, Muhammad Usman Rafique, Zihao He, Pulkit Mehta, Muhammad Ali Jamshed, John M. Cioffi

机构 * Stanford University(斯坦福大学) The University of Oklahoma(俄克拉荷马大学) Emory University(埃默里大学) Purdue University(普渡大学) UC Riverside(加州大学河滨分校) UC Berkeley(加州大学伯克利分校) Khyber Medical University(克希伯医学大学) Universtat Pompeu Fabra(庞培法华大学) Zoox(Zoox公司) Meta Google DeepMind(谷歌DeepMind) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Glasgow(格拉斯哥大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文探讨了大规模大语言模型的根本限制,提出统一框架分析计算、信息和学习的基础限制,并提供缓解方法。

Comments Submitted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09828 2026-01-27 cs.CV cs.LG cs.RO 70%

DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception

DGFusion:基于深度的传感器融合用于鲁棒的语义感知

Tim Broedermannn, Christos Sakaridis, Luigi Piccinelli, Wim Abbeloos, Luc Van Gool

机构 * Computer Vision Laboratory, ETH Zurich(计算机视觉实验室,苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 DGFusion通过整合深度信息提升多模态传感器融合,实现自动驾驶中的鲁棒语义感知。

Comments Code and models are available at https://github.com/timbroed/DGFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23118 2026-01-27 cs.CV 70%

Quantizing Space and Time: Fusing Time Series and Images for Earth Observation

量化空间与时间:融合时间序列和图像用于地球观测

Gianfranco Basile, Johannes Jakubik, Benedikt Blumenstiel, Thomas Brunschwiler, Juan Bernabe Moreno

机构 * IBM Research Europe(IBM欧洲研究院) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种任务无关的多模态融合框架,通过时间序列和图像的统一表示空间提升地球观测任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17073 2026-01-27 cs.LG cs.CV stat.ML 70%

Attention-Based Variational Framework for Joint and Individual Components Learning with Applications in Brain Network Analysis

基于注意力的变分框架用于联合和个体成分学习及其在脑网络分析中的应用

Yifei Zhang, Meimei Liu, Zhengwu Zhang

机构 * Department of Biostatistics, Yale School of Public Health(生物统计学系,耶鲁公共卫生学院) Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工大学) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(统计学与运筹学系,北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CM-JIVNet,一种基于注意力的变分框架,用于联合和个体成分学习,以提升脑网络分析中的跨模态重建和行为预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14697 2026-01-22 cs.IR cs.AI 70%

When Text-as-Vision Meets Semantic IDs in Generative Recommendation: An Empirical Study

当文本作为视觉信号与语义ID在生成推荐中相遇:一项实证研究

Shutong Qiao, Wei Yuan, Tong Chen, Xiangyu Zhao, Quoc Viet Hung Nguyen, Hongzhi Yin

机构 * The University of Queensland(昆士兰大学) City University of Hong Kong(香港城市大学) Griffith University(格里菲斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出将文本视为视觉信号,通过OCR技术提升生成推荐中语义ID学习的鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15394 2026-01-21 cs.CV 70%

Doracamom: Joint 3D Detection and Occupancy Prediction with Multi-view 4D Radars and Cameras for Omnidirectional Perception

Doracamom:多视角4D雷达与摄像头联合3D检测与占用预测用于全方位感知

Lianqing Zheng, Jianan Liu, Runwei Guan, Long Yang, Shouyi Lu, Yuanzhe Li, Xiaokai Bai, Jie Bai, Zhixiong Ma, Hui-Liang Shen, Xichan Zhu

机构 * School of Automotive Studies, Tongji University(同济大学汽车学院) Momoni AI Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Chair of Automotive Engineering, Technische Universität Berlin(柏林技术大学汽车工程学系) College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) School of Information and Electrical Engineering, Hangzhou City University(杭州城市学院信息与电气工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Doracamom通过融合多视角4D雷达与摄像头,实现3D物体检测与语义占用预测,提升自动驾驶环境感知能力。

Comments Accepted by IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12697 2026-01-21 cs.CV cs.CG 70%

Fusing in 3D: Free-Viewpoint Fusion Rendering with a 3D Infrared-Visible Scene Representation

在三维中融合:基于三维红外-可见场景表示的自由视角融合渲染

Chao Yang, Deshui Miao, Chao Tian, Guoqing Zhu, Yameng Gu, Zhenyu He

机构 * School of Computer Science(计算机科学学院) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种基于三维红外-可见场景表示的融合框架,通过跨模态调整模块和融合损失确保融合图像保留关键特征,有效解决传统方法在复杂场景下的信息丢失问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16665 2026-01-14 cs.CV 70%

A Diff-Attention Aware State Space Fusion Model for Remote Sensing Classification

一种面向遥感分类的差分注意力感知状态空间融合模型

Wenping Ma, Boyou Xue, Mengru Ma, Chuang Chen, Hekai Zhang, Hao Zhu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出DAS2F-Model,通过差分注意力模块和线性融合模块提升多模态遥感图像分类性能。

Comments After a careful review, we discovered that there were data errors in the paper, which led to the invalidity of the conclusion. To avoid misleading the readers, we have decided to withdraw this article. We appreciate your understanding and support for our work

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10731 2025-12-30 cs.CV 70%

Removal then Selection: A Coarse-to-Fine Fusion Perspective for RGB-Infrared Object Detection

去除后再选择:一种从粗到细融合视角的RGB红外目标检测

Tianyi Zhao, Maoxun Yuan, Feng Jiang, Nan Wang, Xingxing Wei

机构 * Institute of Artificial Intelligence, Hangzhou Innovation Institute, Beihang University, Beijing(北京航空航天大学人工智能研究院,杭州创新院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Beijing Institute of Control and Electronic Technology(北京控制与电子技术研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种从粗到细的融合视角,通过去除冗余信息和动态选择特征来提升RGB红外目标检测的性能。

Comments 11pages, 10figures

Journal ref IEEE Transactions on Intelligent Transportation Systems, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23950 2025-12-23 cs.AI 70%

InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

InterMT:基于人类反馈的多轮交错偏好对齐

Boyuan Chen, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Bowen Dong, Jiayi Zhou, Kaile Wang, Juntao Dai, Xuyao Wang, Wenqi Chen, Qirui Zheng, Wenxin Li, Sirui Han, Yike Guo, Yaodong Yang

机构 * Institute for AI, Peking University(人工智能研究院,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 InterMT通过多轮多模态交互的偏好数据集探索,旨在提升多模态大模型的交互能力,结合人类反馈和专家注释,揭示多轮扩展规律。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18496 2025-12-23 cs.CV 70%

Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models

Adaptive-VoCo: 用于视觉-语言模型的复杂度感知视觉标记压缩

Xiaoyang Guo, Keze Wang

机构 * Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Adaptive-VoCo通过动态压缩视觉标记提升视觉-语言模型的效率与鲁棒性

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13177 2025-12-17 cs.CV cs.RO 70%

MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion

MMDrive: 通过多表示融合超越视觉的交互场景理解

Minghui Hou, Wei-Hsing Huang, Shaofeng Liang, Daizong Liu, Tai-Hao Wen, Gang Wang, Runwei Guan, Weiping Ding

机构 * organization= College of Computer Science Technology, Jilin University , city= Changchun , country= China organization= Georgia Institute of Technology , city= Atlanta , country= USA organization= Qingdao Institute of Software, College of Computer Science Technology, China University of Petroleum (East China) , city= Qingdao , country= China organization= Institute for Math \& AI, Wuhan University , city= Wuhan , country= China organization= University of Michigan, Ann Arbor , country= USA organization= Thrust of Artificial Intelligence, Hong Kong University of Science organization= School of Artificial Intelligence Computer Science, Nantong University , city= Nantong , country= China

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 MMDrive通过融合占用图、LiDAR点云和文本描述,实现超越视觉的三维场景理解,提升自动驾驶的多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12842 2025-12-16 cs.RO cs.AI cs.LG 70%

SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding

SAGA:通过结构化可及性 grounding 实现开放世界移动操作

Kuan Fang, Yuxin Chen, Xinghao Zhu, Farzad Niroui, Lingfeng Sun, Jiuguang Wang

机构 * RAI Institute(RAI研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 SAGA通过结构化可及性接地实现开放世界移动操作,能有效处理多种任务形式并实现零样本执行。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01042 2025-12-16 cs.CV 70%

WCCNet: Wavelet-context Cooperative Network for Efficient Multispectral Pedestrian Detection

WCCNet:小波-上下文协作网络用于高效多光谱行人检测

Xingjian Wang, Li Chai, Jiming Chen, Zhiguo Shi

机构 * the College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院) the College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 WCCNet通过低计算成本的多光谱特征提取与跨模态融合,提升自动驾驶中的行人检测效率与精度。

Comments 35 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08374 2025-12-10 cs.CV 70%

The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss

看不见的偏见:预规范MLLM中规范差异如何导致视觉信息丢失

Bozhou Li, Xinda Xue, Sihan Yang, Yang Shi, Xinlong Chen, Yushuo Guan, Yuanxing Zhang, Wentao Zhang

机构 * Peking University(北京大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Xi’an Jiaotong University(西安交通大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文揭示了预规范MLLM中规范差异导致的视觉信息丢失问题,并提出通过插入层规范层来解决这一问题,提升模型整体能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06447 2025-12-09 cs.CV 70%

Towards Stable Cross-Domain Depression Recognition under Missing Modalities

迈向稳定跨领域抑郁识别下的缺失模态

Jiuyi Chen, Mingkui Tan, Haifeng Lu, Qiuna Xu, Zhihua Wang, Runhao Zeng, Xiping Hu

机构 * School of Future Technology, South China University of Technology(未来技术学院,华南理工大学) PengCheng Laboratory(鹏城实验室) School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学) Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) Guangdong-Hong Kong-Macao Joint Laboratory for Emotional Intelligence and Pervasive Computing(粤港澳大湾区情感智能与泛在计算联合实验室) School of Computer Science and Technology, Guangdong University of Technology(计算机科学与技术学院,广东工业大学) Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出SCD-MLLM框架,通过多源数据适配器和模态感知自适应融合模块,实现稳定跨领域抑郁症识别,提升多模态数据处理的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01946 2025-12-09 cs.CV 70%

3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding

3DRS: MLLMs 需要 3D 意识表示监督以实现场景理解

Xiaohu Huang, Jingjing Wu, Qunyi Xie, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室) Department of Computer Vision Technology (VIS), Baidu Inc.(百度公司计算机视觉技术部)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 3DRS 通过引入预训练 3D 基础模型的监督,提升 MLLM 的 3D 表示能力,从而增强场景理解性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09560 2025-12-05 cs.CV cs.RO 70%

WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization

WeatherPrompt: 多模态表示学习用于全天候无人机视觉地理定位

Jiahao Wen, Hang Yu, Zhedong Zheng

机构 * School of Computer Engineering and Science, Shanghai University, China(上海大学计算机工程与科学学院) Faculty of Science and Technology and Institute of Collaborative Innovation, University of Macau, China(澳门大学科技学院和协同创新研究院)

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 WeatherPrompt通过多模态学习提升无人机在恶劣天气下的视觉地理定位性能,采用无训练天气推理和动态门控机制实现天气不变表示,提升天气鲁棒性和召回率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04005 2025-12-04 cs.CV cs.LG cs.RO 70%

LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving

LargeAD: 大规模跨传感器数据预训练用于自动驾驶

Lingdong Kong, Xiang Xu, Youquan Liu, Jun Cen, Runnan Chen, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 LargeAD通过跨传感器数据预训练提升自动驾驶中的三维场景理解,结合多模态对比学习和时间一致性,实现更鲁棒的感知性能。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02715 2025-12-03 cs.CV 70%

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

GeoViS: 基于地理奖励的视觉搜索用于遥感视觉定位

Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin, Zonghao Guo, Fengxiang Wang, Xue Yang, Kaiwen Wei, Lei Wang

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) National University of Defense Technology(国防科技大学) Shanghai Jiao Tong University(上海交通大学) Chongqing University(重庆大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 GeoViS通过地理奖励机制,实现了遥感影像中的细粒度视觉定位,通过逐步搜索和推理提升小目标检测精度与跨领域泛化能力。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19856 2025-11-26 cs.CV 70%

Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks

时间-视觉语义对齐:一种统一架构,用于从视觉模型中转移空间先验到零样本时间任务

Xiangkai Ma, Han Zhang, Wenzhong Li, Sanglu Lu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 TimeArtist通过时间-视觉语义对齐框架,实现从时间序列到高质量图像生成的跨模态转换,并在零样本时间任务中取得优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10109 2025-11-26 cs.CV 70%

Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion

Dream-IF: 动态相对增强用于图像融合

Xingxin Xu, Bing Cao, Dongdong Li, Qinghua Hu, Pengfei Zhu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Dream-IF通过动态相对增强框架提升图像融合质量,结合主导区域信息实现多模态图像的协同增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16991 2025-11-24 cs.CV 70%

DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction

DReX:纯视觉融合自监督和卷积表示以预测图像复杂度

Jonathan Skaza, Parsa Madinei, Ziqi Wen, Miguel Eckstein

机构 * University of California, Santa Barbara(加州大学圣芭芭拉分校)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 DReX通过融合自监督和卷积表示,实现图像复杂度预测,取得最佳性能并减少参数量。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16435 2025-11-21 cs.CV 70%

Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation

超越视觉线索:利用通用语义作为少样本分割的支持

Jin Wang, Bingfeng Zhang, Jian Pang, Mengyu Liu, Honglong Chen, Weifeng Liu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出语言驱动属性泛化架构,通过多属性增强和多模态对齐提升少样本分割性能,实现新的最佳效果。

详情

展开后加载摘要…

URL PDF HTML 收藏