arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2604.00503 2026-04-08 cs.CV 70%

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

PET-DINO:将视觉线索统一到Grounding DINO中通过提示增强训练

Weifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li, Yuhuan Lin, Hanqiu Deng, Wenbing Tao, Yong Liu, Chengjie Wang

机构 * YouTu Lab, Tencent(腾讯优图实验室) Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技可灵团队)

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 PET-DINO通过提示增强训练策略,统一视觉线索到Grounding DINO中,提升零样本目标检测性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00867 2026-04-02 cs.CV 70%

A 4D Representation for Training-Free Agentic Reasoning from Monocular Laparoscopic Video

一种用于无训练代理推理的4D表示

Maximilian Fehrentz, Nicolas Stellwag, Robert Wiebe, Nicole Thorisch, Fabian Grob, Patrick Remerscheid, Ken-Joel Simmoteit, Benjamin D. Killeen, Christian Heiliger, Nassir Navab

机构 * Computer Aided Medical Procedures, TU Munich(慕尼黑工业大学计算机辅助医疗程序研究所) Hospital of the LMU Munich, Ludwig-Maximilians-Universität (LMU)(慕尼黑大学医院) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出一种基于显式4D表示的框架,利用多模态大语言模型实现无训练的手术代理推理,提升时空理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00601 2026-04-02 cs.CV 70%

KG-CMI: Knowledge graph enhanced cross-Mamba interaction for medical visual question answering

KG-CMI: 基于知识图谱的跨Mamba交互用于医学视觉问答

Xianyao Zheng, Hong Yu, Hui Cui, Changming Sun, Xiangyu Li, Ran Su, Leyi Wei, Jia Zhou, Junbo Wang, Qiangguo Jin

机构 * School of Software, Northwestern Polytechnical University(西北工业大学软件学院) Tianjin Central Hospital of Gynecology Obstetrics(天津市中心妇产科医院) Department of Computer Science and Information Technology, La Trobe University(拉筹伯大学计算机科学与信息技术系) CSIRO Data61(澳大利亚联邦科学与工业研究组织Data61) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Software, Tianjin University(天津大学计算机软件学院) Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University(澳门理工大学应用科学学院人工智能驱动药物发现中心) Department of Cardiology, Tianjin Chest Hospital(天津市胸科医院心内科)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出KG-CMI框架,通过整合医学知识图谱提升医学视觉问答的准确性与多样性,实验表明其在三个数据集上表现优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19225 2026-04-02 eess.IV cs.CV 70%

Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding

统一的医学图像标记器用于自回归合成与理解

Chenglong Ma, Yuanfeng Ji, Jin Ye, Zilong Li, Chenhui Wang, Junzhi Ning, Wei Li, Lihao Liu, Qiushan Guo, Tianbin Li, Junjun He, Hongming Shan

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Stanford University(斯坦福大学) Shanghai AI Laboratory(上海人工智能实验室) ByteDance Seed(字节跳动Seed)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出MedITok,通过两阶段训练框架统一医学图像标记器,利用大规模未配对图像提升重建精度,并结合图像-文本对注入细粒度语义,实现自回归建模在诊断和生成任务中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27781 2026-03-31 cs.CV 70%

GS3LAM: Gaussian Semantic Splatting SLAM

GS3LAM: 基于高斯语义散射的SLAM

Linfei Li, Lin Zhang, Zhong Wang, Ying Shen

机构 * School of Software Engineering, Tongji University(同济大学软件工程学院) Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 GS3LAM通过多模态数据融合实现实时一致的语义密集地图,采用语义高斯场和多模态误差约束联合优化,引入深度自适应尺度正则化和随机采样关键帧映射策略,提升跟踪鲁棒性和渲染质量。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26724 2026-03-31 cs.CV cs.RO 70%

An Annotation-to-Detection Framework for Autonomous and Robust Vine Trunk Localization in the Field by Mobile Agricultural Robots

一种用于田间自主和鲁棒葡萄藤主干定位的标注到检测框架

Dimitrios Chatziparaschis, Elia Scudiero, Brent Sams, Konstantinos Karydis

机构 * Dept. of Environmental Sciences, Univ. of California, Riverside(加州大学河滨分校环境科学系) Gallo(嘉露酒庄)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种标注到检测框架,利用有限标注数据训练多模态检测器,通过跨模态标注转移和早起传感器融合,提升田间葡萄藤主干定位的鲁棒性与准确性。

Comments 7 pages, 6 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26127 2026-03-30 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Finding Distributed Object-Centric Properties in Self-Supervised Transformers

在自监督变压器中发现分布式以对象为中心的属性

Samyak Rawlekar, Amitabh Swain, Yujun Cai, Yiwei Wang, Ming-Hsuan Yang, Narendra Ahuja

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Queensland(昆士兰大学) UC Merced(加州大学默塞德分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文分析了自监督视觉变压器中对象中心属性的分布式编码机制,提出Object-DINO方法通过聚类注意力头提升无监督对象发现和多模态大语言模型的视觉 grounding。

Comments Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26071 2026-03-30 cs.CV cs.LG 70%

MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality

MUST:一种模态特定表示感知的Transformer,用于具有缺失模态的扩散增强生存预测

Kyungwon Kim, Dosik Hwang

机构 * Yonsei University(延世大学) Korea Institute of Science and Technology(韩国科学技术研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 MUST通过分解模态特定和跨模态上下文化组件,提高生存预测的准确性,在缺失模态情况下仍能保持稳健预测。

Comments Accepted to CVPR 2026. 10 pages, 5 figures, supplementary included

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10971 2026-03-30 cs.CV 70%

ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization

ERMoE:基于特征的混合专家架构用于稳定路由和可解释的专业化

Anzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin, Mingxi Cheng, Heng Ping, Tamoghna Chattopadhyay, Sophia I Thomopoulos, Shahin Nazarian, Paul Thompson, Paul Bogdan

机构 * University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 ERMoE通过重新参数化专家在学习的正交特征基上,改进路由稳定性与专家专业化可解释性,无需显式平衡损失,提升ImageNet和跨模态检索任务性能。

Comments Accepted in CVPR2026 Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23508 2026-03-23 cs.CV 70%

Hyperbolic Cycle Alignment for Infrared-Visible Image Fusion

双曲循环对齐用于红外-可见图像融合

Timing Li, Bing Cao, Jiahe Feng, Haifang Cao, Qinghau Hu, Pengfei Zhu

机构 * College of Intelligence and Computing(智能与计算学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于双曲空间的双曲循环对齐网络,通过双路径交叉模态循环对齐框架和双曲层次对比对齐模块,实现更有效的多模态图像对齐与融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17741 2026-03-17 cs.CV 70%

Reasoning to Attend: Try to Understand How <SEG> Token Works

推理以关注:尝试理解<SEG>标记的作用

Rui Qian, Xin Yin, Dejing Dou

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文研究了<SEG>标记在多模态模型中的作用,通过可视化相似性图揭示其在图像-文本对中的语义相似性贡献,并提出READ方法提升模型的推理能力。

Comments This work has been accepted to CVPR 2025, please refer to https://github.com/rui-qian/READ

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13292 2026-03-17 cs.LG cs.AI 70%

Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs

Pragma-VL:迈向多模态大语言模型中安全与助人的务实仲裁

Ming Wen, Kun Yang, Xin Chen, Jingyu Zhang, Dingding Han, Shiwen Cui, Yuedong Xu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Ant Group(蚂蚁集团) Zhejiang University(浙江大学) UCLA(加州大学洛杉矶分校) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 Pragma-VL通过引入一种端到端的对齐算法,解决了多模态大语言模型在安全与助人之间的平衡问题,通过增强视觉风险感知和理论保证的奖励模型,在多数多模态安全基准上提升了性能。

Comments 31 pages, ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12989 2026-03-16 cs.CV cs.CR 70%

Test-Time Attention Purification for Backdoored Large Vision Language Models

针对受污染的大视觉语言模型的测试时间注意力净化

Zhifang Zhang, Bojun Yang, Shuo He, Weitong Chen, Wei Emma Zhang, Olaf Maennel, Lei Feng, Miao Xu

机构 * University of Queensland(昆士兰大学) Southeast University(东南大学) Nanyang Technological University(南洋理工大学) Adelaide University(阿德莱德大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出CleanSight,一种无需训练的测试时间防御方法,通过检测和净化异常的视觉-文本注意力分布来对抗后门攻击,显著优于现有基于像素的净化方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06250 2026-03-09 cs.CV 70%

Hierarchical Collaborative Fusion for 3D Instance-aware Referring Expression Segmentation

层次化协作融合用于3D实例感知指代表达分割

Keshen Zhou, Runnan Chen, Mingming Gong, Tongliang Liu

机构 * The University of Sydney(悉尼大学) The University of Melbourne(墨尔本大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 HCF-RES通过层次化视觉语义分解和渐进多级融合,实现了3D实例感知指代表达分割的高精度与细粒度定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17938 2026-03-09 cs.CL cs.LG 70%

SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization

SPINE:基于熵带正则化的令牌选择性测试时间强化学习

Jianghao Wu, Yasmeen George, Jin Ye, Yicheng Wu, Daniel F. Schmidt, Jianfei Cai

机构 * Monash University(墨尔本大学) Imperial College London(伦敦帝国理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 SPINE通过令牌选择性和熵带正则化提升测试时间推理稳定性与效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04267 2026-03-05 cs.CV 70%

UniLight: A Unified Representation for Lighting

UniLight: 一种统一的光照表示

Zitian Zhang, Iliyan Georgiev, Michael Fischer, Yannick Hold-Geoffroy, Jean-François Lalonde, Valentin Deschaintre

机构 * Université Laval(拉瓦尔大学) Adobe Research(Adobe研究)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 UniLight通过联合潜在空间统一多种光照表示,实现跨模态的光照特征提取与迁移,支持光照检索、环境映射生成和扩散模型中的光照控制。

Comments Project page: https://lvsn.github.io/UniLight

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02560 2026-03-04 cs.CV 70%

CAWM-Mamba: A unified model for infrared-visible image fusion and compound adverse weather restoration

CAWM-Mamba:一种用于红外可见图像融合和复合恶劣天气恢复的统一模型

Huichun Liu, Xiaosong Li, Zhuangfan Huang, Tao Ye, Yang Liu, Haishu Tan

机构 * School of Physics(物理学院) Optoelectronic Engineering, Foshan University, Foshan 528225, China(光学电子工程学院,佛山大学,佛山528225,中国) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology, Foshan 528225, China(粤港澳联合智能微纳光电子技术实验室,佛山528225,中国) School of Mechanical Electronic(机械电子学院) Information Engineering, China University of Mining and Technology, Beijing 100083, China(信息工程学院,中国矿业大学,北京100083,中国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 CAWM-Mamba提出了一种统一模型,用于红外可见图像融合和复合恶劣天气恢复,通过三个关键模块提升多退化场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14337 2026-03-04 eess.IV cs.CV 70%

Unsupervised Deformable Image Registration with Local-Global Attention and Image Decomposition

无监督可变形图像配准与局部-全局注意力及图像分解

Zhengyong Huang, Xingwen Sun, Xuting Chang, Ning Jiang, Yao Wang, Jianfei Sun, Hongbin Han, Yao Sui

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学医学部医学技术研究所) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) Department of Radiology, Peking University Third Hospital(北京大学第三医院放射科) Department of Pediatrics, Peking University First Hospital(北京大学第一医院儿科) Pediatric Epilepsy Center, Peking University First Hospital(北京大学第一医院儿童癫痫中心) School of Biological Science and Medical Engineering, Southeast University(东南大学生物科学与医学工程学院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出LGANet++,通过局部-全局注意力机制和图像分解技术,实现更准确、鲁棒的无监督可变形图像配准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01720 2026-03-03 cs.CV 70%

Preoperative-to-intraoperative Liver Registration for Laparoscopic Surgery via Latent-Grounded Correspondence Constraints

腹腔手术中基于潜在证据的预手术到手术过程肝脏注册

Ruize Cui, Jialun Pei, Haiqiao Wang, Jun Zhou, Jeremy Yuen-Chun Teoh, Pheng-Ann Heng, Jing Qin

机构 * The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出Land-Reg框架,通过显式学习潜在证据的2D-3D地标对应关系,提升腹腔手术中预手术到术中肝脏的跨模态注册精度与可解释性。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02948 2026-03-03 cs.AI cs.LG physics.ao-ph 70%

FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead

风梧:将高技能全球中期天气预报推至10天提前期

Kang Chen, Tao Han, Junchao Gong, Lei Bai, Fenghua Ling, Jing-Jia Luo, Xi Chen, Leiming Ma, Tianning Zhang, Rui Su, Yuanzheng Ci, Bin Li, Xiaokang Yang, Wanli Ouyang

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 FengWu通过多模态和多任务框架提升全球中期天气预报能力,首次实现10.75天提前期的高精度预测。

Comments 12 pages

Journal ref Commun. Earth Environ. 6, 518 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00409 2026-03-03 cs.CV 70%

SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning

SSR:通过结构化场景推理推动空间智能的极限

Yi Zhang, Youya Xia, Yong Wang, Meng Song, Xin Wu, Wenjun Wan, Bingbing Liu, AiXue Ye, Hongbo Zhang, Feng Wen

机构 * Foundation Model Department, Huawei(华为基础模型部门)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 SSR通过结构化场景推理框架,在减少对齐成本的同时,实现了高效的空间智能,优于更大规模模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00046 2026-03-03 cs.LG cs.AI 70%

REMIND: Rethinking Medical High-Modality Learning under Missingness--A Long-Tailed Distribution Perspective

REMIND: 重新思考医疗多模态学习中的缺失性——从长尾分布视角

Chenwei Wu, Zitao Shuai, Liyue Shen

机构 * University of Michigan(密歇根大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 REMIND从长尾分布视角重新思考医疗多模态学习中的高模态缺失问题,提出组专用混合专家架构和分布鲁棒优化策略,有效提升尾部模态组合的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23959 2026-03-02 cs.CV 70%

Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought

通过图像作为连续动作进行思考:数值视觉链式推理

Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Science and Technology of China(中国科学技术大学) Chalmers University of Technology(楚克理工大学) University of Gothenburg(哥德堡大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 NV-CoT通过将图像推理动作空间扩展为连续欧几里得空间,提升MLLMs的定位精度和回答准确性,同时加速训练收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13689 2026-02-17 cs.RO cs.CV 70%

Symmetry-Aware Fusion of Vision and Tactile Sensing via Bilateral Force Priors for Robotic Manipulation

通过双侧力先验实现视觉与触觉感知的对称感知融合用于机器人操作

Wonju Lee, Matteo Grimaldi, Tao Yu

机构 * DexAI, Emergent Business Unit, Analog Devices Inc.(DexAI,新兴业务部,安森通公司)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种基于双侧力先验的跨模态Transformer,通过结构化注意力机制实现视觉与触觉融合,在机器人操作中实现了96.59%的插入成功率,展示了触觉感知的重要性及物理启发正则化的有效性。

Comments Accepted By ICRA2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12441 2026-02-16 cs.CV 70%

Prototype-driven fusion of pathology and spatial transcriptomics for interpretable survival prediction

基于病理学和空间转录组的原型驱动融合用于可解释的生存预测

Lihe Liu, Xiaoxi Pan, Yinyin Yuan, Lulu Shang

机构 * Department of Biostatistics, MD Anderson Cancer Center(生物统计学系,MD安德森癌症中心) Department of Translational Molecular Pathology, MD Anderson Cancer Center(转化分子病理学系,MD安德森癌症中心) The Institute for Data Science in Oncology (IDSO), MD Anderson Cancer Center(肿瘤学数据科学研究所(IDSO),MD安德森癌症中心)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 PathoSpatial通过整合病理学和空间转录组数据,实现可解释的生存预测,提升多模态学习的可解释性和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11494 2026-02-13 cs.CV 70%

Arbitrary Ratio Feature Compression via Next Token Prediction

通过下一个标记预测实现任意比例特征压缩

Yufan Liu, Daoyuan Ren, Zhipeng Zhang, Wenyang Luo, Bing Li, Weiming Hu, Stephen Maybank

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institution of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) People AI, Inc.(People AI公司) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) School of Computer Science and Mathematics, Birkbeck College, University of London(伦敦大学伯克贝克学院计算机科学与数学学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出了一种通过下一个标记预测实现任意比例特征压缩的框架,解决了传统方法在灵活性和通用性上的不足,通过引入混合解决方案和实体关系图约束模块,提升了压缩特征的质量和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12365 2026-02-12 cs.CL cs.DB 70%

Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics

大语言模型的进展:聚焦推理、适应性、效率和伦理

Asifullah Khan, Muhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, Abdul Rehman

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在推理、适应性、效率和伦理方面的进展,探讨了关键技术和挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04397 2026-02-11 cs.CV 70%

Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration

多专家学习框架与状态空间模型用于光学和SAR图像配准

Wei Wang, Dou Quan, Ning Huyan, Chonghua Lv, Shuang Wang, Yunan Li, Licheng Jiao

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China(教育部智能感知与图像理解重点实验室) Hangzhou Institute of Technology, Xidian University(西安电子科技大学杭州学院) School of Computer Science, Xidian University(西安电子科技大学计算机学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ME-SSM框架,通过多专家学习和状态空间模型提升光学与SAR图像配准的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06040 2026-02-06 cs.CV 70%

SwimBird: Eliciting Switchable Reasoning Mode in Hybrid Autoregressive MLLMs

SwimBird: 在混合自回归大语言模型中实现可切换的推理模式

Jintao Tong, Shilin Yan, Hongwei Xue, Xiaojun Tang, Kunyu Shi, Guannan Zhang, Ruixuan Li, Yixiong Zou

机构 * Huazhong University of Science and Technology(华中科技大学) Accio Team, Alibaba Group(阿里集团Accio团队)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 SwimBird通过动态切换三种推理模式,提升多模态大语言模型在视觉密集任务中的性能,同时保持文本推理能力。

Comments Project Page: https://accio-lab.github.io/SwimBird

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01832 2026-02-03 cs.AI 70%

Synesthesia of Vehicles: Tactile Data Synthesis from Visual Inputs

车辆的联觉:从视觉输入合成触觉数据

Rui Wang, Yaoguang Cao, Yuyi Chen, Jianyi Xu, Zhuoyang Li, Jiachen Shang, Shichun Yang

机构 * Dept. of Transportation Science, Beihang Univ.(北京航空航天大学交通运输科学系) State Key Lab of Intelligent Transportation System, Beihang Univ.(北京航空航天大学智能交通系统国家重点实验室) Hangzhou International Innovation Institute, Beihang Univ.(杭州国际创新研究院) China Software Testing Center(Ministry of Industry and Information Technology Software and Integrated Circuit Promotion Center)(中国软件测试中心(工业和信息化部软件与集成电路促进中心))

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出基于视觉输入的触觉数据合成方法,通过跨模态时空对齐和潜在扩散模型提升自动驾驶车辆的触觉感知能力,从而增强安全性能。

详情

展开后加载摘要…

URL PDF HTML 收藏