arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2602.09255 2026-02-13 cs.RO cs.AI 79%

STaR: Scalable Task-Conditioned Retrieval for Long-Horizon Multimodal Robot Memory

STaR:可扩展的任务条件检索用于长时多模态机器人记忆

Mingfeng Yuan, Hao Zhang, Mahan Mohammadi, Runhao Li, Jinjun Shan, Steven L. Waslander

机构 * University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究所) University of Toronto Robotics Institute(多伦多大学机器人研究所) Department of Earth and Space Science(地球与空间科学系) Lassonde School of Engineering(拉索nde工程学院) York University(约克大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 STaR通过任务条件检索算法提升机器人长时多模态记忆的可扩展性和上下文推理能力

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10023 2026-02-11 cs.CL 79%

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

MEVER:基于图的证据检索的多模态和可解释性声明验证

Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang, Dongwon Lee

机构 * University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) University of California San Diego(加州大学圣地亚哥分校) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

AI总结 MEVER通过多模态图检索和解释生成,实现了准确且可解释的声明验证,同时创建了AI领域的科学数据集AIChartClaim。

Comments Accepted to EACL-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09839 2026-02-11 cs.CV 79%

ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge

ARK:一个双轴多模态检索基准,沿推理与知识

Yijie Lin, Guofeng Ding, Haochen Zhou, Haobin Li, Mouxing Yang, Xi Peng

机构 * College of Computer Science, Sichuan University, Chengdu, China(四川大学计算机学院) National Key Laboratory of Fundamental Algorithms and Models for Engineering Simulation, Sichuan University, China(四川省工程仿真基础算法与模型国家重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 ARK基准通过双轴视角评估多模态检索,揭示知识密集型与推理密集型检索间的差距,发现细粒度视觉推理是主要瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08741 2026-02-10 cs.CL 79%

From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding

从行到推理:一种增强检索的多模态框架用于电子表格理解

Anmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin Paul

机构 * Commercial Technology and Innovation Office, PricewaterhouseCoopers U.S.(普华永道美国商业技术与创新办公室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 FRTR提出了一种多模态检索增强生成框架,通过分解电子表格为细粒度嵌入并整合多模态信息,提升了对复杂电子表格的推理能力,在基准测试中实现了显著的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07642 2026-02-10 cs.AI cs.LG 79%

Efficient Table Retrieval and Understanding with Multimodal Large Language Models

基于多模态大语言模型的高效表格检索与理解

Zhuoyan Xu, Haoyang Fang, Boran Han, Bonan Min, Bernie Wang, Cuixiong Hu, Shuai Zhang

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) AWS(亚马逊网络服务)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 TabRAG通过多模态大语言模型实现高效表格检索与理解,显著提升检索召回率和答案准确率。

Comments Published at EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07208 2026-02-10 cs.IR cs.AI 79%

Sequences as Nodes for Contrastive Multimodal Graph Recommendation

序列作为节点的对比多模态图推荐

Bucher Sahyouni, Matthew Vowels, Liqun Chen, Simon Hadfield

机构 * University of Surrey(萨里大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 MuSICRec通过多视图图方法结合协同、序列和多模态信号,提升推荐系统在冷启动和数据稀疏问题上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06654 2026-02-09 cs.IR cs.AI 79%

Multimodal Generative Retrieval Model with Staged Pretraining for Food Delivery on Meituan

具有分阶段预训练的多模态生成检索模型用于美团外卖

Boyu Chen, Tai Guo, Weiyu Cui, Yuqing Li, Xingxing Wang, Chuan Shi, Cheng Yang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种具有分阶段预训练的多模态生成检索模型,通过分阶段任务优化和语义ID利用,提升美团外卖检索性能与商业收益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02537 2026-02-04 cs.CV cs.LG 79%

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

WorldVQA:评估多模态大语言模型的原子视觉世界知识

Runjie Zhou, Youbo Shao, Haoyu Lu, Bowei Xing, Tongtong Bai, Yujie Chen, Jie Zhao, Lin Sui, Haotian Yao, Zijia Zhao, Hao Yang, Haoning Wu, Zaida Zhou, Jinguo Zhu, Zhiqi Huang, Yiping Bao, Yangyang Liu, Y. Charles, Xinyu Zhou

机构 * Moonshot AI

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 WorldVQA通过评估多模态大语言模型的原子视觉知识,建立衡量其事实性和百科全书广度的基准测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01594 2026-02-03 cs.CV 79%

UV-M3TL: A Unified and Versatile Multimodal Multi-Task Learning Framework for Assistive Driving Perception

UV-M3TL: 一种统一且多功能的多模态多任务学习框架用于辅助驾驶感知

Wenzhuo Liu, Qiannan Guo, Zhen Wang, Wenshuo Wang, Lei Yang, Yicheng Qiao, Lening Wang, Zhiwei Li, Chen Lv, Shanghang Zhang, Junqiang Xi, Huaping Liu

机构 * Energy and Transportation Domain, Beijing Institute of Technology(能源与交通领域,北京理工大学) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(智能技术与系统国家重点实验室和清华大学计算机科学与技术系) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学) School of Transportation Science and Engineering and the State Key Lab of Intelligent Transportation System, Beihang University(交通运输科学与工程学院和智能交通系统国家重点实验室,北京航空航天大学) Beijing University of Chemical Technology(北京化工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UV-M3TL通过双分支结构和自适应损失机制,实现多模态多任务学习,提升辅助驾驶感知的性能与多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22610 2026-02-02 cs.LG cs.AI 79%

Local-Global Multimodal Contrastive Learning for Molecular Property Prediction

局部-全局多模态对比学习用于分子性质预测

Xiayu Liu, Zhengyi Lu, Yunhong Liao, Chan Fan, Hou-biao Li

机构 * School of Mathematical Sciences(数学科学学院) University of Electronic Science and Technology of China(电子科技大学) Department of Computer Science and Engineer(计算机科学与工程系) Oakland University(奥克兰大学) Department of Electrical and Computer Engineering(电气与计算机工程系) College of Management Science(管理科学学院) Chengdu University of Technology(成都理工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 LGM-CL通过局部-全局多模态对比学习框架,整合分子结构和化学语义信息,提升分子性质预测的准确性与性能。

Comments 16 pages, 9 figures. Submitted to Briefings in Bioinformatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01372 2026-02-02 cs.IR cs.AI 79%

Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation

结构化频谱推理用于频率自适应多模态推荐

Wei Yang, Rui Zhong, Yiqun Chen, Chi Lu, Peng Jiang

机构 * Kuaishou Technology(快手科技) Renmin University of China(中国人民大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出结构化频谱推理框架,通过频谱分解、可靠性调节、超光谱融合和对比正则化,提升多模态推荐的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15820 2026-01-23 cs.CL 79%

ExDR: Explanation-driven Dynamic Retrieval Enhancement for Multimodal Fake News Detection

ExDR:基于解释的动态检索增强多模态虚假新闻检测

Guoxuan Ding, Yuqing Li, Ziyan Zhou, Zheng Lin, Daren Zha, Jiangnan Li

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) WeChat AI, Tencent(微信AI,腾讯)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 ExDR通过基于解释的动态检索增强生成框架,提升多模态虚假新闻检测的准确性和泛化能力。

Comments 11 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11970 2026-01-21 cs.CV 79%

Real-Time Multi-Modal Embedded Vision Framework for Object Detection Facial Emotion Recognition and Biometric Identification on Low-Power Edge Platforms

面向低功耗边缘平台的实时多模态嵌入视觉框架:用于目标检测、面部情绪识别和生物识别

S. M. Khalid Bin Zahid, Md. Rakibul Hasan Nishat, Abdul Hasib, Md. Rakibul Hasan, Md. Ashiqussalehin, Md. Sahadat Hossen Sajib, A. S. M. Ahsanul Sarkar Akib

机构 * 1,2Department of Mechatronics Engineering, Rajshahi University of Engineering \& Technology 3 ,4,5Department of Internet of Things Robotics Engineering, University of Frontier Technology, Bangladesh 6Department of Computer Science Engineering, Varendra University, Rajshahi, Bangladesh 7Department of Robotics, Robo Tech Valley, Dhaka, Bangladesh Emails: , , 3

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种面向低功耗边缘平台的实时多模态视觉框架,通过自适应调度机制整合目标检测、面部识别和情绪检测,提升边缘设备上的智能感知效率与隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09606 2026-01-15 cs.CV 79%

GRCF: Two-Stage Groupwise Ranking and Calibration Framework for Multimodal Sentiment Analysis

GRCF: 用于多模态情感分析的两阶段组内排序与校准框架

Manning Gao, Leheng Zhang, Shiqin Han, Haifeng Hu, Yuncheng Jiang, Sijie Mai

机构 * South China Normal University(华南师范大学) Sun Yat-sen University(中山大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 GRCF通过两阶段组内排序与校准框架,解决多模态情感分析中点对点回归的不足,提升预测稳定性和相关性对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10337 2026-01-15 cs.AI cs.LG 79%

A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering

一种基于RAG的强化学习课程学习方法:用于多模态问答

Chenliang Zhang, Lin Wang, Yuanyuan Lu, Yusheng Qi, Kexin Wang, Peixu Hou, Wenshi Chen

机构 * Meituan(美团)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种结合课程学习与强化学习的方法,用于多模态问答任务,通过检索增强生成系统在挑战中取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09278 2026-01-15 cs.AI 79%

M$^3$Searcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning

M$^3$Searcher: 模块化多模态信息检索代理与检索导向推理

Xiaohan Yu, Chao Feng, Lang Mei, Chong Chen

机构 * Huawei Cloud BU(华为云业务单元)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 M$^3$Searcher是一种模块化多模态信息检索代理,通过检索导向的多目标奖励机制,提升多模态任务中的事实准确性、推理合理性和检索忠实度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12735 2026-01-14 cs.CV 79%

Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning

通过多模态提示调优对开放词汇目标检测器进行后门攻击

Ankita Raj, Chetan Arora

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 TrAP通过多模态提示调优对开放词汇目标检测器实施后门攻击,利用轻量级提示标记植入恶意行为,提升攻击成功率并改进下游任务性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04297 2026-01-09 cs.LG cs.CV cs.HC cs.IR 79%

ArtCognition: A Multimodal AI Framework for Affective State Sensing from Visual and Kinematic Drawing Cues

ArtCognition:一种多模态AI框架,用于从视觉和运动学绘画线索中感知情绪状态

Behrad Binaei-Haghighi, Nafiseh Sadat Sajadi, Mehrad Liviyan, Reyhane Akhavan Kharazi, Fatemeh Amirkhani, Behnam Bahrak

机构 * Department of Electrical and Computer Engineering, University of Tehran(电信工程系,德黑兰理工大学) Tehran Institute for Advanced Studies, Khatam University(德黑兰高级研究院,凯塔姆大学) Department of Psychology, Allameh Tabataba’i University(心理学系,阿勒梅·塔巴塔比大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 ArtCognition通过融合视觉和运动学绘画线索,提供非侵入性情绪状态评估的新方法,验证了其在心理测量中的潜力。

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03135 2026-01-09 cs.AI 79%

Beyond Retrieval: Improving Evidence Quality for LLM-based Multimodal Fact-Checking

超越检索:提升基于大语言模型的多模态事实核查的证据质量

Haoran Ou, Gelei Deng, Xingshuo Han, Jie Zhang, Han Qiu, Shangwei Guo, Tianwei Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出Aletheia框架,通过改进证据检索策略提升多模态事实核查的证据质量,实验证明其在虚假信息检测中的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18867 2026-01-07 cs.AI 79%

Topological Perspectives on Optimal Multimodal Embedding Spaces

拓扑视角下的最优多模态嵌入空间

Abdul Aziz A. B, A. B Abdul Rahim

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文通过拓扑数据分析比较CLIP和CLOOB的嵌入空间,揭示其模态差距驱动因素和维度坍缩的影响,为多模态模型优化提供新视角。

Comments This manuscript contains substantive technical inaccuracies and an incomplete treatment of the stated topic. Subsequent developments and a reassessment of the problem indicate that the scope and framing of the work do not adequately reflect the current state of research, and the analysis is therefore incomplete and outdated

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01718 2026-01-06 cs.AI 79%

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

Yuan3.0 Flash:面向企业应用的开源多模态大语言模型

YuanLab. ai, :, Shawn Wu, Sean Wang, Louie Li, Darcy Chen, Allen Wang, Jiangang Luo, Xudong Zhao, Joseph Shen, Gawain Ma, Jasper Jia, Marcus Mao, Claire Wang, Hunter He, Carol Wang, Zera Zhang, Jason Wang, Chonly Shen, Leo Zhang, Logan Chen, Qasim Meng, James Gong, Danied Zhao, Penn Zheng, Owen Zhu, Tong Yu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 Yuan3.0 Flash通过RAPO算法提升企业任务性能,实现高效多模态大语言模型开源

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10393 2026-01-06 cs.SE cs.AI 79%

Cross-modal Retrieval Models for Stripped Binary Analysis

用于剥离二进制分析的跨模态检索模型

Guoqiang Chen, Lingyun Ying, Ziyang Song, Daguang Liu, Qiang Wang, Zhiqi Wang, Li Hu, Shaoyin Cheng, Weiming Zhang, Nenghai Yu

机构 * University of Science and Technology of China(中国科学技术大学) QI-ANXIN Technology Research Institute(QI-ANXIN技术研究所)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出BinSeek模型,通过两阶段跨模态检索框架实现对剥离二进制代码的高效检索,显著提升二进制代码与自然语言描述的相关性检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00623 2026-01-05 cs.AI 79%

DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations

DA-DPO:面向减少多模态大语言模型幻觉的高效难度感知偏好优化

Longtian Qiu, Shan Ning, Chuyu Zhang, Jiaxuan Sun, Xuming He

机构 * ShanghaiTech University(上海科技大学) Lingang Laboratory(灵冈实验室) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)

专题命中 跨模态检索 :MLLM(title);multimodal(abstract);分类 cs.AI

AI总结 DA-DPO通过难度感知机制优化多模态大语言模型的偏好学习,有效减少幻觉并提升模型鲁棒性与泛化能力。

Comments Accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13754 2025-12-30 cs.CV 79%

Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval

跨模态全模式细粒度对齐用于文本到图像人物检索

Hao Yin, Xin Man, Feiyu Chen, Jie Shao, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳先进研究所,电子科学与技术大学) University of Electronic Science and Technology of China(电子科学与技术大学) Sichuan Artificial Intelligence Research Institute(四川人工智能研究院)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出FMFA框架,通过显式细粒度对齐和隐式关系推理实现文本到图像人物检索的高精度匹配。

Comments accepted by ACM Transactions on Multimedia Computing Communications and Applications in December 2025

Journal ref ACM Transactions on Multimedia Computing Communications and Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21698 2025-12-29 cs.CR cs.MM eess.IV 79%

Raster Domain Text Steganography: A Unified Framework for Multimodal Secure Embedding

位图域文本隐写术:一种多模态安全嵌入的统一框架

A V Uday Kiran Kandala

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出了一种基于位图域的多模态安全嵌入框架,通过字形扰动和像素计数实现文本、图像、音频和视频等异构数据的隐蔽传输。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11712 2025-12-23 cs.AI 79%

Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization

通过理论一致的对称多模态偏好优化缓解幻觉

Wenqi Liu, Xuemeng Song, Jiaxi Li, Yinwei Wei, Na Zheng, Jianhua Yin, Liqiang Nie

机构 * Shandong University(山东大学) Southern University of Science and Technology(南方科技大学) University of Georgia(佐治亚大学) National University of Singapore(新加坡国立大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 SymMPO通过理论一致的对称多模态偏好优化方法,有效缓解多模态大语言模型中的幻觉问题。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17194 2025-12-22 cs.AI 79%

MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation

MMRAG-RFT: 两阶段强化微调用于可解释的多模态检索增强生成

Shengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu, Yuying Liu, Dong Zhang, Zhiqiang Tian, Shaoyi Du

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 MMRAG-RFT通过两阶段强化微调提升多模态检索增强生成的可解释性,实现更清晰的推理逻辑和更优的生成效果。

Comments This paper was accepted to AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16802 2025-12-19 cs.CL 79%

Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology

多模态检索增强生成在生物医学领域中的增强策略探索:一项评估糖生物学问答的案例研究

Primož Kocbek, Azra Frkatović-Hodžić, Dora Lalić, Vivian Hui, Gordan Lauc, Gregor Štiglic

机构 * University of Maribor, Faculty of Health Sciences(莫拉维亚大学健康科学学院) University of Ljubljana, Medical Factory(卢布尔雅那大学医疗工厂) Genos Ltd(基因公司) Center for Smart Health, School of Nursing The Hong Kong Polytechnic University(智能健康中心护理学院香港理工大学) University of Zagreb, Faculty of Pharmacy(扎格雷布大学药学院) Usher Institute University of Edinburgh(埃德蒙顿大学usher研究所)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

AI总结 本文研究了多模态检索增强生成在生物医学领域中的增强策略,通过实验发现多模态转换和视觉检索在不同模型中均能提升问答准确率。

Comments Will be published in IEEE BigData 2025 proceedings. Contains 10 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23316 2025-12-16 cs.CV 79%

C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

C3-OWD: 一种用于开放世界检测的课程跨模态对比学习框架

Siheng Wang, Zhengdao Li, Yanshu Li, Canran Xiao, Haibo Zhan, Zhengtao Yao, Xuzhi Zhang, Jiale Kang, Linshan Li, Weiming Liu, Zhikang Dong, Jifeng Shen, Junhao Dong, Qiang Sun, Piotr Koniusz

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 C3-OWD提出一种课程跨模态对比学习框架,通过预训练和视觉-语言对齐提升目标检测的鲁棒性和泛化能力。

Comments one of the authors doesn't agree any more

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10419 2025-12-12 cs.CV 79%

TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning

TransLocNet: 跨模态注意力用于航空-地面车辆定位的对比学习

Phu Pham, Damon Conover, Aniket Bera

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

AI总结 TransLocNet通过跨模态注意力和对比学习实现航空-地面车辆定位,显著提升定位精度和鲁棒性。

Comments 8 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏