arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-24 至 2026-03-24 共收录 148 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 17 篇

2603.15093 2026-03-24 eess.SP cs.IT math.IT 82%

Beam Prediction Based on Multimodal Large Language Models

基于多模态大语言模型的波束预测

Tianhao Mao, Le Liang, Jie Yang, Xiao Li, Shi Jin, Geoffrey Ye Li

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出基于多模态大语言模型的波束预测框架,通过融合RGB图像和LiDAR点云等异构数据,提升动态环境下波束预测精度与通信性能,实验表明其在波束准确性和通信性能上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21778 2026-03-24 cs.CV 79%

Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

Scene-VLM:通过视觉-语言模型进行多模态视频场景分割

Nimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi, Asaf Gendler, Ilan Naiman, Erez Yosef, Igor Kviatkovsky

机构 * Ben-Gurion University(本·古里安大学) Amazon Prime Video(亚马逊Prime视频) Tel-Aviv University(特拉维夫大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Scene-VLM,首个基于视觉-语言模型的视频场景分割框架,通过融合视觉与文本信息实现多模态推理,提升场景分割的准确性和可解释性。

Comments Accepted for publication at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20351 2026-03-24 cs.CR cs.AI 79%

MANA: Towards Efficient Mobile Ad Detection via Multimodal Agentic UI Navigation

MANA:通过多模态代理UI导航实现高效的移动广告检测

Yizhe Zhao, Yongjian Fu, Zihao Feng, Hao Pan, Yongheng Deng, Yaoxue Zhang, Ju Ren

机构 * Department of Computer Science and Technology(计算机科学与技术系) University of Southern California(南加州大学) Shanghai Jiao Tong University(上海交通大学) State Key Laboratory of Internet Architecture(互联网架构国家重点实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出MANA框架,通过整合多种信号实现高效移动广告检测,提升准确率和效率,有效识别隐蔽恶意广告。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17611 2026-03-24 nucl-th astro-ph.SR cond-mat.supr-con nucl-ex quant-ph 78%

Evidence for Multimodal Superfluidity of Neutrons

中子的多模超流性证据

Yuan-Zhuo Ma, Georgios Palkanoglou, Joseph Carlson, Stefano Gandolfi, Alexandros Gezerlis, Gabriel Given, Ashe Hicks, Dean Lee, Kevin E. Schmidt, Jiabin Yu

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 研究通过理论和实验证据揭示了中子富集系统中一种新物相——多模超流性,探讨其在不同维度模型中的表现及对中子星结构的影响。

Comments 61 pages, 37 figures; Minor revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20354 2026-03-24 cs.MM cs.AI 73%

Leum-VL Technical Report

Leum-VL 技术报告

Yuxuan He, Chaiming Huang, Yifan Wu, Hongjun Wang, Chenkui Shen, Jifan Zhang, Long Li

机构 * Hainan Sihe Data Technology Co., Ltd.(海南世和数据技术有限公司)

专题命中 视频多模态 :multimodal(abstract);image-text(abstract);分类 cs.AI、cs.MM

AI总结 本文提出SV6D框架,通过六维结构化表示解析视频内容,构建Leum-VL-8B模型,在多个基准测试中取得优异成绩,强调视频AI需关注结构表示而非像素生成。

Comments 27 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18220 2026-03-24 cs.CV cs.AI 62%

UASTrack: A Unified Adaptive Selection Framework with Modality-Customization in Single Object Tracking

UASTrack: 一种具有模态定制化的统一自适应选择框架在单目标跟踪中

He Wang, Tianyang Xu, Zhangyong Tang, Xiao-Jun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing, University of Surrey(萨里大学视觉、语音和信号处理中心)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 UASTrack通过统一自适应选择框架实现多模态跟踪中的模态定制化,有效过滤噪声并提升跟踪性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20422 2026-03-24 cs.CV cs.AI cs.IR 62%

PEARL: Personalized Streaming Video Understanding Model

PEARL:个性化流视频理解模型

Yuanhong Zheng, Ruichuan An, Xiaopeng Lin, Yuxing Liu, Sihan Yang, Huanyu Zhang, Haodong Li, Qintong Zhang, Renrui Zhang, Guopeng Li, Yifan Zhang, Yuheng Li, Wentao Zhang

机构 * Peking University(北京大学) Adobe CASIA Stepfun CUHK(香港中文大学) Zhongguancun Academy(中关村学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出个性化流视频理解任务PSVU,并引入PEARL-Bench基准测试,通过帧级和视频级两种模式评估模型对个性化概念的响应能力,提出无需训练的PEARL策略,实现先进性能。

Comments Arxiv Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20066 2026-03-24 cs.CV 57%

PAUL: Uncertainty-Guided Partition and Augmentation for Robust Cross-View Geo-Localization under Noisy Correspondence

PAUL:基于不确定性的分区与增强以实现在噪声对应下的鲁棒跨视图地理定位

Zheng Li, Xueyi Zhang, Yanming Guo, Yuxiang Xie, Ding Zhaoyun, Siqi Cai, Haizhou Li, Mingrui Lao

机构 * National University of Defense Technology(国防科技大学) Shenzhen Loop Area Institute(深圳河套学院) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学深圳人工智能学院) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出PAUL框架,通过不确定性感知的共增强和证据共训练,解决跨视图地理定位中的噪声对应问题,提升鲁棒性。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20887 2026-03-24 cs.CV 57%

Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning

基于场景图的细粒度SegCaptioning Transformer用于可控视频分割与标注

Xu Zhang, Jin Yuan, BinHong Yang, Xuan Liu, Qianjun Zhang, Yuyi Wang, Zhiyong Li, Hanwang Zhang

机构 * College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) School of Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学) School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算机与人工智能学院,西南交通大学) CRRC Zhuzhou Institute Company Ltd.(中车株洲院有限公司) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出可控视频分割与标注任务,通过Scene Graph-guided Fine-grained SegCaptioning Transformer框架,结合提示引导时间图模块和细粒度掩码-语言解码器,实现用户意图的精准捕捉与多模态输出生成。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05175 2026-03-24 cs.CV 57%

VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice

VideoAuto-R1:通过一次推理、两次回答实现视频自动推理

Shuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen, Lemeng Wu, Zechun Liu, Chenchen Zhu, Zhipeng Cai, Chong Zhou, Haozhe Liu, Ernie Chang, Saksham Suri, Hongyu Xu, Qi Qian, Wei Wen, Balakrishnan Varadarajan, Zhuang Liu, Hu Xu, Florian Bordes, Raghuraman Krishnamoorthi, Bernard Ghanem, Vikas Chandra, Yunyang Xiong

机构 * Meta AI King Abdullah University of Science and Technology (KAUST)(卡布斯大学) Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出VideoAuto-R1框架,通过一次推理两次回答策略提升视频理解效率,实现准确率和效率的双重提升。

Comments Accepted to CVPR 2026. Project page: https://ivul-kaust.github.io/projects/videoauto-r1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05105 2026-03-24 cs.CV 57%

UniLiPs: Unified LiDAR Pseudo-Labeling with Geometry-Grounded Dynamic Scene Decomposition

UniLiPs: 基于几何引导的动态场景分解的统一激光雷达伪标签方法

Filippo Ghilotti, Samuel Brucker, Nahku Saidy, Matteo Matteucci, Mario Bijelic, Felix Heide

机构 * TORC Robotics(TORC机器人公司) Politecnico of Milan(米兰理工学院) Princeton University(普林斯顿大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出UniLiPs方法,通过利用激光雷达扫描的时空几何一致性,将文本和2D视觉模型的线索直接融合到3D中,生成3D语义标签、包围盒和密集点云,验证了其在三个数据集上的鲁棒性。

Journal ref Proceedings of the International Conference on 3D Vision (3DV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12656 2026-03-24 cs.CV 57%

SPKLIP: Aligning Spike Video Streams with Natural Language

SPKLIP:通过自然语言对齐尖峰视频流

Yongchang Gao, Meiling Jin, Zhaofei Yu, Tiejun Huang, Guozhang Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 SPKLIP提出了一种专门用于尖峰视频-语言对齐的架构,通过分层尖峰特征提取器和尖峰-文本对比学习实现多尺度时间动态建模,并在稀疏异步输出上实现高效的少样本学习。

Comments A dataset partitioning error occurred and is being corrected

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10448 2026-03-24 cs.RO 50%

DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control

DiT4DiT:联合建模视频动态与动作以实现通用机器人控制

Teli Ma, Jia Zheng, Zifan Wang, Chunli Jiang, Andy Cui, Junwei Liang, Shuo Yang

机构 * Mondo Robotics(Mondo机器人公司) HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学)

专题命中 视频多模态 :image-text(abstract)

AI总结 本文提出DiT4DiT模型,通过结合视频扩散变换器与动作扩散变换器,实现视频动态与动作的联合建模,提升机器人控制的泛化能力与样本效率。

Comments https://dit4dit.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 18 篇

2511.01946 2026-03-24 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph 88%

COFAP: A Universal Framework for COFs Adsorption Prediction through Designed Multi-Modal Extraction and Cross-Modal Synergy

COFAP:通过设计的多模态提取和跨模态协同的通用COFs吸附预测框架

Zihan Li, Mingyang Wan, Mingyu Gao, Xishi Tai, Zhongshan Chen, Xiangke Wang, Feifan Zhang

机构 * College of Science, College of Information and Electrical Engineering(科学学院,信息与电气工程学院) China Agricultural University(中国农业大学) Qingdao Institute of Software, College of Computer Science and Technology(软件研究所,计算机科学与技术学院) China University of Petroleum (East China)(中国石油大学(华东)) Weifang university(潍坊大学) College of Environmental Science and Engineering(环境科学与工程学院) North China Electric Power University(华北电力大学) College of Science(科学学院)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出COFAP框架,通过深度学习提取多模态结构和化学特征,并利用跨模态注意力机制融合特征,实现高效COFs吸附预测,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19961 2026-03-24 cs.CL cs.IR 83%

Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval

解锁多模态文档智能:从当前成就到视觉文档检索的未来前沿

Yibo Yan, Jiahao Huo, Guanbo Feng, Mingdong Ou, Yi Cao, Xin Zou, Shuliang Liu, Yuanhuiyi Lyu, Yu Huang, Jungang Li, Kening Zheng, Xu Zheng, Philip S. Yu, James Kwok, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Alibaba Cloud Computing(阿里云计算) Hong Kong University of Science and Technology(香港科技大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

AI总结 本文综述了视觉文档检索领域,探讨了多模态大语言模型时代下的方法演进与挑战,提出未来发展方向。

Comments Under review. This version updates the relevant works released before 15 March, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23306 2026-03-24 cs.CV 83%

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

ThinkOmni: 通过指导解码将文本推理提升到多模态场景

Yiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 跨模态检索 :omni-modal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 ThinkOmni提出一种无需训练和数据的框架,通过指导解码将文本推理扩展到多模态场景,实验显示在多个多模态推理基准上取得显著提升。

Comments Accept by ICLR 2026, Code: https://github.com/1ranGuan/thinkomni

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14408 2026-03-24 cs.CV cs.AI 81%

Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection

基于特征重校准的嗅觉-视觉多模态模型用于增强稻米劣变检测

Rongqiang Zhao, Hengrui Hu, Yijing Wang, Mingchun Sun, Jie Liu

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) National Key Laboratory of Smart Farm Technologies and Systems(国家智能农业技术与系统重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于特征重校准的嗅觉-视觉多模态模型,通过改进的特征嵌入构造器和重校准注意力网络,提升稻米劣变检测的准确性和效率,相比传统方法提升11.51%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04846 2026-03-24 cs.CV 79%

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

多范式协同对抗攻击多模态大语言模型

Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对多模态大语言模型的多范式协同对抗攻击方法,通过聚合视觉和语言特征进行联合优化,提升对抗示例的可转移性,实验表明优于现有方法。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 79%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20970 2026-03-24 cs.CV 79%

GraPHFormer: A Multimodal Graph Persistent Homology Transformer for the Analysis of Neuroscience Morphologies

GraPHFormer:一种多模态图持久同调变换器,用于神经科学形态学分析

Uzair Shah, Marco Agus, Mahmoud Gamal, Mahmood Alzubaidi, Corrado Cali, Pierre J. Magistretti, Abdesselam Bouzerdoum, Mowafa Househ

机构 * Hamad Bin Khalifa University(哈马德·本·卡西姆大学) University of Turin(都灵大学) BESE, King Abdullah University of Science and Technology(贝赛,国王阿卜杜勒阿齐兹大学科学与技术学院) University of Wollongong(沃林根大学) Neuroscience Institute Cavalieri Ottolenghi(卡瓦利埃-奥托伦奇神经科学研究所) Université Grenoble-Alpes(格勒诺布尔阿尔卑斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 GraPHFormer通过CLIP式对比学习统一拓扑和图结构分析,利用持久图像编码和树状LSTM编码器,实现对神经形态的高精度识别与分类,优于传统方法。

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06638 2026-03-24 cs.CV cs.AI 73%

StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering

StaR-KVQA:用于隐式知识视觉问答的结构化推理轨迹

Zhihao Wen, Wenkang Wei, Yuan Fang, Xingtong Yu, Hui Zhang, Weicheng Zhu, Xin Zhang

机构 * Ant International, Ant Group(蚂蚁集团国际部,蚂蚁集团) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院) Anhui Provincial Key Laboratory of High Performance Computing(安徽省高性能计算重点实验室)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 StaR-KVQA通过引入双路径结构化推理轨迹提升隐式知识视觉问答的准确性与推理透明度,采用自蒸馏方法构建轨迹增强数据集,无需外部检索工具,在OK-VQA基准上实现11.3%的精度提升。

Comments 8+3+3 pages, code: https://github.com/jianyingzhihe/StaR-KVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21886 2026-03-24 cs.IR cs.CV 70%

ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval

ADaFuSE: 适应性扩散生成图像与文本融合用于交互式文本到图像检索

Zhuocheng Zhang, Xingwu Zhang, Kangheng Liang, Guanxuan Li, Richard Mccreadie, Zijun Long

机构 * Hunan University(湖南大学) University of Glasgow(格拉斯哥大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 ADaFuSE通过引入双分支融合机制,结合自适应门控和语义感知专家混合,提升交互式文本到图像检索的性能,实现更稳健的多模态融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21135 2026-03-24 cs.CV cs.AI 62%

One Pool Is Not Enough: Multi-Cluster Memory for Practical Test-Time Adaptation

一个池子不够:多簇内存用于实用测试时适应

Yu-Wen Tseng, Xingyi Zheng, Ya-Chen Wu, I-Bin Liao, Yung-Hui Li, Hong-Han Shuai, Wen-Huang Cheng

机构 * National Taiwan University(国立台湾大学) National Yang Ming Chiao Tung University(国立阳明交通大学) Hon Hai Research Institute(宏海研究院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多簇内存(MCM)以解决实用测试时适应中单池存储的不足,通过多簇组织提升适应稳定性,实验证明在多个数据集上均取得显著提升。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20818 2026-03-24 cs.CV cs.AI 62%

PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure Matching

PlanaReLoc:通过基于区域的结构匹配实现3D平面原语的相机重定位

Hanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan Shen

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出PlanaReLoc,利用3D平面原语和地图进行轻量级6自由度相机重定位,通过深度匹配和统一嵌入空间实现可靠的跨模态结构对应。

Comments Accepted by CVPR 2026. 20 pages, 15 figures. Code at https://github.com/3dv-casia/PlanaReLoc

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21925 2026-03-24 cs.AI 57%

Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support

基于指南的检索增强生成用于眼科临床决策支持

Shuying Chen, Sen Cui, Zhong Cao

机构 * University of International Business and Economics(国际商务经济大学) Tsinghua University(清华大学) Heidelberg Institute of Global Health(海德堡全球健康研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 本文提出Oph-Guid-RAG系统,通过多模态视觉RAG方法提升眼科临床问答与决策支持,通过可控检索框架和多模态推理提升证据基础和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01049 2026-03-24 cs.CV cs.RO 57%

KeySG: Hierarchical Keyframe-Based 3D Scene Graphs

KeySG:基于关键帧的层次化3D场景图

Abdelrhman Werby, Dennis Rotondi, Fabio Scaparro, Kai O. Arras

机构 * Socially Intelligent Robotics Lab, Institute for Artificial Intelligence University of Stuttgart, Germany(社会智能机器人实验室,人工智能研究所,斯图加特大学,德国)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 KeySG通过层次化3D场景图结构,结合关键帧提取多模态信息,提升复杂环境中的语义推理与规划能力,优于现有方法。

Comments Code and video are available at https://keysg-lab.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13876 2026-03-24 cs.CV 57%

Scene Prior Filtering for Depth Super-Resolution

场景先验过滤用于深度超分辨率

Zhengxue Wang, Zhiqiang Yan, Ming-Hsuan Yang, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang

机构 * PCA Lab, Nanjing University of Science and Technology, China(南京理工大学科学与工程学院实验室) National University of Singapore(新加坡国立大学) University of California, USA, and Yonsei University, South Korea(美国加州大学和韩国延世大学) PCA Lab, Nanjing University, China(南京大学实验室)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出SPFNet网络,利用大尺度模型的表面法线和语义图作为先验信息,通过多模态先验传播和一对一先验嵌入减少纹理干扰并提升边缘表示,实现在真实和合成数据集上的超分辨率性能。

Comments Accepted to IJCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21083 2026-03-24 cs.CV 57%

Hierarchical Text-Guided Brain Tumor Segmentation via Sub-Region-Aware Prompts

基于子区域感知提示的层次化文本引导脑肿瘤分割

Bahram Mohammadi, Ta Duc Huy, Afrouz Sheikholeslami, Qi Chen, Vu Minh Hieu Phan, Sam White, Minh-Son To, Xuyun Zhang, Amin Beheshti, Luping Zhou, Yuankai Qi

机构 * Macquarie University(麦考瑞大学) Adelaide University(阿德莱德大学) Flinders University(弗林德斯大学) University of Sydney(悉尼大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出TextCSP框架,通过文本引导和子区域感知提示提升脑肿瘤分割精度,实验显示在Dice和HD95指标上优于现有方法。

Comments 10 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12985 2026-03-24 cs.LG cs.CV 57%

Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks

角梯度符号法:揭示双曲网络中的漏洞

Minsoo Jo, Dongyoon Yang, Taesup Kim

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种基于双曲空间几何性质的对抗攻击方法,通过分解梯度为径向和角向分量,生成高影响的对抗样本,提升分类和检索任务的欺骗率。

Comments Accepted by AAAI 2026. Code available at: https://github.com/J-Minsoo/AGSM

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(7), 5566-5574, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20513 2026-03-24 cs.IR cs.AI 57%

ReBOL: Retrieval via Bayesian Optimization with Batched LLM Relevance Observations and Query Reformulation

ReBOL:通过批量LLM相关性观察和查询重述进行检索

Anton Korikov, Scott Sanner

机构 * University of Toronto(多伦多大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 ReBOL通过引入多模态贝叶斯优化和查询重述技术,改进检索阶段的召回率和排名质量,优于现有LLM重排序基线。

详情

展开后加载摘要…

URL PDF HTML 收藏