arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2604.14846 2026-04-17 cs.CV cs.AI 79%

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

通过协调视觉模型实现零样本零售盗窃检测:一种无需训练单个模型的低成本替代方案

Haileab Yagersew

机构 * Paza AI

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Paza框架,通过协调多个现有模型实现零样本零售盗窃检测,降低训练成本,利用多信号预过滤减少昂贵视觉语言模型调用次数,提升效率和准确性。

Comments 16 pages, 3 figures, Code to be released at https://github.com/xHaileab/Paza-AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15946 2026-04-17 cs.LG cs.AI cs.CR 79%

Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval

跌入陷阱,获得智慧:通过误判风险模式检索的认知引导有害迷因检测

Wenshuo Wang, Ziyou Jiang, Junjie Wang, Mingyang Li, Jie Huang, Yuekai Huang, Zhiyuan Chang, Feiyan Duan, Qing Wang

机构 * State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) Science and Technology on Integrated Information System Laboratory(集成信息系统技术研究所) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出PatMD方法,通过学习并主动缓解潜在误判风险,识别有害迷因的深层误判风险模式,提升多模态大语言模型的检测能力,实验显示在5项有害检测任务中,F1-score和准确率均显著提升。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12357 2026-04-15 cs.AI cs.CV 79%

ReflectCAP: Detailed Image Captioning with Reflective Memory

ReflectCAP: 基于反思记忆的详细图像描述

Kyungmin Min, Minbeom Kim, Kang-il Lee, Seunghyun Yoon, Kyomin Jung

机构 * IPAI, Seoul National University(首尔国立大学IPAI) Dept. of ECE, Seoul National University(首尔国立大学电子与计算机工程系) Adobe Research(Adobe研究)

专题命中 视觉定位与Grounding :vision-language model(abstract);InternVL(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 ReflectCAP通过反思笔记指导生成详细且准确的图像描述,平衡事实性和覆盖性,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04099 2026-04-14 cs.CV cs.AI 79%

TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection

TARAC:通过时间注意力实时累积连接缓解大型视觉-语言模型中的幻觉

Lei Jiang, Chunzhao Xie, Tongxuan Liu, Yuting Zeng, jinrong Guo, Yunheng Shen, Weizhe Huang, Jing Li, Xiaohua Xu

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出TARAC,一种无需训练的框架,通过动态累积和重注入历史注意力来维持视觉 grounding,实验表明其在减少幻觉和提升感知得分方面优于现有方法,且计算开销低。

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05117 2026-04-08 cs.CV cs.AI cs.CL 79%

Watch Before You Answer: Learning from Visually Grounded Post-Training

在回答前观看:从视觉引导的后训练中学习

Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Etude AI Kolors Team, Kuaishou Technology(快手科技Kolors团队) University of Toronto(多伦多大学) University of Waterloo(滑铁卢大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文发现现有视频理解基准中40-60%的问题可通过文本线索回答,提出VidGround方法通过仅使用视觉引导问题提升VLM性能,实验显示其在后训练中效果优于复杂方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24133 2026-03-30 cs.CV cs.AI 79%

Compositional Image Synthesis with Inference-Time Scaling

基于推理时缩放的图像合成

Minsuk Ji, Sanghyeok Lee, Namhyuk Ahn

机构 * Inha University(仁荷大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种无需训练的框架,结合对象中心方法与自我完善,提升布局忠实度并保持美学质量,通过大语言模型生成布局并注入图像生成过程,利用对象中心视觉-语言模型进行迭代选择,实现更强的场景对齐。

Comments projcet page: https://github.com/gcl-inha/ReFocus

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17307 2026-03-19 cs.CV cs.AI 79%

Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding

Symphony:一种受认知启发的多智能体系统用于长视频理解

Haiyang Yan, Hongyun Zhou, Peng Xu, Xiaoxue Feng, Mengyi Liu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Kuaishou Technology(快手科技) School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学未来技术学院)

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Symphony多智能体系统,通过模拟人类认知模式,提升长视频理解任务的推理能力,实验显示在多个基准测试中表现优异。

Comments Accepted by cvpr2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12382 2026-03-16 cs.CV cs.AI 79%

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

SPARROW:在像素级视频MLLM中学习空间精度和时间参照一致性

Mohamad Alansari, Naufal Suryanto, Divya Velayudhan, Sajid Javed, Naoufel Werghi, Muzammal Naseer

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 SPARROW通过引入目标特定跟踪特征和双提示设计,提升视频像素级理解的空间精度和时间一致性,基于30646个视频和45231对问答对的数据集,实现六个基准测试的显著改进。

Comments Accepted at CVPR 2026; Project page: https://risys-lab.github.io/SPARROW; Repository: https://github.com/RISys-Lab/SPARROW

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20119 2026-02-24 cs.RO cs.AI cs.CV 79%

NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning

NovaPlan: 通过闭环视频语言规划实现零样本长周期操控

Jiahui Fu, Junyu Nan, Lingfeng Sun, Hongyu Li, Jianing Qian, Jennifer L. Barry, Kris Kitani, George Konidaris

机构 * Robotics and AI Institute(机器人与人工智能研究所) Carnegie Mellon University(卡内基梅隆大学) Brown University(布朗大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 NovaPlan通过闭环视频语言规划实现零样本长周期操控,结合高层任务分解与底层物理执行,无需先验演示即可完成复杂装配任务。

Comments 25 pages, 15 figures. Project webpage: https://nova-plan.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11304 2026-01-23 cs.AI cs.CL cs.CV 79%

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

利用实例分割辅助的多模态大语言模型进行智能交通监控

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

机构 * University of Ottawa(渥太华大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文利用多模态大语言模型和实例分割技术,实现高准确率的交通监控系统,提升交通管理效率和安全性。

Comments 6 pages, 7 figures, submitted to 30th IEEE International Symposium on Computers and Communications (ISCC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22351 2026-01-08 cs.CV cs.AI 79%

VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement

VULCAN:工具增强的多智能体用于迭代3D物体排列

Zhengfei Kuang, Rui Lin, Long Zhao, Gordon Wetzstein, Saining Xie, Sanghyun Woo

机构 * Stanford University(斯坦福大学) Google(谷歌) New York University(纽约大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 VULCAN通过引入MCP API、视觉工具和多智能体框架,提升了3D物体排列任务中MLLMs的视觉接地能力与迭代处理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24023 2026-01-01 cs.CV cs.AI 79%

RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations

RSAgent: 通过多轮工具调用学习推理与行动进行文本引导的分割

Xingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li, Lingyi Hong, Mingxi Chen, Kaixun Jiang, Jiyuan Fu, Wenqiang Zhang

机构 * Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(上海智能信息处理关键实验室,计算机科学与人工智能学院,复旦大学) Artificial Intelligence, Fudan University(人工智能,复旦大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 RSAgent通过多轮工具调用实现文本引导分割的推理与行动,采用两阶段框架提升分割性能,达到领域内和领域外基准的最先进水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00294 2025-12-02 cs.CV cs.AI cs.HC 79%

Words into World: A Task-Adaptive Agent for Language-Guided Spatial Retrieval in AR

词语进入世界:一种任务自适应代理用于增强现实中的语言引导空间检索

Lixing Guo, Tobias Höllerer

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种任务自适应AR代理,结合多模态大语言模型和视觉模型,实现语言引导的空间检索,支持复杂查询和真实世界交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20531 2025-10-24 cs.CV cs.AI 79%

Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis

Lixiong Qin, Yang Zhang, Mei Wang, Jiani Hu, Weihong Deng, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 9 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19599 2025-10-23 cs.CV cs.AI 79%

XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography

Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou, Sebastian Otalora, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士) Shanghai Jiao Tong University, China(上海交通大学,中国) Kaiko.AI, Switzerland(Kaiko.AI,瑞士) Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10426 2025-10-14 cs.CV cs.AI 79%

Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs

Suyang Xi, Chenxi Yang, Hong Ding, Yiqing Ni, Catherine C. Liu, Yunhao Liu, Chengqi Zhang

机构 * Emory University(埃默里大学) University of Electronic Science and Technology of China(电子科技大学) University of Illinois Chicago(伊利诺伊大学香槟分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10424 2025-09-16 cs.CV cs.AI 79%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 视觉定位与Grounding :vision language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13111 2025-09-09 cs.CV cs.CL cs.LG 79%

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Erik Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Kai Kang, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch

机构 * Apple(苹果公司)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22864 2025-07-01 cs.CV cs.AI cs.CL 79%

Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval

Li-Cheng Shen, Jih-Kang Hsieh, Wei-Hua Li, Chu-Song Chen

机构 * National Taiwan University(国立台湾大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06176 2025-05-12 cs.GR cs.CV cs.LG 79%

MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills

Niladri Shekhar Dutt, Duygu Ceylan, Niloy J. Mitra

机构 * University College London(伦敦大学学院) Adobe Research UK(Adobe英国研究)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments Accepted at SIGGRAPH 2025 [ACM Transactions on Graphics]; Project website: https://monetgpt.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09720 2025-02-04 cs.CV cs.AI 79%

A Simple Aerial Detection Baseline of Multimodal Language Models

Qingyun Li, Yushi Chen, Xinya Shu, Dong Chen, Xin He, Yi Yu, Xue Yang

专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 4 pages, 1 table, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00258 2024-06-04 cs.CV cs.AI 79%

Artemis: Towards Referential Understanding in Complex Videos

Jihao Qiu, Yuan Zhang, Xi Tang, Lingxi Xie, Tianren Ma, Pengyu Yan, David Doermann, Qixiang Ye, Yunjie Tian

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 19 pages, 14 figures. Code and data are available at https://github.com/qiujihao19/Artemis

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12714 2024-02-06 cs.CV cs.AI 79%

VIGC: Visual Instruction Generation and Correction

Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, Conghui He

专题命中 视觉定位与Grounding :vision-language model(abstract);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI 2024, Project Website: https://opendatalab.github.io/VIGC, Code and Pretrained Model: https://github.com/opendatalab/VIGC

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05546 2023-03-13 cs.CV cs.AI 79%

Weakly-Supervised HOI Detection from Interaction Labels Only and Language/Vision-Language Priors

Mesut Erhan Unal, Adriana Kovashka

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10105 2025-09-17 cs.CV cs.CL 78%

VARCO-VISION-2.0 Technical Report

Young-rok Cha, Jeongho Ju, SunYoung Park, Jong-Hyeon Lee, Younghyun Yu, Youngjune Kim

机构 * NC AI

专题命中 视觉定位与Grounding :VLM(abstract,comments);vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments 19 pages, 1 figure, 14 tables. Technical report for VARCO-VISION-2.0, a Korean-English bilingual VLM in 14B and 1.7B variants. Key features: multi-image understanding, OCR with text localization, improved Korean capabilities

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11426 2026-08-18 cs.RO 版本更新 78%

Grounding Robot Generalization in Training Data via Retrieval-Augmented VLMs

通过检索增强的视觉语言模型实现机器人在训练数据中的泛化

Jensen Gao, Dorsa Sadigh, Sandy Huang, Dhruv Shah

机构 * Stanford University(斯坦福大学) Google DeepMind(谷歌DeepMind) Princeton University(普林斯顿大学)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract)

AI总结 本文提出RADAR框架,通过检索和视觉语言模型分析,评估机器人策略在不同场景下的泛化能力,验证了其在大规模数据集中的有效性。

Comments IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13216 2026-08-14 cs.HC 新提交 78%

CogChat: Knowledge Graph-Augmented Conversational AI with Heterogeneous Graph Transformer for Cognitive Grounding in Design Generation

CogChat:结合知识图谱与异构图变换器的对话式人工智能,用于设计生成中的认知接地

Jiin Choi, Kyung Hoon Hyun

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 本研究提出结合异构图变换器的CogChat框架,以设计师专属动态知识图谱为基础,提升设计对话的上下文保留度与意图解读效果,降低认知负荷,为LLM交互的长期上下文管理提供新方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12894 2026-08-14 cs.CL 新提交 78%

BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

BavGround:巴伐利亚区域文化接地与方言能力基准

Jophin John, Michael Hoffmann, Jan Fillies, Michael A. Hedderich, Barbara Plank

机构 * Stanford University(斯坦福大学) Freie Universität Berlin(柏林自由大学) Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Leibniz Supercomputing Centre (LRZ)(莱布尼茨超级计算中心)

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 该研究推出BavGround基准,评估LLM的巴伐利亚区域文化接地与方言能力,发现多语言模型在巴伐利亚语及源接地问题上表现欠佳,且评估协议会显著影响结论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10368 2026-08-12 cs.HC cs.MM 新提交 78%

Visual-to-Haptic Augmentation in XR: A Wearable Glove for Perceptual Grounding in Multimodal Interaction

XR中的视觉到触觉增强:一种用于多模态交互感知接地的可穿戴手套

Faisal Mohd, Hamdi Elsaddik, Erhan Baturay Onural, Jihong Zhang, Fedwa Laamarti, Abdulmotaleb El Saddik

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 该研究针对XR系统触觉感知利用不足的问题,提出一款可穿戴手套及视觉到触觉映射算法,经20人实验验证,其触觉增强可提升XR交互的真实感与沉浸感,为多模态XR系统提供感知增强层。

Comments 11 pages, 4 figures. Published in the Proceedings of the 1st Workshop on Shaping Future Human Connection: Social Augmentation through XR Technologies (SAXR 2026), April 13, 2026, Barcelona, Spain

Journal ref Proceedings of the 1st Workshop on Shaping Future Human Connection: Social Augmentation through XR Technologies (SAXR 2026), CEUR Workshop Proceedings, Vol. 4226, pp. 252-262, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13023 2026-08-11 cs.SD cs.MM 版本更新 78%

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding

SpotSound: 通过细粒度时间定位增强大音频-语言模型

Luoyi Sun, Xiao Zhou, Zeqian Li, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室) SAI, Shanghai Jiao Tong University(上海交通大学人工智能院)

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 SpotSound通过引入新的训练目标和挑战性基准,提升了大音频-语言模型在时间定位任务中的性能,同时保持了通用任务的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏