arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1566 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1566 篇

2509.17429 2026-01-27 cs.CV 57%

Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration

多尺度时间预测 via 逐步生成和多智能体协作

Zhitao Zeng, Guojian Yuan, Junyuan Mao, Yuxuan Wang, Xiaoshuang Jia, Yueming Jin

机构 * National University of Singapore(新加坡国立大学) Alibaba Group(阿里巴巴集团) Renmin University of China(中国人民大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种多尺度时间预测方法,通过逐步生成和多智能体协作,提升多尺度和多状态预测的准确性和一致性。

Comments 20 pages, 6 figures

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23927 2026-01-26 cs.CV 57%

FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing

FUSAR-KLIP:迈向遥感多模态基础模型

Yi Yang, Xiaokun Zhang, Qingchen Fang, Jing Liu, Ziqi Ye, Rui Li, Li Liu, Haipeng Wang

机构 * Key Laboratory for Information Science of Electromagnetic Waves (MoE), Fudan University(电磁波信息科学重点实验室(MoE),复旦大学) Institute of Zhejiang Laboratory(浙江实验室研究院) College of Electronic Science and Technology, NUDT(电子科学与技术学院,南大学)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

AI总结 FUSAR-KLIP是首个针对SAR图像的多模态基础模型,通过构建大规模数据集、生成结构化文本、设计自洽优化机制和建立统一评估基准,解决遥感图像与通用视觉表示之间的认知不一致问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04634 2026-01-26 cs.CV 57%

Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models

你所要求的是你所得到的吗?探究文本到图像模型中的概念关联

Salma Abdel Magid, Weiwei Pan, Simon Warchol, Grace Guo, Junsik Kim, Mahia Rahman, Hanspeter Pfister

机构 * Department of Computer Science(计算机科学系) Harvard University(哈佛大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

AI总结 本文提出 Concept2Concept 框架,用于审计文本到图像模型中提示与生成内容之间的概念关联,通过可解释的概念和度量标准进行可视化分析。

Journal ref Trans. Mach. Learn. Res, 2835-8856, 2025, https://openreview.net/forum?id=mk1YIkVvTQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15308 2026-01-23 cs.HC cs.AI 57%

When Generative AI Meets Extended Reality: Enabling Scalable and Natural Interactions

当生成式AI遇见扩展现实:实现可扩展和自然的交互

Mingyu Zhu, Jiangong Chen, Bin Li

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

AI总结 本文探讨生成式AI与扩展现实的结合,通过三个用例展示如何通过语言驱动交互和自动化内容生成解决XR在可扩展性和自然交互方面的挑战。

Comments Accepted by IEEE Internet Computing (Oct. 2025); published in IEEE Xplore (Jan. 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14732 2026-01-22 cs.CV cs.CL cs.MM 57%

DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling

DeepMoLM: 利用视觉和几何结构信息进行分子-文本建模

Jing Lan, Hexiao Ding, Hongzhao Chen, Yufeng Jiang, Nga-Chun Ng, Gwing Kei Yip, Gerald W. Y. Cheng, Yunlin Mao, Jing Cai, Liang-ting Lin, Jung Sun Yoo

机构 * Department of Health Technology and Informatics, The Hong Kong Polytechnic University(健康科技与信息学系,香港理工大学) Department of Nuclear Medicine and PET, Hong Kong Sanatorium and Hospital(核医学与PET部,香港疗养院及医院) Department of Diagnostic and Interventional Radiology, Queen Elizabeth Hospital Hong Kong SAR, China(诊断与介入放射学部,香港特别行政区中国女王医院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 DeepMoLM通过双视角框架结合视觉和几何信息,提升分子-文本建模的准确性与物理合理性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13133 2026-01-21 cs.CV 57%

CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks

面向人类的视觉任务的CLIP引导自监督学习

Mingshuang Luo, Ruibing Hou, Bo Chao, Hong Chang, Zimo Liu, Yaowei Wang, Shiguang Shan

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS), and University of Chinese Academy of Sciences(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)及中国科学院大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 CLASP通过CLIP生成多级语义伪标签,并结合Prompt-Controlled MoE模块提升迁移性,实现人类中心视觉任务的高效预训练。

Comments Accepted by TMM (IEEE Transactions on Multimedia), 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09470 2026-01-15 physics.ed-ph cs.AI 57%

Personalized Multimodal Feedback Using Multiple External Representations: Strategy Profiles and Learning in High School Physics

基于多种外部表征的个性化反馈:策略配置与高中物理学习中的学习

Natalia Revenga-Lozano, Karina E. Avila, Steffen Steinert, Matthias Schweinberger, Clara E. Gómez-Pérez, Jochen Kuhn, Stefan Küchemann

机构 * Chair of Physics Education, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich)(物理教育系主任,物理学院,慕尼黑路易斯-马克西姆利安大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 本文研究了多种外部表征与个性化反馈在高中物理学习中的整合效果,发现详细多表征反馈对学习成绩有积极影响,且学习者根据表征能力选择不同反馈策略。

Comments Keywords: Adaptive Feedback, Multimodal Learning, Multiple External Representations, Physics Education, Science Education, Representational Competences, Intelligent Tutoring Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11169 2026-01-13 cs.CL cs.AI 57%

Correcting misinformation on social media with a large language model

利用大型语言模型纠正社交媒体上的虚假信息

Xinyi Zhou, Ashish Sharma, Amy X. Zhang, Tim Althoff

机构 * Paul G. Allen School of Computer Science and Engineering, University of Washington(保罗·G·阿伦计算机科学与工程学院,华盛顿大学) Computer Science Department, Boise State University(计算机科学系,博伊西州立大学) Microsoft Corporation(微软公司)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

AI总结 MUSE通过结合视觉语言模型和网络检索,有效纠正社交媒体上的虚假信息,优于GPT-4和社交媒体用户回复。

Comments 52 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11399 2026-01-09 cs.CL cs.CV 57%

Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction

最短片段,最大显著性:通过关键时刻提取实现长视频摘要

Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种通过关键时刻提取实现长视频多模态摘要的方法,利用轻量级模型和大型语言模型实现高效摘要生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01366 2026-01-06 cs.AI 57%

KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models

KGCE:基于多模态语言模型的跨平台教育代理基准评估知识增强双图评估器

Zixian Liu, Sihao Liu, Yuqi Zhao

机构 * Faculty of the School of Computer Science, Central China Normal University(中央财经大学计算机学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 KGCE提出一种基于多模态语言模型的跨平台教育代理基准评估框架,通过知识库增强和双图评估方法提升对特定学校软件任务的执行效率和评估精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04678 2025-12-30 cs.CV 57%

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

奖励强制:基于奖励分布匹配蒸馏的高效流式视频生成

Yunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Jiapeng Zhu, Hengyuan Cao, Zhipeng Zhang, Xing Zhu, Yujun Shen, Min Zhang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) SIAS-ZJU SJTU(上海交通大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 奖励强制通过EMA-Sink和Re-DMD提升流式视频生成效率与质量,实现23.1 FPS的高性能视频生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21637 2025-12-29 cs.CV 57%

Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints

无需训练的解耦文本引导图像编辑:通过稀疏潜在约束

Mutiara Shabrina, Nova Kurnia Putri, Jefri Satria Ferdiansyah, Sabita Khansa Dewi, Novanto Yudistira

机构 * Department of Informatics Engineering Universitas Brawijaya Malang, Indonesia(信息工程系 乌姆拉大学 马拉邦,印度尼西亚)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种无需训练的解耦文本引导图像编辑方法,通过引入稀疏潜在约束减少属性纠缠问题,提升编辑的可控性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18699 2025-12-29 cs.CV 57%

Affective Image Editing: Shaping Emotional Factors via Text Descriptions

情感图像编辑:通过文本描述塑造情感因素

Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang, Zijian Jia, Si Li, Boxin Shi

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) State Key Laboratory for Multimedia Information Processing and National Engineering Research Center of Visual Technology, School of Computer Science, Peking University, China(多媒体信息处理国家重点实验室和视觉技术国家工程研究中心,北京大学计算机学院)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

AI总结 AIEdiT通过文本描述实现情感图像编辑,利用情感映射器和MLLM生成符合用户情感需求的图像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01907 2025-12-25 cs.CV cs.CL 57%

RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events

RSCC:一种大规模遥感变化描述数据集用于灾害事件

Zhenyuan Chen, Chenxi Wang, Ningyu Zhang, Feng Zhang

机构 * School of Earth Sciences, Zhejiang University(浙江大学地球科学学院) School of Software Technology, Zhejiang University(浙江大学软件学院) Zhejiang Provincial Key Laboratory of Geographic Information Science(浙江省地理信息科学重点实验室) Key Laboratory of Spatio-temporal Information and Intelligent Services (LSIIS), Ministry of Natural Resources of the People’s Republic of China(国家自然资源部空间时空信息与智能服务重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 RSCC数据集通过提供大规模遥感图像对和详细变化描述,提升视觉-语言模型在灾害事件双时间理解中的性能和应用能力。

Comments Accepted by NeurIPS 2025 Dataset and Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19271 2025-12-23 cs.CV 57%

3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory

3SGen: 一种统一的主体、风格和结构驱动的图像生成方法,具有自适应任务特定记忆

Xinyang Song, Libin Wang, Weining Wang, Zhiwei Li, Jianxin Sun, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun

机构 * School of Artificial Intelligence, UCAS(人工智能学院,UCAS) CASIA AntGroup(蚂蚁集团)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

AI总结 3SGen通过统一的框架实现主体、风格和结构驱动的图像生成,采用自适应任务特定记忆模块提升生成质量和跨任务迁移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04540 2025-12-17 cs.CV 57%

VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management

VideoMem: 通过自适应内存管理增强超长视频理解

Hongbo Jin, Qingyuan Wang, Wenhao Zhang, Yang Liu, Sijie Cheng

机构 * School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

AI总结 VideoMem通过自适应内存管理框架,有效提升超长视频理解任务的性能,采用PRPO算法和两个核心模块实现高效训练和长期记忆保留。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00087 2025-12-12 cs.CV 57%

Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

探索多模态课堂数据中教学活动和话语的自动化识别

Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan Kyle Foster, Peter Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文通过多模态分析方法,实现了课堂活动中教学活动和话语的自动化识别,展示了微调模型在视频和 transcripts 上的高准确率,为可扩展的教师反馈系统提供了基础。

Comments This article has been accepted for publication in the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09670 2025-12-11 cs.CV cs.SY eess.SY 57%

An Automated Tip-and-Cue Framework for Optimized Satellite Tasking and Visual Intelligence

一种自动化提示与提示框架用于优化卫星任务分配和视觉智能

Gil Weissman, Amir Ivry, Israel Cohen

机构 * Andrew and Erna Viterbi Faculty of Electrical and Computer Engineering, Technion-Israel Institute of Technology(安德鲁和伊尔纳·维特比电气与计算机工程学院,技术离子理工学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文提出了一种自动化提示与提示框架,用于优化卫星任务分配和视觉智能,通过生成提示和任务并利用人工智能模型处理影像,提升地球观测效率。

Comments Under review at IEEE Transactions on Geoscience and Remote Sensing (TGRS). 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05127 2025-12-04 eess.IV cs.CV q-bio.QM 57%

PixCell: A generative foundation model for digital histopathology images

PixCell:数字病理图像的生成基础模型

Srikar Yellapragada, Alexandros Graikos, Zilinghan Li, Kostas Triaridis, Varun Belagali, Tarak Nath Nandi, Karen Bai, Beatrice S. Knudsen, Tahsin Kurc, Rajarsi R. Gupta, Prateek Prasanna, Ravi K Madduri, Joel Saltz, Dimitris Samaras

机构 * Stony Brook University(石溪大学) Argonne National Laboratory(阿贡国家实验室) The University of Chicago(芝加哥大学) University of Utah(犹他大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 PixCell是首个针对数字病理图像的生成基础模型,通过扩散模型在大规模数据集上训练,实现隐私保护的数据生成和虚拟染色任务,提升病理学研究效率。

Comments Project page - https://histodiffusion.github.io/docs/projects/pixcell

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02713 2025-12-03 cs.AI 57%

Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs

利用本体对齐的知识图谱训练数据归因于图像生成

Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose

机构 * National Centre for Scientific Research ``Demokritos''(国家科学研究中心「德莫克里特」) University of Glasgow(格拉斯哥大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出利用本体对齐的知识图谱方法,通过多模态大语言模型提取图像中的结构化三元组,以追踪生成模型中训练数据的影响,从而提升透明度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02517 2025-12-03 cs.CV 57%

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

SkyMoE:一种用于增强遥感解释的视觉-语言基础模型

Jiaqi Liu, Ronghao Fu, Lang Sun, Haoran Liu, Xiao Yang, Weipeng Zhang, Xu Na, Zhuoran Duan, Bo Yang

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 SkyMoE通过Mixture-of-Experts架构提升遥感多模态多任务处理能力,实现对不同粒度任务的高效适应与优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06263 2025-12-02 cs.CV 57%

OmniSVG: A Unified Scalable Vector Graphics Generation Model

OmniSVG: 一种统一的可扩展矢量图形生成模型

Yiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng, Fukun Yin, Jiaxu Zhang, Liao Wang, Gang Yu, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) StepFun Project(StepFun项目)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 OmniSVG通过统一框架和预训练视觉-语言模型,实现高效多模态SVG生成,提升复杂结构的表达能力,并引入大规模数据集推动SVG合成发展。

Comments 20 pages; Project Page: https://omnisvg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00557 2025-12-02 cs.CV 57%

NeuroVolve: Evolving Visual Stimuli toward Programmable Neural Objectives

NeuroVolve:通过可编程神经目标演化视觉刺激

Haomiao Chen, Keith W Jamison, Mert R. Sabuncu, Amy Kuceyeski

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技) Weill Cornell Medicine(韦尔医学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 NeuroVolve通过可编程神经目标生成视觉刺激,揭示大脑区域间的协同与对抗性调节关系,实现脑引导的图像编辑与首选刺激生成的统一。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21389 2025-11-27 cs.IR cs.AI 57%

FITRep: Attention-Guided Item Representation via MLLMs

FITRep: 通过大语言模型实现的注意力引导的项目表示

Guoxiao Zhang, Ao Li, Tan Qu, Qianlong Xie, Xingxing Wang

机构 * Meituan(美团)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

AI总结 FITRep通过引入注意力引导的白盒表示框架,利用多模态大语言模型实现细粒度项目去重,提升了广告点击率和每千次展示成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19200 2025-11-26 cs.CV 57%

Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?

现代视觉模型能否理解物体与相似物之间的差异?

Itay Cohen, Ethan Fetaya, Amir Rosenfeld

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文研究了现代视觉模型能否区分真实物体与相似物,通过构建RoLA数据集并改进CLIP模型的嵌入空间方向,提升跨模态检索和描述生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17136 2025-11-24 cs.SD cs.AI 57%

Device-Guided Music Transfer

设备引导的音乐转移

Manh Pham Hung, Changshuo Hu, Ting Dang, Dong Ma

机构 * Singapore Management University(新加坡管理大学) University of Melbourne(墨尔本大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

AI总结 DeMT通过提取扬声器频率响应曲线生成设备嵌入,实现跨设备的音乐风格迁移与鲁棒适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17534 2025-11-21 cs.CV cs.CL cs.MM 57%

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

协同强化学习用于统一多模态理解和生成

Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang, Chao Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出CoRL框架,通过协同强化学习提升多模态大语言模型在生成与理解任务上的性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23584 2025-11-21 cs.CV 57%

VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement

VividFace: 高质量和高效的一步扩散用于视频面部增强

Shulian Zhang, Yong Guo, Long Peng, Ziyang Wang, Ye Chen, Wenbo Li, Xiao Zhang, Yulun Zhang, Jian Chen

机构 * South China University of Technology(华南理工大学) Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University of Science and Technology(南京理工大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

AI总结 VividFace通过单步扩散框架和联合训练策略,高效提升视频面部增强的高质量与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13326 2025-11-20 stat.AP cs.AI 57%

TacEleven: generative tactic discovery for football open play

Siyao Zhao, Hao Ma, Zhiqiang Pu, Jingjing Huang, Yi Pan, Shijie Wang, Zhi Ming

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(交叉科学学院,中国科学院大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Shanghai AI Laboratory(上海人工智能实验室) Association de la Jeunesse Auxerroise(亚眠青年协会)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13993 2025-11-19 cs.CV 57%

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments NeurIPS 2025, Project webpage: https://vision.cs.utexas.edu/projects/CrossTrainer/

详情

展开后加载摘要…

URL PDF HTML 收藏