arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1566 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1566 篇

2412.10908 2025-07-23 cs.CV 57%

Do large language vision models understand 3D shapes?

Sagi Eppel

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13684 2025-07-23 cs.CV 57%

FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models

Minghan Li, Chenxi Xie, Yichen Wu, Lei Zhang, Mengyu Wang

机构 * Harvard AI and Robotics Lab, Harvard University(哈佛人工智能与机器人实验室,哈佛大学) Broad Institute(博德研究所) Hong Kong Polytechnic University(香港理工大学) School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) City University of Hong Kong(香港城市大学) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(自然与人工智能研究学院,哈佛大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 24 pages, 14 figures, 16 tables

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15569 2025-07-22 cs.CV 57%

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团) Peking University(北京大学) Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15025 2025-07-22 cs.SE cs.AI 57%

Survey of GenAI for Automotive Software Development: From Requirements to Executable Code

Nenad Petrovic, Vahid Zolfaghari, Andre Schamschurko, Sven Kirchner, Fengjunjie Pan, Chengdng Wu, Nils Purschke, Aleksei Velsh, Krzysztof Lebioda, Yinglei Song, Yi Zhang, Lukasz Mazur, Alois Knoll

机构 * Chair of Robotics, Artificial Intelligence and Real-Time Systems(机器人、人工智能与实时系统教授席)

专题命中 其他VLM :vision language model(abstract);分类 cs.AI

Comments Conference paper accepted for GACLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00302 2025-07-22 cs.LG cond-mat.mtrl-sci 57%

Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework

Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban

机构 * Computer Engineering, Texas A\&M University, College Station, TX 77843, USA College of Science Engineering, Hamad Bin Khalifa University, Doha, Qatar Dept. of Electrical \& Computer Engineering, Texas A\&M University at Qatar, Doha, Qatar Orthotics, Ankara University, Ankara, Turkey

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

Comments Presented at ICML 2025 Workshop on DataWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23765 2025-07-18 cs.CV 57%

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Yun Li, Yiming Zhang, Tao Lin, Xiangrui Liu, Wenxiao Cai, Zheng Liu, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) China University of Geosciences(中国地质大学) Nanyang Technological University(南洋理工大学) BAAI(百度人工智能研究院) Stanford University(斯坦福大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09961 2025-07-15 cs.LG 57%

Text-Driven Causal Representation Learning for Source-Free Domain Generalization

Lihua Zhou, Mao Ye, Nianxin Li, Shuaifeng Li, Jinlin Wu, Xiatian Zhu, Lei Deng, Hongbin Liu, Jiebo Luo, Zhen Lei

机构 * Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China(人工智能与机器人研究中心,香港科学与创新研究所,中国科学院,香港,中国) School of Computer Science and Engineering, University of Electronic Science and Technology of China(计算机科学与工程学院,电子科技大学) Surrey Institute for People-Centred Artificial Intelligence, CVSSP, University of Surrey(以人为中心的人工智能研究所,CVSSP, Surrey大学) School of Electronics and Information Engineering, Shenzhen University(电子与信息工程学院,深圳大学) University of Rochester(罗切斯特大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10563 2025-07-15 cs.CV 57%

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Jiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang, Yubo Wang, Yuansheng Ni, Wang Zhu, Ziyan Jiang, Bohan Lyu, Dongfu Jiang, Xuan He, Yuan Liu, Hexiang Hu, Xiang Yue, Wenhu Chen

机构 * Core Contributors(核心贡献者) Tiger-AI-Lab(虎鲸人工智能实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR 2025 camera-ready version. Project page: https://tiger-ai-lab.github.io/MEGA-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07105 2025-07-10 cs.CV eess.IV 57%

4KAgent: Agentic Any Image to 4K Super-Resolution

Yushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang, Renjie Li, Jian Wang, Yide Zhang, Gengchen Mai, Lihong V. Wang, James Zou, Xiaoyu Wang, Ming-Hsuan Yang, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) Stanford University(斯坦福大学) Snap Inc.(Snap公司) CU Boulder(科罗拉多大学博尔德分校) UT Austin(得克萨斯大学奥斯汀分校) California Institute of Technology(加州理工学院) Topaz Labs(Topaz实验室) UC Merced(加州大学默塞德分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Project page: https://4kagent.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06973 2025-07-10 cs.CV 57%

Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM

Qiyuan Dai, Sibei Yang

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) Sun Yat-sen University(孙中山大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06523 2025-07-10 cs.CV cs.CL cs.GR 57%

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

Liqiang Jing, Viet Lai, Seunghyun Yoon, Trung Bui, Xinya Du

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05344 2025-07-08 cs.CV 57%

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Jiahui Wang, Zuyan Liu, Yongming Rao, Jiwen Lu

机构 * Tsinghua University(清华大学) Tencent Hunyuan X(腾讯混元实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03997 2025-07-04 cs.CV 57%

CAD-Editor: A Locate-then-Infill Framework with Automated Training Data Synthesis for Text-Based CAD Editing

Yu Yuan, Shizhao Sun, Qi Liu, Jiang Bian

机构 * University of Science(科学大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06285 2025-07-04 cs.CV 57%

DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation

Yueming Lyu, Tianwei Lin, Fu Li, Dongliang He, Jing Dong, Tieniu Tan

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CRIPAC, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所CRIPAC) VIS, Baidu Inc.(百度公司VIS)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Code is available at https://github.com/Yueming6568/DeltaEdit

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00907 2025-07-02 cs.CR cs.AI 57%

The Age of Sensorial Zero Trust: Why We Can No Longer Trust Our Senses

Fabio Correa Xavier

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00891 2025-07-02 cs.CL cs.AI 57%

MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes

Yuheng Wang, Xianhe Tang, Pufeng Huang

机构 * Wuhan University(武汉大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00586 2025-07-02 cs.CV 57%

Context-Aware Academic Emotion Dataset and Benchmark

Luming Zhao, Jingwen Xuan, Jiamin Lou, Yonghui Yu, Wenwu Yang

机构 * Zhejiang Gongshang University(浙江工商大学) Zhejiang Yuexiu University(浙江越秀大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05769 2025-07-02 cs.CV 57%

Exploring Text-Guided Single Image Editing for Remote Sensing Images

Fangzhou Han, Lingyu Si, Zhizhuo Jiang, Hongwei Dong, Lamei Zhang, Yu Liu, Hao Chen, Bo Du

机构 * Department of Information Engineering, Harbin Institute of Technology(信息工程系,哈尔滨工业大学) National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences(空间信息集成国家重点实验室,中国科学院软件研究所) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Hubei Luojia Laboratory, National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(湖北珞珈实验室,国家多媒体软件工程技术研究中心,武汉大学计算机学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 17 pages, 18 figures, Accepted by IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21476 2025-06-27 cs.CV 57%

Global and Local Entailment Learning for Natural World Imagery

Srikumar Sastry, Aayush Dhakal, Eric Xing, Subash Khanal, Nathan Jacobs

机构 * Washington University in St. Louis(圣路易斯华盛顿大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19769 2025-06-25 cs.RO cs.AI 57%

TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning

Yuhui Chen, Haoran Li, Zhennan Jiang, Haowei Wen, Dongbin Zhao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17705 2025-06-24 cs.CV 57%

DreamJourney: Perpetual View Generation with Video Diffusion Models

Bo Pan, Yang Chen, Yingwei Pan, Ting Yao, Wei Chen, Tao Mei

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Laboratory of Art and Archaeology Image(艺术与考古图像实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17608 2025-06-24 cs.CV 57%

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

Nikitha SR, Aradhya Neeraj Mathur, Tarun Ram Menta, Rishabh Jain, Mausoom Sarkar

机构 * Media and Data Science Research Lab, Adobe(Adobe媒体与数据科学研究实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted in CVPR 2025 Workshop on What's Next in Multimodal Foundational Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17500 2025-06-24 cs.CV 57%

Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation

Julio Silva-Rodríguez, Fereshteh Shakeri, Houda Bahig, Jose Dolz, Ismail Ben Ayed

机构 * ÉTS Montréal(ÉTS蒙特利尔) Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM)(蒙特利尔大学中心医院研究中心(CRCHUM))

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments MICCAI 2025. Code: https://github.com/jusiro/SS-Text

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11748 2025-06-24 cs.CV 57%

ILIAS: Instance-Level Image retrieval At Scale

Giorgos Kordopatis-Zilos, Vladan Stojnić, Anna Manko, Pavel Šuma, Nikolaos-Antonios Ypsilantis, Nikos Efthymiadis, Zakaria Laskar, Jiří Matas, Ondřej Chum, Giorgos Tolias

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15903 2025-06-23 cs.LG 57%

VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics

Josef Kuchař, Marek Kadlčík, Michal Spiegel, Michal Štefánik

机构 * Kempelen Institute of Intelligent Technologies(智能技术研究所) Language Technology, University of Helsinki(语言技术,赫尔辛基大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05146 2025-06-23 cs.CV cs.CL 57%

CIVET: Systematic Evaluation of Understanding in VLMs

Massimo Rizzoli, Simone Alghisi, Olha Khomyn, Gabriel Roccabruna, Seyed Mahed Mousavi, Giuseppe Riccardi

机构 * Signals and Interactive Systems Lab, University of Trento, Italy(信号与交互系统实验室,特伦托大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15479 2025-06-19 cs.LG 57%

Creating User-steerable Projections with Interactive Semantic Mapping

Artur André Oliveira, Mateus Espadoto, Roberto Hirata, Roberto M. Cesar, Alex C. Telea

机构 * Institute of Mathematics and Statistics, University of São Paulo(数学与统计学研究所,圣保罗大学) Department of Information and Computing Sciences, Utrecht University(信息与计算科学系,乌得勒支大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13956 2025-06-19 cs.CV 57%

Improving LLM Video Understanding with 16 Frames Per Second

Yixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li, Zejun Ma, Chao Zhang

机构 * Tsinghua University(清华大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12826 2025-06-17 cs.CV 57%

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling

Zhihan Zhang, Xiang Pan, Hongchen Wei, Zhenzhong Chen

机构 * School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16458 2025-06-17 cs.RO cs.CV 57%

BiFold: Bimanual Cloth Folding with Language Guidance

Oriol Barbany, Adrià Colomé, Carme Torras

机构 * Institut de Robòtica i Informàtica Industrial, CSIC-UPC(工业机器人与计算机研究所,CSIC-UPC)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at ICRA 2025. Project page at https://barbany.github.io/bifold/

详情

展开后加载摘要…

URL PDF HTML 收藏