arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1566 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1566 篇

2602.00762 2026-02-03 cs.CL cs.HC 50%

WordCraft: Scaffolding the Keyword Method for L2 Vocabulary Learning with Multimodal LLMs

WordCraft: 通过多模态大语言模型 scaffolding 关键词方法用于L2词汇学习

Yuheng Shao, Junjie Xiong, Chaoran Wu, Xiyuan Wang, Ziyu Zhou, Yang Ouyang, Qinyi Tao, Quan Li

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) School of Creativity and Art, ShanghaiTech University(创意与艺术学院,上海科技大学) Shanghai Fengxian Dai Wen Middle School(上海奉贤戴文中学)

专题命中 其他VLM :multimodal large language model(abstract)

AI总结 WordCraft通过多模态大语言模型辅助L2词汇学习,提升关键词方法的使用效果和学习者参与度。

Comments Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI' 26), April 13--17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17536 2026-01-29 cs.CL 50%

Multimodal Conversation Structure Understanding

多模态对话结构理解

Kent K. Chang, Mackenzie Hanh Cramer, Anna Ho, Ti Ti Nguyen, Yilin Yuan, David Bamman

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 其他VLM :multimodal large language model(abstract)

AI总结 本文提出多模态对话结构理解任务及TV-MMPC数据集,揭示对话角色匿名化对模型性能的影响,并分析对话中性别角色的参与差异。

Comments accepted to EACL 2026 main conference; 22 pages, 9 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20622 2026-01-29 cs.HC 50%

SketchDynamics: Exploring Free-Form Sketches for Dynamic Intent Expression in Animation Generation

SketchDynamics: 探索自由形式草图用于动画生成中的动态意图表达

Boyu Li, Lin-Ping Yuan, Zeyu Wang, Hongbo Fu

专题命中 其他VLM :vision-language model(abstract)

AI总结 SketchDynamics通过自由形式草图与AI交互,探索动态意图表达在动画生成中的应用,展示草图如何有效传达运动意图并指导视频生成。

Comments conditionally accepted by CHI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19203 2026-01-28 cs.HC 50%

Before Smelling the Video: A Two-Stage Pipeline for Interpretable Video-to-Scent Plans

在嗅觉之前:一种可解释的视频到香氛计划的两阶段流程

Kaicheng Wang, Kevin Zhongyang Shao, Ruiqi Chen, Sep Makhsous, Denise Wilson

专题命中 其他VLM :vision-language model(abstract)

AI总结 本文提出了一种两阶段流程,通过视觉语言模型和大型语言模型分离视频语义提取与气味推断,验证了语义规划在提升嗅觉媒体体验中的有效性。

Comments In submission of poster as ongoing project

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14874 2026-01-22 cs.RO 50%

HumanoidVLM: Vision-Language-Guided Impedance Control for Contact-Rich Humanoid Manipulation

HumanoidVLM: 用于接触密集型人形机器人工操作的视觉-语言引导阻抗控制

Yara Mahmoud, Yasheerah Yaqoot, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

机构 * Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所)

专题命中 其他VLM :vision-language model(abstract)

AI总结 HumanoidVLM通过视觉-语言模型和检索增强生成模块,实现人形机器人在接触密集场景中的自适应阻抗控制与抓取配置选择。

Comments This paper has been accepted for publication at LBR of HRI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12195 2026-01-13 cs.CL 50%

Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion

浏览与聚焦:通过先验LLM上下文融合理解多模态内容

Ziyue Wang, Chi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu

机构 * Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China(计算机科学与技术系,人工智能研究院,清华大学,北京,中国) Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China(人工智能产业研究院(AIR),清华大学,北京,中国) Institute of Intelligent Computing, Alibaba Group(智能计算研究院,阿里巴巴集团) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国) Jiangsu Collaborative Innovation Center for Language Competence, Jiangsu, China(江苏省语言能力协同创新中心,江苏,中国)

专题命中 其他VLM :multimodal large language model(abstract)

AI总结 本文提出浏览与聚焦两阶段范式,通过融合先验LLM上下文提升多模态内容理解,显著提升多图像场景的性能。

Comments 17 pages, 5 figures

Journal ref ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21863 2025-12-29 cs.IR cs.MM 50%

Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion

冻结的大型视频语言模型用于微视频推荐:特征提取与融合的系统研究

Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du

专题命中 其他VLM :vision-language model(abstract)

AI总结 本文通过系统研究冻结LVLMs的特征提取与融合策略,提出DFF框架,证明中间隐藏状态优于标题表示,ID嵌入融合优于替换,并在微视频推荐中取得最佳性能。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17228 2025-12-22 cs.HC 50%

LUMIA: A Handheld Vision-to-Music System for Real-Time, Embodied Composition

LUMIA: 一种用于实时具身创作的手持视觉到音乐系统

Chung-Ta Huang, Connie Cheng, Vealy Lai

专题命中 其他VLM :vision-language model(abstract)

AI总结 LUMIA通过视觉到音乐的实时具身创作系统,将环境互动与生成式AI结合,实现基于感知的即兴音乐创作。

Comments 6 pages, 15 pages with appendix, NeurIPS 2025 Creative AI track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20056 2025-12-19 cs.CL 50%

Online-PVLM: Advancing Personalized VLMs with Online Concept Learning

在线-PVLM:通过在线概念学习推进个性化VLMs

Huiyu Bai, Runze Wang, Zhuoyun Du, Yiyang Zhao, Fengji Zhang, Haoyu Chen, Xiaoyong Zhu, Bo Zheng, Xuejiao Zhao

机构 * Nanyang Technological University(南洋理工大学) Alibaba Group(阿里巴巴集团) Zhejiang University(浙江大学) City University of Hong Kong(香港城市大学) University of Oulu(奥卢大学)

专题命中 其他VLM :visual language model(abstract)

AI总结 Online-PVLM通过在线概念学习框架,实现个性化VLMs的高效实时适应,并提出OP-Eval基准评估其在现实场景中的性能。

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14359 2025-11-19 cs.HC cs.SE 50%

Towards LLM-Based Usability Analysis for Recommender User Interfaces

Sebastian Lubos, Alexander Felfernig, Damian Garber, Viet-Man Le, Thi Ngoc Trang Tran

专题命中 其他VLM :multimodal large language model(abstract)

Comments The paper was presented at IntRS'25: Joint Workshop on Interfaces and Human Decision Making for Recommender Systems, September 22, 2025, Prague, Czech Republic and is published in the workshop proceedings: https://ceur-ws.org/Vol-4027/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17669 2025-10-20 cs.CL 50%

Towards Human Cognition: Visual Context Guides Syntactic Priming in Fusion-Encoded Models

Bushi Xiao, Michael Bennie, Jayetri Bardhan, Daisy Zhe Wang

机构 * University of Florida(佛罗里达大学)

专题命中 其他VLM :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10637 2025-10-14 cs.RO 50%

High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting

Haoyu Zhao, Cheng Zeng, Linghao Zhuang, Yaxi Zhao, Shengke Xue, Hao Wang, Xingyue Zhao, Zhongyu Li, Kehan Li, Siteng Huang, Mingxiu Chen, Xin Li, Deli Zhao, Hua Zou

机构 * Wuhan University(武汉大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎扑实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学)

专题命中 其他VLM :MLLM(abstract)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06664 2025-10-09 cs.CL 50%

ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory

Yunzhong Xiao, Yangmin Li, Hewei Wang, Yunlong Tang, Zora Zhiruo Wang

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Rochester(罗切斯特大学)

专题命中 其他VLM :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09747 2025-10-07 cs.NE 50%

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu

专题命中 其他VLM :multimodal large language model(abstract)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09047 2025-10-06 cs.CL 50%

Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs

Yaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术学院 – 以色列理工学院) UC Berkeley(加州大学伯克利分校)

专题命中 其他VLM :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04952 2025-09-30 cs.CL cs.SE 50%

ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation

Chenchen Zhang, Yuhang Li, Can Xu, Jiaheng Liu, Ao Liu, Changzhi Zhou, Ken Deng, Dengpeng Wu, Guanhua Huang, Kejiao Li, Qi Yi, Ruibin Xiong, Shihui Hu, Yue Zhang, Yuhao Jiang, Zenan Xu, Yuanxing Zhang, Wiggin Zhou, Chayse Zhou, Fengzong Lian

机构 * Tencent Hunyuan Team(腾讯文脉团队)

专题命中 其他VLM :MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04415 2025-09-23 cs.CL 50%

MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind

Emilio Villa-Cueva, S M Masrur Ahmed, Rendi Chevi, Jan Christian Blaise Cruz, Kareem Elzeky, Fermin Cristobal, Alham Fikri Aji, Skyler Wang, Rada Mihalcea, Thamar Solorio

机构 * MBZUAI University of Houston(德克萨斯大学休斯顿分校) McGill University(麦吉尔大学) University of Michigan(密歇根大学)

专题命中 其他VLM :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24456 2025-09-23 cs.CL 50%

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofia Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio

机构 * MBZUAI(人工智能研究所)

专题命中 其他VLM :vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15211 2025-09-19 cs.CL 50%

What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques

Petros Stylianos Giouroukis, Dimitris Dimitriadis, Dimitrios Papadopoulos, Zhenwen Shao, Grigorios Tsoumakas

机构 * Aristotle University of Thessaloniki(亚里士多德大学)

专题命中 其他VLM :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14171 2025-09-19 cs.CL 50%

AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity

Yifan Liu, Wenkuan Zhao, Shanshan Zhong, Jinghui Qin, Mingfu Liang, Zhongzhan Huang, Wushao Wen

机构 * Sun Yat-sen University(中山大学) Guangdong University of Technology(广东工业大学) Northwestern University(西北大学)

专题命中 其他VLM :multimodal large language model(abstract)

Comments Accepted by EMNLP 2025 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16405 2025-09-18 cs.MM 50%

EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment

Lancheng Gao, Ziheng Jia, Yunhao Zeng, Wei Sun, Yiming Zhang, Wei Zhou, Guangtao Zhai, Xiongkuo Min

专题命中 其他VLM :MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11065 2025-09-16 cs.SE cs.PL 50%

ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch

Yuan Si, Daming Li, Hanyuan Shi, Jialu Zhang

专题命中 其他VLM :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15241 2025-09-16 math.OC math.DS 50%

Differential Stochastic Variational Inequalities with Parametric Optimization

Xiaojun Chen, Jian Guo, Guan Wang

专题命中 其他VLM :multimodal large language model(abstract)

Comments 35 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04077 2025-09-12 cs.CL cs.SD eess.AS 50%

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

机构 * National Taiwan Normal University(台湾国立台湾师范大学)

专题命中 其他VLM :multimodal large language model(abstract)

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21801 2025-09-01 cs.IR 50%

DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors

Zhuoxing Wei, Qingchen Xie, Qi Liu

专题命中 其他VLM :MLLM(abstract)

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20660 2025-08-29 eess.AS cs.SD 50%

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

Ruifan Deng, Yitian Gong, Qinghui Gao, Luozhijie Jin, Qinyuan Cheng, Zhaoye Fei, Shimin Li, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 其他VLM :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02173 2025-08-05 cs.HC 50%

EchoLadder: Progressive AI-Assisted Design of Immersive VR Scenes

Zhuangze Hou, Jingze Tian, Nianlong Li, Farong Ren, Can Liu

专题命中 其他VLM :vision-language model(abstract)

Comments To appear at UIST 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01092 2025-08-05 cs.HC 50%

DescribePro: Collaborative Audio Description with Human-AI Interaction

Maryam Cheema, Sina Elahimanesh, Samuel Martin, Pooyan Fazli, Hasti Seifi

专题命中 其他VLM :multimodal large language model(abstract)

Comments ASSETS 25 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01857 2025-07-03 cs.RO 50%

TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

Yuhao Lin, Yi-Lin Wei, Haoran Liao, Mu Lin, Chengyi Xing, Hao Li, Dandan Zhang, Mark Cutkosky, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Stanford University, USA(斯坦福大学) Imperial College London, UK(伦敦帝国理工学院)

专题命中 其他VLM :MLLM(abstract)

Comments Project Page: https://isee-laboratory.github.io/TypeTele

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18711 2025-06-24 cs.HC 50%

LLM-enhanced Interactions in Human-Robot Collaborative Drawing with Older Adults

Marianne Bossema, Somaya Ben Allouch, Aske Plaat, Rob Saunders

专题命中 其他VLM :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏