V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
专题命中 其他VLM :vision-language model(abstract)
Comments Accepted by IJCNN 2025
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 其他VLM :vision-language model(abstract)
Comments Accepted by IJCNN 2025
专题命中 其他VLM :multimodal large language model(abstract)
Comments Accepted by Interspeech 2025
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Pengcheng Laboratory(鹏城实验室) ; Shanghai Jiao Tong University(上海交通大学) ; University of Cambridge(剑桥大学)
专题命中 其他VLM :multimodal large language model(abstract)
Comments Accepted in ACL 2025 (Main)
机构 * New York University(纽约大学) ; University of South Florida(佛罗里达州立大学) ; Google LLC(谷歌公司) ; University of California, Santa Barbara(加州大学圣巴巴拉分校) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 其他VLM :multimodal large language model(abstract)
机构 * Fudan University(复旦大学) ; University of Southern California(南加州大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 其他VLM :multimodal large language model(abstract)
专题命中 其他VLM :vision-language model(abstract)
专题命中 其他VLM :multimodal large language model(abstract)
Comments NAACL 2025 Main
专题命中 其他VLM :multimodal large language model(abstract)
Comments Accepted as a CHI 2025 Full Paper
专题命中 其他VLM :multimodal large language model(abstract)
专题命中 其他VLM :multimodal large language model(abstract)
Comments The paper is accepted by WWW2025 Resource Track
专题命中 其他VLM :vision-language model(abstract)
专题命中 其他VLM :vision-language model(abstract)
专题命中 其他VLM :multimodal large language model(abstract)
Comments Published as a conference paper at CHI 2025. Project page: https://yewon-kim.com/amuse
专题命中 其他VLM :vision-language model(abstract)
Comments Accepted by NAACL2025. Our code is available at https://github.com/XiaomengWang-AI/Typographic-Attacks-in-a-Multi-Image-Setting
专题命中 其他VLM :vision-language model(abstract)
Comments Accepted to NAACL 2025
专题命中 其他VLM :multimodal large language model(abstract)
Comments Accepted to HRI 2025
专题命中 其他VLM :vision-language model(abstract)
Comments This work was presented at the WARN, Weighing the Benefits of Autonomous Robot Personalization, workshop at the 33rd IEEE RO-MAN 2024 conference
专题命中 其他VLM :multimodal large language model(abstract)
Comments preprint
专题命中 其他VLM :multimodal large language model(abstract)
Comments 11 pages, 2 figures, 2 tables
专题命中 其他VLM :vision-language model(abstract)
专题命中 其他VLM :MLLM(abstract)
Comments 22 pages, 10 figures
专题命中 其他VLM :vision language model(abstract)
Comments Accepted at IROS 2024
专题命中 其他VLM :multimodal large language model(abstract)
Comments Accepted to EMNLP 2024 Findings
专题命中 其他VLM :visual language model(abstract)
Comments CIKM 2024 Workshop on Industrial Recommendation Systems
专题命中 其他VLM :vision-language model(abstract)
Comments 5 pages
专题命中 其他VLM :vision-language model(abstract)
Comments 6 pages
专题命中 其他VLM :vision-language model(abstract)
专题命中 其他VLM :MLLM(abstract)
专题命中 其他VLM :vision-language model(abstract)
专题命中 其他VLM :MLLM(abstract)