ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
ViMoNet: 一种多模态视觉-语言框架,用于从运动和视频中理解人类行为
Rajan Das Gupta, Lei Wei, Md Yeasin Rahat, Nafiz Fahad, Abir Ahmed, Liew Tze Hui
机构
*
Department of Computer Science, American International University–Bangladesh (AIUB)(美国国际大学-孟加拉国计算机科学系)
;
Faculty of Psychology, Shinawatra University(信武大学心理学系)
;
Faculty of Information Science and Technology, Multimedia University(多媒体大学信息科学与技术系)
;
Department of Information Technology, Washington University of Science & Technology(华盛顿科学与技术大学信息科技系)
CommentsWe found problems in the code while rechecking our implementation. These issues led to noticeable numerical discrepancies, making some of the reported results and conclusions potentially unreliable. Therefore, we request to withdraw this submission
Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
探索多模态课堂数据中教学活动和话语的自动化识别
Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan Kyle Foster, Peter Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Sichuan University(四川大学)
;
National University of Singapore(新加坡国立大学)
;
Shenzhen University(深圳大学)
MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
MMT-ARD: 多模态多教师对抗蒸馏用于鲁棒视觉-语言模型
Yuqi Li, Junhao Dong, Chuanguang Yang, Shiping Wen, Piotr Koniusz, Tingwen Huang, Yingli Tian, Yew-Soon Ong
机构
*
The City University of New York, CUNY(纽约城市大学)
;
Nanyang Technological University(南洋理工大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Technology Sydney(悉尼技术大学)
;
Data61, CSIRO(CSIRO数据61研究所)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu
机构
*
MBZUAI
;
Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室)
;
King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
;
University of Copenhagen(哥本哈根大学)
;
Tsinghua University(清华大学)
机构
*
Institute of Trustworthy Embodied AI(可信具身人工智能研究院)
;
Fudan University(复旦大学)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
Columbia University(哥伦比亚大学)
Descriptive Image-Text Matching with Graded Contextual Similarity
Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn
专题命中
图文多模态
:image-text(title,abstract);分类 cs.CV
CommentsThis version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission