Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
通过文本引导增强弱监督多模态视频异常检测
Shengyang Sun, Jiashen Hua, Junyi Feng, Xiaojin Gong
机构
*
School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院)
;
Alibaba Cloud(阿里云)
;
College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)
机构
*
School of Computer Science & Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)
;
School of Computing, National University of Singapore(计算学院,新加坡国立大学)
机构
*
The University of New South Wales, Sydney, New South Wales, Australia(新南威尔士大学)
;
The University of Sydney, Sydney, New South Wales, Australia(悉尼大学)
;
School of Software and Microelectronics, Peking University, Zibo, Shandong, China(北京大学软件与微电子学院)
;
Shandong University of Technology, Beijing, China(山东理工大学)
Comments19 pages excluding reference. Conditionally accepted to ACM CHI 2026
Journal refIn Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26), April 13-17, 2026, Barcelona, Spain. ACM, New York, NY, USA, 26 pages
机构
*
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT)
;
Shanghai Jiao Tong University(上海交通大学)
;
Ocean University of China(中国海洋大学)
Multi-Modal Zero-Shot Prediction of Color Trajectories in Food Drying
多模态零样本预测食品干燥中的颜色轨迹
Shichen Li, Ahmadreza Eslaminia, Chenhui Shao
机构
*
Department of Mechanical Science and Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA(机械科学与工程系,伊利诺伊大学厄巴纳-香槟分校)
;
Department of Mechanical Engineering, University of Michigan, Ann Arbor, MI, USA(机械工程系,密歇根大学安娜堡分校)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
SAT:多模态语言模型的动态空间能力训练
Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A. Plummer, Ranjay Krishna, Kuo-Hao Zeng, Kate Saenko
机构
*
Boston University(波士顿大学)
;
University of Washington(华盛顿大学)
;
Allen Institute for AI(人工智能研究院)
;
Microsoft Research(微软研究院)
;
New York University(纽约大学)
机构
*
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Zhongguancun Academy, Beijing(中关村学院)
;
China Meteorological Administration(中国气象局)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi
机构
*
Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学)
;
The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院))
;
Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)