Understanding Multimodal Complementarity for Single-Frame Action Anticipation
理解单帧动作预测中的多模态互补性
Manuel Benavent-Lledo, Konstantinos Bacharidis, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez
机构
*
Department of Computer Technology, University of Alicante(阿尔瓦雷斯大学计算机技术系)
;
Institute of Computer Science, FORTH(福蒂研究所)
;
Computer Science Department, University of Crete(克里特大学计算机科学系)
;
Department of Management, Science and Technology, Hellenic Mediterranean University(希腊地中海大学管理、科学与技术系)
机构
*
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院)
;
College of Computer Science and Technology, Harbin Engineering University(哈尔滨工程大学计算机科学与技术学院)
;
Singapore University of Technology and Design (SUTD), Singapore(新加坡科技设计大学)
;
Continental Automotive Singapore Pte. Ltd.
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
China Mobile (Jiangxi) Virtual Reality Technology Co., Ltd.(中国移动(江西)虚拟现实技术有限公司)
;
School of Computer and Information Technology, Beijing Jiaotong University(北京交通大学计算机与信息学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
LVAgent: 通过多轮动态协作的MLLM代理实现长视频理解
Boyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu, Peng Li, Yali Wang
机构
*
Shenzhen Key Lab of Computer Vision and Pattern Recognition, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳计算机视觉与模式识别重点实验室,深圳先进技术研究院,中国科学院)
;
Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China(人工智能产业研究院(AIR),清华大学,北京,中国)
;
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China(计算机科学与技术系,人工智能研究院,清华大学,北京,中国)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shanghai Jiao Tong University(上海交通大学)
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
CoVAR: 通过多模态扩散生成视频与动作用于机器人操作
Liudi Yang, Yang Bai, George Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen, Ziyuan Liu, Abhinav Valada
机构
*
University of Freiburg(弗赖堡大学)
;
Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Technical University of Munich(慕尼黑技术大学)
;
Huawei Heisenberg Research Center (Munich)(华为海森堡研究中心)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
感知、对话,然后适应:面向开放世界视频识别的多模态基础模型知识迁移
Boyu Chen, Siran Chen, Kunchang Li, Qinglin Xu, Yu Qiao, Yali Wang
机构
*
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
the School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Shanghai AI Laboratory(上海人工智能实验室)
Multi-Modal Graph Convolutional Network with Sinusoidal Encoding for Robust Human Action Segmentation
多模态图卷积网络与正弦编码用于鲁棒的人体动作分割
Hao Xing, Kai Zhe Boey, Yuankai Wu, Darius Burschka, Gordon Cheng
机构
*
Institute for Cognitive Systems(认知系统研究所)
;
Chair of Media Technology(媒体技术教授职位)
;
Machine Vision and Perception Group(机器视觉与感知小组)
;
School of Computation, Information and Technology(计算、信息与技术学院)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
瞬间动作预见:多模态线索能替代视频到何种程度?
Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez
机构
*
Universidad de Alicante(阿利坎特大学)
;
Foundation for Research and Technology-Hellas(希腊基础研究与技术基金会)
;
University of Crete(克里特大学)
;
Hellenic Mediterranean University(希腊地中海大学)