机构
*
National University of Defense Technology(国防科技大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学深圳人工智能学院)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
基于场景图的细粒度SegCaptioning Transformer用于可控视频分割与标注
Xu Zhang, Jin Yuan, BinHong Yang, Xuan Liu, Qianjun Zhang, Yuyi Wang, Zhiyong Li, Hanwang Zhang
机构
*
College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院)
;
School of Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学)
;
School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算机与人工智能学院,西南交通大学)
;
CRRC Zhuzhou Institute Company Ltd.(中车株洲院有限公司)
;
Nanyang Technological University(南洋理工大学)
Hyper-STTN: Hypergraph Augmented Spatial-Temporal Transformer Network for Trajectory Prediction
Hyper-STTN:基于超图的时空变换网络用于轨迹预测
Weizheng Wang, Baijian Yang, Sungeun Hong, Wenhai Sun, Byung-Cheol Min
机构
*
School of Applied and Creative Computing, Purdue University(普渡大学应用与创意计算学院)
;
Department of Computer Science and Department of Intelligent Systems Engineering, Indiana University Bloomington(印第安纳大学布卢明顿分校计算机科学系和智能系统工程系)
;
Department of Applied Artificial Intelligence, Sungkyunkwan University(成均馆大学应用人工智能系)
机构
*
East China normal University(东中国正常大学)
;
Sun Yat-sen University(孙中山大学)
;
OPPO AI Center(OPPO AI中心)
;
Shenzhen Institutes of Advanced Technology(深圳先进技术研究所)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
MagicWorld: 向交互视频世界探索的长时稳定迈进
Guangyuan Li, Bo Li, Jinwei Chen, Xiaobin Hu, Lei Zhao, Peng-Tao Jiang
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
vivo BlueImage Lab, vivo Mobile Communication Co., Ltd.(vivo蓝影实验室,vivo移动通信有限公司)
;
National University of Singapore(新加坡国立大学)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
VideoReasonBench: 大规模语言模型能否进行以视觉为中心的复杂视频推理?
Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu, Lin Sui, Xinhao Li, Yan Zhong, Y. Charles, Xinyu Zhou, Xu Sun
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
;
Moonshot AI
;
Nanjing University(南京大学)
;
School of Mathematical Sciences, Peking University(数学科学学院,北京大学)
机构
*
School of Computer Science and Technology, Wuhan University of Science and Technology(武汉科技大学计算机科学与技术学院)
;
Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System(湖北省智能信息处理与实时工业系统重点实验室)
;
Harbin Institute of Technology Zhengzhou Research Institute(哈尔滨工业大学郑州研究所)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
;
Huawei Technologies Ltd(华为技术有限公司)