VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
VideoTemp-o3:在智能视频思考中协调时间定位与视频理解
Wenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu, Qile Su, Tianke Zhang, Haonan Fan, Changyi Liu, Kaiyu Jiang, Jiankang Chen, Kaiyu Tang, Bin Wen, Fan Yang, Tingting Gao, Han Li, Yinwei Wei, Xuemeng Song
机构
*
Shandong University(山东大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beihang University(北航)
;
Southern University of Science and Technology(南方科技大学)
机构
*
Hong Kong Baptist University(香港 Baptist 大学)
;
S-Lab, Nanyang Technological University(S 实验室,南洋理工大学)
;
GVC Lab, Great Bay University(GVC 实验室,Great Bay 大学)
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Tencent(腾讯)
;
Tsinghua University(清华大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Film Academy(北京电影学院)
;
Stanford University(斯坦福大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Singapore Institute of Technology(新加坡理工学院)
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
CoMoGen: 基于掩码引导的视频生成的可控运动动力学与交互
Adil Meric, Lin Geng Foo, Mert Kiray, Benjamin Busam, Rishabh Dabral, Christian Theobalt
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Max Planck Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息研究所,萨尔兰信息校园)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Obsphera
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
EgoInteract: 用于交互理解和预测的合成自我中心视频生成
Rosario Leonardi, Francesco Ragusa, Daniele Materia, Alessandro Passanisi, James Fort, Jakob Engel, Giovanni Maria Farinella
机构
*
Department of Mathematics and Computer Science, University of Catania(卡塔尼亚大学数学与计算机科学系)
;
Next Vision s.r.l.(Next Vision公司)
;
Reality Labs Research, Meta(Meta现实实验室)
机构
*
State Key Laboratory of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)
;
School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学)
;
JIUTIAN Research(JIUTIAN研究机构)
;
Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)
;
Zhejiang Sci-Tech University(浙江科技学院)
;
The University of British Columbia(不列颠哥伦比亚大学)
CommentsWe substantially updated the previous version "Diffusion Models for Tabular Data: Challenges, Current Progress, and Future Directions" by including flow matching models for tabular data
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
解耦时空适配器用于细粒度羽毛球动作定位
Tianyu Wang, Junjie Wu, Jingquan Gao, Shishuo Li
机构
*
School of Economics and Management, Beihang University(北京航空航天大学经济管理学院)
;
Key Laboratory of Data Intelligence and Management, Beihang University, Ministry of Industry and Information Technology(信息产业部北京航空航天大学数据智能与管理重点实验室)