CommentsThis paper has been accepted to the Late-Breaking Results (LBR) track of the 28th International Conference on Multimodal Interaction (ICMI 2026)
ID-VTG: Image-Disambiguated Video Temporal Grounding
ID-VTG:基于图像消歧的视频时间定位
Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents
迈向通用具身智能:整合大语言模型、知识库与推理能力以构建下一代AI智能体
Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao
机构
*
College of Mechanical Engineering, Chongqing University of Technology(重庆理工大学机械工程学院)
;
James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院)
;
School of Energy and Power, Jiangsu University of Science and Technology(江苏科技大学能源与动力学院)
;
Magnesium Research Center, Kumamoto University(熊本大学镁研究中心)
;
Department of Information and Communication Engineering, Nagoya University(名古屋大学信息与通信工程系)
;
State Key Laboratory of Fluid Power and Mechatronic Systems, Zhejiang University(浙江大学流体动力与机电系统国家重点实验室)