ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
ForensicZip: 更多令牌更好但并非必要在取证视觉-语言模型中
Yingxin Lai, Zitong Yu, Jun Wang, Linlin Shen, Yong Xu, Xiaochun Cao
机构
*
Great Bay University(大湾大学)
;
Shenzhen University(深圳大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络安全科学与技术学院)
专题命中
视觉定位与Grounding
:vision-language model(title);multimodal large language model(abstract);分类 cs.CV
Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time
用眼睛倾听:跨时空的自体视觉共指基准测试
Weijie Zhou, Xuantang Xiong, Zhenlin Hu, Xiaomeng Zhu, Chaoyang Zhao, Honghui Dong, Zhengyou Zhang, Ming Tang, Jinqiao Wang
机构
*
Beijing Jiaotong University(北京交通大学)
;
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences (CASIA)(基础模型研究中心、自动化研究所、中国科学院(CASIA))
;
Tencent Robotics X(腾讯机器人X)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST)(计算机科学与工程系、香港科学与技术大学(HKUST))
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
T2SGrid: 视频时间定位的时空网格化
Chaohong Guo, Yihan He, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long
机构
*
South China University of Technology(南方科技大学)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室)
;
Bytedance Inc(字节跳动公司)
机构
*
Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(电子科技大学复杂系统建模与仿真实验室,计算机科学与技术学院)
;
Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江省空间信息感知与传输重点实验室,电子科技大学)
;
Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系)
Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
Invert4TVG: 一种具有逆向任务的时序视频定位框架,以保持动作理解能力
Zhaoyu Chen, Hongnan Lin, Yongwei Nie, Fei Ma, Xuemiao Xu, Fei Yu, Chengjiang Long
机构
*
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室)
;
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
;
ByteDance Inc.(字节跳动公司)