Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
从博弈视角重新思考弱监督视频时间定位
Xiang Fang, Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen, Jianfeng Dong, Keke Tang, Pan Zhou, Yu Cheng, Daizong Liu
机构
*
Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室)
;
Hubei Engineering Research Center on Big Data Security(大数据安全工程研究中心)
;
School of Cyber Science and Engineering(网络安全科学与工程学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
University of Central Florida(佛罗里达中央大学)
;
Zhejiang Gongshang University(浙江工商大学)
;
Guangzhou University(广州大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos
你可以比看见更早定位:一种用于压缩视频中时序句子定位的高效流程
Xiang Fang, Daizong Liu, Pan Zhou, Guoshun Nan
机构
*
The Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,网络安全科学与工程学院,华中科技大学)
;
Peking University(北京大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
VideoTemp-o3:在智能视频思考中协调时间定位与视频理解
Wenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu, Qile Su, Tianke Zhang, Haonan Fan, Changyi Liu, Kaiyu Jiang, Jiankang Chen, Kaiyu Tang, Bin Wen, Fan Yang, Tingting Gao, Han Li, Yinwei Wei, Xuemeng Song
机构
*
Shandong University(山东大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beihang University(北航)
;
Southern University of Science and Technology(南方科技大学)
机构
*
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China(合肥工业大学计算机科学与信息工程学院)
;
Wuhan University, Wuhan, China(武汉大学)
;
Lab for Intelligence and visiON (LION)(智能视觉实验室)
机构
*
Chongqing University(重庆大学)
;
Tianjin University(天津大学)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院MAIS)
;
Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局高性能计算研究所)
;
Chongqing National Data AI Research Institute, AI Research Lab(重庆国家数据AI研究院,AI研究实验室)
机构
*
University of Macau(澳门大学)
;
UESTC(电子科技大学)
;
Purdue University(普渡大学)
;
McGill University(麦吉尔大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Washington(华盛顿大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
机构
*
Macquarie University(麦考瑞大学)
;
Nanjing University of Science and Technology(南京理工大学)
;
The University of Sydney(悉尼大学)
;
National University of Singapore(新加坡国立大学)
Generating Findings for Jaw Cysts in Dental Panoramic Radiographs Using a GPT-Based VLM: A Preliminary Study on Building a Two-Stage Self-Correction Loop with Structured Output (SLSO) Framework
机构
*
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China(复旦大学计算机科学与人工智能学院,上海,中国)
;
Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理重点实验室)
Intelligent Healthcare Imaging Platform: A VLM-Based Framework for Automated Medical Image Analysis and Clinical Report Generation
智能医疗影像平台:基于视觉语言模型的自动化医学图像分析与临床报告生成框架
Samer Al-Hamadani
机构
*
Automated Manufacturing Department/Al-Khwarizmi College of Engineering/ University of Baghdad/Gilgamesh University(自动化制造部门/阿尔·卡瓦尔齐米工程学院/巴格达大学/吉尔伽美什大学)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
HaloProbe: 基于贝叶斯的方法检测和缓解视觉-语言模型中的物体幻觉
Reihaneh Zohrabi, Hosein Hasani, Akshita Gupta, Mahdieh Soleymani Baghshah, Anna Rohrbach, Marcus Rohrbach
机构
*
Multimodal AI Lab, Technical University of Darmstadt(达姆施塔特工业大学多模态人工智能实验室)
;
Department of Computer Engineering, Sharif University of Technology(谢里夫理工大学计算机工程系)