Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization
通过群体相对策略优化在结构因果模型中 grounding 多跳推理
Yunhan Bu, Quan Zhang, Huaping Zhang, Guotong Geng, Chunxiao Gao, Askar Hamdulla, Juan Wang, Qiuchi Li, Baohua Zhang, Shuai Lei, Yunbo Cao, Zhunchen Luo
机构
*
School of Computer Science and Technology, Xinjiang University, Urumqi, China(新疆大学计算机科学与技术学院,乌鲁木齐,中国)
;
Beijing Institute of Technology, Beijing, China(北京理工大学,北京,中国)
;
Military Science Information Research Center, Academy of Military Science, Beijing, China(军事科学院军事科学信息研究中心,北京,中国)
机构
*
School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院,中国)
;
Alibaba Group(阿里巴巴集团)
;
School of Computer Science and Engineering, School of Intelligence Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China(东南大学计算机科学与工程学院、智能科学与工程学院以及新一代人工智能技术及其交叉应用关键实验室,中国)
;
Wangxuan Institute of Computer Technology, National Key Laboratory for Multimedia Information Processing, Peking University, China(北京大学王轩计算机技术研究所、多媒体信息处理国家重点实验室,中国)
;
University of Copenhagen, Denmark(丹麦哥本哈根大学)
机构
*
Multimodal Language Department(多模态语言部门)
;
Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所)
;
Department of Linguistics(语言学系)
;
Boğaziçi University(博多伊奇大学)
;
Donders Institute for Brain Cognition and Behaviour(多纳尔斯脑认知与行为研究所)
;
Radboud University(拉德堡德大学)
;
Department of Linguistics and Communication(语言学与沟通系)
;
University of Birmingham(伯明翰大学)
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
从街景到视觉网络:利用视觉语言模型映射城市地标可见性
Zicheng Fan, Kunihiko Fujiwara, Pengyuan Liu, Fan Zhang, Filip Biljecki
机构
*
organization= Department of Architecture, National University of Singapore , country= Singapore
;
organization= Research \& Development Institute, Takenaka Corporation , country= Japan
;
organization= Urban Analytics Subject Group, Urban Studies \& Social Policy Division, University of Glasgow , country= United Kingdom
;
organization= Institute of Remote Sensing
;
GIS, Peking University , country= China
;
organization= Department of Real Estate, National University of Singapore , country= Singapore
专题命中
视觉定位与Grounding
:vision-language model(title);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.CV
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
OmniVTG:一种大规模数据集和开放世界视频时间定位的训练范式
Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng, Yang Liu
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Central Media Technology Institute, Huawei Technologies Ltd.(华为技术有限公司中央媒体技术研究所)
;
PKU-WUHAN Institute for Artificial Intelligence, Peking University(北京大学武汉人工智能研究所)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV