Commentshis is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: https://doi.org/10.1109/SUMMA68668.2025.11302248
Journal ref2025 7th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA), Lipetsk, Russian Federation, 2025, pp. 799-804
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
从自身学习点击位置:面向GUI定位的在线自蒸馏
Yan Zhang, Daiqing Wu, Huawen Shen, Can Ma, Yu Zhou
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Dexmal
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Zhongguancun Academy(中关村学院)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Tsinghua University(清华大学)
;
Beijing Innovation Center of Humanoid Robotics Co., Ltd.(北京人形机器人创新中心有限公司)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学多媒体信息处理国家重点实验室,计算机学院)
Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation
沙盒计划,开放世界导航:学习物理基础的抽象经验以实现具身导航
Zhixuan Shen, Jiawei Du, Ziyu Guo, Han Luo, Lilan Peng, Joey Tianyi Zhou, Haonan Luo, Tianrui Li
机构
*
School of Computing and Artificial Intelligence, Southwest Jiaotong University, China(计算机与人工智能学院,西南交通大学,中国)
;
Centre for Frontier AI Research A*STAR, Singapore(前沿人工智能研究A*STAR中心,新加坡)
;
School of Computer Science, University of Leeds, UK(计算机科学学院,利兹大学,英国)
Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
Med-StepBench:一种用于评估医学视觉-语言模型幻觉的分层推理框架
Minh Khoi Nguyen, Dai Lam Le, Amir Reza Jafari, Tuan Dung Nguyen, Mai Hong Son, Mai Huy Thong, Quang Huy Nguyen, Thanh Trung Nguyen, Reza Farahbakhsh, Noel Crespi, Phi Le Nguyen
机构
*
AI4LIFE, Hanoi University of Science and Technology, Vietnam(AI4LIFE,越南科学与技术大学)
;
SAMOVAR, Télécom SudParis, Institut Polytechnique de Paris, France(SAMOVAR,法国电信南巴黎学院,巴黎理工学院)
;
Military Central Hospital, Vietnam(越南108军中心医院)
Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models
在操作设计域内运行:基于视觉-语言模型的零样本感知
Berkehan Ünal, Hauke Dierend, Dren Fazlija, Christopher Plachetka
机构
*
Volkswagen Aktiengesellschaft(大众汽车股份有限公司)
;
L3S Research Center(莱比锡大学汉诺威研究中心)
;
Faculty of Information Technology(信息科技学院)
;
MOIA GmbH(MOIA公司)
;
Motor AI GmbH(Motor AI公司)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
C-CoT:基于视觉-语言模型的反事实链式推理用于安全自动驾驶
Kefei Tian, Yuansheng Lian, Kai Yang, Xiangdong Chen, Shen Li
机构
*
College of Transportation, Tongji University(同济大学交通运输学院)
;
Department of Civil Engineering, Tsinghua University(清华大学土木工程系)
;
School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)
;
Department of Civil and Environmental Engineering, National University of Singapore(新加坡国立大学土木与环境工程系)
机构
*
South China University of Technology(华南理工大学)
;
Institute for Super Robotics (Huangpu)(机器人研究所(黄埔))
;
Shanghai Jiao Tong University(上海交通大学)
;
Changsha University of Science and Technology(长沙理工大学)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
PromptGuard: 基于软提示的文本到图像模型不安全内容审查
Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao, Xiaojun Jia, Yihao Huang, Wei Dong, Yang Liu, Bo Li
机构
*
Department of Computer Science, University of Maryland(计算机科学系,马里兰大学)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机科学与数据科学学院)
;
Kahlert School of Computing, The University of Utah(犹他大学Kahlert计算学院)
;
Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校Siebel计算与数据科学学院)
Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning
锚定特征间隙:跨模态谱稳定化以实现样本高效表征学习
Nikhil J. Dhinagar, Vidhi Chhatbar, Chirag Jagad, Pavithra Senthilkumar, Sophia I. Thomopoulos, Mahir H. Khan, Sook-Lei Liew, the ENIGMA-Stroke Recovery Working Group, Paul M. Thompson
机构
*
Imaging Genetics Center, Mark & Mary Stevens Neuroimaging & Informatics Institute, Keck School of Medicine, University of Southern California(影像基因中心,马克与玛丽史蒂文斯神经影像与信息学研究所,凯克医学院,南加州大学)
;
Neuroscience Graduate Program, Mark & Mary Stevens Neuroimaging & Informatics Institute, Chan Division of Occupational Science & Occupational Therapy, Biomedical Engineering, University of Southern California(神经科学研究生项目,马克与玛丽史蒂文斯神经影像与信息学研究所,查恩职业科学与职业治疗 division,生物医学工程,南加州大学)