DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
DGSeg:用于推理分割的语义-空间引导预测的动态门控
Ruizhe Zeng, Siyu Cao, Lu Zhang, Zhiyong Liu
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Nanjing Artificial Intelligence Research of IA(中国科学院自动化所南京人工智能研究院)
机构
*
Nanjing University(南京大学)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Tsinghua University(清华大学)
;
Zhejiang University(浙江大学)
;
University of Washington(华盛顿大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Australian National University(澳大利亚国立大学)
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computing and Information Technology, Great Bay University(大湾区大学计算与信息科技学院)
;
Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息处理重点实验室)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
;
Tencent Youtu Lab(腾讯优图实验室)
;
School of Psychology, Shanghai Jiao Tong University(上海交通大学心理学院)
;
Faculty of Applied Sciences, Macao Polytechnic University(澳门理工学院应用科学学院)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
Department of Computer Science and Engineering, Dhaka International University(达卡国际大学计算机科学与工程系)
;
Department of Computer Science and Engineering, Bangladesh University of Engineering and Technology(孟加拉国工程技术大学计算机科学与工程系)
Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
ByteDance(字节跳动)
;
USTC(中国科学技术大学)
;
Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室)
;
Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部互动技术与体验系统重点实验室)
;
Zhongguancun Academy(中关村科学城)
Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning
按需查看:多模态推理中视觉证据获取的认知调度框架
Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun, Yidong Chen, Jiayi Ji, Qian Chen, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception(多媒体可信感知实验室)
;
Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China(教育部高效计算实验室,厦门大学,361005,中国)
;
Sino-Russian ResearchCenter for Digital Economy(中俄数字经济研究中心)
;
Institute of Artificial Intelligence, Xiamen University, China(人工智能研究院,厦门大学,中国)
;
School of Informatics, Xiamen University, China(信息学院,厦门大学,中国)
;
School of Information Engineering, Xiamen Ocean Vocational College, Xiamen 361102, China(信息工程学院,厦门海洋职业技术学院,厦门361102,中国)
CommentsPreprint. Accepted at NeurIPS 2025 Workshops on SPACE in Vision, Language, and Embodied AI (SpaVLE) as Oral, Embodied World Models for Decision Making (EWM), Aligning Reinforcement Learning Experimentalists and Theorists (ARLET), and Scaling Environments for Agents (SEA)
Explainable Novel Category Discovery in Semantic Concept Space
语义概念空间中的可解释新类别发现
Ifrat Ikhtear Uddin, Yang Zhou, KC Santosh, Longwei Wang
机构
*
Department of Computer Science, University of South Dakota(南达科他大学计算机科学系)
;
Department of Computer Science and Software Engineering, Auburn University(奥本大学计算机科学与软件工程系)
CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models
CAC-VLA:用于视觉-语言-动作模型的上下文门控动作条件调节
Yifu Xiong, Wenhao Yu, Jiaxuan Lin, Bojun Zou, Jiahao Li, Lu Zhang, Yanyong Zhang, Jianmin Ji
机构
*
University of Science and Technology of China (USTC)(中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
CommentsAccepted to ACL 2026 System Demonstrations. 11 pages, 5 figures, 8 tables
Journal refProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 829-839, 2026
机构
*
Department of Computing, Imperial College London(伦敦帝国理工学院计算系)
;
Department of Earth Science & Engineering, Imperial College London(伦敦帝国理工学院地球科学与工程系)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Department of Computer Science, Heriot-Watt University(赫瑞瓦特大学计算机科学系)
;
School of Computer Science, Northumbria University(诺森比亚大学计算机科学学院)
;
College of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件学院)
;
School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(深圳大学广东省智能信息处理重点实验室)
;
School of Engineering and Design, Hunan Normal University(湖南师范大学工程与设计学院)
;
Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)
Audio-Visual Speech Enhancement: Architectural Design and Deployment Strategies
音频-视觉语音增强:架构设计与部署策略
Anis Hamadouche, Haifeng Luo, Mathini Sellathurai, Amir Hussain, Tharm Ratnarajah
机构
*
School of Engineering & Physical Sciences, Heriot-Watt University(赫瑞斯泰学院,赫瑞斯泰大学)
;
College of Engineering Department of Electrical and Computer Engineering, San Diego State University(工程学院电子与计算机工程系,圣地亚哥州立大学)
;
SDAIA-KFUPM Joint Research Centre for Artificial Intelligence, King Fahd University of Petroleum and Minerals(SDAIA-KFUPM人工智能联合研究中心,国王法赫德石油与矿物大学)