arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

2026-07-15 至 2026-07-15 共收录 8
2607.13017 2026-07-15 cs.RO cs.CV 新提交

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

FlowWAM:光流作为世界动作模型的统一动作表示

Yixiang Chen, Peiyan Li, Yuan Xu, Qisen Ma, Jiabing Yang, Kai Wang, Jianhua Yang, Dong An, He Guan, Gaoteng Liu, Jianlou Si, Jun Huang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FiveAges(无) MBZUAI(无) Alibaba Group(阿里巴巴集团)

AI总结 研究针对世界动作模型控制中动作表示难题,提出FlowWAM双流扩散框架,以光流为统一动作表示。该框架可实现WAMs两种模式,能利用无动作标签视频预训练,实验表明在操纵和世界建模任务中表现优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12640 2026-07-15 cs.AI cs.CL 新提交

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

小语言和视觉语言模型网络智能体中GRPO的学习率门控失败:一个可控的无效结果及其机制

Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang

机构 * Monash University(莫纳什大学) University of Chinese Academy of Sciences(中国科学院大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Pusan National University(釜山国立大学)

AI总结 研究4B到8B规模小语言和视觉语言模型网络智能体中GRPO的效果,通过控制变量实验发现其在已掌握任务上无明显提升,解释了学习率导致失败的机制,表明该耦合与模型规模相关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12557 2026-07-15 cs.CV 新提交

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

用于长视频理解中事件感知视觉分配的高斯混合模型

Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Bing Li, Weiming Hu

机构 * Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(中国科学院自动化所多模态信息超智能安全北京市重点实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化所多模态人工智能系统国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hello Group(未知(保留英文)) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

AI总结 针对长视频理解中视觉分配问题,提出GMM-EVA方法,利用高斯混合模型建模事件级结构,采用差异化分配策略,在多个长视频基准实验中显著优于均匀采样,以约一半视觉令牌预算达可比性能,凸显高效性。

Comments accepted at PRCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12429 2026-07-15 cs.CV 新提交

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

不仅仅是你在哪里:从跨视图定位中学习语义、结构和几何

Mao Chen, Xiangkai Zhang, Zhiyong Liu, Chuankai Liu, Xu Yang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Beijing Aerospace Control Center(北京航天飞行控制中心)

AI总结 研究如何在极端视角变化下通过跨视图定位让模型学习语义、结构和几何。提出CROSS框架,克服现有方法局限,经实验验证该框架在跨视图定位中性能最优,能有效学习多方面内容。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11914 2026-07-15 cs.NE cs.AI 新提交

Burst Spiking Neural Networks

突发脉冲神经网络

Jiahong Zhang, Sijun Shen, Man Yao, Han Xu, Mingqiang Huang, Yonghong Tian, Bo Xu, Guoqi Li

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Peng Cheng Laboratory(鹏城实验室) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

AI总结 研究SNN的准确性 - 鲁棒性问题,提出基于突发增强脉冲神经元和动态权重约束机制的BuSNN,通过理论分析和实验表明其在准确性、鲁棒性及低功耗方面优势显著,推进了SNN在相关应用中的可行性。

Comments 18 pages, 21 figures, 1 supplementary material PDF, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11638 2026-07-15 cs.RO 版本更新

DA-Nav: Direction-Aware City-Scale Vision-Language Navigation

DA-Nav:方向感知的城市规模视觉语言导航

Ye Yuan, Kehan Chen, Xinqiang Yu, Wentao Xu, Heng Wang, Libo Huang, Chuanguang Yang, Yan Huang, Jiawei He, Zhulin An

机构 * School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所人工智能安全国家重点实验室) National Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) XYZ Embodied AI(XYZ具身人工智能)

AI总结 研究城市规模户外导航难题,提出DA-Nav框架,利用商业导航工具方向指示,经思维链推理实现轨迹恢复,引入ReDA数据集。实验表明其在未见环境成功率高,优于现有方法,还能适应多种机器人实现稳定户外导航。

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17670 2026-07-15 cs.CV cs.LG q-bio.QM q-bio.TO 版本更新

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割

Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juchler, Johannes C. Paetzold, Rami Al-Maskari, Luciano Höher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstetter, Florian Kofler, Fernando Navarro, Martin J. Menten, Ivan Ezhov, Daniel Rueckert, Iris N. Vos, Ynte M. Ruigrok, Birgitta K. Velthuis, Hugo J. Kuijf, Pengcheng Shi, Wei Liu, Ting Ma, Maximilian R. Rokuss, Yannick Kirchhoff, Fabian Isensee, Klaus Maier-Hein, Chengcheng Zhu, Huilin Zhao, Philippe Bijlenga, Julien Hämmerli, Catherine Wurster, Laura Westphal, Jeroen Bisschop, Elisa Colombo, Hakim Baazaoui, Hannah-Lea Handelsmann, Andrew Makmur, James Hallinan, Amrish Soundararajan, Benedikt Wiestler, Jan S. Kirschke, Evamaria O. Riedel, Roland Wiest, Emmanuel Montagnon, Laurent Letourneau-Guillon, Kwanseok Oh, Dahye Lee, Orhun Utku Aydin, Adam Hilbert, Jana Rieger, Dimitrios Rallios, Satoru Tanioka, Alexander Koch, Dietmar Frey, Abdul Qayyum, Moona Mazher, Steven Niederer, Nico Disch, Julius C. Holzschuh, Dominic LaBella, Francesco Galati, Daniele Falcetta, Maria A. Zuluaga, Chaolong Lin, Haoran Zhao, Zehan Zhang, Minghui Zhang, Xin You, Hanxiao Zhang, Guang-Zhong Yang, Yun Gu, Sinyoung Ra, Jongyun Hwang, Hyunjin Park, Junqiang Chen, Marek Wodzinski, Henning Müller, Nesrin Mansouri, Florent Autrusseau, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Joaquim Salvi, Adrià Casamitjana, Xavier Lladó, Uma Maria Lal-Trehan Estrada, Valeriia Abramova, Luca Giancardo, Arnau Oliver, Paula Casademunt, Adrian Galdran, Matteo Delucchi, Oscar Camara, Jialu Liu, Haibin Huang, Yue Cui, Zehang Lin, Yusheng Liu, Shunzhi Zhu, Tatsat R. Patel, Adnan H. Siddiqui, Vincent M. Tutino, Maysam Orouskhani, Huayu Wang, Mahmud Mossa-Basha, Yuki Sato, Sven Hirsch, Susanne Wegener, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA Institute for Tissue Engineering School of Computation, Information Technology, Technical University of Munich, Germany Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany Munich Center for Machine Learning, Munich, Germany Department of Computing, Imperial College London, London, UK Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands Department of Neurology Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China Peng Cheng Laboratory, Shenzhen, China Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany Faculty of Mathematics Computer Science, Heidelberg University, Germany Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany Data Science School for Health, Karlsruhe/Heidelberg, Germany Learning Group, Department of Radiation Oncology, Heidelberg University Hospital Department of Radiology, University of Washington, Seattle, WA, USA Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland Department of Neurology, University Hospital of Zurich, Zurich, Switzerland Department of Physiology, University of Toronto, Canada Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland Department of Diagnostic Imaging, National University Hospital, Singapore University of Chicago, USA Department of Diagnostic Interventional Neuroradiology, University Hospital Berne University of Berne, Berne, Switzerland Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada DEEPNOID Inc., Seoul, South Korea Department of Artificial Intelligence, Korea University, Seoul, South Korea Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany Lung Institute, Faculty of Medicine, Imperial College London, London, UK Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA Institute of Medical Technology, Peking University Health Science Center, Beijing, China Hangzhou Genlight MedTech Co., Ltd., China Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China Department of Automation, Shanghai Jiao Tong University, Shanghai, China Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea Department of Electrical Computer Engineering, Sungkyunkwan University, Seoul, South Korea Shanghai MediWorks Precision Instruments Co., Ltd., China Institute of Informatics, HES-SO Valais-Wallis, Switzerland Department of Measurement Electronics, AGH University of Krakow, Poland Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France Research Institute of Computer Vision Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA Physense, BCN-Medtech, Department of Communication Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain Department of Mathematical Modeling Machine Learning, University of Zurich, Zurich, Switzerland Laboratory of Brain Atlas Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China School of Computer Information Engineering, Xiamen University of Technology, Xiamen, China Vascular Research Center, University at Buffalo, NY, USA Department of Pathology Anatomical Sciences, University at Buffalo, NY, USA Department of Neurosurgery, University at Buffalo, NY, USA LPIXEL Inc., Tokyo, Japan

AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。

Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435

Journal ref NEJM AI 2026;3(8)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16157 2026-07-15 cs.CV 版本更新

Fast and Accurate Image Restoration and Generation with Rank Enhanced Linear Attention

基于秩增强线性注意力的快速准确图像恢复与生成

Yuang Ai

机构 * University of Chinese Academy of Sciences(中国科学院大学)

AI总结 研究旨在解决Transformer自注意力二次复杂度问题,提出秩增强线性注意力(RELA)及高效视觉Transformer(LAformer),通过集成深度卷积丰富特征表示,经实验验证LAformer在图像恢复和生成任务中性能优且计算高效。

Comments Code: https://github.com/shallowdream204/LAformer

详情

展开后加载摘要…

URL PDF HTML 收藏