arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

2026-03-10 至 2026-03-10 共收录 13
2603.08227 2026-03-10 cs.CV

SRNeRV: A Scale-wise Recursive Framework for Neural Video Representation

SRNeRV:一种用于神经视频表示的分层递归框架

Jia Wang, Jun Zhu, Xinfeng Zhang

机构 * School of Computer Science and Technology, University of Chinese Academy of Sciences(计算机科学与技术学院,中国科学院大学)

AI总结 SRNeRV通过分层递归框架实现神经视频表示的参数高效共享,提升率失真性能。

Comments Accepted by IEEE ISCAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22555 2026-03-10 cs.LG cs.AI

Autoregressive Visual Decoding from EEG Signals

从EEG信号进行自回归视觉解码

Sicheng Dai, Hongwang Xiao, Shan Yu, Qiwei Ye

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artifcial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启智技术国家重点实验室) Beijing Academy of Artificial Intelligence(北京人工智能研究院) National Key Laboratory for Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)

AI总结 AVDE通过自回归生成框架和对比学习提升EEG信号的视觉解码效率,实现高效且可解释的脑机接口应用。

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07691 2026-03-10 cs.RO cs.CV

RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation

RoboPCA:基于姿态的从人类示范中学习空间可及性的方法用于机器人操作

Zhanqi Xiao, Ruiping Wang, Xilin Chen

机构 * Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室、计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 RoboPCA通过联合预测接触区域和姿态,提升机器人操作中从人类示范中学习空间可及性的性能。

Comments Accepted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10416 2026-03-10 cs.CV cs.AI

Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction

超越端点:面向向量化的非公路网络提取的路径中心推理

Wenfei Guan, Jilin Mei, Tong Shen, Xumin Wu, Shuo Wang, Chen Min, Yu Hu

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)

AI总结 本文提出MaGRoad,一种基于路径中心的非公路道路网络提取方法,通过构建大规模数据集和改进模型结构,在野外环境中实现了更鲁棒的道路提取。

Comments This revision improves clarity and consistency throughout the paper. We refine terminology to more precisely describe the vertex extraction optimization, add motivational context to the edge feature encoding section, and clarify the overall inference pipeline. We also add an Acknowledgments section

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17234 2026-03-10 cs.MM cs.AI cs.CV

Taming Modality Entanglement in Continual Audio-Visual Segmentation

驯服持续音频-视觉分割中的模态纠缠

Yuyang Hong, Qi Yang, Tao Zhang, Zili Wang, Zhaojin Fu, Kun Ding, Bin Fan, Shiming Xiang

机构 * School of Artificial Intelligence, UCAS(人工智能学院,UCAS) MAIS, Institute of Automation(自动化研究所MAIS) School of Intelligent Science and Technology, University of Science and Technolog Beijing(智能科学与技术学院,北京理工大学)

AI总结 本文提出CAVS任务和CMR框架,通过解决多模态语义漂移和共现混淆问题,提升持续音频-视觉分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10238 2026-03-10 cs.CV

MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation

MTVCraft: 4D运动分词用于任意角色动画

Yanbo Ding, Xirui Hu, Zhizhi Guo, Yan Zhang, Xinrui Wang, Zhixiang He, Chi Zhang, Yali Wang, Xuelong Li

机构 * Shenzhen Key Laboratory of Computer Vision and Pattern Recognition, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China(深圳计算机视觉与模式识别重点实验室,深圳先进技术研究院,中国科学院,深圳,中国) Institute of Artificial Intelligence (TeleAI), China Telecom, Beijing, China(人工智能研究所(TeleAI),中国电信,北京,中国) School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China(计算机科学与技术学院,西安交通大学,西安,中国) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京,中国) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)

AI总结 MTVCraft通过直接建模4D运动序列,实现任意角色动画,提升运动控制灵活性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06925 2026-03-10 cs.CV

Small Target Detection Based on Mask-Enhanced Attention Fusion of Visible and Infrared Remote Sensing Images

基于可见与红外遥感图像的掩码增强注意力融合的小目标检测

Qianqian Zhang, Xiaolong Jia, Ahmed M. Abdelmoniem, Li Zhou, Junshe An

机构 * National Space Science Center, Chinese Academy of Sciences(中国科学院国家空间科学中心) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦大学玛丽女王学院电子工程与计算机科学学院) School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文学与空间科学学院)

AI总结 本文提出ESM-YOLO+,通过掩码增强注意力融合和结构表示增强方法,实现可见与红外图像中小目标的高效检测,提升检测精度并降低模型复杂度。

Comments The manuscript has been submitted to the journal and is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06920 2026-03-10 cs.CV

DLRMamba: Distilling Low-Rank Mamba for Edge Multispectral Fusion Object Detection

DLRMamba: 低秩Mamba用于边缘多光谱融合目标检测

Qianqian Zhang, Leon Tabaro, Ahmed M. Abdelmoniem, Junshe An

机构 * National Space Science Center, Chinese Academy of Sciences(中国科学院国家空间科学中心) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系) School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦大学Queen Mary电子工程与计算机科学学院) School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文与空间科学系)

AI总结 DLRMamba通过低秩SS2D和结构感知知识蒸馏,提升边缘多光谱融合目标检测的效率与精度。

Comments Has been submitted to the IEEE TGRS journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06670 2026-03-10 cs.CV cs.AI

calibfusion: Transformer-Based Differentiable Calibration for Radar-Camera Fusion Detection in Water-Surface Environments

calibfusion: 基于Transformer的可微校准方法用于水表面环境的雷达-摄像头融合检测

Yuting Wan, Liguo Sun, Jiuwu Hao, Pin LV

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 CalibFusion通过端到端学习隐式外校准,提升水表面环境下雷达-摄像头融合检测的鲁棒性和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23681 2026-03-10 cs.CV

QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification

QuantSparse: 通过模型量化和注意力稀疏化全面压缩视频扩散变换器

Weilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Yulun Zhang, Michele Magno, Yongjun Xu

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ETH Zürich(苏黎世联邦理工学院) City College of New York, City University of New York, USA(纽约城市学院,纽约城市大学,美国) Shanghai Jiao Tong University(上海交通大学)

AI总结 QuantSparse通过结合模型量化和注意力稀疏化,实现视频扩散变换器的高效压缩,提升性能并减少存储与推理时间。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21302 2026-03-10 cs.CV

Quantized Visual Geometry Grounded Transformer

量化视觉几何基础变换器

Weilun Feng, Haotong Qin, Mingqiang Wu, Chuanguang Yang, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Yulun Zhang, Michele Magno, Yongjun Xu

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) ETH Zürich(苏黎世联邦理工学院) City College of New York, City University of New York, USA(纽约城市学院,纽约城市大学,美国) Shanghai Jiao Tong University(上海交通大学)

AI总结 本文提出QuantVGGT,首个针对大规模VGGT的量化框架,通过双平滑细粒度量化和噪声过滤多样化采样技术,在保持高精度的同时实现显著的内存减少和加速。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04016 2026-03-10 cs.CV

S$^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation

S$^2$Q-VDiT: 精确量化视频扩散变换器与显著数据和稀疏令牌蒸馏

Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Han Yang, Yuqi Li, Zhulin An, Libo Huang, Michele Magno, Yongjun Xu

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ETH Zürich(苏黎世联邦理工学院)

AI总结 S$^2$Q-VDiT通过显著数据选择和稀疏令牌蒸馏提升视频扩散变换器的量化性能,实现无损压缩和加速推理。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00329 2026-03-10 cs.LG cs.AI math.OC stat.ML

OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack

OTAD: 一种基于最优传输的鲁棒模型用于无偏对抗攻击

Kuo Gai, Sicong Wang, Shihua Zhang

机构 * Shanghai Institute for Mathematics and Interdisciplinary Sciences(上海数学与交叉科学研究院) Academy of Mathematics and Systems Science(数学系统科学学院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 OTAD通过结合最优传输理论与对抗防御方法,提出一种新的模型以提高深度学习系统的鲁棒性和安全性。

Comments 15 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏