arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Nanjing University(南京大学)

2026-05-28 至 2026-05-28 共收录 17
2605.28809 2026-05-28 cs.CV cs.LG

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

AREA: 基于CLIP的类增量学习中的属性提取与聚合

Zhen-Hao Xie, Yu-Cheng Shi, Da-Wei Zhou

机构 * State Key Laboratory of Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

AI总结 提出AREA方法,通过主测地线分析稳定属性提取、轻量级任务专家和变分信息瓶颈正则化稳定属性聚合,并利用最优传输进行推理,以解决CLIP类增量学习中的灾难性遗忘问题。

Comments Accepted to ICML 2026. Code is available at https://github.com/LAMDA-CL/ICML2026-AREA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28480 2026-05-28 eess.AS cs.SD

Audio-Mind: An Auditable Agentic Framework for Audio Understanding

Audio-Mind: 一种可审计的音频理解智能体框架

Yucheng Wang, Jing Peng, Hanqi Li, Chenghao Wang, Wenming Tu, Yu Xi, Zhaokai Sun, Kai Yu, Shuai Wang

机构 * School of Intelligence Science and Technology, Nanjing University, China(南京大学智能科学与技术学院) Department of Computer Science, ETH Zürich, Switzerland(苏黎世联邦理工学院计算机科学系) X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, China(上海交通大学计算机科学学院X-LANCE实验室) School of Automation Science and Engineering, Xi’an Jiaotong University, China(西安交通大学自动化科学与工程学院) School of Computer Science, Northwestern Polytechnical University, China(西北工业大学计算机科学学院)

AI总结 提出Audio-Mind框架,通过条件性证据获取动态结合强前端与规划器引导的工具使用,解决音频理解中智能体证据获取的时机问题,在MMAR和MSU-Bench上分别达到80.4%和82.8%的准确率,并生成可审计的推理轨迹。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28433 2026-05-28 cs.CL

Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning

角色与轨道:多智能体结构化推理中保持契约的角色演化

Ling-Yue Ge, Lan-Zhe Guo

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

AI总结 提出SERO框架,通过契约保持的角色演化机制(信用引导检索、保护终端聚合器、条件验证器修复、上下文赌博机控制器)解决多智能体系统中角色漂移和契约破坏问题,在真实推理基准上验证有效性。

Comments 33 pages, 23 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27894 2026-05-28 cs.CV

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

面向不完整多模态输入的统一视觉-语言模型

Xiang Fang, Wanlong Fang, Changshuo Wang, Keke Tang, Daizong Liu, Siyi Wang, Wei Ji

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院) Guangzhou University(广州大学) Wuhan University(武汉大学) Nanjing University(南京大学)

AI总结 针对视频-语言模型在传感器失效导致模态不完整数据下的训练-测试不一致问题,提出首个统一的不完整视频-语言模型作为即插即用模块,提升多模态任务性能。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27582 2026-05-28 cs.RO cs.CV

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation

Uni-LaViRA:面向统一具身导航的语言-视觉-机器人动作翻译

Hongyu Ding, Sizhuo Zhang, Ziming Xu, Jinwen Guo, Hongxiu Liu, Xingzhi Cheng, Zixuan Chen, Haifei Qi, Duo Wang, Hao Xu, Jieqi Shi, Yifan Zhang, Jing Huo, Jian Cheng, Yang Gao, Jiebo Luo

机构 * Nanjing University(南京大学) Beihang University(北京航空航天大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) BMW (Nanjing) Information Technology Co., Ltd.(宝马(南京)信息技术有限公司) University of Rochester(罗切斯特大学)

AI总结 提出Uni-LaViRA统一智能体架构,通过语言-视觉-机器人动作翻译结构,结合待办列表记忆和二次机会回溯机制,在零训练下实现四类导航任务和四种真实机器人的零样本泛化,性能匹配或超越近期训练式导航基础模型。

Comments Project page: https://xetroubadour.github.io/Uni-LaViRA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27365 2026-05-28 cs.CV cs.AI cs.LG cs.RO

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

LocateAnything: 基于并行框解码的快速高质量视觉定位

Shihao Wang, Shilong Liu, Yuanguo Kuang, Xinyu Wei, Yangzhou Liu, Zhiqi Li, Yunze Man, Guo Chen, Andrew Tao, Guilin Liu, Jan Kautz, Lei Zhang, Zhiding Yu

机构 * The Hong Kong Polytechnic University(香港理工大学) Princeton University(普林斯顿大学) Nanjing University(南京大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出并行框解码(PBD)方法,将边界框和点作为原子单元单步解码,结合大规模数据集LocateAnything-Data,实现高效统一的目标定位与检测,在保持高精度同时显著提升解码吞吐量。

Comments fix github link

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09367 2026-05-28 cs.CV

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

EpiAgent: 一种以智能体为中心的古铭文修复系统

Shipeng Zhu, Ang Chen, Na Nie, Pengfei Fang, Min-Ling Zhang, Hui Xue

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其跨学科应用关键实验室(东南大学),教育部,中国) Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China(计算机网络与信息集成关键实验室(东南大学),教育部,中国) Nanjing University Museum, Nanjing University, China(南京大学博物馆,南京大学,中国) The China Centre for Linguistic and Strategic Studies, Nanjing University, China(中国语言战略研究中心,南京大学,中国)

AI总结 提出基于智能体的EpiAgent系统,通过分层规划与LLM协调多模态分析、历史经验和专用工具,实现灵活自适应的古铭文修复,在真实退化铭文上取得更优修复质量和泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17354 2026-05-28 cs.CV cs.GR

PocketGS: On-Device Training of 3D Gaussian Splatting for High Perceptual Modeling

PocketGS: 用于高感知建模的3D高斯泼溅设备端训练

Wenzhi Guo, Guangchi Fang, Shu Yang, Bing Wang

机构 * Hong Kong Polytechnic University(香港理工大学) Nanjing University(南京大学)

AI总结 提出PocketGS,通过三个协同设计的算子(G、I、T)在移动设备上实现3D高斯泼溅的高效训练,在严格资源约束下保持高保真重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01766 2026-05-28 cs.RO

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

神经隐式动作场:从离散路点到连续函数的视觉-语言-动作模型

Haoyun Liu, Jianzhuang Zhao, Xinyuan Chang, Tianle Shi, Chuanzhang Meng, Jiayuan Tan, Feng Xiong, Tong Lin, Dongjie Huo, Mu Xu, SongLin Dong, Zhiheng Ma, Yihong Gong, Sheng Zhong

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Faculty of Computility Microelectronics, Shenzhen University of Advanced Technology(深圳大学计算微电子学院) Guangdong Provincial Key Laboratory of Computility Microelectronics(广东省计算微电子重点实验室) Amap, Alibaba Group(阿里集团Amap) Shenzhen University(深圳大学) Xi'an Jiaotong University(西安交通大学) Beijing University of Chemical Technology(北京化工大学)

AI总结 针对视觉-语言-动作模型预测离散动作路点与物理运动连续性不匹配的问题,提出神经隐式动作场(NIAF),通过将动作表示从离散路点重构为连续函数,实现任意时间分辨率的连续动作流形合成,支持解析求导和显式速度监督,提升控制平滑性和物理合理性。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11564 2026-05-28 cs.CV

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

LUVE:基于双频专家的潜在级联超高分辨率视频生成

Chen Zhao, Jiawei Chen, Hongyu Li, Zhuoliang Kang, Shilin Lu, Xiaoming Wei, Kai Zhang, Jian Yang, Ying Tai

机构 * Nanjing University(南京大学) Nanyang Technological University(南洋理工大学)

AI总结 提出LUVE框架,通过三阶段潜在级联架构(低分辨率运动生成、潜在空间上采样、高分辨率内容精炼)结合双频专家,解决超高分辨率视频生成中的运动建模、语义规划和细节合成难题。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01990 2026-05-28 cs.LG cs.AI

SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning

SAME: 用于多模态持续指令调优的稳定混合专家模型

Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye, De-Chuan Zhan, Da-Wei Zhou

机构 * State Key Laboratory of Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室) School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院)

AI总结 针对多模态持续指令调优中专家路由漂移和专家漂移问题,提出稳定混合专家模型(SAME),通过正交子空间分解路由动态和曲率感知缩放更新专家,实现无重放的状态最优性能。

Comments Accepted to ICML 2026. Code is available at https://github.com/LAMDA-CL/Prism

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03048 2026-05-28 cs.CV cs.AI cs.CC

On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning

关于Transformer图像嵌入在非可解空间推理中的内在限制

Siyi Lyu, Quan Liu, Feng Yan

机构 * School of Electronic Science and Engineering, Nanjing University, Nanjing, China(电子科学与工程学院,南京大学,南京,中国)

AI总结 本文通过将空间理解形式化为群同态问题,证明恒定深度Transformer由于TC⁰复杂度限制,无法在单次前向传播中捕获非可解群(如SO(3))的空间结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16483 2026-05-28 cs.CV

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

FasterVAR:视觉自回归模型的即插即用加速

Senmao Li, Kai Wang, Salman Khan, Fahad Shahbaz Khan, Jian Yang, Yaxing Wang

机构 * PCA Lab, VCIP, College of Computer Science, Nankai University(南开大学计算机学院、VCIP、PCA实验室) Program of Computer Science, City University of Hong Kong (Dongguan), China(香港城市大学(东莞)计算机系,中国) City University of Hong Kong, HK SAR, China(香港城市大学,香港特别行政区,中国) Mohamed bin Zayed University of Artificial Intelligence, UAE(阿联酋Mohamed bin Zayed人工智能大学) Linkoping University, Sweden(林地平大学,瑞典) PCA Lab, School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院、PCA实验室) College of Artificial Intelligence, Jilin University(吉林大学人工智能学院)

AI总结 针对VAR模型在大尺度步骤计算复杂度高的问题,提出一种基于阶段感知的即插即用加速框架FasterVAR,通过保留早期关键步骤并剪枝或近似后期细节步骤,实现最高3.4倍加速且几乎无性能损失。

Comments Accepted at ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10185 2026-05-28 cs.CL cs.AI cs.MA

Auditing medical multi-agent AI reveals risks of false consensus

审计医疗多智能体AI揭示虚假共识风险

Yinghao Zhu, Lei Gu, Zixiang Wang, Haoran Sang, Dehao Sui, Wen Tang, Lan Mi, Yasha Wang, Junyi Gao, Liang Yao, Tianfan Fu, Ewen Harrison, Lequan Yu, Liantao Ma

机构 * National Engineering Research Center for Software Engineering, Peking University(北京大学软件工程国家工程研究中心) School of Computing and Data Science, The University of Hong Kong(香港大学计算机与数据科学学院) Department of Nephrology, Peking University Third Hospital(北京大学第三医院肾内科) Key Laboratory of Carcinogenesis and Translational Research (Ministry of Education), Department of Lymphoma, Peking University Cancer Hospital & Institute(教育部癌症发生与转化研究重点实验室、北京大学肿瘤医院淋巴瘤科) Department of Automation, Tsinghua University(清华大学自动化系) Centre for Medical Informatics, The University of Edinburgh(爱丁堡大学医学信息学中心) Health Data Research UK(英国健康数据研究机构) Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李科贤医学院) State Key Laboratory for Novel Software Technology, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室、计算机科学学院)

AI总结 本研究提出MedAgentAudit框架,通过专家验证的审计流程诊断医疗多智能体系统中的协作失败模式,发现虚假共识、权威偏差等系统性风险。

Comments Code and Data: https://github.com/MedX-PKU/MedAgentAudit

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18271 2026-05-28 cs.CV

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

ObjFiller3D:将3D物体修复扩展到密集多视图一致性

Haitang Feng, Xinkai Chen, Jie Liu, Jie Tang, Gangshan Wu, Beiqi Chen, Jianhuang Lai, Guangcong Wang

机构 * Nanjing University(南京大学) Great Bay University(大湾大学) Harbin Institute of Technology(哈尔滨工业大学) Sun Yat-sen University(中山大学)

AI总结 提出ObjFiller-3D方法,通过联合优化密集采样视图的时序生成、语义感知补全和循环一致性3D编码,实现高质量且一致的多视图3D物体修复与编辑。

Comments Project page: https://objfiller3d.github.io/ Code: https://github.com/objfiller3d/ObjFiller-3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09449 2026-05-28 cs.CV

RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration

RASR: 面向实际参考图像复原的检索增强超分辨率

Jiaqi Yan, Shuning Xu, Xiangyu Chen, Dell Zhang, Jiantao Zhou, Jie Tang, Gangshan Wu, Jie Liu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) State Key Laboratory of Internet of Things for Smart City(物联网智慧城市国家重点实验室) Department of Computer Science and Information Science, University of Macau(澳门大学计算机科学与信息科学系)

AI总结 提出检索增强超分辨率(RASR)范式,通过自动检索参考图像实现实际场景下的参考超分辨率,并构建基准数据集RASR-Flickr30及基线模型RASRNet。

Comments Accepted at ISCAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04924 2026-05-28 cs.CV eess.IV

Inter-event Interval Microscopy for Event Cameras

事件相机的帧间间隔显微术

Changqing Su, Yanqin Chen, Zihan Lin, Zhen Cheng, You Zhou, Bo Xiong, Zhaofei Yu, Tiejun Huang

机构 * National Key Laboratory for Multimedia Information Processing(国家多媒体信息处理重点实验室) Westlake Laboratory of Life Sciences and Biomedicine(西湖生命科学与生物医学实验室) School of Automation(自动化学院) Department of Automation(自动化系) Nanjing University Medical School(南京大学医学院)

AI总结 提出基于事件相机的帧间间隔显微术(IEIM),通过量化连续事件的时间间隔实现静态和动态场景的强度重建,在荧光显微镜中实现高时空分辨率和动态范围。

详情

展开后加载摘要…

URL PDF HTML 收藏