arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

2026-05-11 至 2026-05-11 共收录 16
2605.08043 2026-05-11 cs.CV cs.AI

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation

SCOPE:结构分解与条件技能协调用于复杂图像生成

Tianfei Ren, Zhipeng Yan, Yiming Zhao, Zhen Fang, Yu Zeng, Guohui Zhang, Hang Xu, Xiaoxiao Ma, Shiting Huang, Ke Xu, Wenxuan Huang, Lionel Z. Wang, Lin Chen, Zehui Chen, Jie Huang, Feng Zhao

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知 MOE 实验室,中国科学技术大学) The Hong Kong Polytechnic University(香港理工大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出SCOPE框架,通过结构化规范和条件技能协调解决复杂图像生成中语义承诺的持续跟踪问题,通过Gen-Arena基准验证其有效性,取得优于基线的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07692 2026-05-11 cs.AI

GASim: A Graph-Accelerated Hybrid Framework for Social Simulation

GASim:一种用于大规模社会模拟的图加速混合框架

Xuan Zhou, Yanhui Sun, Hantao Yao, Allen He, Yongdong Zhang, Wu Liu

机构 * University of Science and Technology of China(中国科学技术大学) BASIS International School(BASIS国际学校)

AI总结 本文提出GASim框架,通过图优化内存和图消息传递提升大规模社会模拟效率,实现9.94倍速度提升并降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07677 2026-05-11 cs.IR cs.AI cs.CL

TRACE: Tourism Recommendation with Accountable Citation Evidence

TRACE:基于可问责引用证据的旅游推荐

Zixu Zhao, Sijin Wang, Yu Hou, Yuanyuan Xu, Yufan Sheng, Xike Xie, Wenjie Zhang, Won-Yong Shin, Xin Cao

机构 * UNSW Sydney(新南威尔士大学悉尼分校) University of Adelaide(阿德莱德大学) Yonsei University(延世大学) USTC(中国科学技术大学)

AI总结 TRACE通过多轮对话结合评论引用和显式拒绝机制,解决旅游推荐中信任、可验证性和适应性问题,提出三项能力差距并验证评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07378 2026-05-11 cs.LG

Zero-Shot Neural Network Evaluation with Sample-Wise Activation Patterns

基于样本激活模式的零样本神经网络评估

Yameng Peng, Andy Song, HaythamM. Fayek, Vic Ciesielski, Xiaojun Chang

机构 * School of Computing Technologies, RMIT University(计算技术学院,皇家墨尔本理工大学) Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学)

AI总结 本文提出SWAP及其衍生指标SWAP-Score,用于评估神经网络性能,克服了传统零样本指标的局限性,展现强预测能力。

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence. This article is a journal extension of arXiv:2403.04161

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15719 2026-05-11 cs.AI

Harnessing Pre-Resolution Signals for Future Prediction Agents

利用预决议信号为未来预测代理

Chuyang Wei, Maohang Gao, Zhixin Han, Kefei Chen, Yu Zhuang, Haoxiang Guan, Yanzhi Zhang, Yilin Cheng, Xiren Zhou, Huanhuan Chen, Jian Li, Jiyan He, Yu Shi, Yitong Duan, Shuxin Zheng

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国) IIIS, Tsinghua University(清华大学人工智能研究院)

AI总结 本文提出通过预决议信号提升未来预测性能,通过Milkyway代理在未决问题多次回顾中提取信息,改进预测并增强后续准确性。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07492 2026-05-11 cs.CV

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

文档解析距离解决还有多远?PureDocBench:一个跨干净、退化和现实场景的可追溯基准测试

Zhiheng Li, Zongyang Ma, Jiaxian Chen, Jianing Zhang, Zhaolong Su, Yutong Zhang, Zhiyin Yu, Ruiqi Liu, Xiaolei Lv, Bo Li, Jun Gao, Ziqi Zhang, Chunfeng Yuan, Bing Li, Weiming Hu

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学技术大学) NWPU(西北工业大学) JLU(吉林大学) USTC(中国科学技术大学) PKU(北京大学) HelloGroup ShanghaiTech(上海科技大学) Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京多模态信息超级智能安全重点实验室)

AI总结 本文提出PureDocBench,一个可追溯的基准测试,覆盖10个领域、66个子类和1475页文档,包含清洁、数字退化和现实退化三种版本。评估40种模型发现,文档解析尚未解决,最佳模型得分仅74分,公式识别仍是瓶颈,通用VLM在退化下表现不如专用模型。

Comments 42 pages, 20 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07396 2026-05-11 cs.LG cs.AI

Rubric-based On-policy Distillation

基于评分标准的在线蒸馏

Junfeng Fang, Zhepei Hong, Mao Zheng, Mingyang Song, Gengsheng Li, Houcheng Jiang, Dan Zhang, Haiyun Guo, Xiang Wang, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) Tencent(腾讯)

AI总结 本文提出ROPD框架,利用教师生成的响应而非教师日志进行在线蒸馏,实现更灵活的黑盒兼容方案,实验显示在多数场景下性能优于基于日志的蒸馏方法,样本效率提升达10倍。

Comments Preprint. Code is available at https://github.com/Peregrine123/ROPD_official

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07315 2026-05-11 cs.CL

LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification

LaTER:通过潜在探索和显式验证实现高效的推理

Xuan Li, Yining Wang, Yuchen Liu, Guanjun Liu, Delai Qiu, Shengping Liu, Jiaen Liang, Wei Huang, Jun Yu, Junnan Zhu

机构 * University of Science and Technology of China(科学技术大学) Unisound AI Technology Co., Ltd(Unisound AI技术有限公司) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

AI总结 LaTER通过连续潜在空间的有限探索和显式CoT验证提升推理效率,在多个基准测试中减少16%-32%的token使用量,同时提升准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07299 2026-05-11 cs.CV cs.AI

EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

EgoPro-Bench:基于眼动视频流的个性化主动交互基准测试

Dongchuan Ran, Linyu Ou, Xueheng Li, Wenwen Tong, Chenxu Guo, Hewei Guo, Kaibing Wang, Lewei Lu

机构 * Beijing Institute of Technology(北京理工大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出EgoPro-Bench,通过模拟用户画像生成多样化意图,构建高保真人机交互数据,评估主动交互能力,提升多模态大语言模型的意图理解与交互时机识别能力。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06882 2026-05-11 cs.AI

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem

LLMs在最简单的长链推理任务上表现如何:等价类问题的实证研究

Chun Zheng, Lianlong Wu, Bingqian Li, Lvting Liu, Yi Zhou

机构 * University of Science and Technology of China(中国科学技术大学) University of Oxford(牛津大学)

AI总结 本文评估了LLMs在等价类问题上的表现,发现非推理模型无法解决,而推理模型虽更优但仍难以完全解决,且问题难度随连接概率和变量数变化而变化。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06676 2026-05-11 cs.LG cs.CL

LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

LKV:端到端学习头部预算和令牌选择以LLM KV缓存淘汰

Enshuai Zhou, Yifan Hao, Chao Wang, Rui Zhang, Di Huang, Jiaming Guo, Xing Hu, Zidong Du, Qi Guo, Yunji Chen

机构 * University of Science and Technology of China(中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, CAS, Beijing, China(中国科学院计算技术研究所过程器重点实验室,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)

AI总结 本文提出LKV,通过端到端可微优化问题实现KV缓存压缩,学习任务优化全局预算和内在KV重要性,提升长上下文推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05693 2026-05-11 cs.AI cs.LG

Saliency-Aware Regularized Quantization Calibration for Large Language Models

具有显著性感知的正则化量化校准用于大语言模型

Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu, Baihua He, Xinyu Zhang, Harrison Bo Hua Zhu, Wenlong Chen, Li Zeng, Zhuo Sun

机构 * University of Science and Technology of China(中国科学技术大学) University College London(伦敦大学学院) Shanghai University of Finance and Economics(上海财经大学) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) University of Copenhagen(哥本哈根大学) Imperial College London(伦敦帝国学院) Technical University of Denmark(丹麦技术大学) Peking University(北京大学)

AI总结 本文提出SARQC,通过引入显著性感知正则化改进量化校准,提升大语言模型的泛化能力,实验表明在密集和专家混合模型中提升了困惑度和零样本准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16579 2026-05-11 cs.LG cs.AI

EviDep: Trustworthy Multimodal Depression Estimation via Disentangled Evidential Learning

EviDep:通过解耦证据学习实现可信的多模态抑郁估计

Fangyuan Liu, Sirui Zhao, Zeyu Zhang, Jinyang Huang, Feng-Qi Cui, Bin Luo, Meng Li, Tong Xu, Enhong Chen

机构 * School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) Department of Psychiatry, The First Affiliated Hospital of University of Science and Technology of China, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学附属第一医院精神科,生命科学与医学学院)

AI总结 本文提出EviDep框架,通过解耦证据学习量化抑郁严重程度及不确定性,提升多模态抑郁估计的可信度和不确定性校准能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18636 2026-05-11 cs.CV

Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering

注意力稀疏性是输入稳定的:通过离线稀疏性分析和在线QK共聚类实现训练自由的视频生成稀疏注意力

Jiayi Luo, Jiayu Chen, Jiankun Wang, Cong Wang, Hanxin Zhu, Qingyun Sun, Chen Gao, Zhibo Chen, Jianxin Li

机构 * SKLCCSE, School of Computer Science and Engineering, Beihang University(软件学院,北京航空航天大学) Beihang University(北京航空航天大学) School of Computer Science, Peking University(北京大学计算机学院) the State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) BNRist, Tsinghua University(北京理工大学,清华大学) Zhongguancun Academy(中关村学院)

AI总结 本文提出SVOO框架,通过离线层间稀疏性分析和在线双向共聚类实现训练自由的视频生成稀疏注意力,解决传统方法中层异质性和查询-键耦合问题,提升生成质量与速度的平衡。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15356 2026-05-11 eess.IV cs.AI

Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing

Q-Probe:通过上下文感知代理探测扩展图像质量评估至高分辨率

Xiang Li, Xueheng Li, Yu Wang, Xuanhua He, Zhangchi Hu, Weiwei Yu, Chengjun Xie

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Institute of Intelligent Machines, Chinese Academy of Sciences(中国科学院智能 Machines 研究所)

AI总结 Q-Probe通过上下文感知探测方法解决高分辨率图像质量评估中的局部退化捕捉问题,提出Vista-Bench基准和三阶段训练框架,实现高分辨率下的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20909 2026-05-11 cs.CV eess.IV

Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation

Dino U-Net:利用基础模型的高保真密集特征进行医学图像分割

Haoyue Li, Yifan Gao, Feng Yuan, Xiaosong Wang, Xin Gao

机构 * School of Biomedical Engineering (Suzhou), Division of Life Science and Medicine, University of Science and Technology of China, Hefei, China(生物医学工程学院(苏州),生命科学与医学系,中国科学技术大学,合肥,中国) Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou, China(苏州生物医学工程与技术研究所,中国科学院,苏州,中国) Shanghai Innovation Institute, Shanghai, China(上海创新研究院,上海,中国) Medical School of Tianjin University, Tianjin, China(天津大学医学院,天津,中国) Jinan Guoke Medical and Technology Development Co., Ltd., Pharmaceutical Valley New Drug Creation Platform, Jinan, China(济南国科医药科技发展有限公司,药谷新药创制平台,济南,中国)

AI总结 本文提出Dino U-Net,通过融合DINOv3模型的语义特征与低层空间细节,提升医学图像分割精度,实验表明其在多种影像模态中均优于现有方法。

Comments MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏