arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

University of Science and Technology of China(中国科学技术大学)

2026-03-02 至 2026-03-02 共收录 14
2602.24283 2026-03-02 cs.LG cs.AI cs.CL

Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation

驯服动量:通过低秩近似重新思考优化器状态

Zhengbo Wang, Jian Liang, Ran He, Zilei Wang, Tieniu Tan

机构 * University of Science and Technology of China(中国科学技术大学) NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

AI总结 LoRA-Pre通过低秩近似优化器状态,提升预训练和微调效率,实现内存节省与性能提升

Comments Camera-ready version. Accepted as Oral at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24231 2026-03-02 cs.LG

Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference

自适应组合实验设计:决策与推断的帕累托最优性

Hongrui Xie, Junyu Cao, Kan Xu

机构 * University of Science and Technology of China(中国科学技术大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Arizona State University(亚利桑那州立大学)

AI总结 本文提出 MixCombKL 和 MixCombUCB 算法,通过帕累托最优性在组合多臂老虎机中实现 regret 最小化与统计功效的平衡。

Comments 30 pages, 3 figure, AISTATS 2026 accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24041 2026-03-02 cs.CV

Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation

仔细观察:多模态大语言模型中的自适应视觉增强以缓解幻觉

Xingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang, Zhicai Wang, Beier Zhu, Hanwang Zhang

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(信息与电子技术联合实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学)

AI总结 本研究提出自适应视觉增强框架AIR,通过减少冗余标记和选择性整合补丁来缓解多模态大语言模型中的幻觉问题。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24027 2026-03-02 cs.CV cs.MM

GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models

GuardAlign: 多模态大语言模型中的测试时安全性对齐

Xingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang, Yin Zhang, Xiang Wang, Xiangnan He

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(摩埃关键实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) Tianjin University(天津大学)

AI总结 GuardAlign通过OT增强的安全检测和跨模态注意力校准,有效提升多模态大语言模型在测试时的安全性,减少不安全响应率并提升任务表现。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23981 2026-03-02 cs.LG cs.AI

Intrinsic Lorentz Neural Network

内禀洛伦兹神经网络

Xianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu Sebe

机构 * University of Science and Technology of China(中国科学技术大学) University of Trento(特伦托大学) Peking University(北京大学)

AI总结 内禀洛伦兹神经网络通过全内禀双曲架构提升几何决策性能,结合陀螺归一化和双曲距离计算,实现优于现有方法的性能和效率。

Comments Published in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23959 2026-03-02 cs.CV

Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought

通过图像作为连续动作进行思考:数值视觉链式推理

Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Science and Technology of China(中国科学技术大学) Chalmers University of Technology(楚克理工大学) University of Gothenburg(哥德堡大学)

AI总结 NV-CoT通过将图像推理动作空间扩展为连续欧几里得空间,提升MLLMs的定位精度和回答准确性,同时加速训练收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23699 2026-03-02 cs.CV cs.CL

HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit

HiDrop:通过晚期注入、凹形金字塔剪枝和早期退出实现MLLM中的层次视觉令牌减少

Hao Wu, Yingqi Fan, Jinyang Dai, Junlong Tong, Yunpu Ma, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Munich Center for Machine Learning, LMU Munich(慕尼黑大学机器学习中心,慕尼黑大学)

AI总结 HiDrop通过晚期注入、凹形金字塔剪枝和早期退出机制,实现多模态大语言模型中视觉令牌的高效减少,提升训练效率并保持性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23676 2026-03-02 cs.CV

Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering

通过语义解耦潜在引导抑制放射科报告生成中的先验比较幻觉

Ao Li, Rui Liu, Mingjie Li, Sheng Liu, Lei Wang, Xiaodan Liang, Lina Yao, Xiaojun Chang, Lei Xing

机构 * University of New South Wales(新南威尔士大学) Australian Artificial Intelligence Institute, University of Technology Sydney(澳大利亚人工智能研究所,技术悉尼大学) Stanford University(斯坦福大学) School of Computing and Information Technology of University of Wollongong Australia(沃林根澳大利亚大学计算与信息科技学院) Sun Yat-sen University(中山大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出语义解耦潜在引导方法,通过正交化技术减少放射科报告生成中的历史幻觉,提升临床准确性与报告忠实度。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23648 2026-03-02 cs.RO

FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic Manipulation

FAVLA:一种力适应的快速-慢速VLA模型用于接触丰富的机械臂操作

Yao Li, Peiyuan Tang, Wuyang Zhang, Chengyang Zhu, Yifan Duan, Weikai Shi, Xiaodong Zhang, Zijiang Yang, Jianmin Ji, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Xi'an Jiaotong University(西安交通大学) Central South University(中南大学)

AI总结 FAVLA通过解耦慢感知规划与快速接触感知控制,提升接触丰富任务中的反应性和成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21644 2026-03-02 cs.RO

DAGS-SLAM: Dynamic-Aware 3DGS SLAM via Spatiotemporal Motion Probability and Uncertainty-Aware Scheduling

DAGS-SLAM:通过时空运动概率和不确定性感知调度实现动态感知的3DGS SLAM

Li Zhang, Yu-An Liu, Xijia Jiang, Conghao Huang, Danyang Li, Yanyong Zhang

机构 * School of Mathematics, Hefei University of Technology(合肥工业大学数学学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)

AI总结 DAGS-SLAM通过时空运动概率和不确定性感知调度实现动态感知的3DGS SLAM,提升实时定位与密集重建的鲁棒性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24038 2026-03-02 cs.CV cs.MA

Enhancing CLIP Robustness via Cross-Modality Alignment

通过跨模态对齐增强CLIP鲁棒性

Xingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao, Hanwang Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学)

AI总结 COLA通过跨模态对齐提升CLIP对抗鲁棒性,有效缓解对抗扰动导致的特征不一致问题,提升零样本分类性能。

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22353 2026-03-02 cs.LG cs.AI

Context and Diversity Matter: The Emergence of In-Context Learning in World Models

上下文与多样性至关重要:世界模型中情境学习的出现

Fan Wang, Zhiyuan Chen, Yuxuan Zhong, Sunjian Zheng, Pengtao Shao, Bo Yu, Shaoshan Liu, Jianan Wang, Ning Ding, Yang Cao, Yu Kang

机构 * Shenzhen Institute of Artificial Intelligence and Robotics for Society(深圳人工智能与机器人社会研究院) University of Science and Technology of China(中国科学技术大学) Anhui Province Key Laboratory of Intelligent Low-Carbon Information Technology and Equipment(安徽省智能低碳信息技术与设备重点实验室)

AI总结 本文研究了世界模型中情境学习的机制,揭示了环境识别和学习的核心作用,并探讨了长上下文和多样化环境对学习效果的影响。

Journal ref 2026 International Conference on Learning Representations (ICLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01728 2026-03-02 cs.CV

Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion

Shuffle Mamba:基于随机洗牌的态空间模型用于多模态图像融合

Ke Cao, Xuanhua He, Tao Hu, Chengjun Xie, Man Zhou, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Intelligent Machines(智能机器研究所) Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院) Intelligent Agriculture Engineering Laboratory of Anhui Province, Institute of Intelligent Machines(安徽省智能农业工程实验室,智能机器研究所)

AI总结 Shuffle Mamba通过引入随机洗牌策略和逆洗牌,解决多模态图像融合中固定扫描策略带来的偏见问题,提升融合质量。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00320 2026-03-02 cs.LG

TimeMAE: Self-Supervised Representations of Time Series with Decoupled Masked Autoencoders

TimeMAE:解耦掩码自编码器的时序自监督表示

Mingyue Cheng, Xiaoyu Tao, Zhiding Liu, Qi Liu, Hao Zhang, Rujiao Zhang, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

AI总结 TimeMAE通过解耦掩码自编码器和语义单元提升,提升时间序列自监督表示的性能,尤其在数据稀缺和迁移学习场景中表现优异。

Comments Accepted by WSDM'26

详情

展开后加载摘要…

URL PDF HTML 收藏