arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

共收录 2226
2603.00551 2026-03-03 cs.PF cs.AR cs.LG

GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning

GCL-Sampler: 通过图对比学习发现采样GPU模拟中的核相似性

Jiaqi Wang, Jingwei Sun, Jiyu Luo, Han Li, Guangzhong Sun

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 GCL-Sampler通过图对比学习发现GPU模拟中的核相似性,实现高保真度和显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00412 2026-03-03 cs.CV

PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models

PointAlign:用于3D视觉-语言模型的特征级对齐正则化

Yuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi Fan

机构 * University of Science and Technology of China(中国科学技术大学) Fuzhou University(福州大学) Fudan University(复旦大学) Nanjing University(南京大学)

AI总结 PointAlign通过特征级对齐正则化提升3D视觉-语言模型的几何信息保留与任务性能

Comments CVPR 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10575 2026-03-03 cs.CV

UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation

UniFlow:一种统一的像素流标记器用于视觉理解和生成

Zhengrong Yue, Haiyu Zhang, Xiangyu Zeng, Boyu Chen, Chenting Wang, Shaobin Zhuang, Lu Dong, Yi Wang, Limin Wang, Yali Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) Shenzhen Key Lab of Computer Vision and Pattern Recognition, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳计算机视觉与模式识别重点实验室,深圳先进技术研究院,中国科学院) Nanjing University(南京大学) University of Science and Technology of China(中国科学技术大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

AI总结 UniFlow是一种统一的像素流标记器,通过灵活适配视觉编码器和轻量级解码器,在视觉理解和生成任务中实现了性能的双赢。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07940 2026-03-03 cs.CV cs.AI cs.CL cs.LG cs.MM

TTOM: Test-Time Optimization and Memorization for Compositional Video Generation

TTOM:测试时优化与记忆化用于组合视频生成

Leigang Qu, Ziyang Wang, Na Zheng, Wenjie Wang, Liqiang Nie, Tat-Seng Chua

机构 * National University of Singapore(国立新加坡大学) University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

AI总结 TTOM通过测试时优化与记忆化机制,提升组合视频生成的跨模态对齐能力,实现高效且可扩展的实时生成。

Comments ICLR 2026 Camera-ready. Project page: https://ttom-t2v.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22611 2026-03-03 cs.LG cs.AI

Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning

分位数优势估计:稳定LLM推理的RLVR

Junkang Wu, Kexin Huang, Jiancan Wu, An Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 通过分位数优势估计方法,稳定RLVR训练过程,提升LLM推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19877 2026-03-03 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph physics.comp-ph

Advancing Universal Deep Learning for Electronic-Structure Hamiltonian Prediction of Materials

推进通用深度学习用于材料电子结构哈密顿量预测

Shi Yin, Zujian Dai, Xinyang Pan, Lixin He

机构 * Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(1 人工智能研究所,合肥国家综合科学中心) Laboratory of Quantum Information, University of Science and Technology of China(2 量子信息实验室,中国科学技术大学) Hefei National Laboratory, University of Science and Technology of China(3 合肥国家实验室,中国科学技术大学)

AI总结 NextHAM通过神经E(3)对称性和表达性修正方法,提升材料电子结构哈密顿量预测的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12813 2026-03-03 cs.RO cs.SY eess.SY

Bridging Perception and Planning: Towards End-to-End Planning for Signal Temporal Logic Tasks

连接感知与规划:面向信号时序逻辑任务的端到端规划

Bowen Ye, Junyue Huang, Yang Liu, Xiaozhen Qiao, Xiang Yin

机构 * School of Automation & Intelligent Sensing, Shanghai Jiao Tong University(自动化与智能感知学院,上海交通大学) University of Minnesota, Twin Cities(明尼苏达大学,双城分校) School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学)

AI总结 本文提出了一种端到端的STL规划器,通过结构感知的混合专家模型,实现对信号时序逻辑任务的高效规划与执行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03516 2026-03-03 cs.CV

Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

绘画比思考更容易:文本到图像模型能否铺垫,却无法主导?

Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang, Rui Chen, Jiarong Ou, Xin Tao, Pengfei Wan, Xiaojuan Qi, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) The University of Hong Kong(香港大学)

AI总结 本文提出T2I-CoReBench基准测试,用于评估文本到图像模型的组合与推理能力,揭示现有模型在高组合场景和推理任务中的局限性。

Comments Accepted to ICLR 2026. Project Page: https://t2i-corebench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17568 2026-03-03 cs.CR cs.AI cs.SD eess.AS

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

JALMBench: 对音频语言模型中越权攻击的基准测试

Zifan Peng, Yule Liu, Zhen Sun, Mingchen Li, Zeren Luo, Jingyi Zheng, Wenhan Dong, Xinlei He, Xuechao Wang, Yingjie Xue, Shengmin Xu, Xinyi Huang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) State Key Laboratory of Internet Architecture, Tsinghua University(清华大学互联网架构国家重点实验室) University of North Texas(北卡罗来纳州立大学) University of Science and Technology of China(中国科学技术大学) Fujian Normal University(福建师范大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

AI总结 JALMBench通过评估11,316个文本样本和245,355个音频样本,分析LALMs对越权攻击的安全性,揭示模态和架构对安全性的影响,强调需专门设计的防御方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07392 2026-03-03 cs.CV

SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models

SPEED:可扩展、精确和高效的扩散模型概念擦除

Ouxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang, Yanbin Hao, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学)

AI总结 SPEED通过直接编辑模型参数,高效精准地擦除扩散模型中的多个概念,同时保护非目标概念的质量。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00194 2026-03-03 cs.CV cs.AI cs.CR

SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models

SKeDA:一种面向文本到视频扩散模型的生成水印框架

Yang Yang, Xinze Zou, Zehua Ma, Han Fang, Weiming Zhang

机构 * School of Electronic and Information Engineering, Anhui University(安徽大学电子与信息工程学院) Anhui Province Key Laboratory of Digital Security and the CAS Key Laboratory of Electromagnetic Space Information, University of Science and Technology of China(安徽省数字安全重点实验室和中国科学院电磁空间信息重点实验室,中国科学技术大学) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

AI总结 SKeDA是一种专为文本到视频扩散模型设计的生成水印框架,通过分布保持采样和差分注意机制提升水印的鲁棒性和可靠性。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00037 2026-03-03 cs.LG cs.AI

StaTS: Spectral Trajectory Schedule Learning for Adaptive Time Series Forecasting with Frequency Guided Denoiser

StaTS: 基于频域引导去噪器的自适应时间序列预测的频谱轨迹调度学习

Jintao Zhang, Zirui Liu, Mingyue Cheng, Xianquan Wang, Zhiding Liu, Qi Liu

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

AI总结 StaTS通过交替更新学习噪声调度和去噪器,提升时间序列预测的结构保持和异质恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24283 2026-03-02 cs.LG cs.AI cs.CL

Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation

驯服动量:通过低秩近似重新思考优化器状态

Zhengbo Wang, Jian Liang, Ran He, Zilei Wang, Tieniu Tan

机构 * University of Science and Technology of China(中国科学技术大学) NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

AI总结 LoRA-Pre通过低秩近似优化器状态,提升预训练和微调效率,实现内存节省与性能提升

Comments Camera-ready version. Accepted as Oral at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24231 2026-03-02 cs.LG

Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference

自适应组合实验设计:决策与推断的帕累托最优性

Hongrui Xie, Junyu Cao, Kan Xu

机构 * University of Science and Technology of China(中国科学技术大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Arizona State University(亚利桑那州立大学)

AI总结 本文提出 MixCombKL 和 MixCombUCB 算法,通过帕累托最优性在组合多臂老虎机中实现 regret 最小化与统计功效的平衡。

Comments 30 pages, 3 figure, AISTATS 2026 accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24041 2026-03-02 cs.CV

Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation

仔细观察:多模态大语言模型中的自适应视觉增强以缓解幻觉

Xingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang, Zhicai Wang, Beier Zhu, Hanwang Zhang

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(信息与电子技术联合实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学)

AI总结 本研究提出自适应视觉增强框架AIR,通过减少冗余标记和选择性整合补丁来缓解多模态大语言模型中的幻觉问题。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24027 2026-03-02 cs.CV cs.MM

GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models

GuardAlign: 多模态大语言模型中的测试时安全性对齐

Xingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang, Yin Zhang, Xiang Wang, Xiangnan He

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(摩埃关键实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) Tianjin University(天津大学)

AI总结 GuardAlign通过OT增强的安全检测和跨模态注意力校准,有效提升多模态大语言模型在测试时的安全性,减少不安全响应率并提升任务表现。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23981 2026-03-02 cs.LG cs.AI

Intrinsic Lorentz Neural Network

内禀洛伦兹神经网络

Xianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu Sebe

机构 * University of Science and Technology of China(中国科学技术大学) University of Trento(特伦托大学) Peking University(北京大学)

AI总结 内禀洛伦兹神经网络通过全内禀双曲架构提升几何决策性能,结合陀螺归一化和双曲距离计算,实现优于现有方法的性能和效率。

Comments Published in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23959 2026-03-02 cs.CV

Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought

通过图像作为连续动作进行思考:数值视觉链式推理

Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Science and Technology of China(中国科学技术大学) Chalmers University of Technology(楚克理工大学) University of Gothenburg(哥德堡大学)

AI总结 NV-CoT通过将图像推理动作空间扩展为连续欧几里得空间,提升MLLMs的定位精度和回答准确性,同时加速训练收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23699 2026-03-02 cs.CV cs.CL

HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit

HiDrop:通过晚期注入、凹形金字塔剪枝和早期退出实现MLLM中的层次视觉令牌减少

Hao Wu, Yingqi Fan, Jinyang Dai, Junlong Tong, Yunpu Ma, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Munich Center for Machine Learning, LMU Munich(慕尼黑大学机器学习中心,慕尼黑大学)

AI总结 HiDrop通过晚期注入、凹形金字塔剪枝和早期退出机制,实现多模态大语言模型中视觉令牌的高效减少,提升训练效率并保持性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23676 2026-03-02 cs.CV

Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering

通过语义解耦潜在引导抑制放射科报告生成中的先验比较幻觉

Ao Li, Rui Liu, Mingjie Li, Sheng Liu, Lei Wang, Xiaodan Liang, Lina Yao, Xiaojun Chang, Lei Xing

机构 * University of New South Wales(新南威尔士大学) Australian Artificial Intelligence Institute, University of Technology Sydney(澳大利亚人工智能研究所,技术悉尼大学) Stanford University(斯坦福大学) School of Computing and Information Technology of University of Wollongong Australia(沃林根澳大利亚大学计算与信息科技学院) Sun Yat-sen University(中山大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出语义解耦潜在引导方法,通过正交化技术减少放射科报告生成中的历史幻觉,提升临床准确性与报告忠实度。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23648 2026-03-02 cs.RO

FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic Manipulation

FAVLA:一种力适应的快速-慢速VLA模型用于接触丰富的机械臂操作

Yao Li, Peiyuan Tang, Wuyang Zhang, Chengyang Zhu, Yifan Duan, Weikai Shi, Xiaodong Zhang, Zijiang Yang, Jianmin Ji, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Xi'an Jiaotong University(西安交通大学) Central South University(中南大学)

AI总结 FAVLA通过解耦慢感知规划与快速接触感知控制,提升接触丰富任务中的反应性和成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21644 2026-03-02 cs.RO

DAGS-SLAM: Dynamic-Aware 3DGS SLAM via Spatiotemporal Motion Probability and Uncertainty-Aware Scheduling

DAGS-SLAM:通过时空运动概率和不确定性感知调度实现动态感知的3DGS SLAM

Li Zhang, Yu-An Liu, Xijia Jiang, Conghao Huang, Danyang Li, Yanyong Zhang

机构 * School of Mathematics, Hefei University of Technology(合肥工业大学数学学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)

AI总结 DAGS-SLAM通过时空运动概率和不确定性感知调度实现动态感知的3DGS SLAM,提升实时定位与密集重建的鲁棒性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24038 2026-03-02 cs.CV cs.MA

Enhancing CLIP Robustness via Cross-Modality Alignment

通过跨模态对齐增强CLIP鲁棒性

Xingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao, Hanwang Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学)

AI总结 COLA通过跨模态对齐提升CLIP对抗鲁棒性,有效缓解对抗扰动导致的特征不一致问题,提升零样本分类性能。

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22353 2026-03-02 cs.LG cs.AI

Context and Diversity Matter: The Emergence of In-Context Learning in World Models

上下文与多样性至关重要:世界模型中情境学习的出现

Fan Wang, Zhiyuan Chen, Yuxuan Zhong, Sunjian Zheng, Pengtao Shao, Bo Yu, Shaoshan Liu, Jianan Wang, Ning Ding, Yang Cao, Yu Kang

机构 * Shenzhen Institute of Artificial Intelligence and Robotics for Society(深圳人工智能与机器人社会研究院) University of Science and Technology of China(中国科学技术大学) Anhui Province Key Laboratory of Intelligent Low-Carbon Information Technology and Equipment(安徽省智能低碳信息技术与设备重点实验室)

AI总结 本文研究了世界模型中情境学习的机制,揭示了环境识别和学习的核心作用,并探讨了长上下文和多样化环境对学习效果的影响。

Journal ref 2026 International Conference on Learning Representations (ICLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07687 2026-03-02 eess.IV cs.CV

FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis

FermatSyn: 基于改进双向Mamba的多模态医学图像合成方法

Feng Yuan

机构 * USTC(中国科学技术大学) SII(上海信息研究所)

AI总结 FermatSyn通过改进的双向Mamba结合Fermat螺旋扫描策略,解决多模态医学图像合成中全局一致性与局部细节的平衡问题,提升合成图像质量与临床应用价值。

Comments MICCAI 2026(under view)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01728 2026-03-02 cs.CV

Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion

Shuffle Mamba:基于随机洗牌的态空间模型用于多模态图像融合

Ke Cao, Xuanhua He, Tao Hu, Chengjun Xie, Man Zhou, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Intelligent Machines(智能机器研究所) Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院) Intelligent Agriculture Engineering Laboratory of Anhui Province, Institute of Intelligent Machines(安徽省智能农业工程实验室,智能机器研究所)

AI总结 Shuffle Mamba通过引入随机洗牌策略和逆洗牌,解决多模态图像融合中固定扫描策略带来的偏见问题,提升融合质量。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00320 2026-03-02 cs.LG

TimeMAE: Self-Supervised Representations of Time Series with Decoupled Masked Autoencoders

TimeMAE:解耦掩码自编码器的时序自监督表示

Mingyue Cheng, Xiaoyu Tao, Zhiding Liu, Qi Liu, Hao Zhang, Rujiao Zhang, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

AI总结 TimeMAE通过解耦掩码自编码器和语义单元提升,提升时间序列自监督表示的性能,尤其在数据稀缺和迁移学习场景中表现优异。

Comments Accepted by WSDM'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23085 2026-02-27 quant-ph cs.LG

Q-Tag: Watermarking Quantum Circuit Generative Models

Q-Tag: 量子电路生成模型的水印技术

Yang Yang, Yuzhu Long, Han Fang, Zhaoyun Chen, Zhonghui Li, Weiming Zhang, Guoping Guo

机构 * School of Electronic and Information Engineering, Anhui University(安徽大学电子与信息工程学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥国家科学中心人工智能研究所) School of Computing, National University of Singapore(新加坡国立大学计算机学院) School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络科学与技术学院) Laboratory of Quantum Information, University of Science and Technology of China(中国科学技术大学量子信息实验室) Origin Quantum Computing Technology Company(起源量子计算技术有限公司)

AI总结 Q-Tag提出一种集成于量子电路生成模型的水印技术,通过生成过程嵌入所有权信号以保护电路版权,同时保持电路保真度。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22716 2026-02-27 cs.CV cs.AI

SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs

SoPE: 基于球坐标的位置嵌入:增强3D大视觉-语言模型的空间感知

Guanting Ye, Qiyan Zhao, Wenhao Yu, Liangyu Yuan, Mingkai Li, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Qing Jiang, Ka-Veng Yuen

机构 * University of Macau(澳门大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiaotong University(上海交通大学) Hefei University of Technology(合肥工业大学) National University of Singapore(新加坡国立大学)

AI总结 SoPE通过基于球坐标的位置嵌入提升3D LVLMs的空间感知能力,结合多尺度频率混合策略,增强几何表示的一致性和表达性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21743 2026-02-27 cs.CV

Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization

通过难度感知分组归一化增强多模态大语言模型推理

Jinghan Li, Junfeng Fang, Jinda Lu, Yuan Wang, Xiaoyan Guo, Tianyu Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出难度感知分组归一化方法,通过感知复杂度和推理不确定性表征样本,提升多模态大语言模型的推理稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏