arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Nanyang Technological University(南洋理工大学)

共收录 2393
2602.01594 2026-02-03 cs.CV

UV-M3TL: A Unified and Versatile Multimodal Multi-Task Learning Framework for Assistive Driving Perception

UV-M3TL: 一种统一且多功能的多模态多任务学习框架用于辅助驾驶感知

Wenzhuo Liu, Qiannan Guo, Zhen Wang, Wenshuo Wang, Lei Yang, Yicheng Qiao, Lening Wang, Zhiwei Li, Chen Lv, Shanghang Zhang, Junqiang Xi, Huaping Liu

机构 * Energy and Transportation Domain, Beijing Institute of Technology(能源与交通领域,北京理工大学) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(智能技术与系统国家重点实验室和清华大学计算机科学与技术系) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学) School of Transportation Science and Engineering and the State Key Lab of Intelligent Transportation System, Beihang University(交通运输科学与工程学院和智能交通系统国家重点实验室,北京航空航天大学) Beijing University of Chemical Technology(北京化工大学)

AI总结 UV-M3TL通过双分支结构和自适应损失机制,实现多模态多任务学习,提升辅助驾驶感知的性能与多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01574 2026-02-03 cs.CV

SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models

SGHA-Attack:语义引导的层次对齐用于视觉-语言模型的可迁移定向攻击

Haobo Wang, Weiqi Luo, Xiaojun Jia, Xiaochun Cao

机构 * Guangdong Province Key Lab of Information Security Technology, and School of Computer Science and Engineering, Sun Yat-sen University(广东信息安全技术重点实验室,计算机科学与工程学院,中山大学) Nanyang Technological University(南洋理工大学) School of Cyber Science and Technology, Shenzhen Campus, Sun Yat-sen University(网络安全科学与技术学院,深圳校区,中山大学)

AI总结 SGHA-Attack通过语义引导的层次对齐框架,提升视觉-语言模型的定向转移攻击效果,增强跨异构模型的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01558 2026-02-03 cs.LG

How Implicit Bias Accumulates and Propagates in LLM Long-term Memory

隐性偏见如何在LLM长期记忆中积累和传播

Yiming Ma, Lixu Wang, Lionel Z. Wang, Hongkun Yang, Haoming Sun, Xin Xu, Jiaqi Wu, Bin Chen, Wei Dong

机构 * International Research Center for Artificial Intelligence, Harbin Institute of Technology (Shenzhen), Shenzhen, China Chongqing Research Institute of Harbin Institute of Technology, Chongqing, China College of Computing Data Science, Nanyang Technological University, Singapore Department of Management Marketing, The Hong Kong Polytechnic University, Hong Kong Haide College, Ocean University of China, Qingdao, China School of Computer Science, The University of Sheffield, Sheffield, UK Department of Automation, Tsinghua University, Beijing, China

AI总结 本文研究了LLM长期记忆中隐性偏见的积累与传播,提出动态内存标记方法以缓解偏见问题。

Comments Under review, and the first two authors contribute equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20377 2026-02-03 cs.RO eess.SP

RF-MatID: Dataset and Benchmark for Radio Frequency Material Identification

RF-MatID: 用于射频材料识别的数据集和基准

Xinyan Chen, Qinchun Li, Ruiqin Ma, Jiaqi Bai, Li Yi, Jianfei Yang

机构 * MARS Lab, Nanyang Technological University(南洋理工大学MARS实验室) Ibaraki University(岩手大学)

AI总结 RF-MatID是一个大规模射频数据集,用于细粒度材料识别,包含16个细粒度类别和5个超类,覆盖4至43.5 GHz的频率范围,通过多设置多协议基准推动算法发展和实际应用。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11891 2026-02-03 cs.CL cs.AI

Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents

Mobile-Bench-v2: 一种更现实和全面的基于VLM的移动代理基准测试

Weikai Xu, Zhizheng Jiang, Yuxuan Liu, Pengzhi Gao, Wei Liu, Jian Luan, Yuanchun Li, Yunxin Liu, Bin Wang, Bo An

机构 * Nanyang Technological University(南洋理工大学) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallagher人工智能学院) MiLM Plus, Xiaomi Inc.(小米公司 MiLM Plus) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究所)

AI总结 Mobile-Bench-v2通过引入基于槽位的指令生成方法,构建了一个更现实和全面的基准测试,以评估基于VLM的移动代理在处理噪声、多路径任务和主动交互方面的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01284 2026-02-03 cs.MM cs.CV cs.HC

Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection

看见、听见与认知:深度伪造视频检测中的多模态策略

Chen Chen, Dion Hoe-Lian Goh

机构 * Nanyang Technological University(南洋理工大学)

AI总结 研究探讨了多模态策略在深度伪造视频检测中的应用,通过分析人类识别过程中的线索组合,提出了提升媒体素养的指导方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01092 2026-02-03 cs.RO

Failure-Aware Bimanual Teleoperation via Conservative Value Guided Assistance

面向故障意识的双臂遥控操作:通过保守价值引导的辅助

Peng Zhou, Zhongxuan Li, Jinsong Wu, Jiaming Qi, Jun Hu, David Navarro-Alarcon, Jia Pan, Lihua Xie, Shiyao Zhang, Zeqing Zhang

机构 * School of Advanced Engineering, Great Bay University(先进工程学院,大湾大学) School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Department of Mechanical Engineering, The Hong Kong Polytechnic University(机械工程系,香港理工大学) College of Mechanical and Electrical Engineering, Northeast Forestry University(机械电子工程学院,东北林业大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电气与电子工程学院,南洋理工大学)

AI总结 本文提出一种基于保守价值学习的双臂遥控操作框架,通过保守成功分数和学习执行者提供辅助,提升任务成功率并降低操作员负担。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01085 2026-02-03 cs.RO

Estimating Force Interactions of Deformable Linear Objects from their Shapes

从形状估计变形线性物体的力相互作用

Qi Jing Chen, Shilin Shan, Timothy Bretl, Quang-Cuong Pham

机构 * Nanyang Technological University, School of Mechanical and Aerospace Engineering(南洋理工大学机械与航空航天工程学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Eureka Robotics

AI总结 本文提出了一种基于线缆形状信息的力估计方法,无需额外传感器即可准确识别变形线性物体的外部力相互作用。

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01040 2026-02-03 cs.RO

Learning Adaptive Cross-Embodiment Visuomotor Policy with Contrastive Prompt Orchestration

学习适应性跨主体视觉运动策略与对比提示协调

Yuhang Zhang, Chao Yan, Jiaxi Yu, Jiaping Xiao, Mir Feroskhan

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) College of Automation Engineering, Nanjing University of Aeronautics and Astronautics(南京航空航天大学自动化工程学院)

AI总结 CAPO通过对比提示学习和自适应提示协调,提升跨主体视觉运动策略的适应性和样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00707 2026-02-03 cs.AI

Self-Guard: Defending Large Reasoning Models via enhanced self-reflection

Self-Guard:通过增强的自我反思防御大型推理模型

Jingnan Zheng, Jingjun Xu, Yanzhen Luo, Chenhang Cui, Gelei Deng, Zhenkai Liang, Xiang Wang, An Zhang, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Southern University of Science and Technology(南方科技大学) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学)

AI总结 Self-Guard通过增强的自我反思机制,有效解决大型推理模型的安全合规问题,实现安全与实用性的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20042 2026-02-03 cs.CV cs.AI

Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval

超越视觉:基于多模态检索的上下文丰富图像描述

Nguyen Lam Phu Quy, Pham Phu Hoa, Tran Chi Nguyen, Dao Sy Duy Minh, Nguyen Hoang Minh Ngoc, Huynh Trung Kiet

机构 * University of Science - VNUHCM(越南胡志明市大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出一种多模态检索方法,通过整合外部文本知识生成更丰富的上下文图像描述,提升视觉-文本理解能力。

Comments 7 pages, 5 figures. System description for the EVENTA Grand Challenge (Track 1) at ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04048 2026-02-03 cs.CV cs.CL cs.CY

Stable Signer: Hierarchical Sign Language Generative Model

Stable Signer: 层级化手语生成模型

Sen Fang, Yalin Feng, Hongbin Zhong, Yanxin Zhang, Dimitris N. Metaxas

机构 * Rutgers University(罗格斯大学) Nanyang Technological University(南洋理工大学) Georgia Institute of Technology(佐治亚理工学院) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 本文提出Stable Signer模型,通过层级生成端到端任务提升手语视频生成质量,采用SLUL和SLP-MoE模块实现高效生成。

Comments 12 pages, 7 figures. More Demo at https://stablesigner.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22505 2026-02-03 cs.HC cs.AI cs.CL cs.CY stat.AP

Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational Theory

人工智能伴侣对心理健康的影响:通过社交媒体准实验、用户视角和关系理论进行三角验证

Yunhao Yuan, Jiaxun Zhang, Talayeh Aledavood, Renwen Zhang, Koustuv Saha

机构 * Aalto University(阿尔托大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Nanyang Technological University(南洋理工大学)

AI总结 本文通过社交媒体准实验、用户访谈和关系理论,研究人工智能伴侣对心理健康的影响,发现其在提供情感支持的同时也存在依赖风险,提出设计建议以平衡心理社会效益。

Journal ref Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08396 2026-02-03 cs.CV

CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation

CoDi: 一致主体和多姿态文本到图像生成

Zhanxin Gao, Beier Zhu, Liang Yao, Jian Yang, Ying Tai

机构 * Nanjing University(南京大学) Nanyang Technological University(南洋理工大学) Vipshop

AI总结 CoDi通过两阶段策略实现文本到图像生成中主体一致性和姿态多样性的平衡,提升视觉表现和性能。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18355 2026-02-03 cs.RO

Robotic Manipulation of a Rotating Chain with Bottom End Fixed

机器人操控固定底端的旋转链

Qi Jing Chen, Shilin Shan, Quang-Cuong Pham

机构 * Nanyang Technological University, School of Mechanical and Aerospace Engineering(南洋理工大学机械与航空航天工程学院) Eureka Robotics

AI总结 本文提出了一种稳定且一致的机器人操控策略,用于转换固定底端旋转链的形状,通过同胚性质实现不同旋转模式的转换,应用于钻井串和纱线纺纱操作的安全性和效率提升。

Comments 6 pages, 5 figures

Journal ref Proc. IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02524 2026-02-03 cs.RO

Versatile Behavior Diffusion for Generalized Traffic Agent Simulation

多功能行为扩散用于通用交通代理模拟

Zhiyu Huang, Zixu Zhang, Ameya Vaidya, Yuxiao Chen, Chen Lv, Jaime Fernández Fisac

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系) NVIDIA Research(NVIDIA研究)

AI总结 多功能行为扩散通过生成多代理互动场景,提升交通模拟的灵活性和实用性,为自动驾驶系统测试提供基础工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00911 2026-02-03 cs.RO

Accurate Simulation and Parameter Identification of Deformable Linear Objects using Discrete Elastic Rods in Generalized Coordinates

基于广义坐标离散弹性杆的可变形线性物体的精确模拟与参数识别

Qi Jing Chen, Timothy Bretl, Quang-Cuong Pham

机构 * Nanyang Technological University, School of Mechanical and Aerospace Engineering(南洋理工大学机械与航空航天工程学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Eureka Robotics

AI总结 本文提出了一种基于广义坐标的离散弹性杆模型,用于提高可变形线性物体在MuJoCo中的仿真精度和参数识别效率。

Comments 7 pages, 6 figures

Journal ref Proc. IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00205 2026-02-03 cs.LG cs.CV

Reducing Class-Wise Performance Disparity via Margin Regularization

通过边缘正则化减少类别间的性能差异

Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) Hefei University of Technology(合肥工业大学) Singapore Management University(新加坡管理学院) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出MR$^2$方法,通过动态调整边缘来减少分类中类别间的性能差异,提升困难类别性能的同时不损害易类别表现。

Comments To appear in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00154 2026-02-03 cs.CR cs.AI

ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models

ReasoningBomb: 通过诱导病态长推理实现隐蔽的拒绝服务攻击

Xiaogeng Liu, Xinyan Wang, Yechao Zhang, Sanjay Kariyappa, Chong Xiang, Muhao Chen, G. Edward Suh, Chaowei Xiao

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Nanyang Technological University(南洋理工大学) NVIDIA(NVIDIA公司) University of California, Davis(加州大学戴维斯分校) Cornell University(康奈尔大学)

AI总结 ReasoningBomb通过生成短自然提示诱导大型推理模型进入病态长推理,实现隐蔽的拒绝服务攻击,具有高放大率、隐蔽性和可优化性。

Comments Pre-print. Code is available at https://github.com/SaFo-Lab/ReasoningBomb

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00075 2026-02-03 cs.LG cs.MS

Dimensional Peeking for Low-Variance Gradients in Zeroth-Order Discrete Optimization via Simulation

通过模拟实现的维度窥视:用于零阶离散优化中低方差梯度的方法

Philipp Andelfinger, Wentong Cai

机构 * Nanyang Technological University(南洋理工大学)

AI总结 本文提出一种通过模拟实现的维度窥视方法,用于减少零阶离散优化中的梯度方差,提升优化性能。

Comments Accepted at ACM SIGSIM PADS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00016 2026-02-03 cs.CL cs.AI

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

PTCBENCH: 对LLM系统中人格特质情境稳定性的基准测试

Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang, Junjie Wang, Qiang Hu, Chao Shen, Xiaofei Xie

机构 * Singapore Management University(新加坡管理大学) Tianjin University(天津大学) Nanyang Technological University(南洋理工大学) Xi’an Jiaotong University(西安交通大学)

AI总结 PTCBENCH通过评估LLM在不同情境下的人格一致性,揭示外部事件对LLM人格和推理能力的影响,为构建更稳健的心理学对齐AI系统提供新视角。

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07666 2026-02-03 cs.CV cs.AI

iPEAR: Iterative Pyramid Estimation with Attention and Residuals for Deformable Medical Image Registration

iPEAR: 基于注意力和残差的迭代金字塔估计用于形变医学图像配准

Heming Wu, Di Wang, Tai Ma, Peng Zhao, Yubin Xiao, Zhongke Wu, Xing-Ce Wang, Xuan Wu, You Zhou

机构 * College of Software, Jilin University(吉林大学软件学院) Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly, Nanyang Technological University(老年活跃生活卓越研究中⼼(NTU-UBC), 新加坡国立大学) College of Computer Science(计算机科学学院) Technology, Zhejiang University(技术, 浙江大学) Key Laboratory of Symbolic Computation(符号计算重点实验室) Knowledge Engineering of Ministry of Education, College of Computer Science(教育部知识工程重点实验室, 计算机科学学院) Technology, Jilin University(技术, 吉林大学) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

AI总结 iPEAR通过融合注意力与残差模块和双阶段阈值控制迭代策略,提升医学图像配准的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23081 2026-02-02 cs.CL cs.AI cs.CR

Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures

字符作为大语言模型中的潜在变量:对涌现偏差和条件安全失败的机制解释

Yanghao Su, Wenbo Zhou, Tianwei Zhang, Qiu Han, Weiming Zhang, Nenghai Yu, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学)

AI总结 本研究揭示了大语言模型中字符层面倾向对齐风险的核心机制,指出行为倾向的稳定转变是导致涌现偏差和安全失败的关键因素,而非单纯的能力退化或提示防御。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22790 2026-02-02 cs.AI math.ST stat.TH

Conditional Performance Guarantee for Large Reasoning Models

基于条件的大型推理模型性能保证

Jianguo Huang, Hao Zeng, Bingyi Jing, Hongxin Wei, Bo An

机构 * Nanyang Technological University, Singapore(南洋理工大学) Southern University of Science(南方科技大学) The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))

AI总结 本文提出G-PAC和C-PAC推理框架,通过分组策略实现组条件风险控制,提升异质环境下推理效率并降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22738 2026-02-02 cs.CV

StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing

StreamSense: 基于选择性视觉-语言模型路由的流式社交任务检测

Han Wang, Deyi Ji, Lanyun Zhu, Jiebo Luo, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) University of Rochester(罗切斯特大学)

AI总结 StreamSense通过轻量级编码器与选择性路由结合VLM专家,提升流式社交任务检测的准确性和效率。

Comments 10 pages, 4 figures, The Web Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09133 2026-02-02 cs.AI cs.LG math.ST stat.TH

On the Provable Performance Guarantee of Efficient Reasoning Models

关于高效推理模型的可证明性能保证

Hao Zeng, Jianguo Huang, Bingyi Jing, Hongxin Wei, Bo An

机构 * Nanyang Technological University, Singapore(新加坡南洋理工大学) Southern University of Science(南方科技大学) The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))

AI总结 本文提出PAC推理方法,通过动态切换思考与非思考模式,在用户指定容忍度下控制性能损失,实现高效推理与性能保障。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04347 2026-02-02 cs.CL cs.LG

Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models

揭示后门:通过梯度-注意力异常评分对预训练语言模型进行可解释防御

Anindya Sundar Das, Kangjie Chen, Monowar Bhuyan

机构 * Umeå University(乌梅学院) Nanyang Technological University(南洋理工大学)

AI总结 本文提出通过梯度-注意力异常评分对预训练语言模型进行可解释防御,有效降低后门攻击成功率。

Comments 17 pages total (9 pages main text + 6 pages appendix + references), 16 figures. Preprint version; the final camera-ready version may differ. Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04825 2026-02-02 cs.LG cs.AI

Graph Mining under Data scarcity

图数据稀少下的图挖掘

Appan Rakaraddi, Lam Siew-Kei, Mahardhika Pratama, Marcus de Carvalho

机构 * Nanyang Technological University Singapore(南洋理工大学新加坡) University of South Australia Australia(澳大利亚南澳大利亚大学)

AI总结 本文提出一种适用于通用GNN框架的不确定性估计器,通过端到端训练提升少样本图节点分类性能,无需元学习架构。

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12683 2026-02-02 cs.LG cs.CV math.ST stat.TH

TorchCP: A Python Library for Conformal Prediction

TorchCP: 一个用于置信预测的 Python 库

Jianguo Huang, Jianqing Song, Xuanning Zhou, Bingyi Jing, Hongxin Wei

机构 * Southern University of Science and Technology(南方科技大学) Nanyang Technological University(南洋理工大学) Nanjing University(南京大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 TorchCP 是一个基于 PyTorch 的库,旨在将先进的置信预测算法集成到深度学习技术中,提升大规模数据中的不确定性量化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22718 2026-02-02 cs.AI cs.LG

A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization

后退一步:前缀重要性比率稳定了策略优化

Shiye Lei, Zhihao Cheng, Dacheng Tao

机构 * The University of Sydney(悉尼大学) ByteDance(字节跳动) Nanyang Technological University(南洋理工大学)

AI总结 本文提出MinPRO目标,通过使用非累积的最小token级比率替代累积前缀比率,以稳定LLM在非策略性条件下的训练和优化。

详情

展开后加载摘要…

URL PDF HTML 收藏