arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

共收录 1842
2601.16987 2026-01-27 cs.CL cs.AI

Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions

通过成对最大分歧竞赛评估奖励模型泛化能力

Shunyang Luo, Peibei Cao, Zhihui Zhu, Kehua Feng, Zhihua Wang, Keyan Ding

机构 * ZJU-UIUC Institute, Zhejiang University(浙江大学ZJU-UIUC研究院) School of Artificial Intelligence, Nanjing University of Information Science and Technology(南京信息工程大学人工智能学院) ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University(浙江大学Hangzhou全球科技创新中心) City University of Hong Kong(香港城市大学)

AI总结 本文提出PMDC方法,通过动态选择高争议测试案例评估奖励模型的泛化能力,并揭示了现有模型的系统性泛化失败问题。

Comments 17 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01078 2026-01-26 cs.AI

SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds

SimWorld:一种用于物理和社会世界中自主代理的开放式真实模拟器

Jiawei Ren, Yan Zhuang, Xiaokang Ye, Lingjun Mao, Xuhong He, Jianzhi Shen, Mrinaal Dogra, Yiming Liang, Ruixuan Zhang, Tianai Yue, Yiqing Yang, Eric Liu, Ryan Wu, Kevin Benavente, Rajiv Mandya Nagaraju, Muhammad Faayez, Xiyan Zhang, Dhruv Vivek Sharma, Xianrui Zhong, Ziqiao Ma, Tianmin Shu, Zhiting Hu, Lianhui Qin

机构 * UCSD(加州大学圣地亚哥分校) UVA(弗吉尼亚大学) UIUC(伊利诺伊大学香槟分校) JHU(约翰·霍普金斯大学) Purdue(Purdue 大学) PolyU USC(美国南加州大学) UMich(密歇根大学)

AI总结 SimWorld是一个基于Unreal Engine 5构建的开放式真实模拟器,旨在开发和评估LLM/VLM代理在复杂物理和社会环境中的能力,通过多代理配送任务验证其推理模式与局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12150 2026-01-26 cs.RO cs.AI cs.LG

HEIGHT: Heterogeneous Interaction Graph Transformer for Robot Navigation in Crowded and Constrained Environments

HEIGHT:用于拥挤和受限环境中机器人导航的异构交互图变压器

Shuijing Liu, Haochen Xia, Fatemeh Cheraghi Pouria, Kaiwen Hong, Neeloy Chakraborty, Zichao Hu, Joydeep Biswas, Katherine Driggs-Campbell

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 HEIGHT通过异构交互图变压器提升机器人在拥挤和受限环境中的导航性能,有效捕捉时空异构交互,提高路径安全性和效率。

Comments Accepted to IEEE Transactions of Automation Science and Engineering (T-ASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09693 2026-01-23 cs.CV

CropCraft: Complete Structural Characterization of Crop Plants From Images

CropCraft: 作物植物的完整结构表征从图像

Albert J. Zhai, Xinlei Wang, Kaiyuan Li, Zhao Jiang, Junxiong Zhou, Sheng Wang, Zhenong Jin, Kaiyu Guan, Shenlong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Minnesota Twin Cities(明尼苏达大学双城分校)

AI总结 CropCraft通过逆向程序建模优化植物形态参数,实现从图像中完整重建作物3D结构,用于农业监测与模拟应用。

Comments 3DV 2026 (Oral). Project page: https://ajzhai.github.io/CropCraft

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15737 2026-01-23 cs.AI cs.CL

PhysProver: Advancing Automatic Theorem Proving for Physics

PhysProver: 推动物理领域自动定理证明的发展

Hanning Zhang, Ruida Wang, Rui Pan, Wenyuan Wang, Bingxu Meng, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rutgers University(罗格斯大学)

AI总结 PhysProver通过结合可验证语言和强化学习,提升物理领域形式化定理证明的效率和效果。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15560 2026-01-23 cs.CV

Relative Classification Accuracy: A Calibrated Metric for Identity Consistency in Fine-Grained K-pop Face Generation

相对分类准确率:一种用于细粒度K-pop人脸生成中身份一致性的校准度量

Sylvey Lin, Eranki Vasistha

机构 * UIUC(伊利诺伊大学香槟分校)

AI总结 本文提出相对分类准确率(RCA)作为评估细粒度K-pop人脸生成中身份一致性的校准度量,揭示了高视觉质量与严重语义模式崩溃之间的关键权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15412 2026-01-23 cs.HC cs.AI cs.CY

A Checklist for Trustworthy, Safe, and User-Friendly Mental Health Chatbots

可信、安全且用户友好的心理健康聊天机器人清单

Shreya Haran, Samiha Thatikonda, Dong Whi Yoo, Koustuv Saha

机构 * University of Illinois Urbana-Champaign, Urbana, United States(伊利诺伊大学厄巴纳-香槟分校) Indiana University Indianapolis, Indianapolis, United States(印第安纳大学印第安纳波利斯分校)

AI总结 本文提出了一套操作清单,旨在指导开发更可信、安全且用户友好的心理健康聊天机器人,以促进伦理和有效的设计实践。

Journal ref In 28th International Conference on Human-Computer Interaction, Springer LNCS, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15330 2026-01-23 cs.CL cs.AI

ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation

ICPO:用于多轮对话的意涵校准策略优化

Zhebo Wang, Xiaohu Mu, Zijie Zhou, Mohan Li, Wenpeng Xing, Dezhang Kong, Meng Han

机构 * Zhejiang University(浙江大学) Binjiang Institute of Zhejiang University(浙江大学滨江学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) China University of Petroleum (Beijing)(中国石油大学(北京)) Guangzhou University(广州大学)

AI总结 ICPO通过校准模型对指令模糊性的感知,提升多轮对话的鲁棒性和协作性,实现75%的性能提升。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14691 2026-01-23 cs.AI cs.CL

Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation

操纵法官:不忠的推理链可能损害智能体评估

Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Sungryull Sohn, Yunxiang Zhang, Moontae Lee, Hao Peng, Lu Wang, Honglak Lee

机构 * University of Michigan(密歇根大学) LG AI Research(LG人工智能研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文揭示了LLM法官对智能体推理轨迹操纵的脆弱性,表明基于内容的操纵能显著提高假阳性率,凸显了需验证推理与证据的评估机制的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01747 2026-01-23 cs.CR cs.AI cs.CV cs.LG

Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization

为大型视觉-语言模型设计对抗输入使用黑盒优化

Jiwei Guan, Haibo Jin, Haohan Wang

机构 * School of Computing, Macquarie University(麦考瑞大学计算机学院) School of Information Sciences, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校信息科学学院)

AI总结 本文提出了一种基于零阶优化的黑盒劫持攻击方法,针对大型视觉-语言模型实现高成功率的对抗性攻击,揭示了当前模型安全机制的不足。

Comments EACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24257 2026-01-23 cs.CR cs.LG

VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference

VeriLLM:一种轻量级的公开可验证去中心化推理框架

Ke Wang, Zishuo Zhao, Xinyuan Song, Zelin Li, Libin Xia, Chris Tong, Bill Shi, Wenjie Qu, Eric Yang, Lynn Ai

机构 * University of Illinois Urbana-Champaign, Gradient(伊利诺伊大学厄巴纳-香槟分校,Gradient) Emory University(埃默里大学) Ohio State University(俄亥俄州立大学) Peking University(北京大学) National University of Singapore(新加坡国立大学)

AI总结 VeriLLM通过轻量级经验重跑和链上检查实现公开可验证的去中心化LLM推理,提高效率与安全性,减少验证成本。

Comments 18 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22327 2026-01-23 cs.CL cs.CY

NLP for Social Good: A Survey and Outlook of Challenges, Opportunities, and Responsible Deployment

为社会公益服务的NLP:挑战、机遇与负责任部署的综述与展望

Antonia Karamolegkou, Angana Borah, Eunjung Cho, Sagnik Ray Choudhury, Martina Galletti, Pranav Gupta, Oana Ignat, Priyanka Kargupta, Neema Kotonya, Hemank Lamba, Sun-Joo Lee, Arushi Mangla, Ishani Mondal, Fatima Zahra Moudakir, Deniz Nazarova, Poli Nemkova, Dina Pisarevskaya, Naquee Rizwan, Nazanin Sabri, Keenan Samway, Dominik Stammbach, Anna Steinberg, David Tomás, Steven R Wilson, Bowen Yi, Jessica H Zhu, Arkaitz Zubiaga, Anders Søgaard, Alexander Fraser, Zhijing Jin, Rada Mihalcea, Joel R. Tetreault, Daryna Dementieva

机构 * University of Copenhagen(哥本哈根大学) University of Michigan-Ann Arbor(密歇根大学安娜堡分校) ETH Zurich(苏黎世联邦理工学院) University of North Texas(北卡罗来纳州立大学) Sony Computer Science Laboratories - Paris(索尼计算机科学实验室-巴黎) Santa Clara University(圣克拉拉大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Dataminr(DataMinr公司) United Nations Development Programme (UNDP)(联合国开发计划署) University of Maryland, College Park(马里兰大学学院市分校) Max Planck Institute for Intelligent Systems, Tübingen(智能系统马克斯·普朗克研究所,图宾根) Vector Institute(向量研究所) University of Toronto(多伦多大学) University of Washington(华盛顿大学) Queen Mary University of London(伦敦大学玛丽女王学院) IIT Kharagpur(印度理工学院Kharagpur分校) University of California San Diego(加州大学圣地亚哥分校) Princeton University(普林斯顿大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) University of Alicante(阿利坎特大学) University of Michigan-Flint(密歇根大学弗林特分校) University of Southern California(南加州大学) Technical University of Munich(慕尼黑技术大学)

AI总结 本文综述了NLP在社会公益领域的应用现状,指出包容性和AI危害是研究热点,同时呼吁跨学科合作以促进公众福祉。

Comments Accepted to EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14633 2026-01-22 cs.LG

Relational Graph Modeling for Credit Default Prediction: Heterogeneous GNNs and Hybrid Ensemble Learning

关系图建模用于信用违约预测:异构GNN与混合集成学习

Yvonne Yang, Eranki Vasistha

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文提出异构图神经网络与混合集成学习方法,通过构建大规模异构图整合借款人与交易数据,提升信用违约预测的准确性和公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14589 2026-01-22 cs.HC cs.AI cs.CL cs.CY

Designing KRIYA: An AI Companion for Wellbeing Self-Reflection

设计KRIYA:一种用于幸福感自我反思的AI伴侣

Shanshan Zhu, Wenxuan Song, Jiayue Melissa Shi, Dong Whi Yoo, Karthik S. Bhat, Koustuv Saha

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校) Drexel University(德雷塞尔大学)

AI总结 KRIYA是一种通过自我反思功能帮助用户理解个人幸福感数据的AI伴侣,旨在减少表现焦虑,提升用户对健康数据的反思性理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14266 2026-01-22 cs.LG cs.CL cs.CR

GCG Attack On A Diffusion LLM

对扩散语言模型的GCG攻击

Ruben Neyroud, Sam Corley

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文研究了GCG攻击对扩散语言模型LLaDA的影响,评估了多种攻击变体并探讨了模型的鲁棒性及对抗分析策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14235 2026-01-22 astro-ph.IM astro-ph.CO cs.AI cs.LG stat.ML

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

人工智能/机器学习在Rubin LSST暗能量科学合作中的机遇

LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz, Matthew R. Becker, Biswajit Biswas, Rahul Biswas, Boris Bolliet, Adam S. Bolton, Clecio R. Bom, Raphaël Bonnet-Guerrini, Alexandre Boucaud, Jean-Eric Campagne, Chihway Chang, Aleksandra Ćiprijanović, Johann Cohen-Tanugi, Michael W. Coughlin, John Franklin Crenshaw, Juan C. Cuevas-Tello, Juan de Vicente, Seth W. Digel, Steven Dillmann, Mariano Javier de León Dominguez Romero, Alex Drlica-Wagner, Sydney Erickson, Alexander T. Gagliano, Christos Georgiou, Aritra Ghosh, Matthew Grayling, Kirill A. Grishin, Alan Heavens, Lindsay R. House, Mustapha Ishak, Wassim Kabalan, Arun Kannawadi, François Lanusse, C. Danielle Leonard, Pierre-François Léget, Michelle Lochner, Yao-Yuan Mao, Peter Melchior, Grant Merz, Martin Millon, Anais Möller, Gautham Narayan, Yuuki Omori, Hiranya Peiris, Laurence Perreault-Levasseur, Andrés A. Plazas Malagón, Nesar Ramachandra, Benjamin Remy, Cécile Roucelle, Jaime Ruiz-Zapatero, Stefan Schuldt, Ignacio Sevilla-Noarbe, Ved G. Shah, Tjitske Starkenburg, Stephen Thorp, Laura Toribio San Cipriano, Tilman Tröster, Roberto Trotta, Padma Venkatraman, Amanda Wasserman, Tim White, Justine Zeghal, Tianqing Zhang, Yuanyuan Zhang

机构 * Université Paris Cité, CNRS, CEA, Astroparticule et Cosmologie, F-75013 Paris, France Department of Physics, University of Michigan, Ann Arbor, MI 48109, USA Leinweber Institute of Theoretical Physics, University of Michigan, Ann Arbor, MI 48109, USA Argonne National Laboratory, 9700 South Cass Avenue, Lemont, IL 60439, USA Cavendish Astrophysics, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK SLAC National Accelerator Laboratory, Menlo Park, CA 94025, USA Department of Computer Science, University of Milan, Milan, Italy Université Paris Cité, CNRS, Astroparticule et Cosmologie, F-75013 Paris, France Université Paris-Saclay, CNRS/IN2P3, IJCLab, 91405 Orsay, France Department of Astronomy Astrophysics, University of Chicago, Chicago, IL 60637, USA Kavli Institute for Cosmological Physics, University of Chicago, Chicago, IL 60637, USA NSF-Simons AI Institute for the Sky (SkAI), 172 E. Chestnut St., Chicago, IL 60611, USA Fermi National Accelerator Laboratory, P.O. Box 500, Batavia, IL 60510, USA Universit\'e Clermont-Auvergne, CNRS, LPCA, 63000 Clermont-Ferrand, France Kavli Institute for Particle Astrophysics Cosmology, Stanford University, Stanford, CA 94305, USA Department of Physics, Stanford University, 382 Via Pueblo Mall, Stanford, CA 94305, USA Engineering Faculty, Universidad Autonoma de San Luis Potosi, Zona Universitaria, San Luis Potosi, 78290, Mexico Stanford Artificial Intelligence Laboratory, Stanford University, Stanford, CA 94305, USA Kavli Institute of Cosmological Physics, University of Chicago, Chicago, IL 60637, USA The NSF AI Institute for Artificial Intelligence Center for Astrophysics Harvard \& Smithsonian, 60 Garden Street, Cambridge, MA 02138, USA Department of Physics Kavli Institute for Astrophysics Space Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Institut de Física d'Altes Energies (IFAE), The Barcelona Institute of Science Institute of Astronomy Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge, CB3 0HA, UK Imperial Centre for Inference Cosmology (ICIC), Imperial College London, Blackett Laboratory, Prince Consort Road, London SW7 2AZ, UK Data Science Institute, The University of Chicago, Chicago, IL 60615, USA Department of Physics, The University of Texas at Dallas, Richardson, TX 75080, USA Department of Physics, Duke University, Durham, NC 27708, USA Université Paris-Saclay, Université Paris Cité, CEA, CNRS, AIM, F-91191 Gif-sur-Yvette, France School of Mathematics, Statistics Physics, Newcastle University, Newcastle upon Tyne, NE1 7RU, United Kingdom Department of Astrophysical Sciences, Princeton University, Princeton, NJ 08544, USA Astronomy, University of the Western Cape, Bellville, Cape Town, 7535, South Africa Astronomy, University of Utah, Salt Lake City, UT 84112, USA Department of Astrophysical Sciences, Princeton University, Peyton Hall, Princeton, NJ 08544, USA Department of Astronomy, University of Illinois Urbana Champaign, 1002 W. Green St., Urbana, IL, 61801, USA Institute for Particle Physics Astrophysics, ETH Zürich, Wolfgang-Pauli-Strasse 27, CH-8093 Zurich, Switzerland Swinburne University of Technology, Hawthorn, Victoria 3122, Australia Ciela - Montr\'eal Institute for Astrophysical Data Analysis Mila - Quebec Artificial Intelligence Institute, Montréal, QC H2S 3H1, Canada Advanced Research Computing Centre, University College London, 90 High Holborn, London WC1V 6LJ, UK Finnish Centre for Astronomy with ESO (FINCA), University of Turku, FI-20014 Turku, Finland Department of Physics, P.O. Box 64, University of Helsinki, FI-00014 Helsinki, Finland Astronomy, Northwestern University, Evanston, IL, USA Center for Interdisciplinary Exploration Research in Astrophysics, Northwestern University, Evanston, IL, USA Scientific Data Science, International School for Advanced Study, Via Bonomea 265, I-34136 Trieste, Italy Department of Statistics, University of Michigan, Ann Arbor, MI 48109, USA PITT PACC, University of Pittsburgh, Pittsburgh, PA 15260, USA NSF NOIRLab, 950 N. Cherry Ave., Tucson, AZ 85719, USA

AI总结 本文探讨了AI/ML在LSST暗能量科学合作中的应用机遇,强调了大规模贝叶斯推断、物理指导方法和主动学习等关键方法学优先事项,并讨论了新兴技术在重塑工作流程中的潜力。

Comments 84 pages. This is v1.0 of the DESC's white paper on AI/ML, a collaboration document that is being made public but which is not planned for submission to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14209 2026-01-21 cs.LG cs.AI cs.CL

InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning

InT:自我提出干预使LLM推理中的信用分配成为可能

Matthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang, Amrith Setlur, Aviral Kumar

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 InT通过自我提出干预实现LLM推理中的细粒度信用分配,提升模型在数学推理任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08778 2026-01-21 cs.AI cs.DB

Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards

广泛注释错误破坏文本到SQL基准测试和排行榜

Tengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel Kang

机构 * University of Illinois (UIUC)(伊利诺伊大学香槟分校)

AI总结 本研究发现文本到SQL基准测试中广泛存在的注释错误显著影响了代理性能和排行榜排名,可能误导研究方向和部署选择。

Comments 18 pages, 14 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03515 2026-01-21 cs.RO cs.AI cs.LG cs.SY eess.SY stat.AP

Can the Waymo Open Motion Dataset Support Realistic Behavioral Modeling? A Validation Study with Naturalistic Trajectories

Waymo开放运动数据集能否支持真实的行为建模?一项与自然轨迹相结合的验证研究

Yanlin Zhang, Sungyong Chung, Nachuan Li, Dana Monzer, Hani S. Mahmassani, Samer H. Hamdar, Alireza Talebpour

机构 * Department of Civil and Environmental Engineering, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校土木与环境工程系) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Northwestern University Transportation Center(西北大学交通中心) George Washington University(乔治·华盛顿大学)

AI总结 本研究通过对比自然主义数据与Waymo数据集,发现其无法准确反映真实自动驾驶行为,需谨慎使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13183 2026-01-21 cs.CL

OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand

OpenExempt:法律推理的诊断基准及自定义基准框架

Sergio Servantez, Sarah B. Lawsky, Rajiv Jain, Daniel W. Linna, Kristian Hammond

机构 * Northwestern University(西北大学) Adobe Research(Adobe研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 OpenExempt通过动态生成法律推理任务和解决方案,提供一个用于诊断评估的基准和框架,揭示模型在复杂推理中的性能差异。

Comments 25 pages, 9 Figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12758 2026-01-21 cs.CL cs.AI cs.LG

VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

VISPA:通过自动价值选择和激活实现多元对齐

Shenyan Zheng, Jiayou Zhong, Anudeex Shetty, Heng Ji, Preslav Nakov, Usman Naseem

机构 * University of Waterloo(滑铁卢大学) University of Melbourne(墨尔本大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) MBZUAI(马克斯·普朗克智能系统研究所) Macquarie University(麦考瑞大学)

AI总结 VISPA通过自动价值选择和激活实现多元对齐,适用于多种模型和场景,提供了一种可扩展的语言模型对齐方法。

Comments WIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13855 2026-01-21 cs.CL cs.AI

Harnessing Consistency for Robust Test-Time LLM Ensemble

利用一致性提升鲁棒性测试时LLM集成

Zhichen Zeng, Qi Yu, Xiao Lin, Ruizhong Qiu, Xuying Ning, Tianxin Wei, Yuchen Yan, Jingrui He, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 CoRE通过利用模型一致性提升LLM集成的鲁棒性,通过token和model级别的一致性改进集成性能。

Comments 18 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02106 2026-01-21 physics.flu-dyn cs.AI cs.LG gr-qc physics.comp-ph

Resolving Turbulent Magnetohydrodynamics: A Hybrid Operator-Diffusion Framework

解析湍流磁流体动力学:一种混合运算-扩散框架

Semih Kacmaz, E. A. Huerta, Roland Haas

机构 * National Center for Supercomputing Applications, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校国家超级计算中心) Department of Physics, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校物理系) Data Science and Learning Division, Argonne National Laboratory(阿贡国家实验室数据科学与学习部门) Department of Computer Science, The University of Chicago(芝加哥大学计算机科学系) Department of Astronomy, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校天文学系) Department of Physics and Astronomy, University of British Columbia(不列颠哥伦比亚大学物理与天文学系)

AI总结 该研究提出混合运算-扩散框架,结合PINOs与生成扩散模型,实现对高雷诺数MHD湍流的高精度模拟与预测。

Comments 16 pages, 6 figures, 1 table. Content synced with the published version

Journal ref Mach. Learn.: Sci. Technol. 6 (2025) 035057

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03833 2026-01-21 gr-qc astro-ph.IM cs.AI

Sequence modeling of higher-order wave modes of binary black hole mergers

二体黑洞并合高阶波模式的序列建模

Victoria Tiki, Kiet Pham, Eliu Huerta

机构 * Learning Division, Argonne National Laboratory(Argonne国家实验室学习部) Department of Physics, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校物理系) NCSA, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校NCSA) School of Physics and Astronomy, University of Minnesota(明尼苏达大学物理与天文学学院) Department of Computer Science, The University of Chicago(芝加哥大学计算机科学系) Department of Astronomy, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校天文学系)

AI总结 本文提出基于transformer的模型,用于高精度建模二体黑洞并合的高阶引力波模式,实现非线性动力学的快速准确预测。

Comments 32 pages, 2 appendices, 17 figures

Journal ref Class. Quantum Grav. 43 (2026) 015009

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12208 2026-01-21 cs.CL

CoReflect: Conversational Evaluation via Co-Evolutionary Simulation and Reflective Rubric Refinement

CoReflect:通过共进化模拟与反思性评分表细化进行对话评估

Yunzhe Li, Richie Yueqi Feng, Tianxin Wei, Chin-Chia Hsu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 CoReflect通过共进化模拟与反思性评分表细化,实现对话系统的自适应评估,提升测试用例复杂性和评分表精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11560 2026-01-21 cs.IR cs.AI cs.LG

DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research

DeepEvidence: 通过深度知识图谱研究赋能生物医学发现

Zifeng Wang, Zheng Chen, Ziwei Yang, Xuan Wang, Qiao Jin, Yifan Peng, Zhiyong Lu, Jimeng Sun

机构 * Keiji AI Institute of Scientific and Industrial Research, Osaka University(大阪大学科学工业研究所) Bioinformatics Center, Institute for Chemical Research, Kyoto University(京都大学化学研究所生物信息中心) Division of Intramural Research, National Library of Medicine, National Institutes of Health(国家卫生研究院生物医学图书馆内部研究部) Department of Population Health Sciences, Weill Cornell Medicine(韦尔·科恩医学中心流行病学与健康科学系) School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机与数据科学学院)

AI总结 DeepEvidence通过深度知识图谱研究框架,系统化地连接异构生物医学资源,提升科学发现的效率和证据综合能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11559 2026-01-21 cs.AI cs.CL cs.LG

MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?

MIMIC-RD: 能否在真实临床环境中让大语言模型对罕见病进行差异性诊断?

Zilal Eiz AlDin, John Wu, Jeffrey Paul Fung, Jennifer King, Mya Watts, Lauren ONeill, Adam Richard Cross, Jimeng Sun

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois College of Medicine(伊利诺伊大学医学学院)

AI总结 MIMIC-RD通过直接映射临床文本实体到Orphanet,评估LLM在真实临床环境下的罕见病差异性诊断能力,发现现有模型表现不佳,揭示了临床需求与现有能力之间的差距。

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08626 2026-01-21 cs.CL

How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction

大语言模型对顺序敏感性如何?OrderProbe用于确定性结构重建

Yingjie He, Zhaolu Kang, Kehan Jiang, Qianyuan Zhang, Jiachen Qian, Chunlei Meng, Yujie Feng, Yuan Wang, Jiabao Dou, Aming Wu, Leqi Zheng, Pengxiang Zhao, Jiaxin Liu, Zeyu Zhang, Lei Wang, Guansu Wang, Qishi Zhan, Xiaomin He, Meisheng Zhang, Jianyuan Ni

机构 * Peking University(北京大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) City University of Hong Kong(香港城市大学) Fudan University(复旦大学) The Hong Kong Polytechnic University(香港理工大学) Tsinghua University(清华大学) Zhejiang University(浙江大学) University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校) Marquette University(马凯特大学) Juniata College(朱尼阿特学院)

AI总结 研究通过OrderProbe基准评估大语言模型对输入顺序的敏感性,发现即使在前沿模型上,结构重建仍面临挑战,且语义能力与结构鲁棒性存在脱节。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21046 2026-01-21 cs.AI

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

自我进化代理的综述:何时、何地、如何进化以实现人工超级智能

Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Qihan Ren, Cheng Qian, Zhenhailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, Mengdi Wang

机构 * Princeton University(普林斯顿大学) Princeton AI Lab(普林斯顿人工智能实验室) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Pennsylvania State University(宾夕法尼亚州立大学) University of Michigan(密歇根大学) Oregon State University(俄勒冈州立大学) The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) University of California, Santa Barbara(加州大学圣芭芭拉分校) University of California San Diego(加州大学圣地亚哥分校) University of Edinburgh(爱丁堡大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文综述了自我进化代理的现状,探讨了进化机制、适应方法及挑战,为实现人工超级智能提供路线图。

Comments 77 pages, 9 figures, Transactions on Machine Learning Research (01/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05938 2026-01-21 cs.LG cs.AI hep-ex hep-ph hep-th

Uncertainty Quantification From Scaling Laws in Deep Neural Networks

深度神经网络中从缩放定律量化不确定性

Ibrahim Elsharkawy, Yonatan Kahn, Benjamin Hooberman

机构 * Department of Physics, University of Illinois Urbana-Champaign, Urbana, IL, USA(伊利诺伊大学厄巴纳-香槟分校物理系) Department of Physics, University of Toronto, Toronto, ON, Canada(多伦多大学物理系) Vector Institute, Toronto, ON, Canada(向量研究所)

AI总结 本文研究了深度神经网络中通过缩放定律量化不确定性的方法,发现测试损失的方差与均值比值在足够大的训练集下与网络宽度无关。

Comments 18+3 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏