arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Cambridge(剑桥大学)

共收录 1283
2507.02861 2026-03-20 cs.CV cs.AI cs.GR

LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans

LiteReality:从RGB-D扫描生成图形-ready的3D场景重建

Zhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao, Matthias Nießner, Joan Lasenby

机构 * University of Cambridge(剑桥大学) The University of Hong Kong(香港大学) Technical University of Munich(慕尼黑技术大学)

AI总结 LiteReality通过结构化场景图解析扫描数据,生成紧凑且逼真的3D虚拟场景,支持图形管线所需的关键特性,适用于AR/VR、游戏、机器人和数字孪生等应用。

Comments Project Page: https://litereality.github.io; Video: https://www.youtube.com/watch?v=ecK9m3LXg2c&feature=youtu.be Camera-Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17840 2026-03-19 cs.CV

Video Understanding: From Geometry and Semantics to Unified Models

视频理解:从几何与语义到统一模型

Zhaochong An, Zirui Li, Mingqiao Ye, Feng Qiao, Jiaang Li, Zongwei Wu, Vishal Thengane, Chengzu Li, Lei Li, Luc Van Gool, Guolei Sun, Serge Belongie

机构 * Department of Computer Science(计算机科学系) University of Copenhagen(哥本哈根大学) College of Computer Science(计算机科学学院) Nankai University(南开大学) School of Computer and Communication Sciences(计算机与通信科学学校) EPFL(苏黎世联邦理工学院) Department of Computer Science & Engineering(计算机科学与工程系) Washington University in St. Louis(圣路易斯华盛顿大学) Computer Vision Lab(计算机视觉实验室) University of Würzburg(乌尔姆大学) Computer Science Research Centre(计算机科学研究中心) University of Surrey(萨里大学) School of Electrical, Computer and Telecommunications Engineering(电气、计算机和电信工程学院) University of Wollongong(沃林根大学) Language Technology Lab(语言技术实验室) University of Cambridge(剑桥大学) School of Artificial Intelligence(人工智能学院) Beijing Institute of Technology(北京理工大学) Institute for Computer Science(计算机科学研究所) INSAIT

AI总结 本文综述了视频理解的发展,从低层几何理解到高层语义理解和统一模型,探讨了时间动态和视觉上下文建模的重要性,并总结了当前研究趋势和挑战。

Comments A comprehensive survey of video understanding, spanning low-level geometry, high-level semantics, and unified understanding models

Journal ref Machine Intelligence Research 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17781 2026-03-19 cs.AI

Facts as First Class Objects: Knowledge Objects for Persistent LLM Memory

事实作为一等对象:用于持久LLM记忆的知识对象

Oliver Zahn, Simran Chana

机构 * Independent Researcher(独立研究者) University of Cambridge(剑桥大学)

AI总结 本文提出知识对象(KOs)作为持久LLM记忆的解决方案,通过对比上下文记忆和KOs,在多跳推理中KOs表现更优,且能有效应对容量限制、压缩损失和目标漂移等挑战。

Comments 26 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17652 2026-03-19 cs.RO cs.CV

VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs

VectorWorld: 通过向量图上的扩散流实现高效的流式世界模型

Chaokang Jiang, Desen Zhou, Jiuming Liu, Kevin Li Sun

机构 * University of Cambridge, Cambridge, United Kingdom(剑桥大学)

AI总结 VectorWorld通过向量图上的扩散流实现高效的流式世界模型,解决了自动驾驶政策闭环评估中的初始化不匹配、采样延迟和运动可行性问题,提升了地图结构精度和闭环运行稳定性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16301 2026-03-19 cs.RO

OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding

OGScene3D: 增量式开放词汇3D高斯场景图映射用于场景理解

Siting Zhu, Ziyun Lu, Guangming Wang, Chenguang Huang, Yongbo Chen, I-Ming Chen, Wolfram Burgard, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) University of Cambridge(剑桥大学) University of Technology Nuremberg(纽伦堡技术大学) Nanyang Technological University(南洋理工大学) Department of Automation, Key Laboratory of System Control and Information Processing of Ministry of Education, State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai Jiao Tong University(自动化系,教育部系统控制与信息处理重点实验室,航空系统集成与航空系统-of-系统综合国家重点实验室,上海导航与定位基于服务重点实验室,上海交通大学)

AI总结 OGScene3D通过增量式3D高斯场景图映射,实现开放词汇场景理解,结合高斯表示与层次优化策略,提升语义一致性与长期优化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10204 2026-03-19 cs.SE cs.LG

Code Roulette: How Prompt Variability Affects LLM Code Generation

代码掷骰子:提示变化如何影响LLM代码生成

Andrei Paleyes, Radzim Sendyka, Diana Robinson, Christian Cabrera, Neil D. Lawrence

机构 * University of Cambridge(剑桥大学)

AI总结 研究探讨提示变化对LLM代码生成质量的影响,提出评估流程以量化模型对输入变化的敏感性,通过实验验证方法有效性。

Comments Extended version of the paper accepted to LLM4Code @ ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16713 2026-03-18 cs.SD

Evaluating Latent Space Structure in Timbre VAEs: A Comparative Study of Unsupervised, Descriptor-Conditioned, and Perceptual Feature-Conditioned Models

评估在音色VAE中的潜在空间结构:对无监督、描述符条件和感知特征条件模型的比较研究

Joseph Cameron, Alan Blackwell

机构 * Department of Computer Science \& Technology, University of Cambridge\ , United Kingdom

AI总结 本文比较了三种音乐音色生成VAE的潜在空间结构,发现基于感知特征的条件模型更紧凑且具有鉴别力,优于无监督和离散描述符条件模型。

Comments 5 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16682 2026-03-18 cs.SD

A Semantic Timbre Dataset for the Electric Guitar

用于电吉他的情感音色数据集

Joseph Cameron, Alan Blackwell

机构 * Department of Computer Science \& Technology, University of Cambridge\ , United Kingdom

AI总结 本文提出一个电吉他音色数据集,通过19个语义音色描述符和对应幅度标注,支持音色控制和语义音频生成,验证了数据集的有效性。

Comments 5 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16435 2026-03-18 cs.CL

VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization

VQKV:通过向量量化实现高保真和高比的缓存压缩

Yixuan Wang, Qingyu Shi, Jiayu Zhou, Dianbo Liu, Ziwei He, Zhouhan Lin

机构 * LUMIA Lab(LUMIA实验室) School of Artificial Intelligence(人工智能学院) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Cambridge(剑桥大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出VQKV方法,利用向量量化实现高压缩比和高重建保真度的KV缓存压缩,实现在有限资源环境下提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16086 2026-03-18 cs.RO cs.AI cs.CV cs.SD

Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation

迈向视觉-声音-语言-动作范式:用于以声音为中心的操作的HEAR框架

Chang Nie, Tianchen Deng, Guangming Wang, Zhe Liu, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and Shanghai Key Laboratory of Navigation and Location Based Services(自动化与智能感知学院,上海交通大学,导航与基于位置的服务重点实验室) Department of Engineering, Cambridge University(工程系,剑桥大学)

AI总结 本文提出HEAR框架,通过整合声音、视觉、语言和本体感知,解决实时声音中心操作中的关键声音遗漏问题,强调因果持续性和显式时间学习的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15994 2026-03-18 cs.AI

Selective Memory for Artificial Intelligence: Write-Time Gating with Hierarchical Archiving

选择性记忆用于人工智能:具有层次归档的写时门控

Oliver Zahn, Simran Chana

机构 * Independent Researcher(独立研究者) University of Cambridge(剑桥大学)

AI总结 本文提出写时门控机制,通过复合显著性评分筛选知识对象,保持版本链以保存先前状态,实验证明其在噪声环境下优于传统方法,尤其在干扰比例增加时表现更优。

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21637 2026-03-18 cs.CV

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

CARE:一种具有自适应区域建模的分子引导基础模型用于全切片图像分析

Di Zhang, Zhangpeng Gong, Xiaobo Pang, Jiashuai Liu, Junbo Lu, Hao Cui, Jiusong Ge, Zhi Zeng, Kai Yi, Yinghua Li, Si Liu, Tingsong Yu, Haoran Wang, Mireia Crispin-Ortuzar, Weimiao Yu, Chen Li, Zeyu Gao

机构 * Xi’an Jiaotong University(西安交通大学) University of Cambridge(剑桥大学) KingMed(康方生物) BGI Research(贝登基因研究院) A ⋆ STAR

AI总结 CARE通过自适应区域建模和分子引导,提升全切片图像分析的性能,实现对病理区域的精准识别与分类,优于现有基础模型。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12272 2026-03-18 cond-mat.mtrl-sci cs.CE cs.LG math.CO

From structure mining to unsupervised exploration of atomic octahedral networks

从结构挖掘到原子八面体网络的无监督探索

R. Patrick Xian, Ryan J. Morelock, Ido Hadar, Charles B. Musgrave, Christopher Sutton

机构 * Department of Engineering, University of Cambridge(剑桥大学工程系) Department of Chemical and Biological Engineering, University of Colorado Boulder(科罗拉多大学波尔得分校化学与生物工程系) The Institute of Chemistry, Casali Center for Applied Chemistry, and the Center for Nanoscience and Nanotechnology, The Hebrew University of Jerusalem(耶路撒冷希伯来大学化学系、应用化学Casali中心和纳米科学与技术中心) Department of Chemistry and Biochemistry, University of South Carolina(南卡罗来纳大学化学与生物化学系)

AI总结 本文提出利用无监督机器学习自动化分析原子八面体网络的几何解析与分类,通过两个数据集验证了其在发现氧化态变化和揭示八面体连接规则中的有效性。

Comments updated version, incl. three supporting information files

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15091 2026-03-17 math.NA cs.LG cs.NA math.DS math.OC

Trustworthy Koopman Operator Learning: Invariance Diagnostics and Error Bounds

可信的Koopman算子学习:不变性诊断与误差界

Gustav Conradie, Nicolas Boullé, Jean-Christophe Loiseau, Steven L. Brunton, Matthew J. Colbrook

机构 * Centre for Mathematical Sciences, University of Cambridge, UK(剑桥大学数学科学中心,英国) Department of Mathematics, Imperial College London, UK(伦敦帝国理工学院数学系,英国) Department of Mechanical Engineering, University of Washington, USA(美国华盛顿大学机械工程系)

AI总结 本文提出了一种统一的后验方法,用于验证Koopman近似的可信度并改进其性能,通过主角度分解和多步误差界,提供可靠的谱分析和预测保障。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13663 2026-03-17 cs.LG

PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers

PDE-SSM:一种用于扩散变换器中空间混合的谱状态空间方法

Eshed Gal, Moshe Eliasof, Siddharth Rout, Eldad Haber

机构 * University of British Columbia, Vancouver, BC Canada(不列颠哥伦比亚大学,加拿大温哥华,BC省) University of Cambridge, Cambridge, United Kingdom(剑桥大学,英国剑桥)

AI总结 本文提出PDE-SSM,通过学习可扩散-反应偏微分方程替代自注意力机制,解决视觉变换器在生成建模中的二次成本和弱空间归纳偏置问题,实现高效且可扩展的替代方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13466 2026-03-17 eess.IV cs.CV

Open World MRI Reconstruction with Bias-Calibrated Adaptation

开放世界MRI重建中的偏置校准适应

Jiyao Liu, Shangqi Gao, Lihao Liu, Junzhi Ning, Jinjie Wei, Junjun He, Xiahai Zhuang, Ningsheng Xu

机构 * Fudan University, Shanghai, China(复旦大学,上海,中国) University of Cambridge, Cambridge, United Kingdom(剑桥大学,剑桥,英国) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)

AI总结 本文提出BiasRecon框架,通过最小干预原则实现开放世界MRI重建,通过频率引导先验校准、基于分数的去噪和自适应正则化,实现鲁棒适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09677 2026-03-16 cs.AI

The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs

持续扩展的幻觉回报:衡量LLM的长周期执行

Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, Jonas Geiping

机构 * University of Cambridge(剑桥大学) Institute for AI, University of Stuttgart(斯图加特大学人工智能研究所) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所) University of Southampton(南安普顿大学) Tübingen AI Center(图宾根人工智能中心)

AI总结 本文通过分析LLM在长周期任务中的执行能力,揭示了持续扩展LLM并非必然导致回报递减,而是执行能力的提升能显著改善长周期任务表现。

Comments Published at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12543 2026-03-16 cs.LG cs.AI

CALF: Communication-Aware Learning Framework for Distributed Reinforcement Learning

CALF:面向分布式强化学习的通信感知学习框架

Carlos Purves, Pietro Lio'

机构 * University of Cambridge(剑桥大学)

AI总结 本文提出CALF框架,通过在仿真中训练考虑通信延迟等网络条件的策略,减少实际部署中的性能差距,验证了通信约束建模对实现鲁棒现实执行的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10873 2026-03-12 cs.LG q-bio.GN

SNPgen: Phenotype-Supervised Genotype Representation and Synthetic Data Generation via Latent Diffusion

SNPgen: 通过潜在扩散模型生成表型监督的合成基因型

Andrea Lampis, Michela Carlotta Massi, Nicola Pirastu, Francesca Ieva, Matteo Matteucci, Emanuele Di Angelantonio

机构 * DEIB, Politecnico di Milano(米兰理工学院DEIB部门) Politecnico di Milano(米兰理工学院) Health Data Science Centre, Human Technopole(人类技术pole健康数据科学中心) Genomics Research Centre, Human Technopole(人类技术pole基因组研究中心) MOX - Department of Mathematics, Politecnico di Milano(米兰理工学院数学系) Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级护理系)

AI总结 SNPgen通过条件潜在扩散模型生成表型监督的合成基因型,实现统计忠实性和下游任务实用性之间的平衡,与全基因组PRS方法在预测性能上相当。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09517 2026-03-11 cs.CL cs.LG

You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases

你无需那样说:从忠实的同义词中学习的潜意识学习

Isaia Gisler, Zhonghao He, Tianyi Qiu

机构 * ETH Zürich(苏黎世联邦理工学院) University of Cambridge(剑桥大学) Peking University(北京大学)

AI总结 研究通过自然语言同义词的潜意识学习,发现即使内容与教师偏好矛盾,学生模型仍能继承偏好,揭示了训练数据中潜在影响的隐蔽性。

Comments Accepted for Spotlight presentation at EACL 2026 SRW. 5 pages, 2 figures, plus appendix. Equal supervision by Zhonghao He and Tianyi Qiu

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09200 2026-03-11 cs.AI cs.CL cs.CY cs.LG

The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness

推理陷阱——逻辑推理作为情境意识的机制路径

Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary

机构 * MARS 4.0 Fellowship, Cambridge AI Safety Hub(CAISH), University of Cambridge(MARS 4.0 Fellow,剑桥人工智能安全中心(CAISH),剑桥大学) AWS Generative AI Innovation Center, Amazon Web Services, USA(亚马逊生成AI创新中心,亚马逊网络服务,美国) Google, USA(谷歌,美国) Stanford University(斯坦福大学) Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国)

AI总结 本文提出RAISE框架,揭示逻辑推理能力提升与情境意识升级的机制路径,并提出安全原则与测试方法以应对潜在风险。

Comments Accepted at ICLR 2026 Workshop on Logical Reasoning of Large Language Models. 21 Pages. Position Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08219 2026-03-10 cs.LG

Wiener Chaos Expansion based Neural Operator for Singular Stochastic Partial Differential Equations

基于Wiener混沌展开的神经算子用于奇异随机偏微分方程

Dai Shi, Luke Thompson, Andi Han, Peiyan Hu, Junbin Gao, José Miguel Hernández-Lobato

机构 * University of Cambridge(剑桥大学) University of Sydney(悉尼大学) Chinese Academy of Science(中国科学院)

AI总结 本文提出基于Wiener混沌展开和FiLM的神经算子,用于高效求解奇异随机偏微分方程,尤其在Φ^4_2和Φ^4_3模型中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06967 2026-03-10 cs.LG cs.AI

Noisy PDE Training Requires Bigger PINNs

带有噪声的PDE训练需要更大的PINN

Sebastien Andre-Sloan, Anirbit Mukherjee, Matthew Colbrook

机构 * The University of Manchester(曼彻斯特大学) The University of Cambridge(剑桥大学)

AI总结 研究提出PINNs在噪声数据下训练需更大网络规模以降低经验风险,并通过实验验证了这一结论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00296 2026-03-10 cs.RO cs.AI cs.CV cs.LG

From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models

从像素到谓词:通过预训练视觉-语言模型学习符号世界模型

Ashay Athalye, Nishanth Kumar, Tom Silver, Yichao Liang, Jiuguang Wang, Tomás Lozano-Pérez, Leslie Pack Kaelbling

机构 * MIT(麻省理工学院) Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) RAI Institute(RAI研究院)

AI总结 通过预训练视觉-语言模型学习符号世界模型,以实现复杂机器人领域中长周期决策制定的零样本泛化。

Comments A version of this paper appears in the official proceedings of RA-L, Volume 11, Issue 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09787 2026-03-10 cs.LG cs.AI stat.CO stat.ML

BNEM: A Boltzmann Sampler Based on Bootstrapped Noised Energy Matching

BNEM:基于Bootstrap噪声能量匹配的玻尔兹曼采样器

RuiKang OuYang, Bo Qiang, José Miguel Hernández-Lobato

机构 * University of Cambridge(剑桥大学) University of Washington(华盛顿大学)

AI总结 BNEM通过基于噪声能量匹配的Bootstrap技术,在分子动力学等应用中实现高效且鲁棒的采样性能。

Comments Camera-ready version for TMLR (03/2026)

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08082 2026-03-10 cs.LG

Tiny Autoregressive Recursive Models

微型自回归递归模型

Paulius Rauba, Claudio Fanconi, Mihaela van der Schaar

机构 * University of Cambridge(剑桥大学)

AI总结 本文提出自回归TRM,评估其在小型自回归任务上的表现,发现两步细化基线表现强劲,但完整架构未见显著性能提升。

Journal ref ICLR 2026 Workshop RSI Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07779 2026-03-10 cs.CL cs.GL cs.LG

Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems

数据难度的扩展:通过在新鲜且具有挑战性的问题上进行强化学习来改进编码模型

Zongqian Li, Tengchao Lv, Shaohan Huang, Yixuan Su, Qinzheng Sun, Qiufeng Yin, Ying Xin, Scarlett Li, Lei Cui, Nigel Collier, Furu Wei

机构 * Microsoft Research(微软研究院) University of Cambridge(剑桥大学)

AI总结 通过强化学习在新鲜且具有挑战性的问题上改进编码模型,MicroCoder数据集通过系统化数据处理和难度扩展实现了显著的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07777 2026-03-10 cs.LG cs.CL cs.GL

Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models

突破训练瓶颈:为编码模型提供有效且稳定的强化学习

Zongqian Li, Shaohan Huang, Zewen Chi, Yixuan Su, Lexin Zhou, Li Dong, Nigel Collier, Furu Wei

机构 * Microsoft Research(微软研究院) University of Cambridge(剑桥大学) Princeton University(普林斯顿大学)

AI总结 本文提出MicroCoder-GRPO方法,通过改进的组相对策略优化解决编码模型训练瓶颈,实现性能提升并发布相关数据集和评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07233 2026-03-10 cs.LG cs.IR

Retrieval-Augmented Generation for Predicting Cellular Responses to Gene Perturbation

增强检索生成用于预测细胞对基因扰动的响应

Andrea Giuseppe Di Francesco, Andrea Rubbi, Pietro Liò

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) ISTI-CNR(意大利国家研究委员会信息科学与技术研究所) University of Cambridge(剑桥大学) Wellcome Sanger Institute(wellcome桑格研究所)

AI总结 PT-RAG通过细胞类型感知的可微检索增强生成,提升预测细胞对基因扰动响应的性能。

Comments Accepted at ICLR 2026 Workshop: Generative AI in Genomics. 25 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01865 2026-03-10 cs.CL

CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation

CyclicJudge: 有效缓解基于大语言模型的评估中的判断偏差

Ziyi Zhu, Olivier Tieleman, Alexey Bukhtiyarov, Jinghong Chen

机构 * Slingshot AI Department of Engineering, University of Cambridge(工程系,剑桥大学)

AI总结 CyclicJudge通过轮换判断者分配策略,有效消除大语言模型评估中的判断偏差,提升评估的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏