What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct
什么是AI阿谀奉承?对一个碎片化概念的分类和专家调查
Meryl Ye, Lujain Ibrahim, Jessica Y. Bo, Myra Cheng, Ida Mattsson, Daniel Vennemeyer, Robert Kraut, Steve Rathje
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
University of Oxford(牛津大学)
;
University of Toronto(多伦多大学)
;
Stanford University(斯坦福大学)
;
University of Cincinnati(克里夫兰医学中心大学)
;
New York University(纽约大学)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
Specification-Driven Generation and Evaluation of Discrete-Event World Models via the DEVS Formalism
通过DEVS形式化方法驱动的离散事件世界模型生成与评估
Zheyu Chen, Huiteng Zhuang, Zhuohuan Li, Chuanhao Li
机构
*
Zhili College, Tsinghua University(清华大学紫光学院)
;
School of Transportation Science and Engineering, Beihang University(北航交通科学与工程学院)
;
Department of Industrial Engineering, Tsinghua University(清华大学工业工程系)
An AI system to help scientists write expert-level empirical software
一种帮助科学家编写专家级经验软件的AI系统
Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Jake Garrison, Renee Johnston Anton Kast, Cory Y. McLean, Peter Norgaard, Zahra Shamsi, David Smalling, James Thompson, Subhashini Venugopalan, Brian P. Williams, Chujun He, Sarah Martinson, Martyna Plomecka, Lai Wei, Yuchen Zhou, Qian-Ze Zhu, Matthew Abraham, Erica Brand, Anna Bulanova, Jeffrey A. Cardille, Chris Co, Scott Ellsworth, Grace Joseph, Malcolm Kane, Ryan Krueger, Johan Kartiwa, Dan Liebling, Jan-Matthis Lueckmann, Paul Raccuglia, Xuefei, Wang, Katherine Chou, James Manyika, Yossi Matias, John C. Platt, Lizzie Dorfman, Shibl Mourad, Michael P. Brenner
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究)
;
Google Platforms and Devices(谷歌平台与设备)
;
Massachusetts Institute of Technology(麻省理工学院)
;
School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
AI总结
本文提出Empirical Research Assistance (ERA)系统,利用大型语言模型和树搜索技术,自动创建高质量的科学软件,以加速计算实验的开发,从而提高科研效率。
When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering
当案例变得稀少时:一个用于偏离指南临床问答的检索基准
Doeun Lee, Muge Zhang, Yi Yu, Ashish Manne, Stephen Koesters, Frank Wen, Brady Buchanan, Lynda Villagomez, Oluwatoba Moninuola, James Lim, Kathryn Tobin, Andrew Srisuwananukorn, Ping Zhang, Sachin Kumar
机构
*
The Ohio State University(俄亥俄州立大学)
;
The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医学中心)
;
University of Chicago Medical Center(芝加哥大学医学中心)
专题命中
评测与基准
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
VeriScale:对抗性测试套件缩放用于可验证代码生成
Yifan Bai, Xiaoyang Liu, Zihao Mou, Guihong Wang, Jian Yu, Shuhan Xie, Yantao Li, Yangyu Zhang, Jingwei Liang, Tao Luo
机构
*
School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院)
;
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院)
;
School of Mathematics, Jilin University(吉林大学数学学院)
;
School of Mathematical Sciences, Tongji University(同济大学数学科学学院)
;
Zhiyuan College, Shanghai Jiao Tong University(上海交通大学紫阳学院)
;
School of Future Technology, South China University of Technology(华南理工大学未来技术学院)
;
Institute of Natural Sciences, Shanghai Jiao Tong University(上海交通大学自然科学研究院)
;
MOE-LSC, CMA-Shanghai, Shanghai Jiao Tong University(上海交通大学MOE-LSC、CMA-上海)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
STRUCTSENSE:一种任务无关的代理框架,用于结构化信息提取,具有人机协同评估和基准测试
Tek Raj Chhetri, Yibei Chen, Puja Trivedi, Dorota Jarecka, Saif Haobsh, Patrick Ray, Lydia Ng, Satrajit S. Ghosh
机构
*
McGovern Institute for Brain Research, Massachusetts Institute of Technology, Cambridge, MA, USA(麦戈文脑科学研究所,麻省理工学院,马萨诸塞州剑桥市)
;
Fylo Labs Inc., New York, NY, USA(Fylo实验室公司,纽约州纽约市)
;
Allen Institute for Brain Science, Seattle, WA, USA(艾伦脑科学研究所,华盛顿州西雅图市)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results
机构
*
Department of Big Data, Chungbuk National University, Cheongju-si, South Korea(大数据系, Chungbuk国立大学,韩国Cheongju市)
;
Department of Computer Science, Chungbuk National University, Cheongju-si, South Korea(计算机科学系, Chungbuk国立大学,韩国Cheongju市)
;
BigDataLabs Co., Ltd. Department of Management Information Systems, Chungbuk National University, South Korea(BigDataLabs公司 管理信息系, Chungbuk国立大学,韩国)
Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization
支持接近增强的扩散估计用于离线黑盒优化
Yonghan Yang, Ye Yuan, Zipeng Sun, Linfeng Du, Bowei He, Haolun Wu, Can Chen, Xue Liu
机构
*
MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩擦 bin Zayed 大学)
;
McGill University(麦吉尔大学)
;
Mila - Quebec AI Institute(Mila - 加拿大AI研究所)
;
Amazon AGI(亚马逊人工智能实验室)
机构
*
New York University Abu Dhabi(纽约大学阿布扎克校区)
;
Nanyang Technological University(南洋理工大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Harvard University(哈佛大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Beijing University of Technology(北京理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)