GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing
GenPT:通过生成式投射测试实现超越自我报告的可靠LLM心理测量
Ming Wang, Shuang Wu, Bixuan Wang, Lu Lin, Yuxin Chen, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang, Yufan Sun
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算机与信息学院)
;
Mental Health Education Center, Northeastern University(东北大学心理健康教育中心)
;
School of Psychology, Northeast Normal University(东北师范大学心理学系)
;
Faculty of psychology, Southwest University(西南大学心理学系)
;
School of Sociology and Psychology, Central University of Finance and Economics(中央财经大学社会学与心理学学院)
;
College of Arts, Northeastern University(东北大学艺术学院)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
EgoDyn-Bench:评估面向自动驾驶的视觉中心基础模型中的自我运动理解
Finn Rasmus Schäfer, Yuan Gao, Dingrui Wang, Thomas Stauner, Stephan Günnemann, Mattia Piccinini, Sebastian Schmidt, Johannes Betz
机构
*
Professorship of Autonomous Vehicle Systems, Technical University of Munich, Munich, Germany(自动驾驶车辆系统教授职位,慕尼黑技术大学,德国慕尼黑)
;
Bayerische Motoren Werke AG, Munich, Germany(巴伐利亚发动机有限公司,德国慕尼黑)
;
Data Analytics and Machine Learning Group, Technical University of Munich, Munich, Germany(数据分析与机器学习小组,慕尼黑技术大学,德国慕尼黑)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
ROK-FORTRESS:衡量地缘政治转译对国家安全和公共安全的影响
Michael S. Lee, Yash Maurya, Drew Rein, Bert Herring, Jonathan Nguyen, Kyungho Song, Udari Madhushani Sehwag, Jiyeon Cho, Kaustubh Deshpande, Yeongkyun Jang, Jiyeon Joo, Minn Seok Choi, Evi Fuelle, Christina Q. Knight, Joseph Brandifino, Max Fenkell
机构
*
Scale AI
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
CommentsAccepted for publication at the IEEE International Conference on LLM-Aided Design, 2026, to be held: Time: July 30-31, 2026 Location: Stanford University, Stanford, CA
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
PluraMath:将数学推理评估扩展到资源丰富语言之外
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan Özer, Nikola Selic, Subhankar Swain, Tsedeniya Kinfe Temesgen, Galit Bary Weisberg, Alexander Fraser
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Munich Data Science Institute(慕尼黑数据科学研究所)
;
Applied AI Institute(应用人工智能研究所)
;
Charles University(查尔斯大学)
;
Nazarbayev University(纳扎尔拜大学)
;
Inria(法国国家信息与自动化研究所)
;
Indian Institute of Technology, Kharagpur (IIT Kharagpur)(印度Kharagpur理工学院)
;
German University of Digital Science(德国数字科学大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments28 pages, 4 figures. Pre-registered cross-provider evaluation harness and per-regime results at github.com/canonic-canonic/canonic-pub. Construction claims resolve to commands run at the evidence-window-close ref (see Appendix C)