Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
十年视觉语言人工智能模型中准确性和视觉认知错误的演变
Shravan Murlidaran, Miguel P. Eckstein
机构
*
Psychological & Brain Sciences, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校心理与脑科学系)
;
Department of Computer Science, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校计算机科学系)
;
Department of Electrical and Computer Engineering, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校电气与计算机工程系)
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
一种用于人工智能辅助鉴别诊断的面向安全的假设-演绎框架
Fan Ma, Mauro Giuffrè, Donald Wright, Kent McCann, Mark Iscoe, Lingfei Qian, Mingyang Jiang, Chi Wing Ng, Na Hong, Huan He, Cathy Shyr, Qingyu Chen, Lee Schwamm, Lucila Ohno-Machado, Hua Xu
机构
*
Yale School of Medicine, Yale University(耶鲁大学医学院,耶鲁大学)
;
Università degli Studi di Trieste(的里雅斯特大学)
;
Vanderbilt University(范德堡大学)
A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving
一个用于自动驾驶的具有大语言模型注释的高风险驾驶场景知识增强数据集
Heye Huang, Jingguang Li, Zhiyuan Zhou, Paul Liang, Mingyu Wu, Kitae Jang, Jianqiang Wang
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Fudan University(复旦大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Tsinghua University(清华大学)
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
一种评估智能体人工智能自主模型发现的实验设计方法
Hao He, Xueying Liu, Chris J. Kuhlman, Xinwei Deng
机构
*
Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工学院)
;
Department of Statistical Science, Baylor University(统计科学系,贝勒大学)
;
Advanced Research Computing, Virginia Tech(高级研究计算,弗吉尼亚理工学院)
Comments7 pages. To be published in the proceedings of 41st International Conference on Automated Software Engineering (ASE '26), October 12-16, 2026, Munich, Germany (Industry Showcase Track)
Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems
自主信息寻求:智能推荐系统路线图
Xinyu Lin, Yashar Deldjoo, Sunhao Dai, Honghui Bao, Xiaopeng Ye, Fatemeh Nazary, Wenjie Wang, Tommaso Di Noia, Jun Xu, Tat-Seng Chua
机构
*
National University of Singapore(新加坡国立大学)
;
Polytechnic University of Bari(巴里理工大学)
;
Renmin University of China(中国人民大学)
;
University of Science and Technology of China(中国科学技术大学)
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models
关于强化学习中奖励函数对大语言模型置信度校准的有效性
Chee Heng Tan, Zhuoyi Lin, Mehul Motani, Wee Sun Lee
机构
*
School of Computing National University of Singapore(新加坡国立大学计算机学院)
;
Institute of Advanced Intelligence and Computing(高级智能与计算研究所)
;
Agency for Science, Technology and Research(科技研究局)
;
Department of Electrical & Computer Engineering(电气与计算机工程系)
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
HCSU:用于细粒度历史书法风格理解的数据集和基准测试
Yinsheng Yao, Yan Liu, Chen Ye
机构
*
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
The Key Laboratory of Embedded System and Service Computing, Ministry of Education(教育部嵌入式系统与服务计算重点实验室)
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
用于开源情报和网络调查的智能与生成式人工智能:分类法、评估、挑战及未来方向
Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
机构
*
School of Computer Science and Mathematics, Keele University(基尔大学计算机科学与数学学院)
;
Cybersecurity Institute, University of Liverpool(利物浦大学网络安全研究所)
;
Department of Applied Computing IICL, University of Wales Trinity Saint David(威尔士特里尼达大学应用计算系)
;
Deakin Cyber Research and Innovation Hub, Deakin University(德金大学网络安全研究与创新中心)
;
Department of Information Systems and Cyber Security, The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校信息系与网络安全系)
机构
*
Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
University of Manchester(曼彻斯特大学)
;
Institute of Metal Research, Chinese Academy of Sciences(中国科学院金属研究所)
;
China University of Geosciences(中国地质大学)
机构
*
South China University of Technology(华南理工大学)
;
The University of Hong Kong(香港大学)
;
Snap Inc(Snap公司)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
EduArt:评估大型语言模型艺术史知识的教育级基准
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
University of Bologna(博洛尼亚大学)
;
Villa i Tatti – The Harvard University Center for Italian Renaissance Studies(哈佛大学意大利文艺复兴研究中心(I Tatti))
;
University of Copenhagen(哥本哈根大学)