机构
*
Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University(数据科学与人工智能系,香港理工大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation
通过细粒度事实核查和领域特定适应减轻医疗保健大语言模型中的幻觉
Musarrat Zeba, Abdullah Al Mamun, Kishoar Jahan Tithee, Debopom Sutradhar, Mohaimenul Azam Khan Raiaan, Saddam Mukta, Reem E. Mohamed, Md Rafiqul Islam, Yakub Sebastian, Mukhtar Hussain, Sami Azam
机构
*
Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统实验室)
;
Department of Computer Science and Engineering(计算机科学与工程系)
;
Department of Data Science and Artificial Intelligence(数据科学与人工智能系)
;
Department of Software Engineering(软件工程系)
;
Faculty of Science and Information Technology(科学与信息技术学院)
;
Faculty of Science and Technology(科学与技术学院)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
重新审视印度语言机器翻译和摘要细粒度评估的度量可靠性
Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri Koto
机构
*
Sharif University of Technology(谢里夫理工学院)
;
Vellore Institute of Technology(韦洛雷理工学院)
;
IIT Kharagpur(印度理工学院达卡分校)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
FieldWorkArena:面向真实作业任务的代理AI基准测试
Jun Takahashi, Atsunori Moteki, Akiyoshi Uchida, Shoichi Masui, Fan Yang, Kanji Uchino, Yueqi Song, Yonatan Bisk, Graham Neubig, Ikuo Kusajima, Yasuto Watanabe, Hiroyuki Ishida, Koki Nakagawa, Shan Jiang
机构
*
Fujitsu Limited(富士通株式会社)
;
Fujitsu Research of America(富士通美国研究部)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Master’s Student, The University of Tokyo(东京大学硕士研究生)
;
Agent Research Collective(代理研究集体)
CommentsRevised version. Refined ontological modeling of legislative events (adopted F27/E64 joint typing over E11). Introduced technical distinctions for bitemporal modeling in legal knowledge graphs and enriched the critical analysis of related standards in Section 2
Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation
通过表征地标和Nyström近似的可解释自监督学习
Maedeh Zarvandi, Michael Timothy, Theresa Wasserer, Debarghya Ghoshdastidar
机构
*
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Technical University of Munich, TUM School of Computation, Information and Technology(慕尼黑技术大学,TUM计算、信息与技术学院)
Comments34 pages, 2 figures, 5 tables. v2 substantially revises Secs. 5-8 and the conclusion, adds a measured instance of factor-structure instability, and corrects a claim in Sec. 3.3 that endpoint alignment suffices for composition. Supersedes the withdrawn arXiv:2510.15236. Code: https://github.com/BrettRey/benchmark-inference-composition
Semantic Drift in Bug Resolution: How Behavioral Signals Propagate from Reports to Tests and Patches
错误修复中的语义漂移:行为信号如何从报告传播到测试和补丁
Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang, Paweł Borsukiewicz, Liang Xiao, Lingfeng Bao, Anil Koyuncu, Jacques Klein, David Lo, Tegawendé F. Bissyandé
Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection
剖析与剪枝:增强AI生成图像检测的鲁棒性
Dahye Kim, Jaehyun Choi, Hyun Seok Seong, Seongho Kim, Donghun Lee, Sungwon Yi, Jang-Ho Choi
机构
*
Korea AI Safety Institute (AISI), ETRI, Seongnam, South Korea(韩国人工智能安全研究所(AISI)、ETRI、Seongnam韩国)
;
Department of Artificial Intelligence, Sungkyunkwan University, Suwon, South Korea(人工智能系,全州大学,Suwon韩国)