机构
*
Salesforce AI Research(Salesforce AI研究)
;
National University of Singapore(国立新加坡大学)
;
Nanyang Technological University(南洋理工大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
A*STAR, Singapore(新加坡A*STAR)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Dehao Dai, Ding Ma, Dou Liu, Kerui Geng, Yiqing Wang
机构
*
University of California San Diego(加州大学圣迭戈分校)
;
Georgia Institute of Technology(佐治亚理工学院)
;
New York University(纽约大学)
;
Tulane University(路易斯安那州立大学)
;
Southern Methodist University(南方 Methodist 大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
Attribution Quality in AI-Generated Content:Benchmarking Style Embeddings and LLM Judges
AI生成内容中的归因质量:基准测试风格嵌入与LLM裁判
Misam Abbas
机构
*
IEEE International Conference on Data Mining Workshops(IEEE国际数据挖掘会议研讨会)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结
本文通过Human AI Parallel Corpus评估风格嵌入与LLM裁判在AI生成内容归因中的表现,发现LLM裁判在小说和学术写作中表现更优,而嵌入在口语和剧本对话中更占优势,强调归因问题的多维性。
CommentsAccepted for publication at the 2025 IEEE ICDM Workshop on "Grounding Documents with Reasoning, Agents, Retrieval, and Attribution". This is author submitted version. Not yet published
Journal refProc. IEEE Int. Conf. Data Mining Workshops (ICDMW), pp. 1713-1720, 2025
IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans
IOSVLM:一种用于从牙科内窥扫描进行统一牙科诊断的3D视觉-语言模型
Huimin Xiong, Zijie Meng, Tianxiang Hu, Chenyi Zhou, Yang Feng, Zuozhu Liu
机构
*
ZJU-UIUC Institute, Zhejiang University, Haining, 314400, China(浙大浙ICU研究所,浙江大学,海宁,314400,中国)
;
Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Hangzhou, 310058, China(口腔医院,口腔医学院,浙江大学医学院,杭州,310058,中国)
;
Angelalign Research Institute, Angel Align Inc., Shanghai, 200011, China(天使对齐研究院,天使对齐公司,上海,200011,中国)
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
正则化潜在动态预测是行为基础模型的强基线
Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White
机构
*
Department of Computing Science, University of Alberta, Canada(阿尔伯塔大学计算机科学系)
;
Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能 chair)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)
DUCTILE: Agentic LLM Orchestration of Engineering Analysis in Product Development Practice
DUCTILE:工程分析中产品开发实践的代理LLM协调
Alejandro Pradas-Gomez, Arindam Brahma, Ola Isaksson
机构
*
Product Development Division, Department of Mechanical Engineering, Chalmers University of Technology, Gothenburg, Sweden(产品开发部门,机械工程系,楚姆勒斯技术大学,哥德堡,瑞典)
;
GKN Aerospace Engine Systems, Trollhättan, Sweden(GKN航空发动机系统,特罗尔哈塔纳,瑞典)
On Theoretically-Driven LLM Agents for Multi-Dimensional Discourse Analysis
关于理论驱动的LLM代理在多维话语分析中的应用
Maciej Uberna, Michał Wawer, Jarosław A. Chudziak, Marcin Koszowy
机构
*
Laboratory of The New Ethos, Warsaw University of Technology, Poland(新伦理实验室,华沙理工大学,波兰)
;
Faculty of Electronics and Information Technology, Warsaw University of Technology, Poland(电子与信息技术学院,华沙理工大学,波兰)
Comments8 pages, 4 figures, 3 tables. This is the accepted version of the paper presented at the 18th International Conference on Agents and Artificial Intelligence (ICAART 2026), Marbella, Spain
Journal refProceedings of the 18th International Conference on Agents and Artificial Intelligence (ICAART 2026)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
SWE-QA-Pro:一个代表性的基准和可扩展的训练配方用于仓库级代码理解
Songcheng Cai, Zhiheng Lyu, Yuansheng Ni, Xiangchao Chen, Baichuan Zhou, Shenzhe Zhu, Yi Lu, Haozhe Wang, Chi Ruan, Benjamin Schneider, Weixu Zhang, Xiang Li, Andy Zheng, Yuyu Zhang, Ping Nie, Wenhu Chen
机构
*
University of Waterloo(滑铁卢大学)
;
University of Toronto(多伦多大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
McGill University & MILA(麦吉尔大学及MILA)
;
Verdent AI, Inc.(Verdent AI公司)
专题命中
评测与基准
:large language model(abstract);language model(abstract);SFT(abstract);分类 cs.CL、cs.AI
AOI: Turning Failed Trajectories into Training Signals for Autonomous Cloud Diagnosis
AOI: 将失败轨迹转化为自主云诊断的训练信号
Pei Yang, Wanyi Chen, Asuka Yuxi Zheng, Xueqian Li, Xiang Li, Haoqin Tu, Jie Xiao, Yifan Pang, Dongdong Zhang, Fuqiang Li, Alfred Long, Lynn Ai, Eric Yang, Bill Shi
机构
*
Gradient
;
Soochow University(苏州大学)
;
UC Santa Cruz
;
Georgia Institute of Technology(佐治亚理工学院)
;
University College London(伦敦大学学院)
;
WeJoy
;
ByteDance(字节跳动)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
ReGAIN: Retrieval-Grounded AI Framework for Network Traffic Analysis
ReGAIN:基于检索的网络流量分析人工智能框架
Shaghayegh Shajarian, Kennedy Marsh, James Benson, Sajad Khorsandroo, Mahmoud Abdelsalam
机构
*
Computer Science(计算机科学)
;
North Carolina A\&T State University(北卡罗来纳A&T州立大学)
;
Institute for Cyber Security(网络安全研究所)
;
University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
CRIMSON:一种基于临床的LLM评估指标用于生成放射学报告评估
Mohammed Baharoon, Thibault Heintz, Siavash Raissi, Mahmoud Alabbad, Mona Alhammad, Hassan AlOmaish, Sung Eun Kim, Oishi Banerjee, Pranav Rajpurkar
机构
*
Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院)
;
Department of Radiation Oncology, Mass General Brigham(放射肿瘤科,麻省总医院与布里奇沃特医院)
;
King Fahad Hospital, Al-Ahsa Health Cluster, Al Hofuf, Saudi Arabia(国王法赫德医院,阿尔阿萨卫生集群,阿尔霍夫,沙特阿拉伯)
;
Ras-Tanura General Hospital, Ministry of Health, Eastern Province, Saudi Arabia(拉斯-坦纳医院,卫生部,东部省,沙特阿拉伯)
;
Department of Medical Imaging, King Abdulaziz Medical City, Ministry of National Guard, Riyadh, Saudi Arabia(医学影像科,国王阿卜杜勒-阿齐兹医疗城,国民卫队部,利雅得,沙特阿拉伯)