AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
AgentStream:自进化大语言模型智能体在流式任务下的表现如何?
Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan, Qi Zhang, Tieniu Tan
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
Microsoft(微软公司)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Nanjing University(南京大学)
专题命中
评测与基准
:LLM(title,summary_cn);large language model(abstract);language model(abstract);foundation model(abstract)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
MEDIC:对LLM在临床应用中的安全性和实用性领先指标的综合评估
Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan
机构
*
M42
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Given, When, Then, Again: Mining Subscenario Refactoring Candidates in Behaviour-Driven Test Suites with ML Classifiers and LLM-Judge Baselines
在行为驱动软件测试套件中挖掘子场景重构机会:ML分类器和LLM-判断基线
Ali Hassaan Mughal, Noor Fatima, Muhammad Bilal
机构
*
Independent Researcher(独立研究者;应用MBA(数据分析),德克萨斯韦斯利安大学)
;
Applied MBA (Data Analytics), Texas Wesleyan University(独立研究者;计算机工程学士,国立科学与技术大学(NUST))
;
Independent Researcher(独立研究者;管理硕士,慕尼黑技术大学)
;
B.E. Computer Engineering, National University of Sciences and Technology (NUST)
;
Independent Researcher
;
M.Sc. Management, Technical University of Munich
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data
随机规则森林 (RRF): 基于LLM生成问题的可解释且可控集成方法用于从非结构化数据预测成功
Ben Griffin, Aaron Ontoyin Yin, Diego Vidaurre, Ugur Koyluoglu, Joseph Ternasky, Fuat Alican, Yigit Ihlamur
机构
*
University of Oxford, United Kingdom Aarhus University, Aarhus, Denmark Centre de Recerca Matem\`atica, Barcelona, Spain Oliver Wyman, New York, United States Vela Research, San Francisco, United States
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)
A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback
用于评估手术反馈质量的多智能体LLM框架
Rafal Kocielnik, J. Everett Knudsen, Steven Y. Cen, Jasmine Lin, Cherine H. Yang, Atharva Deo, Ujjwal Pasupulety, Peter Wager, Anima Anandkumar, Andrew J. Hung
机构
*
Computing + Mathematical Sciences, California Institute of Technology(加州理工学院计算与数学科学系)
;
Department of Urology, Cedars-Sinai(塞斯医疗中心泌尿科)
;
Keck School of Medicine, University of Southern California(美国南加州大学凯克医学院)
Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models
使用多智能体语言模型自动检测和分类自然音频日记中的妄想相关内容
Feng Chen, Justin Tauscher, Changye Li, Meliha Yetisgen, Alex Cohen, Adam Kuczynski, Angelina Pei-Tzu Tsai, Benjamin Buck, Dror Ben-Zeev, Trevor Cohen
机构
*
Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, USA(生物医学信息学与医学教育系,华盛顿大学,西雅图,华盛顿州,美国)
;
Department of Psychiatry and Behavioral Sciences, University of Washington, Seattle, WA, USA(精神病学与行为科学系,华盛顿大学,西雅图,华盛顿州,美国)
;
Department of Psychology, Louisiana State University, Baton Rouge, LA, USA(心理学系,路易斯安那州立大学,巴吞鲁日,路易斯安那州,美国)
;
Department of Psychiatry, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA(精神病学系,北卡罗来纳大学教堂山分校,教堂山,北卡罗来纳州,美国)
专题命中
评测与基准
:LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract);foundation model(abstract)
Wentao Zhang, Qi Zhang, Mingkun Xu, Mu You, Henghua Shen, Zhongzhi He, Keyan Jin, Derek F. Wong, Tao Fang
机构
*
Business School, Shandong University of Technology(山东理工大学商学院)
;
Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)
;
Guangdong Institute of Intelligent Science and Technology(广东智能科学与技术研究院)
;
Macau Millennium College(澳门 millennium 学院)
;
Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)
CommentsThis work is an expanded version of our prior paper published in the IEEE ICASSP 2026 conference arXiv:2512.24947, from 4 to 20+ pages, presenting a well-structured and principled framework, extensive experiments, and deeper insights. Tao Fang is the corresponding author
Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation
超越参数:大型语言模型中上下文丰富技术的综述:从上下文提示到因果检索增强生成
Prakhar Bansal, Shivangi Agarwal
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);prompting(title);分类 cs.CL、cs.AI