VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
VitaBench 2.0:评估长期用户交互中的个性化与主动型代理
Yuxin Chen, Yi Zhang, Zhengzhou Cai, Yaorui Shi, Zhiyuan Yao, Chenhang Cui, Jingnan Zheng, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Xiang Wang, An Zhang, Tat-Seng Chua
机构
*
National University of Singapore(新加坡国立大学)
;
Meituan(美团)
;
University of Science and Technology of China(中国科学技术大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Zhejiang University(浙江大学)
机构
*
McGill University(麦吉尔大学)
;
Mila - Quebec AI Institute(魁北克人工智能研究所)
;
University of Cambridge(剑桥大学)
;
MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学)
;
University of Toronto(多伦多大学)
;
Salesforce
Comments41 pages, 7 figures, 7 tables. Preliminary cJSON-only evaluation (N=5 main, N=3 ablation; descriptive statistics, no significance claims). Code and 25-run artifacts at https://github.com/Qiao-Zhiyi/fuzz_agent (tag paper01-arxiv-v1). Venue-version Stage-1 pilot on libxml2, sqlite3, openssl_x509 currently in flight; v2 will report those results
Comments18 pages, 4 figures, 14 tables; includes appendices with verbatim prompts, example session, and full ablation tables; prepared by the LLM Suite Engineering Team, JP Morgan Chase & Co