机构
*
School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)
;
Melbourne School of Psychological Sciences, The University of Melbourne(墨尔本大学墨尔本心理科学学院)
;
LILT
专题命中
评测与基准
:language model(title,abstract);LLM(abstract_cn);large language model(abstract);分类 cs.CL
CommentsImproved benchmark design and reproducibility, replaced restricted datasets with public benchmarks in primary analyses, and added sensitivity analyses supporting the interpretation of model scaling and evaluation protocol effects in molecular prediction
TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs
TABVERSE:大语言模型与视觉语言模型中跨格式表格理解的基准测试
Momina Ahsan, Sarfraz Ahmad, Ming Shan Hee, Roy Ka-Wei Lee, Preslav Nakov
机构
*
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)
专题命中
评测与基准
:LLM(summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
PACT: Learning Diverse Diagnostic Strategies via Privileged Synthesis and Branch Consensus
PACT: 通过特权合成与分支共识学习多样化诊断策略
Gen Li, Yuanze Hu, Zhichao Yang, Qingchen Yu, Jianwei Lv, Yue Guo, Yujing Liu, Faguo Wu, Hongwei Zheng, Xiandong Li, Bo Yuan, Yifan Sun, Zhaoxin Fan
机构
*
Beihang University(北京航空航天大学)
;
Baidu(百度)
;
ByteDance(字节跳动)
;
Beijing Academy of Blockchain and Edge Computing(北京区块链与边缘计算研究院)
;
Renmin University of China(中国人民大学)
机构
*
Saarland University(萨尔大学)
;
Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)
;
University of Cambridge(剑桥大学)
;
University of Edinburgh(爱丁堡大学)
;
Zhejiang University(浙江大学)
;
Tencent YouTu Lab(腾讯优图实验室)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
机构
*
Northeastern University(东北大学)
;
University of Notre Dame(Notre Dame 大学)
;
University of Waterloo(滑铁卢大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Adobe(Adobe公司)
;
Microsoft Research Asia(微软亚洲研究院)
机构
*
The Hong Kong University of Science(香港科学与技术大学)
;
Inner Mongolia University(内蒙古大学)
;
Beihang University(北京航空航天大学)
;
Queen Mary University of London(伦敦玛丽女王大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
National University of Singapore(新加坡国立大学)
;
University of Surrey(萨里大学)
;
University of Rochester(罗切斯特大学)
;
Independent Researcher(独立研究者)
专题命中
评测与基准
:LLM(abstract_cn);large language model(abstract);language model(abstract);instruction tuning(abstract)
Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs
测试黑箱:面向消费者的健康大语言模型独立评估的结构性障碍
Rahul Gorijavolu, Kaushik Madapati, Pritika Vig, Rawan Abulibdeh, Nikhil Jaiswal, Mahri Kadyrova, Zeamanuel Hailu Tesfaye, Charles Senteio, Paula Maurutto, Leo Anthony Celi
机构
*
Massachusetts Institute of Technology(麻省理工学院)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
Toronto General Hospital, University Health Network(多伦多综合医院,大学健康网络)
;
McGill University(麦吉尔大学)
;
University of Toronto(多伦多大学)
;
Independent Researcher(独立研究者)
;
Rutgers University(罗格斯大学)
;
Beth Israel Deaconess Medical Center(贝斯以色列女执事医疗中心)
;
Harvard T.H. Chan School of Public Health(哈佛大学陈曾熙公共卫生学院)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics
支持向量评分准则:弥合自生成与人工评分准则之间的差距
Mengyuan Sun, Yu Li, Zhuohao Yu, Shikun Zhang, Wei Ye
机构
*
National Engineering Research Center for Software Engineering, Peking University(北京大学软件工程国家工程研究中心)
;
University of Science and Technology of China(中国科学技术大学)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL
CommentsWe have curated a paper list on RAG security in https://github.com/TreeAI-Lab/Awesome-RAG-Security, and we warmly welcome authors who wish to have their new work included to contact us via email