URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
URSA:用于功利性逆合成评估的化学感知基准测试
Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Anton Morgunov, Arkadii Lin, Maksim Kuznetsov, Rim Shayakhmetov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
机构
*
Insilico Medicine AI Limited(Insilico Medicine AI有限公司)
;
Independent researcher(独立研究者)
;
Insilico Medicine Canada Inc.(Insilico Medicine加拿大公司)
;
Insilico Medicine Hong Kong Ltd.(Insilico Medicine香港有限公司)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
GuideMe:流视频中的多域任务指导与干预
Fang Liu, Jinpeng Chen, Ke Xu, Yuhao Liu, Huankang Guan, Xudong Lu, Bo Yang, Gerhard Hancke, Rui Liu, Rynson W. H. Lau
机构
*
City University of Hong Kong(香港城市大学)
;
Huawei Research(华为研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong (Dongguan)(香港城市大学(东莞))
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract)
CommentsThis paper is withdrawn due to significant methodological errors in the experimental design that fundamentally affect the validity of the results. The errors are not correctable within the current framework, and the conclusions can no longer be supported. We apologize for any inconvenience caused to readers
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Renmin University of China(中国人民大学)
;
Sun Yat-Sen University(中山大学)
;
Chinese Academy of Sciences, Institute of Automation(中国科学院自动化研究所)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Advanced Topic Modeling Techniques for Categorizing Software Vulnerabilities
用于软件漏洞分类的高级主题建模技术
Utkarsh Tiwari, Spoorthi M, Anirudh S, Nidhin Prabhakar T.
机构
*
Department of Computer Science & Engineering, Amrita School of Computing, Bengaluru, Amrita Vishwa Vidyapeetham, India(计算机科学与工程系,Amrita计算学院,班加罗尔,Amrita世界大学,印度)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
Comments10 pages, 10 figures. Accepted at the 16th International Conference on Computing, Communication and Networking Technologies (ICCCNT 2025), July 6-11, 2025, IIT Indore, Madhya Pradesh, India. IEEE proceedings
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
TeachObs:多模态教学观察与模型评估的人工验证基准
Yeil Jeong, Youngjin Yoo, Jiyoung Bae, Seobin Sohn, Hyejin Han, Jinseo Lee, Howard Scott, Unggi Lee
机构
*
Indiana University Bloomington(印第安纳大学布卢明顿分校)
;
Pai Chai University(培才大学)
;
Seoul National University(首尔国立大学)
;
Ewha Womans University(成均馆大学)
;
University of Wolverhampton(沃尔夫汉普顿大学)
;
Korea University Sejong Campus(韩国大学世宗校区)
TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking
TrendFact:自动事实核查中热点感知的基准
Xiaocheng Zhang, Xi Wang, Yifei Lu, Jianing Wang, Zhuangzhuang Ye, Mengjiao Bao, Peng Yan, Xiaohong Su
机构
*
Harbin Institute of Technology(哈尔滨工业大学)
;
National University of Defense Technology(国防科学技术大学)
;
Northeastern University(东北大学)
;
East China Normal University(华东师范大学)
;
Beihang University(北京航空航天大学)
;
Tsinghua University(清华大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
机构
*
Nanjing University(南京大学)
;
Australian National University(澳大利亚国立大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Nvidia(英伟达)