fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery
fmxcoders: 分层掩码交叉编解码器用于跨层特征发现
Andreas D. Demou, Panagiotis Koromilas, James Oldfield, Yannis Panagakis, Mihalis A. Nicolaou
机构
*
The Cyprus Institute(塞浦路斯研究所)
;
University of Athens(雅典大学)
;
University of Oxford(牛津大学)
;
Archimedes AI/Athena Research Center(Archimedes AI/ Athena 研究中心)
;
University of Cyprus(塞浦路斯大学)
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment
SymptomAI:迈向日常症状评估的对话式AI代理
Joseph Breda, Fadi Yousif, Beszel Hawkins, Marinela Cotoi, Miao Liu, Ray Luo, Po-Hsuan Cameron Chen, Mike Schaekermann, Samuel Schmidgall, Xin Liu, Girish Narayanswamy, Samuel Solomon, Maxwell A. Xu, Xiaoran Fan, Longfei Shangguan, Anran Wang, Bhavna Daryani, Buddy Herkenham, Cara Tan, Mark Malhotra, Shwetak Patel, John B. Hernandez, Quang Duong, Yun Liu, Zach Wasson, Dimitrios Antos, Bob Lou, Matthew Thompson, Jonathan Richina, Anupam Pathak, Nichole Young-Lin, Jake Sunshine, Daniel McDuff
机构
*
Google Research(谷歌研究)
;
Google DeepMind(谷歌DeepMind)
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
EmoS:一种高保真多模态基准,用于细粒度流式情感理解
Pengze Guo, Jingxi Liang, Zhiwen Xie, Qifeng Wang, Derek F. Wong
机构
*
NLP(自然语言处理)
;
CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学学院,澳门大学CT实验室)
;
School of Computer Science, Central China Normal University(Central China Normal University计算机科学学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL
CommentsWe voluntarily withdraw this manuscript. Extensive post-submission testing shows the method lacks the originally reported generality and effectiveness. The benchmark metrics originally designed are inadequate for assessing existing model editing algorithms. To avoid misleading the community, we have decided to withdraw this paper and will not release an updated version.
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
AgentCollabBench: 评估好代理为何会成为差合作者
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia, Tanzila Khan, Kainat Raisa Hossain, Nehaa Shri, Shubhrangshu Debsarkar, Humayra Tasnim, Gour Gupal Talukder Shawon, Debjoty Mitra, Sumaiya Ahmed Rani, Al Jami Islam Anik, Al Nafeu Khan
机构
*
University of Utah(犹他大学)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
University of Dhaka(达卡大学)
;
Vellore Institute of Technology(韦洛雷理工学院)
;
University of Virginia(弗吉尼亚大学)
;
Rajshahi University of Engineering and Technology(拉贾加赫尔工程与技术大学)
;
Shahjalal University of Science and Technology(沙赫jalal科学与技术大学)
;
BRAC University(BRAC大学)
;
Islamic University of Technology(伊斯兰技术大学)
;
Comilla University(科摩拉大学)
Comments3 Tables and 4 Figures. Companion data descriptor to the Unified Galaxy HI Rotation Curve Corpus v7 (Zenodo 10.5281/zenodo.19563417). v7: corrected Harris [Fe/H] description to spectroscopic metallicity (Carretta et al. scale) following communication from E. Carretta (INAF Bologna); added Kruijssen et al. (2019) as recommended ages source for v1.4
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
CDS4RAG:基于循环双序列的RAG超参数优化
Pengzhou Chen, Tao Chen
机构
*
School of Computer Science and Engineering, UESTC, Chengdu, China(电子科技大学计算机科学与工程学院,成都,中国)
;
IDEAS Lab, University of Birmingham, Birmingham, UK(伯明翰大学IDEAS实验室,英国布里斯托尔)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
AssayBench: 一个基于实验层面的虚拟细胞基准,用于LLMs和代理
Edward De Brouwer, Carl Edwards, Alexander Wu, Jenna Collier, Graham Heimberg, Xiner Li, Meena Subramaniam, Ehsan Hajiramezanali, David Richmond, Jan-Christian Hütter, Sara Mostafavi, Gabriele Scalia
ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
ThreatCore:一种用于显式和隐式威胁检测的基准测试
Davide Bruni, Carlo Bardazzi, Maurizio Tesconi
机构
*
Computer Science Department, University of Pisa, Italy(比萨大学计算机科学系)
;
Institute of Informatics and Telematics, National Research Council, Italy(意大利国家研究委员会信息与电信学研究院)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
任务感知校准:在大语言模型中的可证明最优解码
Tim Tomov, Dominik Fuchsgruber, Rajeev Verma, Stephan Günnemann
机构
*
School of Computation, Information & Technology, Technical University of Munich(慕尼黑技术大学计算、信息与技术学院)
;
Munich Data Science Institute(慕尼黑数据科学研究所)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
University of Amsterdam(阿姆斯特丹大学)
Timo Stoll, Chendi Qian, Ben Finkelshtein, Ali Parviz, Darius Weber, Fabrizio Frasca, Hadar Shavit, Antoine Siraudin, Arman Mielke, Marie Anastacio, Erik Müller, Maya Bechler-Speicher, Michael Bronstein, Mikhail Galkin, Holger Hoos, Mathias Niepert, Bryan Perozzi, Jan Tönshoff, Christopher Morris
机构
*
RWTH Aachen University(亚琛RWTH大学)
;
University of Oxford(牛津大学)
;
Mila – Quebec AI Institute(魁北克AI研究所)
;
Technion - Israel Institute of Technology(技术学院-以色列理工学院)
;
ETAS Research University of Stuttgart(斯图加特大学ETAS研究所)
;
Google Research(谷歌研究)
;
Microsoft Research(微软研究院)