Shrinking the Generation-Verification Gap with Weak Verifiers
缩小生成-验证差距的弱验证器
Jon Saad-Falcon, E. Kelly Buchanan, Mayee F. Chen, Tzu-Heng Huang, Brendan McLaughlin, Tanvir Bhathal, Shang Zhu, Ben Athiwaratkun, Frederic Sala, Scott Linderman, Azalia Mirhoseini, Christopher Ré
机构
*
Stanford University(斯坦福大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
Together AI
Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs
推理错误在大语言模型(LLMs)残差流轨迹中具有特定区域与方向
Hamed Damirchi, Ignacio Meza De la Jara, Damith Ranasinghe, Yuhang Liu, Javen Shi
机构
*
Australian Institute for Machine Learning(澳大利亚机器学习研究所)
;
Adelaide University(阿德莱德大学)
;
Naval Group Pacific(太平洋海军集团)
;
Responsible AI Research Centre(负责任人工智能研究中心)
Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
;
Wuhan University(武汉大学)
;
Sun Yat-sen University(中山大学)
;
Xidian University(西安电子科技大学)
;
Southeast University(东南大学)
SemDINO: Foundation Prior-Guided Cross-Temporal Semantic Alignment Network for Remote Sensing Change Detection
SemDINO: 一种基于DINOv3的跨时间语义对齐变化检测网络
Xinyu Tong, Meihua Zhou, Jinxiao Sun, Zaiyan Zhang, Hongruixuan Chen, Lei Wang
机构
*
Xinjiang Institute of Ecology and Geography, Chinese Academy of Sciences(中国科学院新疆生态与地理研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Computer Science, Xiangtan University(湘潭大学计算机科学学院)
;
College of Information and Communication Engineering, Harbin Engineering University(哈尔滨工程大学信息与通信工程学院)
CommentsThis paper has already been accepted and presented at the IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026) from July 31 to August 2, 2026, in Seattle, WA, USA
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
BenHalluEval:孟加拉语大语言模型的多任务幻觉评估框架
Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Ishmam Tashdeed, Md Taukir Azam Chowdhury
机构
*
Department of Computer Science and Engineering, Islamic University of Technology(伊斯兰科技大学计算机科学与工程系)
;
Department of Computer Science and Engineering, University of California(加州大学计算机科学与工程系)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);prompting(abstract)
Commentsv2: substantially revised and retitled; withdraws the variance-compression and transgender-suppression results and adds a pre-registered decoding control. 21 pages, 8 figures, 7 tables