How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
辅助推理如何释放VLM中的GUI定位能力
Weiming Li, Yan Shao, Jing Yang, Yujing Lu, Ling Zhong, Yuhan Wang, Min Yu, Tongxiao Ruan, Manni Duan
机构
*
Zhejiang Lab(浙江实验室)
;
Hangzhou Research and Development Center(杭州研发中心)
;
China Mobile(中国移动)
;
Innovation Center of Yangtze River Delta(长江三角洲创新中心)
;
Zhejiang University(浙江大学)
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
CodeGraphVLP:代码规划器与语义图状态的结合用于非马尔可夫视觉-语言-动作模型
Khoa Vo, Sieu Tran, Taisei Hanyu, Yuki Ikebe, Duy Nguyen, Nghi D. Q. Bui, Minh Vu, Anthony Gunderman, Chase Rainwater, Anh Nguyen, Ngan Le
机构
*
University of Arkansas(亚拉巴马大学)
;
Max Planck Research School for Intelligent Systems and the University of Stuttgart(马克斯·普朗克智能系统研究学校和斯图加特大学)
;
Center of AI Research, VinUniversity(Vin大学人工智能研究中心)
;
TU Wien(维也纳技术大学)
;
University of Liverpool(利物浦大学)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
重新思考大语言模型作为评判器:通过语义能力不对称利用小语言模型进行表示作为评判器
Zhuochun Li, Yong Zhang, Ming Li, Yuelyu Ji, Yiming Zeng, Ning Cheng, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao, Daqing He
机构
*
Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
;
University of Pittsburgh(匹兹堡大学)
;
University of Maryland, College Park(马里兰大学学院公园分校)
;
University of Connecticut(康涅狄格大学)
专题命中
评测与基准
:LLM(title,abstract);language model(title,abstract);small language model(title);large language model(abstract)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
打破镜像:基于激活的LLM评估者自我偏好缓解方法
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer
机构
*
University of Virginia(弗吉尼亚大学)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Carnegie Mellon University(卡内基梅隆大学)
;
School of Computer Science(计算机科学学院)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);preference optimization(abstract)
Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
面向情感的提示方法用于共情LLM辅导
Shuangquan Feng, Laura Fleig, Ruisen Tu, Philip Chi, Edmund Bu, Melinda Ozel, Junhua Ma, Teng Fei, Virginia R. de Sa
机构
*
Neurosciences Graduate Program, University of California San Diego(加州大学圣地亚哥分校神经科学研究生项目)
;
Department of Cognitive Science, University of California San Diego(加州大学圣地亚哥分校认知科学系)
;
Department of Computer Science and Engineering, University of California San Diego(加州大学圣地亚哥分校计算机科学与工程系)
;
Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系)
;
Department of Mathematics, University of California San Diego(加州大学圣地亚哥分校数学系)
;
Face the FACS
;
Halıcıoğlu Data Science Institute, University of California San Diego(Halıcıoğlu数据科学研究所)
专题命中
评测与基准
:LLM(title,title_cn);prompting(title);large language model(abstract);language model(abstract)
CommentsAccepted at the ICML 2026 workshops "Statistical Frameworks for Uncertainty in Agentic Systems" and "Combining Theory and Benchmarks: Towards a Virtuous Cycle to Understand and Guarantee Foundation Model Performance". 13 pages, 9 figures
机构
*
University of Pittsburgh(匹兹堡大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of Notre Dame(诺特丹大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Washington(华盛顿大学)
;
Allen Institute for Artificial Intelligence(人工智能研究院)
;
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);post-training(abstract)
Culturally uneven urban perception in large language models
大型语言模型通过文化不平等的基线感知城市
Rong Zhao, Wanqi Liu, Zhizhou Sha, Nanxi Su, Yecheng Zhang, Ying Long
机构
*
Centre for Advanced Spatial Analysis (CASA), UCL, London, UK(高级空间分析中心(CASA),伦敦大学学院,英国)
;
School of Architecture, Tsinghua University, Beijing, China(清华大学建筑学院,北京,中国)
;
Department of Computer Science, UT Austin, Austin, TX, USA(得克萨斯大学奥斯汀分校计算机科学系,奥斯汀,德克萨斯,美国)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);prompting(abstract)
Comments74 pages, 2 figures, 4 tables. Hybrid systematic survey and conceptual framework on LLM evaluation and AI-safety failures, synthesizing 373 primary studies (2018-2026). Introduces the EvalSafetyGap framework (Instability Decomposition, Alignment Trilemma) and reports an exploratory ten-model audit. Submitted as a review/survey article; not currently under consideration elsewhere
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学)
;
School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
The Chinese University of Hong Kong, Shenzhen(香港大学深圳校区)
;
Southern University of Science and Technology(南方科技大学)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
大型语言模型的红队框架:以忠实性评估为例
Abrar Alotaibi, Raed Mughus, Moataz Ahmed
机构
*
King Fahd University of Petroleum & Minerals(法赫德国王石油矿产大学)
;
Imam Abdulrahman Bin Faisal University(伊玛目阿卜杜勒拉赫曼·本·费萨尔大学)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence(沙特数据与人工智能局-法赫德国王石油矿产大学人工智能联合研究中心)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI