Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
超越正确性:混合思维多模态大语言模型(MLLM)的响应行为基准测试与对齐
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Large Language Model Department, Tencent(腾讯大语言模型部) ; University of Electronic Science and Technology of China(电子科技大学) ; Hong Kong University of Science and Technology(香港科技大学) ; Zhongguancun Academy(中关村学院)
专题命中 视觉定位与Grounding :MLLM(title_cn,summary_cn);grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 该研究针对混合思维 MLLM 的思维与非思维模式响应错位问题,构建 PatternEval 基准并开发 PatternRL 方法,可减轻跨模式错位且任务性能损失极小。
Comments 8 tables and 6figures