SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity
SSMNBench: 通过单视图充分性与多视图必要性诊断基于图像的跨视角人-物理解
Tianchen Guo, Chen Liu, Ling Chen, Xin Yu
机构
*
The University of Queensland(昆士兰大学)
;
Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所)
;
University of Technology Sydney(悉尼科技大学)
;
Follow Me AI Pty LTD(Follow Me AI有限公司)
专题命中
视觉问答
:MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV
机构
*
Centennial High School, Frisco, Texas, USA(Centennial High School, Texas, USA)
;
Lebanon Trail High School, Frisco, Texas, USA(Lebanon Trail High School, Texas, USA)
;
West Windsor-Plainsboro High School, Princeton Junction, New Jersey, USA(West Windsor-Plainsboro High School, New Jersey, USA)
;
Algoverse AI Research, Palo Alto, California, USA(Algoververse AI Research, California, USA)
专题命中
视觉推理
:VLM(title_cn);vision-language model(abstract);vision language model(abstract);visual reasoning(abstract)
Comments14 pages, 4 figures, 8 tables. Presented at the 39th Conference on Neural Information Processing Systems Workshop: VLM4RWD. Presented at the 43th International Conference on Machine Learning Workshops: ICML 2026 CTB, ICML 2026 FAGEN, ICML 2026 EMM-QA. Authors Aahana Basappa and Pranay Goel contributed equally. Code: https://github.com/AahanaB24/AMVICC, Data: https://doi.org/10.5281/zenodo.17646068
TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
TriViewBench: 多视图结构推理中受控复杂度缩放
Yu-Yang Chen, Lan-Zhe Guo
机构
*
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
专题命中
视觉推理
:visual reasoning(abstract);visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract_cn)
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
物理问题场景图:文本到视频生成中物理合理性的细粒度评估
Atin Pothiraj, Jaemin Cho, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
AI2(艾伦人工智能研究所)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
PhaseWin: An Efficient Search Algorithm for Faithful Visual Attribution
PhaseWin:一种用于忠实视觉归因的高效搜索算法
Zihan Gu, Junchi Zhang, Li Liu, Xiaochun Cao, Hua Zhang
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
;
Shanghai Center for Mathematical Sciences, Fudan University(复旦大学上海数学中心)
;
College of Electronic Science and Technology, National University of Defense Technology(国防科技大学电子科学学院)
;
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区网络空间安全学院)