Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
学习聚焦与精确裁剪:一种带有信息缺口和接地损失的强化学习框架用于多模态大语言模型
Xuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen, Yao Liu, Yue Wu, Tao Gong, Qi Chu, Nenghai Yu
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络空间安全学院)
;
Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院)
;
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
通过强化学习提升多图像接地在MLLMs中的推理能力
Bob Zhang, Haoran Li, Tao Zhang, Jianan Li, Cilin Yan, Xikai Liu, Jiayin Cai, Yanbin Hao
机构
*
Xiaohongshu Inc.(小红书公司)
;
University of Science and Technology of China(中国科学技术大学)
;
Wuhan University(武汉大学)
;
Technical University of Munich(慕尼黑工业大学)
;
Hefei University of Technology(合肥工业大学)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
UVLab, Department of Computer Science, University of Warwick(华威大学计算机科学系UVLab)
;
Department of Automation, University of Cambridge(剑桥大学自动化系)
;
Department of Computer Science, The University of Sheffield(谢菲尔德大学计算机科学系)
FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios
FORGE:面向制造场景的细粒度多模态评估
Xiangru Jian, Hao Xu, Wei Pang, Xinjian Zhao, Chengyu Tao, Qixin Zhang, Xikun Zhang, Chao Zhang, Guanzhi Deng, Alex Xue, Juan Du, Tianshu Yu, Garth Tarr, Linqi Song, Qiuzhuang Sun, Dacheng Tao
机构
*
University of Waterloo(滑铁卢大学)
;
University of Sydney(悉尼大学)
;
Singapore Management University(新加坡管理大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Hunan University(湖南大学)
;
Nanyang Technological University(南洋理工大学)
;
Royal Melbourne Institute of Technology(皇家墨尔本理工大学)
;
City University of Hong Kong(香港城市大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
City University of Hong Kong Shenzhen Research Institute(香港城市大学深圳研究院)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
机构
*
Vermont Artificial Intelligence Lab, Department of Computer Science, University of Vermont(佛蒙特大学计算机科学系佛蒙特人工智能实验室)
;
Intelligent Machines Lab, Department of Artificial Intelligence, Information Technology University(信息技术大学人工智能系智能机器实验室)
;
Institute of Artificial Intelligence, University of Central Florida(中佛罗里达大学人工智能研究所)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments11 pages, 6 figures. This is the author's version of the article that appeared at the IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR) 2026
FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding
FashionStylist: 一个增强专家知识的多模态数据集用于时尚理解
Kaidong Feng, Zhuoxuan Huang, Huizhong Guo, Yuting Jin, Xinyu Chen, Yue Liang, Yifei Gai, Li Zhou, Yunshan Ma, Zhu Sun
机构
*
Yanshan University(燕山大学)
;
Central South University(中南大学)
;
Zhejiang University(浙江大学)
;
Southwest University(西南大学)
;
Singapore Management University(新加坡管理大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping
CLASP: 闭环异步空间感知用于开放词汇桌面物体抓取
Yiran Ling, Wenxuan Li, Siying Dong, Yize Zhang, Xiaoyao Huang, Jing Jiang, Ruonan Li, Jie Liu
机构
*
Harbin Institute of Technology(哈尔滨工业大学)
;
National Key Laboratory of Smart Farm Technologies and Systems(智慧农场技术与系统全国重点实验室)
;
Peng Cheng Laboratory(鹏城实验室)
;
Northeastern University(东北大学)
CommentsThe original title, "Abstract Argumentation with Subargument Relations," has been replaced by "Subargument Argumentation Frameworks: Separating Direct Conflict from Structural Dependency"
Diffusion-CAM: Faithful Visual Explanations for dMLLMs
扩散-CAM:面向dMLLMs的可信视觉解释
Haomin Zuo, Yidi Li, Luoxiao Yang, Xiaofeng Zhang
机构
*
Department of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知系)
;
Sun Yat-sen University(中山大学)
;
Northwestern University(西北大学)
;
Technion - Israel Institute of Technology(以色列理工学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.AI
Wei Chen, Qibin Zhao, John Paisley, Junmei Yang, Delu Zeng
机构
*
School of Mathematics, South China University of Technology(华南理工大学数学学院)
;
Tensor Learning Team, RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心张量学习团队)
;
Department of Electrical Engineering, Columbia University(哥伦比亚大学电气工程系)
;
School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息工程学院)