Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality
跨模态掩码组合概念建模以增强视觉-语言组合性
Wei Li, Zhen Huang, Xinmei Tian
机构
*
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学,教育部脑启发智能感知与认知重点实验室)
;
Independent Researcher(独立研究员)
机构
*
Shanghai Key Lab of Intelligent Information Processing, Fudan University(复旦大学上海智能信息处理重点实验室)
;
School of Computer Science, Fudan University(复旦大学计算机科学技术学院)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
Youtu Lab, Tencent(腾讯优图实验室)
;
Meta AI
;
Shanghai AI Laboratory(上海人工智能实验室)
M*: A Modular, Extensible, Serving System for Multimodal Models
M*: 一个模块化、可扩展的多模态模型服务系统
Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang
机构
*
Stanford University(斯坦福大学)
;
University of Washington(华盛顿大学)
;
Carnegie Mellon University(卡内基梅隆大学)
Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences
顺序很重要:LVLMs作为图像序列时间推理的评判者
Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes
机构
*
University of Bologna(博洛尼亚大学)
;
NOVA School of Science and Technology(NOVA科技学院)
;
NOVA Laboratory for Computer Science and Informatics(NOVA计算机科学与信息实验室)
机构
*
Shandong University(山东大学)
;
City University of Hong Kong(香港城市大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Shandong Jianzhu University(山东建筑大学)
机构
*
South China University of Technology(华南理工大学)
;
Westlake University(西湖大学)
;
Johns Hopkins University(约翰·霍普金斯大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
机构
*
Dalian University of Technology(大连理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Zhejiang University(浙江大学)
;
WeChat, Tencent Inc.(腾讯微信)
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
推进面向艺术字的场景文本识别:数据集与方法
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Haojie Zhang, Chong Sun, Chen Li, Jing Lyu, Zhineng Chen
机构
*
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所)
;
Shanghai Key Laboratory of Multimodal Embodied AI, Fudan University(复旦大学上海市多模态具身人工智能重点实验室)
;
WeChat Vision, Tencent Inc.(腾讯微信视觉团队)
;
South China University of Technology(华南理工大学)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
S-Lab, Nanyang Technological University(南洋理工大学S实验室)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
University of Science and Technology of China(中国科学技术大学)
;
Stanford University(斯坦福大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
CPII under InnoHK(InnoHK下的CPII)
;
Adobe Research(Adobe研究)