Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能
Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh
机构
*
Computer Science Department, University of California, Los Angeles(加州大学洛杉矶分校计算机科学系)
;
Mechanical and Aerospace Engineering Department, University of California, Los Angeles(加州大学洛杉矶分校机械与航空航天工程系)
专题命中
VLM训练与架构
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi
机构
*
of Computer Science, University of Southern California, Los Angeles, CA, USA
;
Computer Engineering, University of Southern California, Los Angeles, CA, USA
专题命中
VLM训练与架构
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG
CommentsAccepted to IEEE International Conference on Data Mining (ICDM) 2025
Cross-Modal Instructions for Robot Motion Generation
William Barron, Xiaoxiang Dong, Matthew Johnson-Roberson, Weiming Zhi
机构
*
College of Connected Computing, Vanderbilt University(连接计算学院,范德比尔特大学)
;
Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)
;
School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Zeyi Sun, Ziyang Chu, Pan Zhang, Tong Wu, Xiaoyi Dong, Yuhang Zang, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Tsinghua University(清华大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
CPII under InnoHK(创新香港科技促进会)
;
MThreads AI
专题命中
VLM训练与架构
:vision-language model(abstract);vision language model(abstract);分类 cs.CV、cs.AI、cs.LG
CommentsIn Figure 2, the correlation coefficient and the scatter plot do not match. I calculated this correlation using two sets of settings. I used the scatter plot from setting A, but accidentally wrote the correlation coefficient, r, from setting B
NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
Shiyu Liu, Lianlei Shan
机构
*
School of Electrical and Electronic Engineering(电气与电子工程学院)
;
Nanyang Technological University(南洋理工大学)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
VLM训练与架构
:visual language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
Tianyu Chen, Xingcheng Fu, Yisen Gao, Haodong Qian, Yuecen Wei, Kun Yan, Haoyi Zhou, Jianxin Li
机构
*
SKLCCSE, School of Computer Science and Engineering, Beihang University, China(信息与通信工程学院,北京航空航天大学)
;
School of Software, Beihang University, China(软件学院,北京航空航天大学)
;
Key Lab of Education Blockchain and Intelligent Technology, Guangxi Normal University, China(教育区块链与智能技术重点实验室,广西师范大学)
;
Institute of Artificial Intelligence, Beihang University, Beijing, China(人工智能研究院,北京航空航天大学)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
Jiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning, Zhihui Zhu
机构
*
Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学)
;
Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学)
;
Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
Run Luo, Renke Shan, Longze Chen, Ziqiang Liu, Lu Wang, Min Yang, Xiaobo Xia
机构
*
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
National University of Singapore(新加坡国立大学)
;
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室)