ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation
ARP:通过视觉对齐与迭代细化增强机器人操作的量化技能抽象
Yuntian Wang, Zesheng Jia, Yuhui Duan, Qibing Wang, Yang Liu, Song Wang, Siao Liu, Jin Wang
机构
*
School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院)
;
Key Laboratory of General Artificial Intelligence and Large Models in Provincial Universities, Soochow University(苏州大学省高校通用人工智能与大模型重点实验室)
;
College of Mechanical and Electronic Engineering, China Jiliang University(中国计量大学机电工程学院)
;
College of Electronics and Information Engineering, Tongji University(同济大学电子与信息工程学院)
;
Leju Robotics(乐聚机器人)
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Beijing University of Chemical Technology(北京化工大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Beijing Institute of Technology (Zhuhai)(北京理工大学(珠海))
;
Tencent Hy(腾讯(深圳))
;
Peng Cheng Laboratory(鹏城实验室)
Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs
我们的基准测试足够吗?多模态大语言模型持续学习分析
Van-Tuan Tran, Shruthi Gowda, Merim Dzaferagic, Marco Ruffini
机构
*
School of Computer Science and Statistics, Trinity College Dublin(都柏林圣三一学院计算机科学与统计学院)
;
Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)
专题命中
其他VLM
:MLLM(summary_cn,abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG
Vision-Language Models are Fragile Multilingual Associators
视觉-语言模型是脆弱的多语言关联器
Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi, Umapada Pal
机构
*
Manipal University Jaipur(斋浦尔马尼帕尔大学)
;
University of North Carolina Charlotte(北卡罗来纳大学夏洛特分校)
;
University of Salford(索尔福德大学)
;
University of Manchester(曼彻斯特大学)
;
Indian Statistical Institute Kolkata(印度统计研究所加尔各答分所)
机构
*
University College Dublin(都柏林大学学院)
;
School of Computer Science, University College Dublin(都柏林大学学院计算机科学学院)
;
The Third Affiliated Hospital of Southern Medical University(南方医科大学第三附属医院)
;
Dublin City University(都柏林城市大学)
SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models
SAB-LVLM: 面向大型视觉-语言模型的重要性感知二值化
Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han
机构
*
State Key Laboratory of Robotics and Intelligent Systems(机器人学国家重点实验室)
;
Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Shanghai Jiao Tong University(上海交通大学)
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Nanyang Technological University(南洋理工大学)
;
Imperial College London(帝国理工学院)