Beyond Textual Knowledge-Leveraging Multimodal Knowledge Bases for Enhancing Vision-and-Language Navigation
超越文本知识的多模态知识库用于增强视觉-语言导航
Dongsheng Yang, Yinfeng Yu, Liejun Wang
机构
*
School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
;
Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语言认知计算联合国际研究实验室)
机构
*
University of Washington(华盛顿大学)
;
National University of Singapore(新加坡国立大学)
;
Clemson University(克莱姆森大学)
;
Drexel University(德雷塞尔大学)
;
Microsoft Research(微软研究院)
Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
Goal-VLA: 图像生成视觉语言模型作为对象中心世界模型,赋能零样本机器人操作
Haonan Chen, Jingxiang Guo, Bangjun Wang, Tianrui Zhang, Xuchuan Huang, Boren Zheng, Yiwen Hou, Chenrui Tie, Jiajun Deng, Lin Shao
机构
*
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
The HKU Musketeers Foundation Institute of Data Science, The University of Hong Kong(香港大学数据科学研究院)
;
Yuanpei College, Peking University(北京大学元培学院)
;
Department of Automation, Tsinghua University(清华大学自动化系)
机构
*
Guangdong Provincial Key Laboratory of Space-Aerial Networking and Intelligent Sensing, Harbin Institute of Technology, Shenzhen(广东省空天网络与智能感知重点实验室,哈尔滨工业大学(深圳))
;
Peng Cheng Laboratory (PCL), Shenzhen(鹏城实验室)
;
Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计系)
;
School of Computer Science and Engineering, Central South University, Changsha(中南大学计算机科学与工程学院)
机构
*
Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建筑工程学院)
;
Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系)
;
Google(谷歌)
Assessing Vision-Language Models for Perception in Autonomous Underwater Robotic Software
评估用于自主水下机器人软件中的视觉-语言模型感知能力
Muhammad Yousaf, Aitor Arrieta, Shaukat Ali, Paolo Arcaini, Shuai Wang
机构
*
Simula Research Laboratory(Simula研究实验室)
;
Oslo Metropolitan University(奥斯陆城市大学)
;
Mondragon University(蒙德拉贡大学)
;
National Institute of Informatics(国立信息学研究所)
;
Det norske Veritas (DNV)(挪威船级社(DNV))
Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
通过因果发现和双模态安全子空间投影诊断和修复视觉-语言模型中的不安全通道
Jinhu Fu, Yihang Lou, Qingyi Si, Shudong Zhang, Yan Bai, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Huawei Technologies Ltd.(华为技术有限公司)
;
Peking University(北京大学)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models
过度思考导致幻觉:追踪视觉语言模型中的混杂因素传播
Abin Shoby, Ta Duc Huy, Tuan Dung Nguyen, Minh Khoi Ho, Qi Chen, Anton van den Hengel, Phi Le Nguyen, Johan W. Verjans, Vu Minh Hieu Phan
机构
*
Australian Institute for Machine Learning, University of Adelaide(阿德莱德大学澳大利亚机器学习研究所)
;
Hanoi University of Science and Technology(河内科技大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中
幻觉与鲁棒性
:vision language model(title,abstract);分类 cs.CV
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
SECOND:通过选择性和对比性解码缓解视觉-语言模型中的感知幻觉
Woohyeon Park, Woojin Kim, Jaeik Kim, Jaeyoung Do
机构
*
Department of Electrical and Computer Engineering, Seoul National University(首尔大学电气与计算机工程系)
;
Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔大学人工智能跨学科项目)
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
所有权的回声:对抗引导的双注入用于MLLMs中的版权保护
Chengwei Xia, Fan Ma, Ruijie Quan, Yunqiu Xu, Kun Zhan, Yi Yang
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
ImagenWorld: 通过可解释的人类评估对图像生成模型进行压力测试,针对开放性现实任务
Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani, Paria Khoshtab, Wei-Chieh Sun, Parnian Fazel, Zhi Rui Tam, Thomas Chong, Edisy Kin Wai Chan, Donald Wai Tong Tsang, Chiao-Wei Hsu, Ting Wai Lam, Ho Yin Sam Ng, Chiafeng Chu, Chak-Wing Mak, Keming Wu, Hiu Tung Wong, Yik Chun Ho, Chi Ruan, Zhuofeng Li, I-Sheng Fang, Shih-Ying Yeh, Ho Kei Cheng, Ping Nie, Wenhu Chen
机构
*
University of Waterloo(滑铁卢大学)
;
G-G-G
;
Comfy Org
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Independent(独立研究者)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
针对LAION-400M的人为中心标注:审查偏见及其在模型中的转移
Leander Girrbach, Stephan Alaniz, Genevieve Smith, Trevor Darrell, Zeynep Akata
机构
*
Technical University of Munich, Munich Center for Machine Learning (MCML), MDSI(慕尼黑工业大学,慕尼黑机器学习中心(MCML),MDSI)
;
LTCI, Télécom Paris, Institut Polytechnique de Paris, France(法国巴黎理工学院,巴黎电信学院,LTCI)
;
Helmholtz Munich(亥姆霍兹慕尼黑中心)
;
University of California, Berkeley(加州大学伯克利分校)
From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
从探索到利用:一种两阶段熵RLVR方法用于噪声容忍的多模态大语言模型训练
Donglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang, Jinpeng Chen, Wenao Ma, Zhijian Hou, Mengyang Wu, Xiaolei Li, Senkang Hu, Ziyi Guan, Jason Chun Lok Li, Lai Man Po
机构
*
Independent Researcher(独立研究者)
;
The Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong(香港城市大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
University of Hong Kong(香港大学)
专题命中
VLM训练与架构
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
机构
*
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)
;
King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
;
Pengcheng Laboratory(鹏城实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI、cs.LG
Synergizing Discriminative Exemplars and Self-Refined Experience for MLLM-based In-Context Learning in Medical Diagnosis
融合判别示例与自我优化经验的医学诊断基于MLLM的上下文学习
Wenkai Zhao, Zipei Wang, Mengjie Fang, Di Dong, Jie Tian, Lingwei Zhang
机构
*
CAS Key Laboratory of Molecular Imaging, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所分子影像重点实验室)
;
School of Software, North University of China(中北大学软件学院)
专题命中
VLM训练与架构
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV