PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation
PressMimic: 压力引导的动作捕捉与控制用于人形机器人模仿
Yi Lu, Shenghao Ren, Tianyu Xiong, Zhaoxiang Li, Jiaqi Li, He Zhang, Tao Yu, Qiu Shen, Xun Cao
机构
*
School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院)
;
Key Laboratory of Optoelectronic Devices and Systems with Extreme Performances of MOE, Nanjing University(南京大学极端性能光电技术与系统教育部重点实验室)
;
BNRist, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
MKG-RAG-Bench:多模态知识图谱增强生成中的检索基准
Xiaochen Wang, Bao Hoang, Han Liu, Ting Wang, Fenglong Ma
机构
*
The Pennsylvania State University(宾夕法尼亚州立大学)
;
Michigan State University(密歇根州立大学)
;
Dalian University of Technology(大连理工大学)
;
Stony Brook University(石溪大学)
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
通过自主经验探索与事后经验利用赋能GUI智能体任务规划
Tianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
机构
*
Department of Artificial Intelligence, School of Informatics, Xiamen University(厦门大学信息学院人工智能系)
;
Department of Computer Science, Aberystwyth University(阿伯里斯特威斯大学计算机科学系)
Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations
安全关键场景下基于蒸馏视觉语言模型的事件自适应运动规划
Zhenwei Huang, Changsheng You, Shuai Wang, Chao Zhou, Wei Xu, Yi Gong
机构
*
Southern University of Science and Technology(南方科技大学)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
Manifold Tech Limited(曼孚科技股份有限公司)
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
6根手指,1个肾脏:自然对抗性医学图像揭示视觉语言模型的关键弱点
Leon Mayer, Piotr Kalinowski, Caroline Ebersbach, Marcel Knopp, Tim Rädsch, Evangelia Christodoulou, Annika Reinke, Fiona R. Kolbinger, Lena Maier-Hein
机构
*
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部门)
;
Medical Faculty, Heidelberg University(海德堡大学医学院)
;
Faculty of Mathematics and Computer Science, Heidelberg University(海德堡大学数学与计算机科学学院)
;
HIDSS4Health - Helmholtz Information and Data Science School for Health, Karlsruhe/Heidelberg(HIDSS4Health - 哈勃-马克斯信息与数据科学健康学院,卡尔斯鲁厄/海德堡)
;
Helmholtz Imaging, German Cancer Research Center (DKFZ)(哈勃-马克斯成像,德国癌症研究中心(DKFZ))
;
Engineering Faculty, Heidelberg University(海德堡大学工程学院)
;
School of Computation, Information and Technology, TUM(技术大学(TUM)计算、信息与技术学院)
;
Weldon School of Biomedical Engineering, Purdue University(普渡大学韦尔登生物医学工程学院)
;
Department of Visceral, Thoracic and Vascular Surgery, University Hospital and Faculty of Medicine Carl Gustav Carus, TUD Dresden University of Technology(visceral、胸腔和血管外科部门,技术大学(TUD)德累斯顿大学医院和医学院)
;
National Center for Tumor Diseases (NCT), NCT Heidelberg, a partnership between DKFZ and University Hospital Heidelberg(肿瘤疾病国家中心(NCT),海德堡NCT,DKFZ与海德堡大学医院之间的合作)
;
Heidelberg University Hospital, Surgical Clinic, Surgical AI Research Group(海德堡大学医院,外科诊所,外科人工智能研究组)
;
Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE(Mohamed Bin Zayed人工智能大学(MBZUAI),阿布扎赫,阿拉伯联合酋长国)
机构
*
Peking University(北京大学)
;
University of Pennsylvania(宾夕法尼亚大学)
;
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
Virginia Tech(弗吉尼亚理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
De Artificial Intelligence Lab(人工智能实验室)
Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models
跨模态对抗扩散:文本、视觉与视觉-语言模型的攻击、防御与评估融合综述
Abrar Alotaibi, Moataz Ahmed
机构
*
Information and Computer Science Department, King Fahd University of Petroleum & Minerals(国王法赫德石油矿物大学信息与计算机科学系)
;
College of Computer Science and Information Technology, Imam Abdulrahman Bin Faisal University(伊玛目阿卜杜勒拉赫曼·本·法伊塞尔大学计算机科学与信息科技学院)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence, King Fahd University of Petroleum & Minerals(SDAIA-KFUPM人工智能联合研究中心,国王法赫德石油矿物大学)
机构
*
Brain-inspired Cognitive AI Lab, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所类脑认知智能实验室)
;
Beijing Key Laboratory of Safe AI and Superalignment(北京市安全人工智能与超级对齐重点实验室)
;
Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究所)
;
Gaoling School of AI, Renmin University of China(中国人民大学高瓴人工智能学院)
;
University of Chinese Academy of Sciences (UCAS)(中国科学院大学)
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
TOPS:通过构建令牌最优保留集实现高效多模态大语言模型推理的第一性原理视觉令牌剪枝
Tinghao Wang, Yichen Guo, Rui Huang, Zheng Lu, Qizhe Zhang, Chenxi Li, Yuan Zhang, Jiajun Cao, Zhirong Shen, Yaosong Du, Guangyan Gan, Wenya Wang, Lin William Cong, Shanghang Zhang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Nanyang Technological University(南洋理工大学)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
专题命中
VLM训练与架构
:MLLM(title,abstract);LLaVA(summary_cn,abstract);multimodal large language model(abstract);分类 cs.AI
CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs
CheXanatomy: 面向胸部X光片的解剖感知视觉-语言建模
Sergios Gatidis, Curtis Langlotz, Christian Bluethgen
机构
*
Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University(斯坦福大学医学与影像人工智能中心)
;
Department of Radiology, Stanford University(斯坦福大学放射学系)
机构
*
Chinese University of Hong Kong(香港中文大学)
;
Westlake University(西湖大学)
;
Southern Medical University(南方医科大学)
;
Jiangnan University(江南大学)
;
Dalian University of Technology(大连理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Copenhagen(哥本哈根大学)
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
AIDEN:面向视障人士的AI助手设计与初步研究
Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer, Bessie Dominguez-Dager, Alvaro Belmonte-Baeza, Mónica Pina-Navarro, Francisco Morillas-Espejo, Felix Escalona, Miguel Cazorla
机构
*
Institute for Computer Research, University of Alicante(计算机研究所,阿利坎特大学)
;
ETH Zurich(苏黎世联邦理工学院)
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
ReasonCLIP-58M: CLIP的视觉基础常识推理监督
Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto, Shi Qiu, Jamal Bentahar, Naveed Akhtar, Mubarak Shah
机构
*
Khalifa University(卡利法大学)
;
University of Western Australia(西澳大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Melbourne(墨尔本大学)
;
University of Central Florida(佛罗里达中央大学)
专题命中
VLM训练与架构
:LLaVA(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI