Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models
Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang, Hang Yu
机构
*
Anhui Polytechnic University (AHPU)(安徽工程大学)
;
Shanghai University (SHU)(上海大学)
;
Shanghai Jiao Tong University (SJTU)(上海交通大学)
;
Sichuan University (SCU)(四川大学)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
Zahraa Al Sahili, Ioannis Patras, Matthew Purver
机构
*
School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院)
;
Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)
专题命中
视觉推理
:multimodal large language model(title);分类 cs.CV、cs.AI、cs.LG
Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
Houjian Yu, Zheming Zhou, Min Sun, Omid Ghasemalizadeh, Yuyin Sun, Cheng-Hao Kuo, Arnie Sen, Changhyun Choi
机构
*
Department of Electrical and Computer Engineering, Univ. of Minnesota(电气与计算机工程系)
;
Amazon Lab126(亚马逊实验室126)
;
National Tsing Hua University(国立清华大学)
专题命中
视觉推理
:grounding(title,abstract)
CommentsAccepted to 2025 IEEE-RAS 24th International Conference on Humanoid Robots
Robotic Perception with a Large Tactile-Vision-Language Model for Physical Property Inference
Zexiang Guo, Hengxiang Chen, Xinheng Mai, Qiusang Qiu, Gan Ma, Zhanat Kappassov, Qiang Li, Nutan Chen
机构
*
College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国)
;
Sino-German College of Intelligent Manufacturing, Shenzhen Technology University, China(中德智能制造学院,深圳科技大学,中国)
;
Robotics Department, Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, Kazakhstan(机器人系,智能系统与人工智能研究所(ISSAI),纳扎尔巴耶夫大学,哈萨克斯坦)
;
Foundation Robotics Labs, Germany(基础机器人实验室,德国)
专题命中
视觉推理
:vision-language model(title,abstract)
CommentsThis paper has been accepted by the 2025 International Conference on Climbing and Walking Robots (CLAWAR). These authors contributed equally to this work: Zexiang Guo, Hengxiang Chen, Xinheng Mai
机构
*
College of Informatics, Harbin Institute of Technology(信息学院,哈尔滨工业大学)
;
Key Laboratory of Forest and Grassland Fire Risk Prevention, Ministry of Emergency Management, China Fire and Rescue Institute(森林和草原火灾风险预防重点实验室,应急管理部,中国消防救援学院)
;
Southern University of Science and Technology(南方科技大学)
专题命中
视觉推理
:multimodal large language model(title,abstract)
HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
Muhammad Haris Khan, Miguel Altamirano Cabrera, Dmitrii Iarchuk, Yara Mahmoud, Daria Trinitatova, Issatay Tokmurziyev, Dzmitry Tsetserukou
机构
*
Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室、数字工程中心、斯克尔科沃科学与技术研究所)
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
Mohammed Saidul Islam, Raian Rahman, Ahmed Masry, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, Enamul Hoque