CommentsTo appear at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), Vienna, Austria, July 2025, https://2025.aclweb.org/
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Qing Yu, Go Irie, Yixuan Li, Hai Li, Ziwei Liu, Kiyoharu Aizawa
机构
*
The University of Tokyo(东京大学)
;
S-Lab, Nanyang Technological University(南洋理工大学S实验室)
;
Duke University(杜克大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
LY Corporation(LY公司)
;
Tokyo University of Science(东京科学大学)
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
Jihyoung Jang, Minwook Bae, Minji Kim, Dilek Hakkani-Tur, Hyounghun Kim
机构
*
Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院)
;
Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)
;
Artificial Intelligence Graduate School, UNIST(UNIST人工智能研究生院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Comments6 pages, 4, figures. This work has been submitted for publication to the 12th Annual IEEE EMBS International Conference on Neural Engineering (https://neuro.embs.org/2025/)
机构
*
School of Software(软件学院)
;
Dalian University of Technology(大连理工大学)
;
School of Foreign Languages and School of Software(外语学院和软件学院)
;
Faculty of Business Administration(商学院)
;
University of Macau(澳门大学)
;
School of Computing Technologies(计算技术学院)
;
RMIT University(皇家墨尔本理工大学)
BETTY Dataset: A Multi-modal Dataset for Full-Stack Autonomy
Micah Nye, Ayoub Raji, Andrew Saba, Eidan Erlich, Robert Exley, Aragya Goyal, Alexander Matros, Ritesh Misra, Matthew Sivaprakasam, Marko Bertogna, Deva Ramanan, Sebastian Scherer
机构
*
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
;
University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)
;
University of Waterloo(滑铁卢大学)
;
University of Pittsburgh(匹兹堡大学)
Advancing Conversational Diagnostic AI with Multimodal Reasoning
Khaled Saab, Jan Freyberg, Chunjong Park, Tim Strother, Yong Cheng, Wei-Hung Weng, David G. T. Barrett, David Stutz, Nenad Tomasev, Anil Palepu, Valentin Liévin, Yash Sharma, Roma Ruparel, Abdullah Ahmed, Elahe Vedadi, Kimberly Kanada, Cian Hughes, Yun Liu, Geoff Brown, Yang Gao, Sean Li, S. Sara Mahdavi, James Manyika, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Pushmeet Kohli, S. M. Ali Eslami, Joëlle Barral, Adam Rodman, Vivek Natarajan, Mike Schaekermann, Tao Tu, Alan Karthikesalingam, Ryutaro Tanno
机构
*
Google(谷歌)
;
DeepMind(深Mind)
;
Google Research(谷歌研究院)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
Hanlei Zhang, Zhuohang Li, Yeshuang Zhu, Hua Xu, Peiwu Wang, Haige Zhu, Jie Zhou, Jinchao Zhang
机构
*
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院)
;
Kennesaw State University(凯斯西储大学)
REMEMBER: Retrieval-based Explainable Multimodal Evidence-guided Modeling for Brain Evaluation and Reasoning in Zero- and Few-shot Neurodegenerative Diagnosis
Duy-Cat Can, Quang-Huy Tang, Huong Ha, Binh T. Nguyen, Oliver Y. Chén
CommentsThis paper has been accepted for presentation at International Workshop on Spoken Dialogue Systems Technology 2025 (IWSDS 2025) and represents the author's version of the work
MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research
James Burgess, Jeffrey J Nirschl, Laura Bravo-Sánchez, Alejandro Lozano, Sanket Rajan Gupte, Jesus G. Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M. Hasan, Alexandra Johannesson, William D. Leineweber, Malvika G Nair, Ridhi Yarlagadda, Connor Zuraski, Wah Chiu, Sarah Cohen, Jan N. Hansen, Manuel D Leonetti, Chad Liu, Emma Lundberg, Serena Yeung-Levy