Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
视觉-语言模型将头部方向误认为注视方向:非语言对话线索
Zory Zhang, Pinyuan Feng, Bingyang Wang, Tianwei Zhao, Suyang Yu, Qingying Gao, Hokin Deng, Ziqiao Ma, Yijiang Li, Dezhi Luo
机构
*
Brown University(布朗大学)
;
Columbia University(哥伦比亚大学)
;
Emory University(埃默里大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of Washington(华盛顿大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Michigan(密歇根大学)
;
UC San Diego(圣地亚哥大学)
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Bosch Corporate Research(博世企业研究)
;
King Abdullah University of Science and Technology(卡布斯大学)
机构
*
Institute of Trustworthy Embodied AI (TEAI)(可信具身人工智能研究院)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
OpenDriveLab, The University of Hong Kong(OpenDrive实验室,香港大学)
WALL-WM: Carving World Action Modeling at the Event Joints
WALL-WM:在事件关节处雕刻世界动作建模
Shalfun Li, Victor Yao, Charles Yang, Truth Qu, Regis Cheng, Ryan Yu, Howard Lu, Newton Von, Vincent Chen, Yohann Tang, Maeve Zhang, Ellie Ma, Gody Li, Sage Yang, Lorien Shu, J. W. Gao, Ethan Chen, Colin Ye, Yu Sun, Elise Mon, PS Zhang, Neo Li, Lily Li, James Wang, Ping Yang, Chris Pan, Lucy Liang, Hang Su, Roy Gan, Hao Wang, Qian Wang
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
WorldMemArena: 通过动作-世界交互评估多模态智能体记忆
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu, Yepeng Liu, Lin Long, Yichen Guo, Nuo Chen, Zhaotian Weng, Elena Kochkina, Simerjot Kaur, Charese Smiley, Xiaomo Liu, James Zou, Sheng Liu, Yuheng Bu, Songyou Peng, Xin Eric Wang
机构
*
University of California, Santa Barbara(加州大学圣芭芭拉分校)
;
J.P. Morgan Chase(摩根大通)
;
ETH Zurich(苏黎世联邦理工学院)
;
Stanford University(斯坦福大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Carnegie Mellon University(卡内基梅隆大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.CV
AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing
AIGaitor: 面向所有人的隐私保护与无云端运动分析——基于边缘计算
Lauhitya Reddy, Trisha M. Kesar, Hyeokhyen Kwon
机构
*
Department of Biomedical Informatics, Emory University(埃默里大学生物医学信息学系)
;
Department of Rehabilitation Medicine, Emory University(埃默里大学康复医学系)
;
The Wallace H. Coulter Department of Biomedical Engineering, Emory University and Georgia Institute of Technology(埃默里大学和佐治亚理工学院的Wallace H. Coulter生物医学工程系)
机构
*
Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳瓦分校计算机科学与工程系)
;
School of Information Technology, King Mongkut’s Institute of Technology Ladkrabang, Thailand(泰国拉差班国王理工大学信息科技学院)
RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
RadAgent:一种用于胸部CT逐步解读的工具型AI智能体
Mélanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor
机构
*
Department of Biosystems Science and Engineering, ETH Zurich(生物系统科学与工程系,苏黎世联邦理工学院)
;
ETH AI Center, Zurich(ETH人工智能中心,苏黎世)
;
Department of Computer Science, ETH Zurich(计算机科学系,苏黎世联邦理工学院)
;
Faculty of Computer Science and Mathematics, Heidelberg University(计算机科学与数学学院,海德堡大学)
;
Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University(斯坦福大学人工智能在医学和影像中的中心)
;
Department of Radiology, Stanford University(放射科,斯坦福大学)
;
Department of Quantitative Biomedicine, University of Zurich(定量生物医学系,苏黎世大学)
;
Institute of Computer Science, Zurich University of Applied Sciences(应用科学大学计算机科学研究所)
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
使用多片段视频破解多模态大语言模型
Choongwon Kang, Seungjong Sun, Hyunmin Jun, Jang Hyun Kim
机构
*
Department of Applied Artificial Intelligence, Sungkyunkwan University(应用人工智能系,成均馆大学)
;
Department of Human-Artificial Intelligence Interaction, Sungkyunkwan University(人机交互系,成均馆大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI