Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
使用基于语音的多模态大语言模型实现可泛化的认知障碍检测
Yingchao Huang, Xin Wang, Yuhan Su, Shanshan Yao
机构
*
Faculty of Digital Innovation, Arts \& Sciences, Saskatchewan Polytechnic, Regina SK S4S 5X1, Canada
;
School of Basic Medical Sciences, Hebei University, Baoding 071000, China
;
Department of Civil \& Environmental Engineering
;
School of Mining \& Petroleum Engineering, University of Alberta, Edmonton AB T6G 2H5, Canada
专题命中
视觉定位与Grounding
:multimodal large language model(title);分类 cs.LG
Anomalous Frame Detection by Grouping Frame Similarities between Two Videos Computed by Vision-Language Model to Extract Expert Workers' Unique Actions
Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance
用于现场检查和维护中移动操作的部分观察下语言引导抓取
Dilermando Almeida, Juliano Negri, Guilherme Lazzarini, Thiago H. Segreto, Ranulfo Bezerra, Gustavo J. G. Lahr, Ricardo V. Godoy, Marcelo Becker
机构
*
Department of Mechanical Engineering, Federal University of Uberlândia(联邦大学伯南迪利亚机械工程系)
;
Department of Mechanical Engineering, University of São Paulo(圣保罗大学机械工程系)
;
Graduate School of Information Sciences, Tohoku University(东北大学信息科学研究生院)
;
Faculdade Israelita de Ensino e Pesquisa Albert Einstein, Hospital Israelita Albert Einstein(艾伯特·爱因斯坦以色列教学与研究学院,艾伯特·爱因斯坦医院)
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
COMPASS:在统一多模态模型中锚定构图意图引导
Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan
机构
*
University of Edinburgh(爱丁堡大学)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Ant Group(蚂蚁集团)
CommentsAccepted at FOIS 2026 (16th International Conference on Formal Ontology in Information Systems), Vitória, Brazil; to appear in Frontiers in Artificial Intelligence and Applications, IOS Press. 16 pages, 1 figure, 2 tables
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
基于XAI的语音深度伪造检测解释生成:使用免训练多模态大语言模型
Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang, Berrak Sisman, Björn W. Schuller
机构
*
Imperial College London(帝国理工学院)
;
Technical University of Munich(慕尼黑工业大学)
;
University of Southampton(南安普顿大学)
;
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
;
Johns Hopkins University(约翰霍普金斯大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title);分类 cs.AI
Graph Alignment Topology as an Inductive Bias for Grounding Detection
图对齐拓扑作为接地检测的归纳偏置
Paul Landes, Pranav Herur, Adam Cross, Jimeng Sun
机构
*
Department of Pediatrics, University of Illinois College of Medicine Peoria(伊利诺伊大学皮奥里亚医学院儿科部)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机与数据科学学院)
;
Carle Illinois College of Medicine, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校卡莱医学院)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
无需训练的零样本时序动作检测与视觉-语言模型
Chaolei Han, Hongsong Wang, Jidong Kuang, Lei Zhang, Jie Gui
机构
*
Southeast University School of Cyber Science and Engineering(东南大学网络安全科学与工程学院)
;
Southeast University School of Computer Science and Engineering(东南大学计算机科学与工程学院)
;
Nanjing Normal University School of Electrical Engineering and Automation(南京师范大学电气工程与自动化学院)
Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework
通过临床世界模型和技能混合框架在人类认知中奠定临床AI能力
Seyed Amir Ahmad Safavi-Naini, Elahe Meftah, Josh Mohess, Pooya Mohammadi Kazaj, Georgios Siontis, Zahra Atf, Peter R. Lewis, Mauricio Reyes, Girish Nadkarni, Roland Wiest, Stephan Windecker, Christoph Grani, Ali Soroush, Isaac Shiri
机构
*
Department of Cardiology, Inselspital, Bern University Hospital, University of Bern(伯尔尼大学医院心脏病学系,伯尔尼大学)
;
Department of Digital Medicine, Bern University Hospital, University of Bern(伯尔尼大学医院数字医学系,伯尔尼大学)
;
Division of Data-Driven and Digital Medicine (D3M), Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院数据驱动与数字医学部)
;
Clinical Research Development Center, Amir Oncology Teaching Hospital, Shiraz University of Medical Sciences(设拉子医科大学阿米尔肿瘤教学医院临床研究发展中心)
;
Graduate School for Cellular and Biomedical Sciences, University of Bern(伯尔尼大学细胞与生物医学研究生院)
;
Faculty of Business and Information Technology, Ontario Tech University(安大略理工大学商业与信息技术学院)
;
Department of Radiation Oncology, Inselspital, Bern University Hospital and University of Bern(伯尔尼大学医院放射肿瘤学系,伯尔尼大学)
;
ARTORG Center for Biomedical Engineering Research, University of Bern(伯尔尼大学ARTORG生物医学工程研究中心)
;
The Charles Bronfman Institute of Personalised Medicine, Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院查尔斯·布朗夫曼个性化医学研究所)
;
University Institute of Diagnostic and Interventional Neuroradiology, Inselspital, Bern University Hospital, University of Bern(伯尔尼大学医院诊断与介入神经放射学大学研究所,伯尔尼大学)
;
Translational Imaging Center (TIC), Swiss Institute for Translational and Entrepreneurial Medicine(瑞士转化与创业医学研究所转化影像中心)
;
Henry D. Janowitz Division of Gastroenterology, Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院亨利·D·雅诺维茨消化内科)
IoT-Brain: Grounding LLMs for Semantic-Spatial Sensor Scheduling
IoT-Brain:为语义-空间传感器调度 grounding LLMs
Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Junda Lin, Jinke Song
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
机构
*
College of Computer Science and Artificial Intelligence, Shanghai Key Laboratory of Intelligent Information Processing, Fudan University(复旦大学计算机科学与人工智能学院,上海智能信息处理重点实验室)
;
University of Oxford(牛津大学)
;
The Hong Kong University of Science and Technology(香港科技大学)