KG-CMI: Knowledge graph enhanced cross-Mamba interaction for medical visual question answering
KG-CMI: 基于知识图谱的跨Mamba交互用于医学视觉问答
Xianyao Zheng, Hong Yu, Hui Cui, Changming Sun, Xiangyu Li, Ran Su, Leyi Wei, Jia Zhou, Junbo Wang, Qiangguo Jin
机构
*
School of Software, Northwestern Polytechnical University(西北工业大学软件学院)
;
Tianjin Central Hospital of Gynecology Obstetrics(天津市中心妇产科医院)
;
Department of Computer Science and Information Technology, La Trobe University(拉筹伯大学计算机科学与信息技术系)
;
CSIRO Data61(澳大利亚联邦科学与工业研究组织Data61)
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
;
School of Computer Software, Tianjin University(天津大学计算机软件学院)
;
Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University(澳门理工大学应用科学学院人工智能驱动药物发现中心)
;
Department of Cardiology, Tianjin Chest Hospital(天津市胸科医院心内科)
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
BigEarthNet.txt: 一个大规模多传感器图像-文本数据集和地球观测基准
Johann-Ludwig Herzog, Mathis Jürgen Adler, Leonard Hackel, Yan Shu, Angelos Zavras, Ioannis Papoutsis, Paolo Rota, Begüm Demir
机构
*
BIFOLD(柏林智能数据与机器学习研究所)
;
Technische Universität Berlin(柏林工业大学)
;
University of Trento(特伦托大学)
;
National Technical University(雅典国家技术大学)
;
National Observatory of Athens(雅典国家天文台)
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
空间推理并非免费午餐:对LLaVA的受控研究
Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu, Timothy Chung, Drishti Sharma, Akshata A., Kranthi Kiran, Wesley Tam, Bala Krishna S Vegesna
机构
*
Cohere Labs Community(Cohere Labs社区)
;
Indian Institute of Science, Bangalore(印度科学研究所班加罗尔分校)
;
UIUC(伊利诺伊大学厄巴纳-香槟分校)
;
Imperial College London(伦敦帝国学院)
;
Eisai Inc.(卫材公司)
;
EleutherAI
;
Georgia Institute of Technology(佐治亚理工学院)
A 4D Representation for Training-Free Agentic Reasoning from Monocular Laparoscopic Video
一种用于无训练代理推理的4D表示
Maximilian Fehrentz, Nicolas Stellwag, Robert Wiebe, Nicole Thorisch, Fabian Grob, Patrick Remerscheid, Ken-Joel Simmoteit, Benjamin D. Killeen, Christian Heiliger, Nassir Navab
机构
*
Computer Aided Medical Procedures, TU Munich(慕尼黑工业大学计算机辅助医疗程序研究所)
;
Hospital of the LMU Munich, Ludwig-Maximilians-Universität (LMU)(慕尼黑大学医院)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
专题命中
视觉推理
:vision-language model(abstract);grounding(abstract);multimodal large language model(abstract);MLLM(abstract)
A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation
一种用于胸部X光解读的推理增强视觉-语言基础模型
Yabin Zhang, Chong Wang, Yunhe Gao, Jiaming Liu, Maya Varma, Justin Xu, Sophie Ostmeier, Jin Long, Sergios Gatidis, Seena Dehkharghani, Arne Michalson, Eun Kyoung Hong, Christian Bluethgen, Haiwei Henry Guo, Alexander Victor Ortiz, Stephan Altmayer, Sandhya Bodapati, Joseph David Janizek, Ken Chang, Jean-Benoit Delbrouck, Akshay S. Chaudhari, Curtis P. Langlotz
机构
*
Tsinghua University(清华大学)
;
National University of Defense Technology(国防科技大学)
;
Beihang University(北京航空航天大学)
;
Tianjin University(天津大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Peking University(北京大学)
;
Institute of Microelectronics, Chinese Academy of Sciences(中国科学院微电子研究所)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Civil Aviation University of China(中国民航大学)
;
Beijing Institute of Technology(北京理工大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
RoboClaw:一个用于可扩展长周期机器人任务的代理框架
Ruiying Li, Yunlang Zhou, YuYao Zhu, Kylin Chen, Jingyuan Wang, Sukai Wang, Kongtao Hu, Minhui Yu, Bowen Jiang, Zhan Su, Jiayao Ma, Xin He, Yongjian Shen, Yang Yang, Guanghui Ren, Maoqing Yao, Wenhao Wang, Yao Mu
机构
*
AgiBot
;
National University of Singapore(新加坡国立大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
MoE Key Lab of Artificial Intelligence, AI Institute, SJTU(上海交通大学人工智能研究院教育部人工智能重点实验室)
Advancing Multi-Robot Networks via MLLM-Driven Sensing, Communication, and Computation: A Comprehensive Survey
通过MLLM驱动的感知、通信与计算推进多机器人网络:综述
Hyun Jong Yang, Howon Lee, Kyuhong Shim, Jeongho Kwak, Hyunsoo Kim, Donghoon Kim, Khoa Anh Ngo, Sehyun Ryu, Jaehyun Choi, Youbin Kim, Chanjun Moon, Michael Ryoo, Byonghyo Shim
机构
*
Seoul National University(首尔大学)
;
Ajou University(亚洲大学)
;
Sungkyunkwan University(成均馆大学)
;
Korea University(高丽大学)
;
POSTECH(浦项科技大学)
;
Stony Brook University(石溪大学)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
VMAD:视觉增强的多模态大语言模型用于零样本异常检测
Huilin Deng, Hongchen Luo, Wei Zhai, Yang Cao, Yu Kang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Northeastern University(东北大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer
MotionGrounder: 通过扩散变换器实现多物体运动迁移
Samuel Teodoro, Yun Chen, Agus Gunawan, Soo Ye Kim, Jihyong Oh, Munchurl Kim
机构
*
School of Electrical Engineering, Korea Advanced Institute of Science and Technology(韩国科学技术院电气工程学院)
;
Adobe Research(Adobe研究院)
;
CMLab, Chung-Ang University(中央大学CMLab)