机构
*
Department of Biomedical Informatics, Harvard Medical School(哈佛医学学校生物医学信息学系)
;
Department of Computer Science, University of North Carolina(北卡罗来纳大学计算机科学系)
;
Department of Computer Science, Massachusetts Institute of Technology(麻省理工学院计算机科学系)
机构
*
Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, South Korea
;
Department of Electrical
;
Computer Engineering, Seoul National University, Seoul, South Korea
;
Department of Computer Science \& Engineering, Korea University, Seoul, South Korea
;
ISRC, Seoul National University, Seoul, South Korea
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
OmniAD:基于多模态推理的工业异常检测与理解
Shifang Zhao, Yiheng Lin, Lu Han, Yao Zhao, Yunchao Wei
机构
*
Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院)
;
Visual Intelligence + X International Joint Laboratory of the Ministry of Education(教育部视觉智能+X国际合作实验室)
;
Key Laboratory of Noise and Vibration Research, Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所噪声与振动重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
机构
*
Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学)
;
Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学)
;
Department of Electrical and Computer System Engineering, Monash University(电子与计算机系统工程系,莫纳什大学)
机构
*
Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院)
;
School of Computation, Information and Technology, Technical University of Munich, Garching, Germany(慕尼黑技术大学计算与信息学院)
;
Department of Shenyang Institute of Computing Technology, University of Chinese Academy of Sciences, Beijing, China(中国科学院沈阳计算技术研究所部门)
;
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家重点实验室)
;
Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
区域感知多模态大语言模型:基于慢快标记化与伪掩码引导的3D CT报告生成
Sunggu Kyung, Jinyoung Seo, Hyunseok Lim, Dongyeong Kim, Hyungbin Park, Jimin Sung, Jihyun Kim, Wooyoung Jo, Yoojin Nam, Namkug Kim
机构
*
Department of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Republic of Korea(韩国首尔峨山医疗中心蔚山大学医学院融合医学系)
;
University of Ulsan College of Medicine, Seoul, Republic of Korea(韩国首尔蔚山大学医学院)
;
Department of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Republic of Korea(韩国首尔峨山医疗中心蔚山大学医学院放射科与放射学研究所)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
EvoLMM:具有连续奖励的自进化大型多模态模型
Omkar Thawakar, Shravan Venkatraman, Ritesh Thawkar, Abdelrahman Shaker, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Fahad Khan
机构
*
Mohamed bin Zayed University of AI(Mohamed bin Zayed人工智能大学)
;
Aalto University(阿alto大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林肯大学)
机构
*
AnnLab(安实验室)
;
Institute of Semiconductors, Chinese Academy of Sciences(中国科学院半导体研究所)
;
Zhongguancun Academy(中关村学院)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
State Key Laboratory of High Performance Ceramics(高性能陶瓷国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Electronic, Electrical and Communication Engineering(电子电气与通信工程学院)
;
University of ChineseAcademy of Sciences(中国科学院大学)
Now We Know? A Systematic Comparison of TerraMind and THOR
我们现在知道了吗?TerraMind和THOR的系统比较
Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-Børre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
机构
*
University of Münster(明斯特大学)
;
IBM Research(IBM研究院)
;
Norwegian Computing Center(挪威计算中心)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
MMaDA-VLA: 基于统一多模态指令与生成的大型扩散视觉-语言-动作模型
Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang
机构
*
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
East China University of Science and Technology(华东理工大学)
;
Huawei Celia Team(华为Celia团队)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
OpenHelix Robotics