Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports
Razi Mahmood, Diego Machado-Reyes, Joy Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep Kalra, Ge Wang, Pingkun Yan, Tanveer Syeda-Mahmood
机构
*
Rensselaer Polytechnic Institute, NY, USA(罗文学院)
;
IBM Research, Almaden, CA, USA(IBM研究院)
;
Stanford University, CA, USA(斯坦福大学)
;
Massachusetts General Hospital (MGH), Boston, USA(麻省总医院)
专题命中
视觉定位与Grounding
:vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji
机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Institute of AI for Industries(工业人工智能研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
Nadim Barakat, William Lotter
机构
*
Dana-Farber Cancer Institute & Tufts University School of Medicine(达纳-法伯癌症研究所及塔夫茨大学医学院)
;
Dana-Farber Cancer Institute Brigham and Women’s Hospital & Harvard Medical School(达纳-法伯癌症研究所布里特妇女医院及哈佛医学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen
机构
*
College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院)
;
School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学)
;
State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室)
;
College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校)
;
Department of Geography, The University of Hong Kong(地理系,香港大学)
;
Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)
专题命中
视觉定位与Grounding
:vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
机构
*
Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学)
;
Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学)
;
Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations
Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang
机构
*
Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以)
;
NVIDIA(NVIDIA公司)
;
Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
Yufei Zhan, Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)
;
Wuhan AI Research, Wuhan, China(武汉人工智能研究所)
专题命中
视觉定位与Grounding
:vision language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Ser-Nam Lim, Rajiv Ramnath
机构
*
Department of Computer Science(计算机科学系)
;
Engineering, Ohio State University, Ohio, US(工程系,俄亥俄州立大学,俄亥俄,美国)
;
Department of Computer Science, University of Central Florida, Florida, US(计算机科学系,中央佛罗里达大学,佛罗里达,美国)
专题命中
视觉定位与Grounding
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation
Adam Goodge, Xun Xu, Bryan Hooi, Wee Siong Ng, Jingyi Liao, Yongyi Su, Xulei Yang
机构
*
Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息通信研究所,科技研究局(A*STAR),新加坡)
;
School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡)
DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Glasgow(格拉斯哥大学)
;
Boston University(波士顿大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))