Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model
Wild-Drive: 通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划
Zihang Wang, Xu Li, Benwu Wang, Wenkai Zhu, Xieyuanli Chen, Dong Kong, Kailin Lyu, Yinan Du, Yiming Peng, Haoyang Che
机构
*
School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院)
;
Southeast University Nanjing Jiangbei New Area Innovation Research Institute(东南大学南京江滨新区创新研究院)
;
National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology(国防科技大学装备状态感知与智能支撑国家重点实验室)
;
School of Transportation, Shandong University of Science and Technology(山东科技大学交通学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
FiLo++: 通过融合细粒度描述和变形定位实现零/少样本异常检测
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Ming Tang, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,自动化研究所,中国科学院)
;
School of Artifcial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Wuhan AI Research(武汉AI研究所)
;
Peng Cheng Laboratory(鹏城实验室)
;
Guangdong Provincial Key Laboratory of Intellectual Property & Big Data, Guangdong Polytechnic Normal University(广东省知识产权与大数据重点实验室,广东工业大学)
Vision-Language Feature Alignment for Road Anomaly Segmentation
视觉-语言特征对齐用于道路异常分割
Zhuolin He, Jiacheng Tang, Jian Pu, Xiangyang Xue
机构
*
School of Computer Science, Fudan University(复旦大学计算机科学学院)
;
Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑启发智能科学与技术研究院)
MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine
MedGPT-oss: 为生物医学训练一个通用的视觉-语言模型
Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
机构
*
Department of Computer Science and Engineering, Lehigh University(莱斯大学计算机科学与工程系)
;
Department of Computer Science and Engineering, University of Notre Dame(圣母大学计算机科学与工程系)
;
Department of Health Outcomes & Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系)
;
Department of Radiation Oncology, Mayo Clinic(梅奥诊所放射肿瘤科)
;
Research Computing, University of Florida(佛罗里达大学研究计算中心)
;
AI Technology Center, NVIDIA(NVIDIA人工智能技术中心)
CommentsWe present a framework for evaluation of Multi-modal Agents consisting of Voice-to-voice model components viz. Text to Speech (TTS), Retrieval Augmented Generation (RAG) and Speech-to-text (STT)