From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation
从对齐到合成:用于文本到CT生成的对比体 grounding
Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
机构
*
Unit of Artificial Intelligence and Computer Systems, Department of Engineering, Università Campus Bio-Medico di Roma(人工智能与计算机系统单位,工程系,罗马生物医学大学)
;
Department of Diagnostics and Intervention, Biomedical Engineering and Radiation Physics, Umeå University(诊断与干预部门,生物医学工程与放射物理学,乌梅拉大学)
From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation
从语义接地到决策优化:面向长 horizon 无人机视觉语言导航的统一框架
Zeyuan Ma, Jiaxin Chen, Di Huang
机构
*
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
Comments21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval
CommentsThis manuscript is currently under peer review. Copyright may subsequently be transferred to the publisher, after which the availability of this version may be subject to the publisher's policy
A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
外科AI的比较研究:数据、计算和扩展的潜力与局限
Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
机构
*
Center for Applied AI, Chicago Booth(应用人工智能中心,芝加哥商学院)
;
Surgical Data Science Collective(外科数据科学集体)
;
Children’s National Hospital(儿童医学中心)
;
Operations Management & Tolan Center for Healthcare, Chicago Booth(运营管理与托兰医疗中心,芝加哥商学院)
专题命中
视觉定位与Grounding
:vision language model(abstract);分类 cs.CV、cs.AI、cs.LG
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
SARVLM:一种用于SAR图像语义理解的视觉语言基础模型
Qiwei Ma, Xukun Lu, Wang Liu, Puhong Duan, Xudong Kang, Shutao Li
机构
*
School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院)
;
Yuelushan Center for Industrial Innovation(岳麓山创新中心)
;
School of Medical Information Engineering, Jining Medical University(济南医学院医学信息工程学院)