BioGait-VLM: A Tri-Modal Vision-Language-Biomechanics Framework for Interpretable Clinical Gait Assessment
BioGait-VLM:一种多模态视觉-语言-生物力学框架用于可解释的临床步态评估
Erdong Chen, Yuyang Ji, Jacob K. Greenberg, Benjamin Steel, Faraz Arkam, Abigail Lewis, Pranay Singh, Feng Liu
机构
*
Department of Computer Science, Drexel University(德克萨斯大学计算机科学系)
;
Department of Neurological Surgery, Washington University(华盛顿大学神经外科系)
;
University of California, Berkeley(加州大学伯克利分校)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
超越准确率:评估多模态医学推理中的视觉语义
Anas Zafar, Leema Krishna Murali, Ashish Vashist
机构
*
The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
;
Cohere Labs(Cohere实验室)
;
Eisai Inc.(艾伯维公司)
;
Indian Institute of Science, Bangalore(班加罗尔印度科学研究院)
HulluEdit: Single-Pass Evidence-Consistent Subspace Editing for Mitigating Hallucinations in Large Vision-Language Models
HulluEdit: 单次通过证据一致子空间编辑用于缓解大视觉-语言模型中的幻觉
Yangguang Lin, Quan Fang, Yufei Li, Jiachen Sun, Junyu Gao, Jitao Sang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Jiaotong University(北京交通大学)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
面向细粒度图像编辑评估的人类对齐MLLM评判:一个基准、框架和分析
Runzhou Liu, Hailey Weingord, Sejal Mittal, Prakhar Dungarwal, Anusha Nandula, Bo Ni, Samyadeep Basu, Hongjie Chen, Nesreen K. Ahmed, Li Li, Jiayi Zhang, Koustava Goswami, Subhojyoti Mukherjee, Branislav Kveton, Puneet Mathur, Franck Dernoncourt, Yue Zhao, Yu Wang, Ryan A. Rossi, Zhengzhong Tu, Hongru Du
机构
*
University of Virginia(弗吉尼亚大学)
;
Columbia University(哥伦比亚大学)
;
Vanderbilt University(范德比大学)
;
Adobe Research(Adobe研究)
;
Dolby Laboratories(杜比实验室)
;
Cisco Research(思科研究)
;
University of Southern California(南加州大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
University of Oregon(俄勒冈大学)
;
Texas A&M University(德克萨斯大学)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
StreamSense: 基于选择性视觉-语言模型路由的流式社交任务检测
Han Wang, Deyi Ji, Lanyun Zhu, Jiebo Luo, Roy Ka-Wei Lee
机构
*
Singapore University of Technology and Design(新加坡科技设计大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
University of Rochester(罗切斯特大学)
RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
RSGround-R1: 重新思考通过空间推理的遥感视觉定位
Shiqi Huang, Shuting He, Bihan Wen
机构
*
School of Electrical and Electronic Engineering, Nanyang Technological University(电气电子工程学院,南洋理工大学)
;
MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(教育部计算与经济交叉学科重点实验室,上海财经大学)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV
A Training-Free Guess What Vision Language Model from Snippets to Open-Vocabulary Object Detection
无需训练的Guess What视觉语言模型:从片段到开放词汇物体检测
Guiying Zhu, Bowen Yang, Yin Zhuang, Tong Zhang, Guanqun Wang, Zhihao Che, He Chen, Lianlin Li
机构
*
Aerospace and Informatics Domain(航空航天与信息领域)
;
National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing(空间智能信息处理国家级重点实验室)
;
School of Electronic(电子学院)
专题命中
视觉定位与Grounding
:vision language model(title,abstract);VLM(abstract);分类 cs.CV
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
基于三元组的引用范式用于统一视觉解码
Yan Tai, Luhao Zhu, Yunan Ding, Yiying Dong, Guangtao Zhai, Xiaohong Liu, Guodong Guo
机构
*
School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China(上海交通大学计算机科学学院)
;
Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo, China(宁波数字孪生研究院)
;
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai, 200240, China(上海交通大学信息科学与电子工程学院)
专题命中
视觉定位与Grounding
:VLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
释放大型视觉-语言模型的能力以实现道路基础设施的智能感知
Luxuan Fu, Chong Liu, Bisheng Yang, Zhen Dong
机构
*
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS)(信息工程测绘遥感国家重点实验室)
;
Wuhan University(武汉大学)
;
Hubei Luojia Laboratory(湖北珞珈实验室)
专题命中
视觉定位与Grounding
:vision-language model(title);vision language model(abstract);grounding(abstract);分类 cs.CV