SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
SEDualVLN:一种空间增强的双系统用于视觉语言导航
Jingzhi Huang, Junkai Huang, Wenxuan Song, Haoyang Yang, Hailong Huang, Haoang Li, Yi Wang
机构
*
Hong Kong Polytechnic University(香港理工大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构
*
College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
School of Computing Technologies, RMIT University(皇家墨尔本理工学院计算技术学院)
RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing
RAPTOR+: 一种基于视觉的视觉-语言框架,用于提高自动化癌症转诊处理中的临床信任度和可审计性
Sofiat Abioye, Ufaq Khan, Shazad Ashraf, Anusha Jose, Adam Byfield, Lukman Akanbi, Muhammad Bilal
机构
*
Birmingham City University(伯明翰城市大学)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
University Hospitals Birmingham NHS Foundation Trust(伯明翰大学医院 NHS 基础信托)
;
NHS England(英格兰国家卫生服务体系)
CommentsThis manuscript has been withdrawn by the authors. It reproduced the methodology of Gardinazzi et al., arXiv:2410.11042, without citation, and utilized code and data from the associated repository (github.com/RitAreaSciencePark/ZigZagLLMs) without disclosure or violate the MIT License. A revised future version with full attribution may be prepared. For any feedback, please contact Pengcheng Zheng
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
VOLD:通过在线蒸馏将LLM推理能力转移到视觉语言模型
Walid Bousselham, Hilde Kuehne, Cordelia Schmid
机构
*
Tuebingen AI Center(图宾根人工智能中心)
;
University of Tuebingen(图宾根大学)
;
MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
;
Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究院、巴黎-萨克勒大学、École Normale Supérieure、PSL研究大学)
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
多模态还是非多模态:通过主动模态检测的查询自适应音视频人物检索
Erfan Loweimi, Mengjie Qian, Kate Knill, Guanfeng Wu, Chi-Ho Chan, Abbas Haider, Muhammad Awan, Josef Kittler, Hui Wang, Mark Gales
机构
*
University of Cambridge(剑桥大学)
;
Queen's University Belfast(贝尔法斯特女王大学)
;
University of Surrey(萨里大学)
;
Cisco(思科)
;
Southwest Jiaotong University(西南交通大学)
;
Teesside University(泰赛德大学)
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
O_O-VC: 基于合成数据驱动的任意到任意语音转换一一对一对齐
Huu Tuong Tu, Huan Vu, cuong tien nguyen, Dien Hy Ngo, Nguyen Thi Thu Trang
机构
*
VNPT AI, VNPT Group(VNPT AI,VNPT集团)
;
Hanoi University of Science and Technology(河内科学技术大学)
;
Business AI Lab, National Economics University(国家经济大学商业人工智能实验室)
Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting
使用大语言模型和梯度提升的多模态性别歧视识别与表征
Kyriakos Chaviaras, Maria Lymperaiou, Athanasios Voulodimos
机构
*
Artificial Intelligence and Learning Systems Laboratory(人工智能与学习系统实验室)
;
School of Electrical and Computer Engineering(电气与计算机工程学院)
;
National Technical University of Athens(雅典国家技术大学)
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
MemoryCard: 面向长视频问答的主题感知多模态线索压缩
Qing Yang, Pengcheng Huang, Xinze Li, Zhenghao Liu, Yukun Yan, Yu Gu, Ge Yu, Gang Li, Maosong Sun
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Digital China Group(数字中国集团)
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
在预测之前想象:用于视频事件预测的交错潜在视觉推理
Tianxiang Jiang, Linquan Wu, Sheng Xia, Songze Li, Ziang Yan, Haoyu Yang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
City University of Hong Kong(香港城市大学)
;
Nanjing University(南京大学)
;
Fudan University(复旦大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video
LongSpace: 从感知到回忆的视频长程空间记忆探索
Shiqiang Lang, Jing Liu, Haoyang He, Peiwen Sun, Yuanteng Chen, Tao Liu, Lan Yang, Longteng Guo, Honggang Zhang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Zhongguancun Academy(中关村学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
The Chinese University of Hong Kong(香港中文大学)
;
Xi’an Jiaotong University(西安交通大学)
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning
VTI-CoT: 用于视频推理的视觉-文本交织思维链
Shufan Zhang, Ziyue Lin, Bairun Wang, Lei Jin, Xuanding Ding, Xinzhu Ma, Kunlin Yang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Hong Kong(香港大学)
;
Beijing Shanwei Zhixing Technology Co., Ltd.(北京尚维智行科技有限公司)
;
Tsinghua University(清华大学)
;
Beihang University(北航)
Ritabrata Roy Choudhury, Arkajyoti Karmakar, Rudra Pratap Mitra
机构
*
School of Computer Engineering, Kalinga Institute of Industrial Technology(计算机工程学院,凯林加工业技术学院)
;
School of Electronics Engineering, Kalinga Institute of Industrial Technology(电子工程学院,凯林加工业技术学院)