V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
V-ABS:基于动作-观察者的动态视觉推理束搜索
Zhiwei Ning, Xuanang Gao, Jiaxi Cao, Gengming Zhang, Shengnan Ma, Wenwen Tong, Hanming Deng, Jie Yang, Wei Liu
机构
*
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)
;
Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(上海交通大学图像处理与模式识别研究所)
;
SenseTime Research(商汤科技研究院)
;
Institute of Medical Robotics, Shanghai Jiao Tong University(上海交通大学医学机器人研究所)
Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models
面向视觉-语言-动作模型的后门基于所有权验证
Ming Sun, Rui Wang, Xingrui Yu, Lihua Jing, Hangyu Du, Zhenglin Wan, Xu Pan, Ivor Tsang
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
A*STAR Institute of High Performance Computing (A*STAR IHPC)(新加坡A*STAR高性能计算研究所)
;
A*STAR Centre for Frontier AI Research (A*STAR CFAR)(新加坡A*STAR前沿人工智能研究中心)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
College of Design and Engineering, Nanyang Technological University(南洋理工大学设计与工程学院)
;
Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS), Wuhan University(武汉大学测绘遥感信息工程国家重点实验室(LIESMARS))
机构
*
Interdepartmental Program in Computational Biology and Bioinformatics, Yale University(耶鲁大学计算生物学与生物信息学联合计划)
;
Department of Biostatistics, Yale University(耶鲁大学生物统计学系)
;
Broad Institute of MIT and Harvard(哈佛大学与麻省理工学院联合Broad研究所)
;
Center for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡医学学校生物医学数据科学中心)
;
Sport and Exercise Medicine Service, KK Women’s and Children’s Hospital Training Program, Duke-NUS Medical School(杜克-新加坡医学学校KK妇女儿童医院运动与医学服务培训项目)
;
Training Program, Duke-NUS Medical School(杜克-新加坡医学学校培训项目)
;
Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系)
;
Department of Complexity Science and Engineering, The University of Tokyo(东京大学复杂科学与工程系)
;
Center for Advanced Intelligence Project, RIKEN(日本理化学研究所高级智能项目中心)
;
NUS Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所)
;
Department of Biostatistics and Bioinformatics, Duke University(杜克大学生物统计学与生物信息学系)
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
Department of Genetics, Yale University(耶鲁大学遗传学系)
;
Wu Tsai Institute, Yale University(耶鲁大学吴天教授研究所)
;
Department of Biomedical Informatics and Data Science, Yale University(耶鲁大学生物医学信息学与数据科学系)
机构
*
Sev.en Global Investments(Sev.en全球投资)
;
Vrije Universiteit Brussels(布鲁塞尔自由大学)
;
Université Libre de Bruxelles(布鲁塞尔自由大学)
;
Ministry of Finance of the Slovak Republic(斯洛伐克共和国财政部)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
动态跨模态提示生成用于多模态持续指令微调
Tao Hu, Da-Wei Zhou
机构
*
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
;
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth
用于非合作航天器6D位姿估计的跨模态RGB-D融合Transformer
Yongliang Zhen, Bo LÜ, Hang Yang, Xiaotian WU
机构
*
School of Physics, Northeast Normal University(东北师范大学物理学院)
;
Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Science(长春光学精密机械与物理研究所)
Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer
不再有间隙:通过一个分词器实现零间隙多模态整合
Yanan Li, Christina Yi Jin, Yuan Jin, Manli Luo, Tie Xu, Shuai Jiao, Wei He, Qing Zhang
机构
*
Research Center for Frontier Fundamental Studies, Zhejiang Lab(前沿基础研究研究中心,浙江实验室)
;
Research Center for Scientific Data Hub, Zhejiang Lab(科学数据枢纽研究中心,浙江实验室)
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
MulTaBench: 基于文本和图像的多模态表格学习基准测试
Alan Arazi, Eilam Shapira, Shoham Grunblat, Mor Ventura, Elad Hoffer, Gioia Blayer, David Holzmüller, Lennart Purucker, Gaël Varoquaux, Frank Hutter, Roi Reichart
机构
*
Technion – Israel Institute of Technology(技术ion – 以色列理工学院)
;
Prior Labs(Prior实验室)
;
NVIDIA
;
SODA Team, INRIA Saclay, Palaiseau(SODA团队,INRIA萨克莱,帕莱索)
;
University of Freiburg(弗赖堡大学)
;
Probabl
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
机构
*
Dongguan Key Laboratory of Intelligent Equipment and Smart Industry, School of Advanced Engineering, Great Bay University(东莞智能装备与智能制造重点实验室,先进工程学院,大湾大学)
;
Chair of Applied Statistics, Technische Universität Dresden(应用统计学教授职位,德累斯顿技术大学)
;
Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据解析与人工智能中心(ScaDS.AI))
;
College of Automation, Guangdong University of Technology(自动化学院,广东技术大学)
机构
*
Tencent Hunyuan(腾讯文言)
;
University of Maryland, College Park(马里兰大学 College Park 分校)
;
University of Virginia(弗吉尼亚大学)
;
University of North Carolina, Chapel Hill(北卡罗来纳大学 Chapel Hill 分校)
Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning
锚定特征间隙:跨模态谱稳定化以实现样本高效表征学习
Nikhil J. Dhinagar, Vidhi Chhatbar, Chirag Jagad, Pavithra Senthilkumar, Sophia I. Thomopoulos, Mahir H. Khan, Sook-Lei Liew, the ENIGMA-Stroke Recovery Working Group, Paul M. Thompson
机构
*
Imaging Genetics Center, Mark & Mary Stevens Neuroimaging & Informatics Institute, Keck School of Medicine, University of Southern California(影像基因中心,马克与玛丽史蒂文斯神经影像与信息学研究所,凯克医学院,南加州大学)
;
Neuroscience Graduate Program, Mark & Mary Stevens Neuroimaging & Informatics Institute, Chan Division of Occupational Science & Occupational Therapy, Biomedical Engineering, University of Southern California(神经科学研究生项目,马克与玛丽史蒂文斯神经影像与信息学研究所,查恩职业科学与职业治疗 division,生物医学工程,南加州大学)