RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
RAR:用于视觉识别的检索与排序增强多模态大语言模型
Ziyu Liu, Zeyi Sun, Yuhang Zang, Wei Li, Pan Zhang, Xiaoyi Dong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
MThreads, Inc.(MThreads公司)
;
Nanyang Technological University(南洋理工大学)
GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective
GOMA:从图信号平滑视角迈向结构驱动的多模态对齐
Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang
机构
*
School of Airspace Science and Engineering, Shandong University(山东大学 airspace 科学与工程学院)
;
Department of Computer Science, Beijing Institute of Technology(北京理工大学计算机学院)
;
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
ICED: 通过可解释的概念分解实现概念级机器去学习
Shen Lin, Jing Lin, Junhao Dong, Piotr Koniusz, Li Xu
机构
*
Fujian Normal University(福建师范大学)
;
Nanyang Technological University(南洋理工大学)
;
University of New South Wales(新南威尔士大学)
;
Data61 CSIRO(Data61澳大利亚联邦科学与工业研究组织)
EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
EntropyScan: 向通过视觉注意力熵实现LVLMs的模型级后门检测
Xuanyu Ge, Zhongqi Wang, Jie Zhang, Shiguang Shan, Xilin Chen
机构
*
China University of Geosciences(中国地质大学)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构
*
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所国家工程研究中心与人工智能院,中国科学院)
;
School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision
通过合成神经符号监督学习结构化机器人策略
Alessandro Adami, Tommaso Tubaldo, Marco Todescato, Ruggero Carli, Pietro Falco
机构
*
University of Padova, Dept. of Information Engineering(帕多瓦大学信息工程系)
;
Fraunhofer Italia Research(弗劳恩霍夫意大利研究所)
;
Polytechnic of Bari Dept. of Electrical and Information Engineering(巴里理工学院电气与信息工程系)
机构
*
Department of Electronics and Computer Engineering, Thapathali Campus, Institute of Engineering, Tribhuvan University(电子与计算机工程系,Thapathali校区,工程学院,塔波胡万大学)
Driving Through the Network: Performance and Workload Under Latency and Video Impairment
通过网络驾驶:延迟和视频失真下的性能与负载
Ines Trautmannsheimer, Ahmed Azab, Frank Diermeyer
机构
*
Technical University of Munich, School of Engineering and Design, Institute of Automotive Technology and Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑工业大学,工程学院,汽车技术研究所,慕尼黑机器人与机器智能研究所(MIRMI))
Agent4POI: Agentic Context-Conditioned Affordance Reasoning for Multimodal Point-of-Interest Recommendation
Agent4POI: 基于代理的上下文条件化 affordance 推理用于多模态兴趣点推荐
Jinze Wang, Yangchen Zeng, Tiehua Zhang, Lu Zhang, Yuze Liu, Yongchao Liu, Xingjun Ma, Zhu Sun
机构
*
Tongji University(同济大学)
;
Swinburne University of Technology(斯威本理工大学)
;
Southeast University(东南大学)
;
Chengdu University of Information Technology(成都信息工程大学)
;
Fudan University(复旦大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
Integrating chemical structures as treatments improves representations of microscopy images for morphological profiling
将化学结构作为治疗手段提高显微镜图像的表示以进行形态学分析
Yemin Yu, Emre Hayir, Neil Tenenholtz, Lester Mackey, Ying Wei, David Alvarez-Melis, Ava P. Amini, Alex X. Lu
机构
*
Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
;
Microsoft Research(微软研究院)
;
Department of Computer Science, Zhejiang University(浙江大学计算机科学系)
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance(北京多模态数据智能感知与治理重点实验室)