Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
面向可扩展多模态推理的推理对齐感知解耦
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Xin Jin, Zhenguo Li, James T. Kwok, Yu Zhang
机构
*
Southern University of Science and Technology(南方科技大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
Huawei Cloud Project(华为云项目)
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
MARCUS:一种用于心脏诊断和管理的代理式多模态视觉-语言模型
Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe, Akshay Chaudhari, Tahoura Nedaee, Francois Haddad, Michael Salerno, Li Fe-Fei, Ehsan Adeli, Rima Arnaout, Euan A Ashley
机构
*
Division of Cardiology, Department of Medicine, Stanford University(斯坦福大学心脏病学系)
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
Department of Medicine, Radiology, and Pediatrics, UCSF(旧金山大学医学系、放射学与儿科学系)
;
Bakar Institute, UCSF(Bakar研究所,旧金山大学)
;
UCSF–UC Berkeley Joint Program in Computational Precision Health(旧金山大学-伯克利计算精准健康联合计划)
;
Department of Radiology, Stanford University(斯坦福大学放射学系)
;
Department of Psychiatry and Behavioral Sciences, Stanford University(斯坦福大学精神病学与行为科学系)
;
Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
;
Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系)
;
Department of Biology, Stanford University(斯坦福大学生物学系)
Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
通过语义增强的动态对比交互实现高度可迁移的视觉-语言攻击
Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler
机构
*
School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院)
;
Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
PixelVLA:推进视觉-语言-动作模型中的像素级理解
Wenqi Liang, Gan Sun, Yao He, Jiahua Dong, Suyan Dai, Ivan Laptev, Salman Khan, Yang Cong
机构
*
University of Trento(特伦托大学)
;
School of Automation Science and Engineering, South China University of Technology(华南理工大学自动化科学与工程学院)
;
Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学)
;
Australian National University(澳大利亚国立大学)
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家实验室,南京大学,南京,中国)
;
School of Artificial Intelligence, Nanjing University, Nanjing, China(人工智能学院,南京大学,南京,中国)
;
Mila - Quebec AI Institute(魁北克AI研究所)
Comments4 pages, 2 figures. To appear in Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction (HRI '26), Edinburgh, Scotland, March 2026
Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
始终信守承诺:一种针对神经网络的模型无关属性算法
Kevin Lee, Duncan Smith-Halverson, Pablo Millan Arias
机构
*
David R. Cheriton School of Computer Science, University of Waterloo, ON, Canada(滑铁卢大学戴维·R·切里顿计算机科学学院)
;
Scotiabank AML AI Research, Toronto, ON, Canada(多伦多加拿大Scotiabank反洗钱AI研究)
机构
*
School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China(北京理工大学计算机科学与技术学院,北京,中国)
;
Beijing Engineering Research Center of High Volume Language Information Processing and Cloud Computing Applications, Beijing Institute of Technology, Beijing, China(高性能语言信息处理与云计算应用北京工程研究中心,北京理工大学,北京,中国)