Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Nanyang Technological University(南洋理工大学)
;
Imperial College London(帝国理工学院)
Comments8 pages, 1 figure, 4 tables. Copyright 2026 IEEE. This is the accepted manuscript for 2025 IEEE International Conference on Intelligent Transportation Systems (ITSC), not the final published version
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
SciOrch: 学习编排专家大语言模型以解决前沿多模态科学推理任务
Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin
机构
*
Imperial College London(伦敦帝国学院)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
University of Oxford(牛津大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
ScoutVLA: UAV-Centric Active Perception via a Dual-Expert VLA Model for Open-World Embodied Question Answering
ScoutVLA:面向开放世界具身问答的无人机中心主动感知双专家VLA模型
Wenhao Lu, Zhengqiu Zhu, Xiaofeng Wang, Xiaoran Zhang, Yatai Ji, Yong Zhao, Yue Hu, Yingzhen Nie, Jinlong Zhu, Zheng Zhu
机构
*
National Key Laboratory of Digital Intelligent Modeling and Simulation, National University of Defense Technology(国防科技大学数字智能建模与仿真国家重点实验室)
;
GigaAI
机构
*
Shenzhen Loop Area Institute(深圳河套学院)
;
Dalian University of Technology(大连理工大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI
VL2Spike:面向具身AI低功耗视觉感知的VLM脉冲驱动蒸馏
Zinan Liu, Eric Zheng, Soumyaratna Debnath, Hao Shi, Ling Xiao, Lin Wang
机构
*
School of EEE, Nanyang Technological University (NTU)(南洋理工大学电气与电子工程学院)
;
Department of Computer Science, University of Toronto(多伦多大学计算机科学系)
;
Advanced Micro Devices, Inc.(超威半导体公司)
;
State Key Laboratory of Extreme Photonics and Instrumentation, Zhejiang University(浙江大学极端光子学与仪器国家重点实验室)
;
Faculty of Information Science and Technology, Hokkaido University(北海道大学信息科学与技术学院)
机构
*
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
University of California, Berkeley(加州大学伯克利分校)
GraphBEV++: Multi-Modal Feature Alignment for Autonomous Driving
GraphBEV++: 自动驾驶中的多模态特征对齐
Ziying Song, Caiyan Jia, Lin Liu, Shaoqing Xu, Lei Yang, Yadan Luo
机构
*
Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院,交通数据挖掘与具身智能北京市重点实验室)
;
School of Artificial Intelligence (School of Software), Yanshan University(燕山大学人工智能学院(软件学院))
;
University of Macau(澳门大学)
;
Nanyang Technological University(南洋理工大学)
;
The University of Queensland(昆士兰大学)
Attention, not scale, drives human-AI alignment in multimodal language prediction
注意力,而非规模,驱动多模态语言预测中的人机对齐
Viktor Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards, Quitterie Lacome D'Elascombe, Akilles Rechardt, Jeremy I Skipper, Gabriella Vigliocco
机构
*
Psychology and Language Science, Experimental Psychology, University College London, London, UK(心理学与语言科学、实验心理学,伦敦大学学院,伦敦,英国)
;
Google Deepmind, Mountain View, US(谷歌DeepMind,山景城,美国)
;
Princeton Neuroscience Institute, Princeton University, Princeton, NJ, USA(普林斯顿神经科学研究所,普林斯顿大学,普林斯顿,新泽西州,美国)
;
Computer Science Department, Exeter University(计算机科学系,埃克塞特大学)
Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans
传播结构引导:从眼底图像和稀疏OCT扫描合成荧光素血管造影
Tengfei Ma, Ruiqi Wu, Chenran Zhang, Ye Geng, Na Su, Xiangyuan Duanmu, Tao Zhou, Yi Zhou, Wen Fan
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education(教育部新一代人工智能技术及其跨学科应用重点实验室)
;
Tianyuan Honors School, Nanjing Medical University(南京医科大学天元荣誉学院)
;
Nanjing University of Science and Technology(南京理工大学)
;
Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科)
机构
*
East China University of Science and Technology(华东理工大学)
;
Tsinghua University(清华大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
University of Science and Technology of China(中国科学技术大学)
机构
*
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
;
Peking University(北京大学)
;
Independent Researcher(独立研究者)
;
Jinling Clinical Medical College College of Artificial Intelligence Nanjing University of Aeronautics and Astronautics(金陵临床医学院人工智能学院南京航空航天大学)
机构
*
School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学)
;
School of Computer Science, Wuhan University(计算机学院,武汉大学)
;
Great Wall Motor(长城汽车)
;
School of Information and Electronic Engineering, Zhejiang University of Science and Technology(信息电子工程学院,浙江理工大学)
CommentsThis paper is an extended version of the authors' work previously presented at the ICRA conference. To appear in IEEE Transactions on Circuits and Systems for Video Technology. DOl: 10.1109/TCSVT.2026.3701706
EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
EyeMVP: 通过配对CFP-OCT预训练实现OCT启发的眼底表征学习
Zhuo Deng, Ruiheng Zhang, Ziheng Zhang, Weihao Gao, Yitong Li, Qian Wang, Lei Shao, Jiaoyue Dong, Zhixi Zeng, Lijian Fang, Haibo Wang, Xiaobin Lin, Tao Liu, Zhicheng Du, Zhengwei Zhang, Lin Yang, Zheng Gong, Xinyu Zhao, Zhenquan Wu, Fang Li, Zhiguang Zhou, Guoming Zhang, Sun Jing, Han Lv, Wenbin We, Lan Ma
机构
*
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Beijing Tongren Eye Center, Beijing Tongren Hospital, Capital Medical University(首都医科大学附属北京同仁医院北京同仁眼科中心)
;
Liangxiang Hospital of Beijing Fangshan District, Capital Medical University(首都医科大学北京市房山区良乡医院)
;
The Third People's Hospital of Dalian(大连市第三人民医院)
;
National Clinical Research Center for Endocrine and Metabolic Diseases, The Second Xiangya Hospital of Central South University(中南大学湘雅二医院国家内分泌代谢病临床医学研究中心)
;
The Central Hospital of Baoji City(宝鸡市中心医院)
;
Wuxi No.2 People's Hospital, Affiliated Wuxi Clinical College of Nantong University(南通大学附属无锡临床学院无锡市第二人民医院)
;
Shenzhen Eye Hospital, Southern Medical University(南方医科大学深圳眼科医院)
;
Beijing Friendship Hospital, Capital Medical University(首都医科大学附属北京友谊医院)