机构
*
Tsinghua University(清华大学)
;
Peking University(北京大学)
;
Fudan University(复旦大学)
;
Microsoft Research Asia(微软亚洲研究院)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Zhejiang University(浙江大学)
机构
*
CS Department Virginia Tech(弗吉尼亚理工学院计算机科学系)
;
EECS Department MIT(麻省理工学院电子工程与计算机科学系)
;
CS Department Dartmouth College(达特茅斯学院计算机科学系)
;
Statistics Department Virginia Tech(弗吉尼亚理工学院统计学系)
;
CS Department Michigan State University(密歇根州立大学计算机科学系)
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
Vid-LLM:一种基于视频的紧凑型3D多模态大语言模型,具有重建-推理协同效应
Haijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu, Jiaxuan Lin, Jingrong Wang
机构
*
School of Geodesy and Geomatics, Wuhan University(武汉大学测绘学院)
;
Hubei Luojia Laboratory(湖北珞珈实验室)
;
School of Architecture and Urban Planning, Shenzhen University(深圳大学建筑与城市规划学院)
机构
*
Southwestern University of Finance(西南财经大学)
;
Department of Informatics,Universitat Hamburg, Hamburg, Germany(汉堡大学信息学院)
;
University of Electronic Science(电子科技大学)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
;
School of Information and Intelligent Science, Donghua University(信息与智能科学学院,东华大学)
MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence
MLLM-4D: 向基于视觉的空间-时间智能迈进
Xingyilang Yin, Chengzhengxu Li, Jiahao Chang, Chi-Man Pun, Xiaodong Cun
机构
*
University of Macau(澳门大学)
;
Xi'an Jiaotong University(西安交通大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
GVC Lab, Great Bay University(大湾大学GVC实验室)
MMNavAgent: Multi-Magnification WSI Navigation Agent for Clinically Consistent Whole-Slide Analysis
MMNavAgent: 多倍率WSI导航代理用于临床一致的全滑片分析
Zhengyang Xu, Han Li, Jingsong Liu, Linrui Xie, Xun Ma, Xin You, Shihui Zu, Ayako Ito, Xinyu Hao, Hongming Xu, Shaohua Kevin Zhou, Nassir Navab, Peter J. Schüffler
机构
*
Institute of Pathology, Technical University of Munich, Germany(慕尼黑技术大学病理研究所)
;
Munich Data Science Institute (MDSI), Munich, Germany(慕尼黑数据科学研究所)
;
Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心)
;
Computer Aided Medical Procedures (CAMP), TU Munich, Munich, Germany(计算机辅助医疗程序(CAMP),慕尼黑技术大学)
;
Dalian University of Technology(大连理工大学)
;
Cancer Hospital of Dalian University of Technology, Shenyang(大连理工大学肿瘤医院)
;
Department of Human Pathology, Juntendo University Graduate School of Medicine(立命大学医学研究生院人类病理部门)
;
University of Science and Technology of China(中国科学技术大学)
;
Institute of Medical Robotics, Shanghai Jiao Tong University(上海交通大学医学机器人研究所)
;
Northwest University of China(中国西北大学)
机构
*
University of Toronto, Canada(多伦多大学)
;
Vector Institute, Canada(向量研究所)
;
KITE Research Institute, University Health Network, Canada(KITE研究机构)
;
York University, Canada(约克大学)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
AgentVista: 评估在超挑战性现实视觉场景中的多模态代理
Zhaochen Su, Jincheng Gao, Hangyu Guo, Zhenhua Liu, Lueyang Zhang, Xinyu Geng, Shijue Huang, Peng Xia, Guanyu Jiang, Cheng Wang, Yue Zhang, Yi R. Fung, Junxian He
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构
*
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室)
;
Nation Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心)
;
Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所)
;
Xi’an Jiaotong University(西安交通大学)
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
SesaHand: 通过语义和结构对齐可控生成增强3D手重建
Zhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei, Pan Hui, Anyi Rao
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Communication University of China(中国通信大学)