ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
ProMSA: 渐进式多模态搜索智能体用于基于知识的视觉问答
ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, Haoqian Wang
机构
*
OPPO AI Center, OPPO Inc. China(OPPO AI中心,OPPO公司)
;
The Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures
ReScene: 基于多视角图像的结构化室内场景重建
Haoran Xu, Lechao Zhang, Daoguo Dong, Yan Gao, Xin Tan
机构
*
School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
;
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究院)
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
QiYuanLab(启元实验室)
;
Tsinghua University(清华大学)
;
University of Electronic Science and Technology of China(电子科技大学)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
EchoSonar-R: 一种用于超声心动图疾病分类和报告生成的多视图推理增强模型
Darya Taratynova, Ahmed Aly, Numan Saeed, Mohammad Yaqub
机构
*
Division of Computing and Mathematical Sciences, Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE(穆罕默德·本·扎耶德人工智能大学计算与数学科学系,阿布扎比,阿联酋)
Comments6 pages, 5 figures. Published in MobiSys Workshop '26
Journal refIn Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services Workshops (MobiSys Workshop '26), June 21-25, 2026, Cambridge, United Kingdom. ACM, New York, NY, USA, 6 pages