Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA
全决策:一种用于全模态问答的渐进式证据状态智能体系统
Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Weigao Sun, Junhan Shi, Lingrui Mei, Tianming Yang, Steven Hoi
机构
*
Institute of Neuroscience, Chinese Academy of Sciences(中国科学院神经科学研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)
;
Tsinghua University(清华大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Beijing University of Chemical Technology(北京化工大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Beijing Institute of Technology (Zhuhai)(北京理工大学(珠海))
;
Tencent Hy(腾讯(深圳))
;
Peng Cheng Laboratory(鹏城实验室)
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
LEDGERMIND:基于结构化证据账本的溯源约束多模态智能体推理
Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The University of Hong Kong(香港大学)
;
Tsinghua University(清华大学)
;
University of Sussex(萨塞克斯大学)
机构
*
Zhejiang University(浙江大学)
;
Hunan University(湖南大学)
;
Tianjin University(天津大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Jilin University(吉林大学)
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
ProMSA: 渐进式多模态搜索智能体用于基于知识的视觉问答
ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, Haoqian Wang
机构
*
OPPO AI Center, OPPO Inc. China(OPPO AI中心,OPPO公司)
;
The Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
机构
*
School of Information, Computer
;
Communication Technology Sirindhorn International Institute of Technology, Thammasat University Pathum Thani, Thailand 1
机构
*
National Taiwan University(国立台湾大学)
;
Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所)
;
Radboud University(拉德堡德大学)
;
Institut Jean Nicod(让·尼科研究所)
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Nanyang Technological University(南洋理工大学)
;
Imperial College London(帝国理工学院)
Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
基于语义三维高斯溅射的开放词汇移动操作具身多模态定位
Huosen Ou, Dongni Song, Yuncong Wang, Tao Zhou, Yiding Ji
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Midea Group(美的集团)
;
The Hong Kong University of Science and Technology(香港科技大学)