Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
基于语义三维高斯溅射的开放词汇移动操作具身多模态定位
Huosen Ou, Dongni Song, Yuncong Wang, Tao Zhou, Yiding Ji
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Midea Group(美的集团)
;
The Hong Kong University of Science and Technology(香港科技大学)
JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures
JEPA-DNA:通过联合嵌入预测架构夯实基因组基础模型
Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
机构
*
Applied AI Architecture, NVIDIA, Israel(NVIDIA应用人工智能架构,以色列)
;
Worldwide Field Ops, NVIDIA, Israel(NVIDIA全球现场运营,以色列)
;
Developer Programs, NVIDIA, Israel(NVIDIA开发者计划,以色列)
;
Cancer Research Center and Wohl Institute of Translational Medicine, Sheba Medical Center, Tel Hashomer, Israel(癌症研究中心和Wohl转化医学研究所,Sheba医疗中心,Tel Hashomer,以色列)
;
Windreich Department of AI and Human Health, Icahn School of Medicine at Mount Sinai, New York, USA(AI与人类健康风reich部门,Mount Sinai医学中心,纽约,美国)
Comments11 pages, 4 figures. Published in the Proceedings of the 1st Workshop on Shaping Future Human Connection: Social Augmentation through XR Technologies (SAXR 2026), April 13, 2026, Barcelona, Spain
Journal refProceedings of the 1st Workshop on Shaping Future Human Connection: Social Augmentation through XR Technologies (SAXR 2026), CEUR Workshop Proceedings, Vol. 4226, pp. 252-262, 2026
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
EventCoT:用于推理时间定位的以事件为中心的视频思维链
Youngkil Song, Yoonjae Baek, Dongwon Kim, Inho Kim, Dongkeun Kim, Suha Kwak
机构
*
Pohang University of Science and Technology(浦项科技大学)
;
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Handong Global University(韩东国际大学)
On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis
在大型语言模型中自我改进的极限:没有符号模型合成,奇点并不临近
Hector Zenil
机构
*
Algorithmic Dynamics Lab(算法动力实验室)
;
Department of Biomedical Computing(生物医学计算系)
;
School of Biomedical Engineering and Imaging Sciences(生物医学工程与成像科学学院)
;
King’s Institute for AI(国王人工智能研究所)
;
King’s College London(伦敦国王学院)
;
Oxford Immune Algorithmics(牛津免疫算法公司)
;
Oxford University Innovation(牛津大学创新中心)
;
London Institute for Healthcare Engineering(伦敦医疗工程研究所)
机构
*
Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
;
Samueli Electrical and Computer Engineering, UCLA(UCLA Samueli电气与计算机工程系)
;
Department of Computer Science and Informatics, Emory University(埃默里大学计算机科学与信息学系)
;
Mayo Clinic(梅奥诊所)
Situation Graph Prediction for User Perspective Modeling
情境图预测:用户建模的结构化视角推断
Jisung Shin, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
机构
*
Flybits Labs, Creative AI Hub(Flybits实验室、创意人工智能中心)
;
University of Toronto(多伦多大学)
;
Toronto Metropolitan University(多伦多 Metropolitan 大学)
;
MIT Media Lab(MIT媒体实验室)
专题命中
视觉定位与Grounding
:grounding(abstract);分类 cs.AI
AI总结
情境图预测通过结构化视角推断提升用户建模能力,揭示潜在状态推断比表层提取更困难。
CommentsAccepted to PILA 2026: Workshop on Personal Intelligence in the Agentic AI Era, at KDD 2026, 5 pages