Multi-Modal Grounded Planning and Efficient Replanning For Learning Embodied Agents with A Few Examples
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
Comments AAAI 2025 (Project page: https://twoongg.github.io/projects/flare/)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
Comments AAAI 2025 (Project page: https://twoongg.github.io/projects/flare/)
专题命中 多模态Agent :multimodal(title);分类 cs.AI
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments Accepted and presented at ECCV 2024 2nd Workshop on Vision-Centric Autonomous Driving (VCAD) on September 30, 2024. 13 pages, 5 figures
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
Comments Accepted to NeurIPS 2024
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments 5 pages, 1 figure, 1 table, accepted in Embodied AI 2024 Workshop held in conjunction with CVPR 2024
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments The 1st place solution of End-to-end Driving at Scale at the CVPR 2024 Autonomous Grand Challenge
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments Accepted for publication at MICCAI 2024 workshop on AI for Imaging Genomics Learning (AIIG)
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments CO and DHP contributed equally to this work. JSD and ETR are corresponding authors
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments Accepted as full paper in AAMAS 2024
专题命中 多模态Agent :multimodal(title);分类 cs.AI
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
Comments Challenge report doi.org/10.1016/j.cmpb.2023.107561
Journal ref Computer Methods and Programs in Biomedicine, Volume 236, 2023
专题命中 多模态Agent :cross-modal(title);分类 cs.CV
Comments Our project page with videos is at https://RchalYang.github.io/LocoTransformer
专题命中 多模态Agent :multimodal(title);分类 cs.CL
Comments 10 pages; This position paper was presented at the Rethinking the Senses: A Workshop on Multisensory Embodied Experiences and Disability Interactions associated with the ACM CHI Conference on Human Factors in Computing Systems, May 2021
Journal ref ACM CHI Conference on Human Factors in Computing Systems, May 2021
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Journal ref Proceedings of 21st International Radar Symposium (IRS 2020)
专题命中 多模态Agent :multi-modal(title);分类 cs.AI
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Comments The paper has been accepted by IEEE Transactions on Intelligent Transportation Systems 2020
专题命中 多模态Agent :multimodal(title);分类 cs.AI
Comments 16 pages, 13 figures, new experiments, new explanatory figures for intuition and new title
专题命中 多模态Agent :multimodal(title);分类 cs.CV
Journal ref Asian Conference on Computer Vision (ACCV). 2018. 511-526
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
专题命中 多模态Agent :multimodal(title);分类 cs.CV
专题命中 多模态Agent :multimodal(title);分类 cs.CV
具身形态塑造多模态婴儿模型中的翻滚行为
机构 * Frankfurt Institute for Advanced Studies(法兰克福高等研究院) ; Goethe University Frankfurt(法兰克福大学) ; University of New South Wales(新南威尔士大学)
专题命中 多模态Agent :multimodal(title,comments)
AI总结 通过虚拟婴儿MIMo学习仰卧到俯卧翻滚,研究婴儿运动发展中的具身形态变化如何影响行为,发现与真实婴儿一致的发育趋势和协调模式。
Comments 7 pages, 7 figures. Accepted at the 2026 IEEE ICDL Conference. Cite as: L. Philipp, F. M. López, and J. Triesch, "Embodiment Shapes Rolling Behavior in a Multimodal Infant Model", in 2026 IEEE International Conference on Development and Learning (ICDL). IEEE, 2026, pp. 1-7
前沿大语言模型在空间意象推理中的局限性
机构 * Institute of Mathematics and Statistics – University of São Paulo(数学统计研究所 – 圣保罗大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 本研究通过引入外部“意象模块”辅助3D模型旋转任务,发现即使外包整体3D状态维护,前沿模型仍缺乏基础视觉空间原语,导致准确率最高仅62.5%。
Comments 25 pages. v2: Title updated; added a section on object/spatial imagery and propositional reasoning; added new experimental results for the single-object rotation probe
快门之前:3D场景中美学的且可执行的人像摄影规划
机构 * The Hong Kong Polytechnic University(香港理工大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 提出在3D场景中生成人像姿态、相机、照明和曝光方案的方法,通过构建摄影场景图实现美学引导的规划,生成视觉上引人注目且几何与光度可行的人像。
具身人工智能的安全性:风险、攻击与防御综述
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院) ; City University of Hong Kong(香港城市大学) ; Jilin University(吉林大学) ; Singapore Management University(新加坡管理大学) ; Deakin University(德肯大学) ; Tongji University(同济大学) ; Nanyang Technological University(南洋理工大学) ; Chinese Academy of Sciences(中国科学院) ; The University of Melbourne(墨尔本大学) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
AI总结 本文综述了具身AI在感知、认知、规划、行动及交互全流程中的安全风险、攻击与防御方法,提出了多层次分类体系,并指出了多模态感知融合脆弱性、规划不稳定及人机交互可信度等关键挑战。
Comments Survey paper; 75 pages, 4 figures, 18 tables; v2 expands embodied-specific coverage of agentic threats, World Action Model threats, and contextual risk mitigation, with over 100 new references added. Project page: https://x-zheng16.github.io/Awesome-Embodied-AI-Safety/
StarVLA:一种积木式代码库,用于视觉-语言-动作模型开发
机构 * Von Neumann Institute, HKUST(香港科技大学冯·诺依曼研究所)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
AI总结 StarVLA通过模块化架构、可重用训练策略和统一评估接口,解决VLA方法碎片化问题,提升可复现性和跨架构兼容性。
Comments Open-source VLA infra, Technical Report
HippoCamp:在个人电脑上对上下文代理进行基准测试
机构 * S-Lab, Nanyang Technological University, Singapore(新加坡南洋理工大学S-Lab)
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 HippoCamp是一个新的基准,用于评估代理在多模态文件管理中的能力,通过用户为中心的环境建模个体用户档案并搜索大规模个人文件进行上下文感知推理,揭示了当前代理在真实环境中的局限性。
Comments Project Page: https://hippocamp-ai.github.io/
FAPE-IR:面向全场景图像修复的频率感知规划与执行框架
机构 * Tianjin University(天津大学) ; University of Macau(澳门大学) ; City University of Hong Kong(香港城市大学) ; Institute of Artificial Intelligence (TeleAI)(人工智能研究所(TeleAI))
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出FAPE-IR框架,通过频率感知规划与执行模块,结合多模态大语言模型和LoRA-MoE架构,实现统一且可解释的全场景图像修复,实验显示其在七项任务中表现优异。
多智能体系统实现从化学文献中灵活的信息提取
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本研究提出了一种基于多模态大语言模型的多智能体系统,实现了从化学文献中高效提取化学信息,F1分数达76.27%,显著提升信息提取效率。
ACE-Brain-0:空间智能作为通用具身化体系的共享框架
机构 * Shanghai Jiao Tong University(上海交通大学) ; Nanyang Technological University(南洋理工大学) ; The Chinese University of Hong Kong(香港中文大学) ; The University of Hong Kong(香港大学) ; University of Science(科学技术大学) ; Fudan University(复旦大学) ; Xiamen University(厦门大学) ; East China Normal University(华东师范大学) ; Wuhan University(武汉大学) ; Sun Yat-sen University(中山大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL
AI总结 ACE-Brain-0 通过空间智能作为共享框架,统一了自动驾驶、机器人和 UAVs 的具身化任务,采用 SSR 范式和 GRPO 方法实现跨领域泛化和领域精通的平衡。
Comments Code: https://github.com/ACE-BRAIN-Team/ACE-Brain-0 Hugging Face: https://huggingface.co/ACE-Brain/ACE-Brain-0-8B