From Goals, Waypoints & Paths To Long Term Human Trajectory Forecasting
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI
Comments 14 pages, 7 figures (including 2 GIFs)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI
Comments 14 pages, 7 figures (including 2 GIFs)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted at CorRL 2020 (matching camera-ready version)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted ICCV 2019 Workshop
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.MM
Comments 10 pages
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI
Comments 7 pages, 8 figures
Journal ref Third Workshop in Machine Learning in the Planning and Control of Robot Motion at ICRA, 2018
专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.MM
Comments 2018 IEEE International Conference on Image Processing - ICIP'18. arXiv admin note: text overlap with arXiv:1806.02609
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI
Comments Accepted to ICLR 2018
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI
Comments 10 pages, RoboNLP Workshop from ACL Conference
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI
Comments 11 pages, 5 appendix pages, 11 figures, 3 tables, under review as a conference paper at ICLR 2017
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL
Comments 9 pages, manuscript under submission
AutoTool: 面向智能体推理的动态工具选择与集成
机构 * Nanyang Technological University(南洋理工大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL;multi-modal(comments)
AI总结 提出AutoTool框架,通过双阶段优化(SFT+RL轨迹稳定化和KL正则化Plackett-Luce排序)使大语言模型具备动态工具选择能力,在数学、科学、代码和多模态推理等任务上平均提升6.4%-7.7%。
Comments ICML2026; Best Paper Award at ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence
专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.AI
Comments CCC Blue Sky Ideas paper, published at the ACM International Conference on Multimodal Interaction (ICMI '22), November 7-11, 2022
专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.CL
Comments 6 pages, ICLR 2021 Embodied Multimodal Learning Workshop
专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.CV
Comments Accepted to 2019 International MICCAI Brainlesion Workshop -- Multimodal Brain Tumor Segmentation Challenge (BraTS) 2019. arXiv admin note: substantial text overlap with arXiv:1810.11654
专题命中 多模态Agent :multimodal(abstract,journal_ref);分类 cs.CV
Journal ref ICMI 2018 - Workshop at 20th ACM International Conference on Multimodal Interaction, Oct 2018, Boulder, Colorado, United States. pp.1-13
MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制
机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 该研究提出多智能体大语言模型框架MistyPilot,可解释自然语言指令并编排Misty社交机器人技能,经评估其在多项任务上准确率高、方差低,用户反馈积极,代码将公开。
Comments Accepted at the ECCV 2026 ACVR Workshop
ETHOS:面向临床多智能体系统的模块化伦理框架
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。
Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}
ForceU-VLA:用于实体超声扫描的力感知视觉-语言-动作模型
机构 * Faculty of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院) ; School of Information Science and Engineering, Shandong University(山东大学信息科学与工程学院) ; Innovation School of Artificial Intelligence, Hefei University of Technology(合肥工业大学人工智能创新学院)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV
AI总结 本文针对现有实体超声扫描方法的不足,提出力感知视觉-语言-动作模型ForceU-VLA,设计FUSFM与SAMM模块,构建ForceU-VLA-Data数据集,实验证实其可提升超声扫描的接触稳定性与压力调节能力。
视觉语言模型(VLMs)无法规划,但它们能进行形式化吗?
机构 * Drexel University(德雷塞尔大学) ; University of Pennsylvania(宾夕法尼亚大学) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL
AI总结 本研究提出5种VLM作为形式化器的流水线,评估后发现其表现优于端到端规划生成,较弱VLMs的瓶颈为物体关系视觉grounding,较强模型已克服该限制。
AdvDex:通过关节对齐动作与对抗学习从人类演示中学习灵巧操作
机构 * Zhejiang University(浙江大学) ; Shanghai Innovation Institute(上海创新研究院) ; Fudan University(复旦大学) ; Shanghai Jiao Tong University(上海交通大学) ; Paxini Tech(帕西尼科技)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 AdvDex是一种统一视觉-语言-动作框架,通过OmniShare数据集、JAAS动作空间与领域对抗学习,实现从人类和机器人演示中学习灵巧操作,提升跨实体泛化与技能迁移能力。
PlayWorld:基于智能体玩家的长程目标世界模型基准测试
机构 * The Chinese University of Hong Kong(香港中文大学) ; The University of Hong Kong(香港大学) ; Zhejiang University(浙江大学) ; Kuaishou Technology(快手科技)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV
AI总结 该研究针对现有世界模型跨模型公平比较的难题,推出含171个场景的PlayWorld基准,通过多模态智能体玩家从多维度评估9种先进世界模型,发现其长程交互式目标表现仍不可靠。
Comments project page: https://kxding.github.io/project/PlayWorld/
LLM-Advisor: 一种用于多地形高效路径规划的LLM基准
机构 * Graduate School of Information Science and Technology, Hokkaido University(弘前大学信息科学与技术研究生院) ; Department of Information and Communication Engineering, The University of Tokyo(东京大学信息与通信工程系)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 LLM-Advisor利用大型语言模型优化多地形路径规划,提升路径成本效率,适用于现实场景。
Comments This paper has been accepted by IEEE Transactions on Automation Science and Engineering
Lines and Ladders:面向大规模零售价格分类的上下文感知多智能体框架
机构 * Walmart Global Tech(沃尔玛全球科技)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI
AI总结 针对大规模零售商品定价管理难题,提出上下文感知多智能体框架自动化构建Lines and Ladders价格分类,3智能体系统在Lines任务F1达0.83,在多品类数据上表现优异且已投入生产。
Comments 8 pages. Accepted in the Main Conference of IEEE ICMLA 2026
多智能体目标存在性验证与学习型掩码几何优化:2026年第8届LSVOS挑战赛MeViS-Text赛道获胜报告
机构 * Soongsil University(崇实大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV
AI总结 本研究提出SSUPER方案,通过多智能体审计解耦存在性验证、StyleRefiner优化掩码几何,获2026年LSVOS挑战赛MeViS-Text赛道冠军,最终得分0.9081339614。
XCoT-VLA:面向视觉-语言-动作自动驾驶的可执行思维链
机构 * XPeng Inc(小鹏汽车)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 该研究提出XCoT-VLA,用可执行CoT令牌替代自然语言CoT,结合Reason FFN与Control FFN实现轨迹生成,在自动驾驶任务中降低了误差并满足实时规划要求。
GAM-Agent:面向复杂视觉推理的博弈论与感知不确定性感知协作框架
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 该研究提出GAM-Agent博弈论多智能体框架,通过基础感知智能体与关键验证智能体的非零和博弈及不确定性感知协作,在四个视觉推理基准上显著提升了中小及强规模VLM的性能,为可靠可解释多模态推理提供了新路径。
Comments Accepted at NeurIPS 2025. Code available at https://github.com/jushengzhang/Gam-Agent
用于物理一致性开放世界运动规划的能量结构化潜在世界模型与神经时间场
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 针对开放世界运动规划的物理一致性挑战,提出能量结构化潜在世界模型ELWM,结合物理条件神经时间场PC-NTF,显著提升了运动预测精度与导航成功率,降低了碰撞率与残差。
Comments 9 pages, 5 figures
阅读并非推理:弥合视觉-文本压缩中的智能体策略差距
专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI
AI总结 该研究针对视觉-文本压缩导致的智能体策略差距,提出CAPS跨模态智能体策略自蒸馏框架,在SearchQA、ALFWorld等数据集上显著提升性能并大幅降低上下文成本。