arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2777 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2777 篇

2603.17055 2026-03-19 cs.CV 57%

PaAgent: Portrait-Aware Image Restoration Agent via Subjective-Objective Reinforcement Learning

PaAgent:通过主观-客观强化学习实现的面向人物图像修复代理

Yijian Wang, Qingsen Yan, Jiantao Zhou, Duwei Dai, Wei Dong

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院) State Key Laboratory of Internet of Things for Smart City, University of Macau(澳门大学智慧城市物联网国家重点实验室) National-Local Joint Engineering Research Center of Biodiagnosis and Biotherapy, the Second Affiliated Hospital of Xi’an Jiaotong University(西安交通大学生物诊断与生物治疗国家地方联合工程研究中心) College of Information and Control Engineering, Xi’an University of Architecture and Technology(西安建筑科技大学信息与控制工程学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 PaAgent通过结合自进化的人物银行和检索增强生成技术,提升图像修复任务中对复杂场景的感知能力,通过主观-客观强化学习策略优化修复工具选择,实验验证其在多种修复基准上的优越性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13710 2026-03-17 cs.AI 57%

InterventionLens: A Multi-Agent Framework for Detecting ASD Intervention Strategies in Parent-Child Shared Reading

InterventionLens:一种多智能体框架,用于检测自闭症干预策略在父母-儿童共读中的应用

Xiao Wang, Lu Dong, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出InterventionLens,一种多智能体系统,用于自动检测和时间分割共读视频中的护理人员干预策略,实验表明其在ASD-HI数据集上的F1分数达到79.44%,优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12482 2026-03-16 cs.CV 57%

CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning

CalliMaster:通过布局引导的空间规划掌握页面级中文书法

Tianshuo Xu, Tiantian Hong, Zhifei Chen, Fei Chao, Ying-cong Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Faculty of Engineering and IT, University of Technology Sydney(悉尼大学工程与信息学院) Xiamen University(厦门大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文提出CalliMaster框架,通过解耦空间规划与内容合成,解决页面级书法生成中精度与布局的平衡问题,支持可控生成与编辑,扩展至文物修复与鉴定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00016 2026-03-16 cs.RO cs.AI cs.HC 57%

Beyond Static Instruction: A Multi-agent AI Framework for Adaptive Augmented Reality Robot Training

超越静态指令:一种多智能体AI框架用于自适应增强现实机器人训练

Nicolas Leins, Jana Gonnermann-Müller, Malte Teichmann, Sebastian Pokutta

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出多智能体AI框架,用于增强现实机器人训练中的动态适应,通过多模态输入预处理和LLM自主推理实现个性化学习环境调整。

Journal ref Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction (2026) 989-993

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11392 2026-03-13 cs.NI cs.AI 57%

Agentic AI for Embodied-enhanced Beam Prediction in Low-Altitude Economy Networks

面向低空经济网络的具身增强型波束预测的代理AI

Min Hao, Zhizhuo Li, Zirui Zhang, Maoqiang Wu, Han Zhang, Rong Yu

机构 * School of Electronic Science and Engineering, South China Normal University(南方科技大学电子科学与工程学院) School of Intelligent Engineering, Shaoguan university(韶关大学智能工程学院) School of Automation, Guangdong University of Technology(广东工业大学自动化学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于多代理协作推理的混合波束预测模型,通过整合时序建模、视觉编码和多模态融合,提升低空经济网络中UAV通信的波束预测精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09733 2026-03-11 cs.CV cs.MA 57%

FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis

FetalAgents:一种用于胎儿超声图像和视频分析的多智能体系统

Xiaotian Hu, Junwei Huang, Mingxuan Liu, Kasidit Anmahapong, Yifei Chen, Yitong Luo, Yiming Huang, Xuguang Bai, Zihan Li, Yi Liao, Haibo Qu, Qiyuan Tian

机构 * Tsinghua University(清华大学) University of California San Diego(加州大学圣地亚哥分校) West China Second University Hospital, Sichuan University(四川大学华西第二医院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 FetalAgents是一种多智能体系统,通过动态协调视觉专家,实现胎儿超声图像和视频的端到端分析与报告,提供高准确性和流程对齐的解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02083 2026-03-10 cs.RO cs.CV 57%

$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

$π$-StepNFT: 更宽的空间需要更细的步骤在线RL用于基于流的VLAs

Siting Wang, Xiaofeng Wang, Zheng Zhu, Minnan Pei, Xinyu Cui, Cheng Deng, Jian Zhao, Guan Huang, Haifeng Zhang, Jun Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 $π$-StepNFT通过分步负向感知微调方法,在在线强化学习中提升基于流的VLAs的性能和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08113 2026-03-10 cs.CV 57%

SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving

SAMoE-VLA:一种面向自动驾驶的场景自适应混合专家视觉-语言-动作模型

Zihan You, Hongwei Liu, Chenxu Dang, Zhe Wang, Sining Ang, Aoqi Wang, Yan Wang

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) School of Instrument Science and Engineering, Southeast University(仪器科学与工程学院,东南大学) Zhili College, Tsinghua University(紫荆学院,清华大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Department of Automation, University of Science and Technology of China(自动化学院,中国科学技术大学) Department of Automation, University of Science and Technology Beijing(自动化学院,北京科技大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

AI总结 SAMoE-VLA通过场景自适应混合专家机制提升自动驾驶中的视觉-语言-动作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08013 2026-03-10 cs.AI 57%

PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents

PIRA-Bench: 从反应式 GUI 代理到基于 GUI 的主动意图推荐代理的转变

Yuxiang Chai, Shunye Tang, Han Xiao, Rui Liu, Hongsheng Li

机构 * Nankai University(南开大学) Huawei Research(华为研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 PIRA-Bench通过引入主动意图推荐代理基准,推动GUI代理从反应式向主动式转变,评估多模态大语言模型在连续弱监督视觉输入中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10918 2026-03-10 cs.HC cs.CL 57%

CompanionCast: Toward Social Collaboration with Multi-Agent Systems in Shared Experiences

CompanionCast:迈向多智能体系统在共享体验中的社会协作

Yiyang Wang, Chen Chen, Tica Lin, Vishnu Raj, Josh Kimball, Alex Cabral, Josiah Hester

机构 * Georgia Institute of Technology(佐治亚理工学院) Dolby Laboratories, Inc.(杜比实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

AI总结 CompanionCast通过多智能体系统提升共享体验中的社会协作,通过多模态检测、上下文缓存和空间音频增强共在感,实验显示其在体育观看中显著提升社会存在感和情感共享。

Comments Accepted at ACM CHI 2026 Workshop on Human-Agent Collaboration

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11978 2026-03-10 cs.RO cs.AI 57%

Accelerating Robotic Reinforcement Learning with Agent Guidance

通过代理引导加速机器人强化学习

Haojun Chen, Zili Zou, Chengdong Ma, Yaoxiang Pu, Haotong Zhang, Yuanpei Chen, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) PKU-PsiBot Joint Lab(北京大学- PsiBot 联合实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 AGPS通过多模态代理替代人类监督,提升机器人强化学习的样本效率和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06061 2026-03-09 cs.CV cs.RO 57%

Transforming Omnidirectional RGB-LiDAR data into 3D Gaussian Splatting

将全方位RGB-LiDAR数据转换为3D高斯散点

Semin Bae, Hansol Lim, Jongseong Brad Choi

机构 * Department of Computer Science, State University of New York(计算机科学系,纽约州立大学) Department of Mechanical Engineering, State University of New York(机械工程系,纽约州立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种将全方位RGB-LiDAR数据转换为3D高斯散点的重用管道,解决数据处理中的非线性失真和计算开销问题,提升复杂场景的渲染保真度。

Comments This work has been submitted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05270 2026-03-09 cs.RO cs.AI cs.HC cs.MA cs.SY eess.SY 57%

XR-DT: Extended Reality-Enhanced Digital Twin for Safe Motion Planning via Human-Aware Model Predictive Path Integral Control

XR-DT:增强现实增强型数字孪生用于通过人感知模型预测路径积分控制的安全运动规划

Tianyi Wang, Jiseop Byeon, Ahmad Yehia, Yiming Xu, Jihyung Park, Tianyi Zeng, Sikai Chen, Ziran Wang, Junfeng Jiao, Christian Claudel

机构 * Department of Civil, Architectural, and Environmental Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校土木、建筑与环境工程系) School of Architecture, The University of Texas at Austin(德克萨斯大学奥斯汀分校建筑学院) School of Civil and Construction Engineering, Purdue University(普渡大学土木与建设工程学院) Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 XR-DT通过结合增强现实与数字孪生技术,提出HA-MPPI控制模型,实现基于人类行为预测的安全高效人机交互。

Comments 8 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03762 2026-03-05 cs.CV 57%

Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding

如同专家所见:一个知识增强的代理用于开放集细粒度视觉理解

Junhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang, Yitao Luo, Zhanyu Ma

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 KFRA通过知识增强推理代理实现开放集细粒度视觉理解,提升推理准确率和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23141 2026-03-04 cs.CV 57%

Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents

Earth-Agent: 解锁地球观测的全貌与潜力

Peilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang, Xinjie Huo, Jinhua Yu, Wanghan Xu, Wenlong Zhang, Lei Bai, Conghui He, Weijia Li

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV

AI总结 Earth-Agent是一种结合RGB和光谱数据的代理框架,通过多模态工具生态系统实现跨模态、多步骤推理,提升地球观测分析的科学性和应用潜力。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17520 2026-03-04 cs.RO cs.CV 57%

InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation

InstructVLA: 从理解到操作的视觉-语言-动作指令微调

Shuai Yang, Hao Li, Bin Wang, Yilun Chen, Yang Tian, Tai Wang, Hanqing Wang, Feng Zhao, Yiyi Liao, Jiangmiao Pang

机构 * University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 InstructVLA通过视觉-语言-动作指令微调,实现了在多模态推理和动作生成之间的平衡,提升了机器人在复杂任务中的操控性能和泛化能力。

Comments 48 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19661 2026-03-03 cs.CV 57%

CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization

CodeV: 通过工具感知策略优化实现基于图像的代码推理

Xinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li, Jia Liu, Todd C. Hollon, Bryan Wang

机构 * University of Michigan(密歇根大学) The Ohio State University(俄亥俄州立大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 CodeV通过工具感知策略优化提升视觉推理的忠实度,实现更高准确率和更可靠的工具使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00420 2026-03-03 cs.RO cs.AI cs.SY eess.SY 57%

TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot

TMR-VLA:一种用于三腿硅基软机器人磁性运动控制的视觉-语言-动作模型

Ruijie Tang, Chi Kit Ng, Kaixuan Wu, Long Bai, Guankun Wang, Yiming Huang, Yupeng Wang, Hongliang Ren

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 TMR-VLA通过结合视觉、语言和动作模块,实现三腿硅基软机器人的磁性运动控制与导航能力提升。

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24100 2026-03-02 cs.AI cs.LG 57%

Artificial Agency Program: Curiosity, compression, and communication in agents

人工代理计划:代理中的好奇心、压缩与交流

Richard Csaky

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本文提出人工代理计划,旨在通过整合好奇心、压缩与交流机制,构建资源受限的AI代理系统,以提升感知、理解和行动能力并减少人机交互摩擦。

Comments This is a working draft. Feedback and criticism is most welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21112 2026-02-26 cs.RO cs.AI 57%

EO-1: An Open Unified Embodied Foundation Model for General Robot Control

EO-1:一个通用机器人控制的开放统一具身基础模型

Delin Qu, Haoming Song, Qizhi Chen, Zhaoqing Chen, Xianqiang Gao, Dong Wang, Xinyi Ye, Qi Lv, Modi Shi, Guanghui Ren, Cheng Ruan, Maoqing Yao, Haoran Yang, Jiacheng Bao, Bin Zhao, Xuelong Li

机构 * Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) AgiBot Northwestern Polytechnical University(西北工业大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 EO-1是一个统一的具身基础模型,通过交错视觉-文本-动作预训练实现了在多模态推理和机器人控制中的卓越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20773 2026-02-25 cs.CV 57%

Federated Learning for Cross-Modality Medical Image Segmentation via Augmentation-Driven Generalization

通过增强驱动泛化实现跨模态医学图像分割的联邦学习

Sachin Dudda Nagaraju, Ashkan Moradi, Bendik Skarre Abrahamsen, Mattijs Elschot

机构 * Department of Circulation and Medical Imaging, Norwegian University of Science and Technology(循环医学成像系,挪威科学与技术大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种联邦学习方法,通过增强驱动泛化实现跨模态医学图像分割,提升模型泛化能力的同时保护数据隐私。

Comments Submitted to IEEE JBHI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23055 2026-02-25 cs.AI 57%

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

MindPower: 使基于视觉语言的具身智能体具备心智理论推理能力

Ruoxuan Zhang, Qiyun Zheng, Zhiyu Zhou, Ziqi Liao, Siyu Wu, Jian-Yu Jiang-Lin, Bin Wen, Hongxia Xie, Jianlong Fu, Wen-Huang Cheng

机构 * Jilin University(吉林大学) National Taiwan University(台湾大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 MindPower通过整合感知、心理推理、决策和行动,使基于视觉语言的具身智能体具备心智理论推理能力,并在决策和行动生成上超越GPT-4o。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18729 2026-02-24 cs.CV 57%

GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving

GuideFlow: 基于约束的流匹配用于端到端自动驾驶中的规划

Lin Liu, Caiyan Jia, Guanyi Yu, Ziying Song, JunQiao Li, Feiyang Jia, Peiliang Wu, Xiaoshuai Hao, Yadan Luo

机构 * School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通大数据挖掘与具身智能重点实验室) Qcraft Yanshan University(燕山大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) The University of Queensland(昆士兰大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 GuideFlow通过约束流匹配技术,在端到端自动驾驶中实现高效规划,直接强制显式约束并提升物理约束满足能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12707 2026-02-19 cs.LG cs.AI cs.MA 57%

PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI

PLAICraft: 大规模时间对齐的视觉-语音-动作数据集用于具身人工智能

Yingchen He, Christian D. Weilbach, Martyna E. Wojciechowska, Yuxuan Zhang, Frank Wood

机构 * University of British Columbia(不列颠哥伦比亚大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 PLAICraft通过大规模时间对齐的视觉-语音-动作数据集,推动具身人工智能的研究与评估。

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15767 2026-02-18 cs.RO cs.AI cs.HC 57%

Robot-Assisted Social Dining as a White Glove Service

机器人辅助社交用餐作为白手套服务

Atharva S Kashyap, Ugne Aleksandra Morkute, Patricia Alves-Oliveira

机构 * University of Michigan(密歇根大学) Leiden University(莱顿大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 本研究通过参与设计和AI工具,提出机器人辅助社交用餐的白手套服务理念,强调多模态输入、情境敏感行为及角色扩展,以提升残疾人在野外社交用餐中的独立性和尊严。

Comments 20 pages, 9 figures. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15294 2026-02-18 cs.AI 57%

EAA: Automating materials characterization with vision language model agents

EAA: 用视觉语言模型代理自动化材料表征

Ming Du, Yanqi Luo, Srutarshi Banerjee, Michael Wojcik, Jelena Popovic, Mathew J. Cherukara

机构 * Argonne National Laboratory(阿贡国家实验室) Advanced Photon Source(先进光子源) Data Science and Learning Division(数据科学与学习 division) Department of Radiation Oncology(放射肿瘤学系) Northwestern University(西北大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 EAA利用视觉语言模型代理自动化材料表征,通过多模态推理和工具增强操作提升显微镜实验效率,减少操作负担并降低用户专业门槛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14093 2026-02-17 cs.AI cs.LG 57%

GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training

GUI-GENESIS: 自动合成具有可验证奖励的高效环境以实现GUI代理训练

Yuan Cao, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Xin Chen, Ang Li, Gang Cao, Gong Zhi, Hao Yu, Linyi Li, Wei Yang, Tao Xie

机构 * Key Lab of HCST (PKU), MOE SCS, Peking University, Beijing, China Tencent Inc., Shenzheng, China Hong Kong University of Science Department of Computer Science, University of Texas at Dallas, Richardson, USA School of Computing Science, Simon Fraser University, Burnaby, BC, Canada

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 GUI-GENESIS通过自动合成高效GUI训练环境并提供可验证奖励,显著提升了GUI代理的训练效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14048 2026-02-17 cs.RO cs.CV cs.GR 57%

ProAct: A Dual-System Framework for Proactive Embodied Social Agents

ProAct:一种双系统框架用于主动具身社交代理

Zeyi Zhang, Zixi Kang, Ruijie Zhao, Yusen Feng, Biao Jiang, Libin Liu

机构 * Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 ProAct通过双系统框架实现主动具身社交代理,结合低延迟行为系统与慢速认知系统,提升交互的主动性和社会参与度。

Comments Project Page: https://proactrobot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14003 2026-02-17 cs.AI 57%

Prompt-Driven Low-Altitude Edge Intelligence: Modular Agents and Generative Reasoning

基于提示的低空边缘智能:模块化代理与生成推理

Jiahao You, Ziye Jia, Chao Dong, Qihui Wu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 本文提出P2AECF框架,通过模块化代理和生成推理实现灵活、高效和适应的低空边缘智能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10080 2026-02-17 cs.CV 57%

BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals

BEVTraj: 无地图端到端鸟瞰图轨迹预测方法,采用可变形注意力和稀疏目标提案

Minsang Kong, Myeongjun Kim, Sang Gu Kang, Hejiu Lu, Yupeng Zhong, Sang Hun Lee

机构 * Department of Automobile and IT Convergence, Kookmin University(汽车与IT融合系,韩国釜山大学) Department of Automotive Engineering, Kookmin University(汽车工程系,韩国釜山大学) Graduate School of Automobile and Mobility, Kookmin University(汽车与移动研究生院,韩国釜山大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 BEVTraj通过可变形注意力和稀疏目标提案实现无地图端到端鸟瞰图轨迹预测,提升自动驾驶的鲁棒性和灵活性。

Comments Submitted to IEEE Transactions on Intelligent Transportation Systems (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏