arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 8666 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人数据与评测 8666 篇

2503.09626 2025-12-29 cs.SI cs.AI cs.LG 62%

Certainly Bot Or Not? Trustworthy Social Bot Detection via Robust Multi-Modal Neural Processes

确定是机器人还是不是?通过鲁棒多模态神经过程进行可信的社交机器人检测

Qi Wu, Yingguang Yang, hao liu, Hao Peng, Buyun He, Yutong Xia, Yong Liao

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出鲁棒多模态神经过程框架,通过增强多模态神经过程的鲁棒性来检测社交机器人,同时提升不确定性估计能力。

Comments We withdraw this paper due to an error identified in the experimental setup. Specifically, the evaluation protocol described in Section 4 does not correctly reflect the intended experimental design, which may affect the validity of the reported results. To avoid potential misunderstanding by readers, we choose to withdraw this version and revise the work before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02519 2025-12-29 cs.AI cs.LG 62%

Creative Agents: Empowering Agents with Imagination for Creative Tasks

创意代理:通过想象力赋能代理以完成创意任务

Penglin Cai, Chi Zhang, Yuhui Fu, Haoqi Yuan, Zongqing Lu

机构 * School of Computer Science Peking University(北京大学计算机学院)

专题命中 机器人数据与评测 :embodied agent(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出通过增强想象力的创意代理,首次在Minecraft生存模式中实现多样化建筑创建,利用大语言模型和扩散模型提升创造力表现。

Comments The first two authors contribute equally

Journal ref Proceedings of the 41st Conference on Uncertainty in Artificial Intelligence (UAI 2025), PMLR 244:471-496

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06157 2025-12-29 cs.CV cs.AI 62%

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

UrbanVideo-Bench: 基于城市空间视频数据评估视觉-语言模型的具身智能

Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li

机构 * Tsinghua University(清华大学)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.AI、cs.CV

AI总结 UrbanVideo-Bench通过城市空间视频数据评估视频大语言模型的具身智能,揭示其在城市环境中感知、推理和导航能力的局限性,并验证了Sim-to-Real迁移的潜力。

Comments 22 pages

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 32400-32423, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18146 2025-12-23 cs.RO cs.AI cs.MA 62%

On Swarm Leader Identification using Probing Policies

基于探测策略的蜂群领头者识别

Stergios E. Bachoumas, Panagiotis Artemiadis

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种基于探测策略的蜂群领头者识别方法,利用TGR层和S5模型实现高效的领头者识别,并通过仿真和实际实验验证了其在对抗环境中的有效性。

Comments 13 pages, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21052 2025-12-22 cs.CV cs.AI cs.MM 62%

FakeParts: a New Family of AI-Generated DeepFakes

FakeParts: 一种新的AI生成深度伪造家族

Ziyi Liu, Firas Gabetni, Awais Hussain Sani, Xi Wang, Soobash Daiboo, Gaetan Brison, Gianni Franchi, Vicky Kalogeiton

机构 * Hi!PARIS, Institut Polytechnique de Paris, Palaiseau, France(巴黎Hi!实验室,巴黎理工学院,Palaiseau,法国) LIX, École Polytechnique, CNRS, Institut Polytechnique de Paris, Palaiseau, France(LIX实验室,巴黎高等学院,国家科学研究中心,巴黎理工学院,Palaiseau,法国) U2IS, ENSTA Paris, Institut Polytechnique de Paris, Palaiseau, France(U2IS,ENSTA巴黎,巴黎理工学院,Palaiseau,法国)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.CV

AI总结 FakeParts是一种通过局部修改真实视频生成的深度伪造技术,通过引入FakePartsBench基准数据集,揭示了现有检测方法在应对部分修改时的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13107 2025-12-19 cs.CV cs.AI 62%

Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather

基于扩散的多模态3D物体检测在恶劣天气中的修复

Zhijian He, Feifei Liu, Yuwei Li, Zhanpeng Luo, Jintao Cheng, Xieyuanli Chen, Xiaoyu Tang

机构 * School of Xingzhi College, South China Normal University(星智学院,华南师范大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学) School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.AI、cs.CV

AI总结 DiffFusion通过基于扩散的修复和自适应跨模态融合,提升多模态3D物体检测在恶劣天气中的鲁棒性与清洁数据性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15957 2025-12-19 cs.CV cs.AI 62%

Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models

看见即信仰(并预测):基于视觉语言模型的上下文感知多人类行为预测

Utsav Panchal, Yuchen Liu, Luigi Palmieri, Ilche Georgievski, Marco Aiello

机构 * Institute of Architecture of Application Systems, University of Stuttgart, Germany(应用系统建筑研究所,斯图加特大学,德国) Bosch Research, Germany(博世研究,德国)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI、cs.CV

AI总结 CAMP-VLM通过结合视觉语言模型与上下文特征,提升了多人类行为预测的准确性,其在预测精度上比基线模型高66.9%。

Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09677 2025-12-18 cs.CV cs.AI 62%

Benchmarking Gaslighting Negation Attacks Against Reasoning Models

对抗性否定攻击下推理模型的基准测试

Bin Zhu, Hailong Yin, Jingjing Chen, Yu-Gang Jiang

机构 * Singapore Management University, Singapore College of Computer Science(新加坡国立管理学院计算机科学学院) Artificial Intelligence, Fudan University, China(人工智能,复旦大学)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.CV

AI总结 本文评估了三种顶级推理模型在对抗性否定攻击下的表现,发现其准确率显著下降,并提出了GaslightingBench-R基准以进一步研究模型对这类攻击的防御能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14601 2025-12-17 cs.CV cs.AI 62%

FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake Videos

FakeRadar: 探测伪造异常以检测未知深度伪造视频

Zhaolun Li, Jichang Li, Yinqi Cai, Junye Chen, Xiaonan Luo, Guanbin Li, Rushi Lan

机构 * Guilin University of Electronic Technology(桂林电子科技大学) Pengcheng Laboratory(鹏城实验室) Sun Yat-sen University(中山大学) Guangxi Key Laboratory of Image and Graphic Intelligent Processing(广西图像与图形智能处理重点实验室) Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.CV

AI总结 FakeRadar通过动态子聚类建模和异常引导训练,有效检测未知深度伪造视频,提升跨域泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13100 2025-12-16 cs.RO cs.AI 62%

OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning

OXE-AugE: 一种大规模机器人增强的OXE以实现跨躯体政策学习的扩展

Guanhua Ji, Harsha Polavaram, Lawrence Yunliang Chen, Sandeep Bajamahal, Zehan Ma, Simeon Adebola, Chenfeng Xu, Ken Goldberg

机构 * Department of EECS, UC Berkeley(电子工程与计算机科学系,加州大学伯克利分校) GRASP Laboratory, University of Pennsylvania(格拉斯实验室,宾夕法尼亚大学) Department of CS, UT Austin(计算机科学系,得克萨斯大学奥斯汀分校)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.RO、cs.AI

AI总结 OXE-AugE通过引入9种不同机器人躯体增强OXE数据集,提升跨躯体政策学习的泛化能力和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03726 2025-12-16 cs.CV cs.RO 62%

Active 6D Pose Estimation for Textureless Objects using Multi-View RGB Frames

基于多视角RGB图像的无纹理物体6D位姿主动估计

Jun Yang, Wenjie Xue, Sahar Ghavidel, Steven L. Waslander

机构 * University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究学院) Robotics Institute(机器人研究所) Epson Canada Ltd(加拿大爱普生有限公司)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO、cs.CV

AI总结 本文提出基于多视角RGB图像的无纹理物体6D位姿主动估计方法,通过两步分解过程提升准确性和效率,并引入主动感知策略减少位姿不确定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07620 2025-12-16 cs.LG cs.AI 62%

DGTEN: A Robust Deep Gaussian based Graph Neural Network for Dynamic Trust Evaluation with Uncertainty-Quantification Support

DGTEN:一种基于深度高斯的图神经网络,用于具有不确定性量化的动态信任评估

Muhammad Usman, Yugyung Lee

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 DGTEN通过结合不确定性意识的消息传递、时间建模和防御机制,实现了动态信任评估中的鲁棒性和准确性提升。

Comments 15 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05622 2025-12-16 cs.CR cs.AI cs.LG 62%

DATABench: Evaluating Dataset Auditing in Deep Learning from an Adversarial Perspective

DATABench: 从对抗角度评估深度学习中的数据集审计

Shuo Shao, Yiming Li, Mengren Zheng, Zhiyang Hu, Yukun Chen, Boheng Li, Yu He, Junfeng Guo, Dacheng Tao, Zhan Qin

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院) Nanyang Technological University(南洋理工大学) Chongqing University(重庆大学) University of Maryland at College Park(马里兰大学学院公园分校)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 DATABench从对抗角度评估深度学习中的数据集审计,提出新的分类法和攻击策略,揭示现有审计方法在对抗环境下的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21502 2025-12-16 cs.LG cs.AI 62%

Process mining-driven modeling and simulation to enhance fault diagnosis in cyber-physical systems

基于过程挖掘的建模与仿真以增强网络物理系统中的故障诊断

Francesco Vitale, Nicola Dall'Ora, Sebastiano Gaiardelli, Enrico Fraccaroli, Nicola Mazzocca, Franco Fummi

机构 * University of Naples Federico II, Department of Electrical Engineering and Information Technology(那不勒斯费德里科二世大学电气工程与信息科技系) Guglielmo Marconi University, Department of Engineering Sciences(古吉莱奥·马尔科尼大学工程科学系)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于过程挖掘的建模与仿真方法,通过可解释的随机Petri网提升网络物理系统故障诊断的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12128 2025-12-16 cs.CV cs.AI 62%

A Benchmark Dataset for Spatially Aligned Road Damage Assessment in Small Uncrewed Aerial Systems Disaster Imagery

用于小无人 aerial 系统灾害影像中空间对齐道路损坏评估的基准数据集

Thomas Manzini, Priyankari Perali, Raisa Karnik, Robin R. Murphy

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.AI、cs.CV

AI总结 本文提出一个用于小无人 aerial 系统灾害影像中道路损坏评估的基准数据集,并通过空间对齐技术提升模型性能。

Comments 11 pages, 6 figures, 6 tables. To appear AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15351 2025-12-15 cs.AI cs.CV 62%

Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration

Octopus: 六种能力协同的代理多模态推理

Yifu Guo, Zishan Xu, Zhiyuan Yao, Yuquan Lu, Jiaye Lin, Sen Hu, Zhenheng Tang, Huacan Wang, Ronghao Chen

机构 * Sun Yat-sen University(中山大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Tsinghua University(清华大学) Peking University(北京大学) The Hong Kong University of Science and Technology(香港科技大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.CV

AI总结 Octopus通过六种能力协同实现多模态代理推理,有效提升复杂任务中的自主探索与动态能力选择能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09463 2025-12-11 cs.CV cs.AI 62%

Privacy-Preserving Computer Vision for Industry: Three Case Studies in Human-Centric Manufacturing

工业隐私保护计算机视觉:人类导向制造中的三个案例研究

Sander De Coninck, Emilio Gamba, Bart Van Doninck, Abdellatif Bey-Temsamani, Sam Leroux, Pieter Simoens

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.AI、cs.CV

AI总结 本文提出一种隐私保护的计算机视觉框架,通过三个工业案例验证,展示如何在监控中减少隐私风险,为人类导向的AI应用提供指导。

Comments Accepted to the AAAI26 HCM workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07729 2025-12-09 cs.CV cs.AI 62%

Improving action classification with brain-inspired deep networks

用脑启发的深度网络改进动作分类

Aidas Aglinskas, Stefano Anzellotti

机构 * Department of Psychology and Neuroscience(心理学与神经科学系)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.AI、cs.CV

AI总结 本研究提出了一种受大脑领域特异性启发的深度网络架构,通过分别处理身体和背景信息,提升动作识别性能并更接近人类表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07351 2025-12-09 cs.CV cs.AI cs.SD 62%

DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection

DeepAgent: 一种双流多智能体融合用于鲁棒多模态深度伪造检测

Sayeem Been Zaman, Wasimul Karim, Arefin Ittesafun Abian, Reem E. Mohamed, Md Rafiqul Islam, Asif Karim, Sami Azam

机构 * Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统实验室) Department of Computer Science and Engineering(计算机科学与工程系) University of Scholars(学者大学) Faculty of Science and Information Technology(科学与信息技术学院) Faculty of Science and Technology(科学与技术学院)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.CV

AI总结 DeepAgent通过双流多智能体融合方法,提升多模态深度伪造检测的鲁棒性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17705 2025-12-09 cs.RO cs.AI 62%

Towards Autonomous and Safe Last-mile Deliveries with AI-augmented Self-driving Delivery Robots

迈向自主安全的最后-mile配送:基于AI增强的自动驾驶配送机器人

Eyad Shaklab, Areg Karapetyan, Arjun Sharma, Murad Mebrahtu, Mustofa Basri, Mohamed Nagy, Majid Khonji, Jorge Dias

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种基于AI增强的自动驾驶配送机器人系统,用于实现小型城市社区的自主安全最后-mile配送,通过整合优化和现实约束,提升配送效率与客户满意度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05230 2025-12-08 cs.RO cs.AI 62%

Invariance Co-training for Robot Visual Generalization

机器人视觉泛化中的不变性协同训练

Jonathan Yang, Chelsea Finn, Dorsa Sadigh

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 该研究通过引入状态相似性和观测扰动不变性辅助任务,提升机器人在多样视角、光照和干扰物体下的泛化能力,实验表明其性能比现有方法提高了18%。

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17665 2025-12-08 cs.CV cs.RO 62%

Perspective-Invariant 3D Object Detection

视角不变的3D物体检测

Ao Liang, Lingdong Kong, Dongyue Lu, Youquan Liu, Jian Fang, Huaici Zhao, Wei Tsang Ooi

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO、cs.CV

AI总结 本文提出Pi3DET数据集和跨平台适应框架,实现视角不变的3D物体检测,推动非车辆平台的3D检测研究。

Comments ICCV 2025; 54 pages, 18 figures, 22 tables; Project Page at https://pi3det.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06580 2025-12-08 cs.CV cs.AI 62%

Exploring Ordinal Bias in Action Recognition for Instructional Videos

探索教学视频中动作识别的序数偏差

Joochan Kim, Minjoon Jung, Byoung-Tak Zhang

机构 * Korea Institute of Science and Technology(韩国科学技术院) Seoul National University(首尔国立大学)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.AI、cs.CV

AI总结 本文提出两种方法探索教学视频中动作识别的序数偏差问题,通过实验揭示模型在面对非常规动作序列时的脆弱性,强调了重新设计评估策略和开发更通用模型的必要性。

Comments Accepted at SCSL @ ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06963 2025-12-04 cs.LG cs.AI stat.ML 62%

Tree Ensembles for Contextual Bandits

基于树集成的上下文老虎机

Hannes Nilsson, Rikard Johansson, Niklas Åkerblom, Morteza Haghir Chehreghani

机构 * Chalmers University of Technology and University of Gothenburg(查尔姆斯理工大学和哥德堡大学) Volvo Car Corporation(沃尔沃汽车公司)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于树集成的上下文老虎机框架,通过改进不确定性估计方法,在减少遗憾和提升计算效率方面优于传统方法。

Comments The first two authors contributed equally to this work

Journal ref Transactions on Machine Learning Research (TMLR), 2024, https://openreview.net/forum?id=59DCkSGw8S, GitHub: https://github.com/HannesNilsson/tree_ensemble_bandits

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14510 2025-12-04 cs.CV cs.GR cs.LG 62%

Deep-BrownConrady: Prediction of Camera Calibration and Distortion Parameters Using Deep Learning and Synthetic Data

Deep-BrownConrady: 使用深度学习和合成数据预测相机校准和畸变参数

Faiz Muhammad Chaudhry, Jarno Ralli, Jerome Leudet, Fahad Sohrab, Farhad Pakdaman, Pierre Corbani, Moncef Gabbouj

机构 * AILiveSim Ltd.(AILiveSim有限公司) Faculty of Information Technology and Communication Sciences, Tampere University(信息科技与通讯科学学院,塔尔基耶大学)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.CV、cs.LG

AI总结 本文提出使用深度学习和合成数据预测相机校准和畸变参数,基于ResNet架构的模型在合成数据集上训练,以提升真实场景下的校准性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04096 2025-12-03 cs.RO cs.CV 62%

Image-Based Relocalization and Alignment for Long-Term Monitoring of Dynamic Underwater Environments

基于图像的重定位与对齐用于动态水下环境的长期监测

Beverley Gorry, Tobias Fischer, Michael Milford, Alejandro Fontan

机构 * Queensland University of Technology(昆士兰理工大学)

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.CV

AI总结 本文提出了一种结合VPR、特征匹配和图像分割的方法,用于水下环境的长期监测,并引入了首个大规模水下VPR基准测试。

Journal ref Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hangzhou, China, pp. 10749-10756, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01816 2025-12-02 cs.CV cs.AI 62%

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

Envision:基于因果世界过程洞察的统一理解和生成基准测试

Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu, Cheng Tan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.AI、cs.CV

AI总结 Envision通过因果事件进程基准测试,评估统一模型在动态时空一致性上的表现,揭示其在因果叙事一致性上的优势及仍需改进的时空一致性挑战。

Comments 35 pages, 12 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21053 2025-12-02 cs.RO cs.CV 62%

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

AerialMind:迈向无人机场景中的指称多目标跟踪

Chenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun, Haocheng Zhao, Haiyun Jiang, Tao Huang, Henghui Ding, Qing-Long Han

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.CV

AI总结 AerialMind是首个针对无人机场景的RMOT基准,通过COALA框架和HETrack方法提升无人机的自然语言交互能力与多目标跟踪性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01755 2025-12-02 cs.CV cs.RO 62%

3EED: Ground Everything Everywhere in 3D

3EED: 在三维中万物皆 grounded

Rong Li, Yuhao Dong, Tianshuai Hu, Ao Liang, Youquan Liu, Dongyue Lu, Liang Pan, Lingdong Kong, Junwei Liang, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 机器人数据与评测 :embodied agent(abstract);分类 cs.RO、cs.CV

AI总结 3EED提出一个大规模多平台多模态三维 grounding 基准测试,通过提供丰富的户外场景数据和跨平台学习技术,推动语言驱动的三维具身感知研究。

Comments NeurIPS 2025 DB Track; 38 pages, 17 figures, 10 tables; Project Page at https://project-3eed.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00080 2025-12-02 cs.CV cs.RO 62%

Conceptual Evaluation of Deep Visual Stereo Odometry for the MARWIN Radiation Monitoring Robot in Accelerator Tunnels

深度视觉立体视觉里程计的概念评估:用于加速器隧道中MARWIN辐射监测机器人的评估

André Dehne, Juri Zach, Peer Stelldinger

机构 * Faculty of Computer Science and Digital Society(计算机科学与数字社会学院) HAW Hamburg(汉堡应用技术大学)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.RO、cs.CV

AI总结 本文评估了深度视觉立体视觉里程计在加速器隧道中为MARWIN机器人提供自主导航的可能性,通过结合3D几何约束,旨在提高在未知环境中的鲁棒性和灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏