arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-05-01 至 2026-05-01 共收录 57 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人学习 2 篇

2604.27621 2026-05-01 cs.RO cs.CV 88%

Robot Learning from Human Videos: A Survey

从人类视频学习机器人:综述

Junyi Ma, Erhang Zhang, Haoran Yang, Ditao Li, Chenyang Xu, Guangming Wang, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) University of Cambridge(剑桥大学)

专题命中 机器人学习 :robot learning(title);robotics(abstract);embodied AI(abstract);manipulation(abstract)

AI总结 本文综述了通过人类视频学习机器人技能的方法,探讨了政策学习基础、人类视频接口、技能转移层次分类及数据基础,分析了挑战与未来研究方向。

Comments Paper list: https://github.com/IRMVLab/awesome-robot-learning-from-human-videos

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27667 2026-05-01 cs.RO cs.LG 84%

Can Tabular Foundation Models Guide Exploration in Robot Policy Learning?

表基础模型能否指导机器人策略学习中的探索?

Buqing Ou, Frederike Dümbgen

机构 * Department of Mechanical Engineering, Carnegie Mellon University(卡内基梅隆大学机械工程系) Inria, Département d’informatique de l’ENS, CNRS, PSL Research University(法国国家科学研究中心(CNRS)、巴黎综合理工学院(ENS)和PSL研究大学的Inria)

专题命中 机器人学习 :robot policy(title,abstract);robotics(abstract);分类 cs.RO、cs.LG

AI总结 本文提出TFM-S3方法,通过结合局部和全局搜索提升机器人策略学习的全局探索效率,实验显示其在连续控制基准上加速收敛并提升性能。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 机器人操作 12 篇

2602.00937 2026-05-01 cs.RO cs.AI cs.CV cs.LG 90%

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining

CLAMP: 基于对比学习的3D多视角动作条件机器人操控预训练

I-Chun Arthur Liu, Krzysztof Choromanski, Sandy Huang, Connor Schenck

机构 * Google DeepMind(谷歌DeepMind) University of Southern California(南加州大学)

专题命中 机器人操作 :manipulation(title,abstract);robotic(title,abstract);分类 cs.RO、cs.AI、cs.CV;robotics(comments)

AI总结 CLAMP通过3D点云和机器人动作进行预训练,利用对比学习提升机器人操控精度与效率,优于现有基线方法。

Comments Accepted to the Robotics: Science and Systems (RSS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22615 2026-05-01 cs.RO 88%

GazeVLA: Learning Human Intention for Robotic Manipulation

GazeVLA:学习人类意图以促进机器人操作

Chengyang Li, Kaiyi Xiong, Yuan Xu, Lei Qian, Yizhou Wang, Wentao Zhu

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology, Ningbo(宁波东部科技研究院) Peking University(北京大学) ShanghaiTech University(上海科技大学)

专题命中 机器人操作 :manipulation(title,abstract);robotic(title,abstract);分类 cs.RO

AI总结 本文提出GazeVLA框架,通过学习人类意图弥合人机之间的固有差距,提升机器人操作性能。

Comments Project page: https://gazevla.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27175 2026-05-01 cs.RO 79%

Global Sampling-Based Trajectory Optimization for Contact-Rich Manipulation via KernelSOS

基于全局采样的轨迹优化用于接触密集操作的KernelSOS方法

Zhongqi Wei, Frederike Dümbgen

机构 * Department of Mechanical Engineering, Carnegie Mellon University(卡内基梅隆大学机械工程系) Inria, Département d’informatique de l’ENS, CNRS, PSL Research University(法国国家科学研究中心(CNRS)、巴黎高等师范学院(ENS)和PSL研究大学的Inria)

专题命中 机器人操作 :manipulation(title,abstract);分类 cs.RO

AI总结 本文提出Global-MPPI框架,结合全局探索与局部优化,解决接触密集操作中的高维、长时域问题,实验显示其收敛更快且成本更低。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28161 2026-05-01 cs.RO 70%

RopeDreamer: A Kinematic Recurrent State Space Model for Dynamics of Flexible Deformable Linear Objects

RopeDreamer:一种用于柔性可变形线性物体动态的运动学递归状态空间模型

Tim Missal, Lucas Domingues, Berk Guler, Simon Manschitz, Jan Peters, Paula Dornhofer Paro Costa

机构 * Technical University of Darmstadt(德意志技术大学) School of Electrical and Computer Engineering, Universidade Estadual de Campinas (UNICAMP)(坎皮纳斯州立大学电气与计算机工程学院) Instituto de Pesquisas Eldorado(Eldorado研究所) Honda Research Institute Europe GmbH(本田欧洲研究院) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Robotics Institute Germany (RIG)(德国机器人研究所) Centre for Cognitive Science(认知科学研究中心) Artificial Ingelligence Lab, Recod.ai(Recod.ai人工智能实验室)

专题命中 机器人操作 :manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出结合递归状态空间模型与四元数运动链表示的潜变量框架,用于预测柔性可变形线性物体的状态,通过约束物理有效流形减少自交和非物理变形,提升长周期预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27557 2026-05-01 cs.RO 70%

Function-based Parametric Co-Design Optimization of Dexterous Hands

基于函数的参数化双工手优化设计

Mohammad Amin Mirzaee, Harsh Gupta, Wenzhen Yuan

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 机器人操作 :manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出了一种综合参数框架,统一手掌结构、手指运动学、指尖几何和微尺度表面曲率,通过参数化表面变形核直接影响接触交互,提升抓取稳定性任务的仿真与现实优化性能。

Comments 8 pages, 7 figures, https://www.aminmirzaee.com/HandCDO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27414 2026-05-01 cs.CV cs.CR cs.LG 62%

Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis

理解视觉语言模型在自动驾驶中的对抗转移性:跨架构分析

David Fernandez, Pedram MohajerAnsari, Amir Salarpour, Mert D. Pese

机构 * School of Computing, Clemson University, USA(克莱姆森大学计算机学院)

专题命中 机器人操作 :manipulation(abstract);分类 cs.CV、cs.LG

AI总结 本文通过跨架构研究探讨视觉语言模型在自动驾驶中的对抗转移性,评估了三种架构在不同场景下的对抗效果,发现高跨架构有效性。

Comments 9 pages, 2 figures. Accepted at SAE WCX 2026

Journal ref SAE Technical Paper 2026-01-0170, SAE WCX 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28183 2026-05-01 cond-mat.mtrl-sci cond-mat.mes-hall physics.app-ph 50%

Uniaxial strain-driven ferroelastic domain control in LaAlO3

单轴应变驱动的LaAlO3铁电弹性域控制

Matthias Roeper, Robin Buschbeck, Jakob Wetzel, Tobias Ritschel, Anna-Lena Hofmann, Vladyslav Kovtunovych, Mike N. Pionteck, Javier Taboada-Gutiérrez, Alexey B. Kuzmenko, Martina Basini, Vivek Unikandanunni, Iuliia Kiseleva, Jochen Geck, Susanne C. Kehr, Maximilian Lederer, Simone Sanna, Lukas M. Eng, Samuel D. Seddon

专题命中 机器人操作 :manipulation(abstract)

AI总结 通过单轴应变调控LaAlO3单晶的铁电弹性域结构,实现可控的域分布变化,为异质结构中的主动实时编程提供新方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27518 2026-05-01 cs.HC math.OC 50%

lpviz: Interactive Linear Programming Visualization

lpviz: 交互式线性规划可视化

Evan Grand, Michael Klamkin

专题命中 机器人操作 :manipulation(abstract)

AI总结 lpviz是一款基于浏览器的线性规划可视化工具,提供直观界面让用户直接绘制和编辑可行域及目标向量,支持比较多种线性规划算法的性能,且开源免费。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27417 2026-05-01 cond-mat.mes-hall physics.optics 50%

Mobile Exceptional Points Generate Momentum-Space Switching Domains

移动的极点生成动量空间切换域

Jung-Wan Ryu, Chang-Hwan Yi

专题命中 机器人操作 :manipulation(abstract)

AI总结 研究探讨了在周期调制下移动极点如何生成动量空间切换域,通过两能带模型提出带排列不变量,揭示了极点运动对布里渊区带切换行为的影响。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27281 2026-05-01 cs.SD 50%

Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints

声调转换:社会语言学与技术约束驱动的调研

Yurii Halychanskyi, Jianfeng Steven Guo, Volodymyr Kindratenko

机构 * Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学与数据科学学院) National Center for Supercomputing Applications, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校国家超级计算应用中心) Department of East Asian Languages and Cultures, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校东亚语言与文化系)

专题命中 机器人操作 :manipulation(abstract)

AI总结 本文调研了声调转换方法的发展,分析了数据对齐、表征解耦和资源稀缺等挑战,回顾了从早期规则基方法到现代神经网络架构的演变,并探讨了应用需求对声调修改与说话人身份保护的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27202 2026-05-01 cs.CR 50%

Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives

野外间接提示注入:网页中间接提示注入的实证研究

Soheil Khodayari, Xuenan Zhang, Bhupendra Acharya, Giancarlo Pellegrino

专题命中 机器人操作 :manipulation(abstract)

AI总结 研究通过分析网页和HTTP响应中的间接提示注入实例,揭示其在真实环境中的普遍存在性、技术手段及影响,发现大部分指令针对机器而非人类,且存在多种不同的目标和影响。

Comments 18 pages total, 12 pages main content, 8 figures, and 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27041 2026-05-01 econ.GN q-fin.EC q-fin.TR 50%

The Signal Credibility Index for Prediction Markets: A Microstructure-Grounded Diagnostic with Weighted and Time-Varying Extensions

预测市场信号可信度指数:基于微观结构的诊断方法及其加权和时间变化扩展

Maksym Nechepurenko

专题命中 机器人操作 :manipulation(abstract)

AI总结 本文提出信号可信度指数作为独立诊断工具,通过改进的持久性组件、加权柯布-道格拉斯形式、时间变化规格及蒙特卡洛验证,区分微观结构制度,揭示协调可信度而非纯信息内容的两种失效模式。

Comments 19 pages, 5 figures, 5 tables. Companion to arXiv:2604.24147. Replication code: https://github.com/ForesightFlow/signal-credibility-index

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 具身导航 9 篇

2604.27620 2026-05-01 cs.CV 83%

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

SpaAct:基于课程适应的空间激活转换学习用于视觉语言导航

Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Hanbing Li, Lin Zhao, Kailin Lyu, Long Chen, Zhi-Xin Yang, Nanning Zheng

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) The State Key Laboratory of Internet of Things for Smart City(智能城市物联网国家重点实验室) Centre for Artificial Intelligence and Robotics(人工智能与机器人中心) University of Macau(澳门大学) Xiaomi EV(小i EV) School of Automation(自动化学院) Beijing Institute of Technology(北京理工大学) Institute of Automation(自动化研究所) Chinese Academy of Sciences(中国科学院)

专题命中 具身导航 :navigation(title,abstract);embodied agent(abstract);分类 cs.CV

AI总结 SpaAct通过引入空间激活任务和课程适应方法,提升视觉语言导航中动态空间感知能力,实现更高效的导航性能。

Comments Submmited to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13559 2026-05-01 cs.AI 79%

OpAgent: Operator Agent for Web Navigation

OpAgent:网页导航的运算代理

Yuyu Guo, Wenjie Yang, Siyuan Yang, Ziyang Liu, Cheng Chen, Yuan Wei, Yun Hu, Yang Huang, Guoliang Hao, Dongsheng Yuan, Jianming Wang, Xin Chen, Hang Yu, Lei Lei, Peng Di

机构 * Ant Group(蚂蚁集团)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI

AI总结 本文提出OpAgent,通过在线强化学习优化网页导航策略,结合层级多任务微调和混合奖励机制,实现71.6%的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27010 2026-05-01 cs.HC 78%

Quantifying the Cost of Manual Navigation: A Comparison of Gesture-Based Magnification versus Direct Access Reading in Digital Layout-based Documents

量化手动导航的成本:手势放大与直接访问阅读在基于数字布局的文档中的比较

Sebastián Gallardo, Hui-Yin Wu, Dorian Mazauric, Pierre Kornprobst, Monica Di Meo, Stéphanie Baillif, Aurelie Calabrese

专题命中 具身导航 :navigation(title,abstract)

AI总结 研究比较手势放大与直接访问阅读在数字布局文档中的表现,发现大字号版块在阅读速度和目标定位效率上更优,且恢复了自然的阅读策略,同时降低用户负荷并提升偏好。

Journal ref IMX 2026 - International Conference on Interactive Media Experiences, Technological University of the Shannon, Jun 2026, Athlone, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20990 2026-05-01 cs.HC 71%

The Impact of Navigation on Proxemics in an Immersive Virtual Environment with Conversational Agents

导航对沉浸式虚拟环境中人际距离的影响:与对话代理的互动

Rose Connolly, Lauren Buck, Victor Zordan, Rachel McDonnell

专题命中 具身导航 :navigation(title)

AI总结 研究探讨了在沉浸式虚拟环境中,导航方式对人际距离的影响,发现 teleportation 使参与者与对话代理保持更近的距离,且女性与男性在距离上存在差异,自然行走则带来更高的自主感和身体所有权。

Comments for the associated supplementary video, see project page https://connolr3.github.io/TeleportationIPD Accepted for presentation at IEEE VR 2025 and for publication in a special issue of the IEEE Transactions on Visualization and Computer Graphics (IEEE TVCG) file was updated 30.04.2026 to include updated grant information

Journal ref IEEE Transactions on Visualization and Computer Graphics ( Volume: 31, Issue: 5, May 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27450 2026-05-01 cs.RO cs.AI 62%

RAY-TOLD: Ray-Based Latent Dynamics for Dense Dynamic Obstacle Avoidance with TDMPC

RAY-TOLD: 基于射线的任务导向潜在动力学用于密集动态障碍物避障与TDMPC

Seungho Han, Seokju Lee, Jeonguk Kang

机构 * School of Electrical Engineering, Hanyang University(翰阳大学电气工程学院) Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(机械工程系,韩国科学技术院(KAIST)机电系统与控制实验室(MSC实验室)) Samsung Research, Samsung Electronics(三星研究所,三星电子)

专题命中 具身导航 :navigation(abstract);分类 cs.RO、cs.AI

AI总结 本文提出RAY-TOLD,结合物理基础MPPI的鲁棒性与强化学习的长视界,通过LiDAR中心的潜在动力学模型实现动态障碍物避障,提升导航可靠性与安全性。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27383 2026-05-01 eess.IV cs.CV 57%

A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation

一种实时尺度鲁棒网络用于鼻内气管插管中的声带分割

Yang Zhou, Chaoyong Zhang, Ruoyi Hao, Huilin Pan, Yang Zhang, Hongliang Ren

机构 * School of Mechanical Engineering, Hubei University of Technology(湖北工业大学机械工程学院) National Key Laboratory for Novel Software Technology, Department of Computer Science and Technology, Nanjing University(南京大学新型软件技术国家重点实验室,计算机科学与技术系) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) Shun Hing Institute of Advanced Engineering, The Chinese University of Hong Kong(香港中文大学顺安先进工程研究院)

专题命中 具身导航 :navigation(abstract);分类 cs.CV

AI总结 本文提出了一种轻量级多感受野特征提取模块,用于提高鼻内气管插管中声带分割的鲁棒性和精度,实验表明其在三个数据集上表现优异,达到92.9%的mDice分数。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27253 2026-05-01 cs.AI 57%

AutoSurfer -- Teaching Web Agents through Comprehensive Surfing, Learning, and Modeling

AutoSurfer -- 通过全面浏览、学习和建模教学网络代理

Fazle Elahi Faisal, Qianhui Wu, Baolin Peng, Jianfeng Gao

机构 * Microsoft Research(微软研究院)

专题命中 具身导航 :navigation(abstract);分类 cs.AI

AI总结 AutoSurfer通过系统性广度优先探索策略、任务合成引导和轨迹优化,实现全面覆盖网站操作空间,提升网络代理轨迹生成的准确性和多样性。

Comments 21 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17435 2026-05-01 cs.RO 57%

ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

ImagineNav++: 通过场景想象促使视觉语言模型作为具身导航器

Teng Wang, Xinxin Zhao, Wenzhe Cai, Changyin Sun

机构 * School of Automation, Southeast University(东南大学自动化学院)

专题命中 具身导航 :navigation(abstract);分类 cs.RO

AI总结 本文提出ImagineNav++,通过场景想象将视觉语言模型用于无地图导航,利用想象模块生成高探索潜力的视点,并通过选择性聚焦记忆机制实现空间一致性,实验表明其在无地图导航中表现优异。

Comments 17 pages, 10 figures. arXiv admin note: text overlap with arXiv:2410.09874

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15350 2026-05-01 cs.RO cs.NE 57%

Nauplius Optimisation for Autonomous Hydrodynamics

幼体优化用于自主水动力学

Shyalan Ramesh, Scott Mann, Alex Stumpf

机构 * La Trobe University(拉特罗布大学)

专题命中 具身导航 :robotics(abstract);分类 cs.RO

AI总结 本文提出NOAH算法,结合水流感知漂移、不可逆锚定和群体通信,解决水下集群机器人在强流和持续感知中的优化问题。

Comments IEEE Access, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 具身推理 8 篇

2512.24329 2026-05-01 cs.CL 82%

World model inspired sarcasm reasoning with large language model agents

受世界模型启发的大型语言模型代理 sarcasm 推理

Keito Inoshita, Shinnosuke Mizuno

机构 * Faculty of Business and Commerce, Kansai University(大阪 kansai 大学 商业与文理学院) Faculty of Medicine, The University of Tokyo(东京大学 医学部)

专题命中 具身推理 :world model(title,abstract)

AI总结 本文提出 WM-SAR 模型,通过分解语义不一致性和意图进行 sarcasm 推理,实现高可解释性和性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28196 2026-05-01 cs.CV 79%

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++:迈向统一的驾驶世界模型用于3D场景理解和生成

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Mach Drive University of Hong Kong(香港大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出HERMES++,一种统一的驾驶世界模型,整合3D场景理解和未来几何预测。通过BEV表示、LLM增强世界查询和当前到未来链接等设计,提升驾驶场景的生成与理解能力。

Comments Extended version of ICCV 25 paper HERMES, Code: https://github.com/H-EmbodVis/HERMESV2, Project page: https://h-embodvis.github.io/HERMESV2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27935 2026-05-01 cs.RO cs.SY eess.SP eess.SY 79%

Flying by Inference: Active Inference World Models for Adaptive UAV Swarms

通过推断飞行:用于自适应无人机群的主动推断世界模型

Kaleem Arshid, Ali Krayani, Lucio Marcenaro, David Martin Gomez, Carlo Regazzoni

机构 * Department of Engineering and Naval Architecture (DITEN), University of Genoa(工程与 naval 架构系(DITEN),热那亚大学) Intelligent Systems Laboratory, Department of Systems Engineering and Automation, Carlos III University of Madrid(智能系统实验室,系统工程与自动化系,马德里卡洛斯三世大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO

AI总结 本文提出了一种专家引导的主动推断框架,用于自适应无人机群轨迹规划。该方法将多无人机轨迹设计转化为分层概率推断问题,通过遗传算法生成专家示范并学习世界模型,实现高效的轨迹规划与动态调整。

Comments Submitted to IEEE journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27895 2026-05-01 cs.AI 79%

Graph World Models: Concepts, Taxonomy, and Future Directions

图世界模型:概念、分类与未来方向

Jiawei Liu, Senqiao Yang, Mingjun Wang, Yu Wang, Bei Yu

机构 * The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文系统阐述了图世界模型的概念,分类了基于关系归纳偏置的三种类型,并探讨了其未来研究方向与挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28122 2026-05-01 cs.CV cs.LG 62%

Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces

超越高斯瓶颈:基于拓扑对齐的视觉Transformer特征空间编码

Andrew Bond, Ilkin Umut Melanlioglu, Erkut Erdem, Aykut Erdem

机构 * Department of Computer Engineering, Koç University, Istanbul, Turkey(科克大学计算机工程系,伊斯坦布尔,土耳其) Department of Computer Engineering, Hacettepe University, Ankara, Turkey(哈恰塔佩大学计算机工程系,安卡拉,土耳其) KUIS AI Research Center, Istanbul, Turkey(KUIS人工智能研究中心,伊斯坦布尔,土耳其) Department of Electrical and Electronics Engineering, Koç University, Istanbul, Turkey(科克大学电气与电子工程系,伊斯坦布尔,土耳其)

专题命中 具身推理 :world model(abstract);分类 cs.CV、cs.LG

AI总结 本文提出S²VAE框架,通过压缩和表示场景的3D状态,包括相机运动、深度和点结构,以提升视觉模型的几何一致性。实验显示,几何对齐的超球面隐空间在高压缩条件下优于传统高斯瓶颈。

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27899 2026-05-01 cs.AI 57%

Simulating clinical interventions with a generative multimodal model of human physiology

用生成式多模态模型模拟临床干预

Guy Lutsker, Gal Sapir, Jordi Merino, Smadar Shilo, Anastasia Godneva, Eli Meirom, Shie Mannor, Hagai Rossman, Gal Chechik, Eran Segal

机构 * Department of Computer Science and Applied Mathematics, Weizmann Institute of Science(魏茨曼科学研究所计算机科学与应用数学系) Department of Molecular Cell Biology, Weizmann Institute of Science(魏茨曼科学研究所分子细胞生物学系) NVIDIA Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen(诺沃维克基金会基础代谢研究中心,哥本哈根大学) Faculty of Medical and Health Sciences, Tel Aviv University(特拉维夫大学医学与健康科学学院) The Jesse Z and Sara Lea Shafer Institute for Endocrinology and Diabetes, National Center for Childhood Diabetes, Schneider Children’s Medical Center of Israel(杰西Z和索菲亚·李·沙弗内分泌学与糖尿病研究所,以色列儿童糖尿病国家中心,施耐德儿童医学中心) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 具身推理 :world model(abstract);分类 cs.AI

AI总结 本文提出HealthFormer模型,通过训练人类表型项目数据,生成人类生理轨迹,实现对个体生理变化的预测和干预模拟,提升临床风险评分和疾病预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27576 2026-05-01 cs.LO cs.LG 57%

BAss: Symbolic Reasoning in Abstract Dialectical Frameworks

BAss:基于BDD的抽象辩证框架符号推理

Samuel Pastva, Van-Giang Trinh

机构 * Faculty of Informatics, Masaryk University(马萨里克大学信息学院) Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT)(胡志明市技术大学计算机科学与工程学院)

专题命中 具身推理 :world model(abstract);分类 cs.LG

AI总结 BAss基于BDD提出新型分析工具,实现抽象辩证框架的所有可接受、完整和优先解释的全符号计算,优于现有工具并在大规模解空间场景中表现优异,推动系统生物学研究。

详情

展开后加载摘要…

URL PDF HTML 收藏