arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9819 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9146 篇

2602.16109 2026-02-19 cs.CR cs.AI cs.CE 57%

Federated Graph AGI for Cross-Border Insider Threat Intelligence in Government Financial Schemes

联邦图式AGI用于政府金融方案跨境内部威胁情报

Srikumar Nayak, James Walmesley

机构 * Incedo Inc.(Incedo公司) Indian Institute of Technology(印度理工学院) University of Kent(肯特大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 FedGraph-AGI通过整合AGI推理与联邦图学习,实现了隐私保护下的跨境内部威胁检测,准确率高达92.3%。

Comments 35 Pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15384 2026-02-18 cs.AI cs.CL 57%

World-Model-Augmented Web Agents with Action Correction

具有动作修正的世界模型增强型网络代理

Zhouzhou Shen, Xueyu Hu, Xiyun Li, Tianqing Fang, Juncheng Li, Shengyu Zhang

机构 * Zhejiang University(浙江大学) Tencent AI Lab(腾讯AI实验室)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 WAC通过多代理协作和风险感知机制,提升网络代理在复杂任务中的鲁棒性和执行效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10717 2026-02-12 cs.RO 57%

Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation

说、梦、做:学习视频世界模型以驱动指令式机器人操作

Songen Gu, Yunuo Cai, Tianyu Wang, Simo Wu, Yanwei Fu

机构 * Fudan University(复旦大学)

专题命中 VLA模型 :action model(abstract);分类 cs.RO

AI总结 本文提出一种视频条件动作框架,通过生成稳健的视频模型和对抗性蒸馏,提升机器人操作中的预测能力和空间准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10485 2026-02-12 cs.AI 57%

Abstraction Generation for Generalized Planning with Pretrained Large Language Models

基于预训练大语言模型的通用规划抽象生成

Zhenhe Cui, Huaxiang Xia, Hangjun Shen, Kailun Luo, Yong He, Wei Liang

机构 * Hunan University of Science and Technology(湖南科技大学) Dongguan University of Technology(东莞理工学院)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 本文提出利用预训练大语言模型生成通用规划的抽象,并通过自动化调试修正抽象,以提升规划效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08822 2026-02-10 cs.CV 57%

Any-to-All MRI Synthesis: A Unified Foundation Model for Nasopharyngeal Carcinoma and Its Downstream Applications

任意到全部MRI合成:鼻咽癌及其下游应用的统一基础模型

Yao Pu, Yiming Shi, Zhenxi Zhang, Peixin Yu, Yitao Zhuang, Xiang Wang, Hongzhao Chen, Jing Cai, Ge Ren

专题命中 VLA模型 :VLA(abstract);分类 cs.CV

AI总结 本文提出统一基础模型实现任意到全部MRI合成,提升鼻咽癌放疗的准确性和临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06062 2026-02-03 cs.LG cs.DC 57%

Resolving Extreme Data Scarcity by Explicit Physics Integration: An Application to Groundwater Heat Transport

通过显式物理整合解决极端数据稀缺问题:应用于地下水热传输的应用

Julia Pelzer, Corné Verburg, Alexander Heinlein, Miriam Schulte

机构 * Institute for Parallel and Distributed Systems, University of Stuttgart, Stuttgart, Germany(并行与分布式系统研究所,斯图加特大学,斯图加特,德国) Delft Institute of Applied Mathematics, Delft University of Technology (TU Delft), Delft, the Netherlands(代尔夫特应用数学研究所,代尔夫特理工大学(TU Delft),代尔夫特,荷兰)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 本文提出了一种局部-全局卷积神经网络,通过显式物理整合解决极端数据稀缺问题,应用于地下水热传输建模,展示了其在城市尺度和实际地下参数地图上的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22928 2026-01-26 cs.AI cs.CL 57%

Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning

通过基于强化学习的数值推理增强临床试验论文的论文级别推断

Massimiliano Pronesti, Michela Lorandi, Paul Flanagan, Oisin Redmond, Anya Belz, Yufang Hou

机构 * IBM Research Europe - Ireland(IBM欧洲研究院-爱尔兰) Dublin City University(都柏林城市大学) IT:U Interdisciplinary Transformation University Austria(IT:U跨学科转型大学奥地利)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 本研究通过强化学习和数值推理提升临床试验论文的论文级别推断,实现更准确的系统评价自动化。

Comments Accepted at EMNLP 2025 Main Conference. This revision corrects a minor typo in the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12397 2026-01-21 cs.RO 57%

Learning Diverse Skills for Behavior Models with Mixture of Experts

通过专家混合学习行为模型的多样化技能

Wangtian Shen, Jinming Ma, Mingliang Zhou, Ziyang Meng

专题命中 VLA模型 :action model(abstract);分类 cs.RO

AI总结 Di-BM通过专家混合学习方法,提升机器人行为模型在多任务场景下的性能和数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05271 2026-01-12 cs.CL cs.LG 57%

Enhancing Foundation Models in Transaction Understanding with LLM-based Sentence Embeddings

利用基于大语言模型的句子嵌入增强交易理解的基础模型

Xiran Fan, Zhimeng Jiang, Chin-Chia Michael Yeh, Yuzhong Chen, Yingtong Dou, Menghai Pan, Yan Zheng

机构 * Visa Research(Visa研究)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 本文提出一种结合大语言模型生成的句子嵌入与轻量级交易模型的混合框架,以提升交易理解任务的性能和效率。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track (EMNLP 2025), pages 903-911

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04657 2026-01-09 cs.RO cs.HC 57%

Model of Spatial Human-Agent Interaction with Consideration for Others

考虑他人的空间人机交互模型

Takafumi Sakamoto, Yugo Takeuchi

机构 * Graduate School of Science and Technology, Shizuoka University(静冈大学理学技术研究生院)

专题命中 VLA模型 :action model(abstract);分类 cs.RO

AI总结 本文提出了一种考虑他人的空间人机交互模型,通过实验验证了该模型在调节人机交互行为方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21121 2026-01-07 cs.IR cs.AI 57%

Beyond Patch Aggregation: 3-Pass Pyramid Indexing for Vision-Enhanced Document Retrieval

超越补丁聚合:面向视觉增强文档检索的三阶段金字塔索引

Anup Roy, Rishabh Gyanendra Upadhyay, Animesh Rameshbhai Panara, Robin Mills, Aidan Millar

机构 * Inception AI Mubadala

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 VisionRAG是一种无OCR、模型无关的多模态检索系统,通过三阶段金字塔索引提升文档检索效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01196 2026-01-06 cs.RO 57%

EduSim-LLM: An Educational Platform Integrating Large Language Models and Robotic Simulation for Beginners

EduSim-LLM: 一个整合大语言模型和机器人模拟的教育平台,用于初学者

Shenqi Lu, Liangwei Zhang

机构 * Hangzhou Dianzi University Information Engineering College(杭州电子科技大学信息工程学院)

专题命中 VLA模型 :action model(abstract);分类 cs.RO

AI总结 EduSim-LLM通过整合大语言模型与机器人模拟,实现自然语言到机器人行为的转换,提升初学者在人机交互和多机器人协作中的控制与操作能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09921 2025-12-23 cs.CV 57%

FaceShield: Defending Facial Image against Deepfake Threats

FaceShield:防御面部图像 against 深度伪造威胁

Jaehwan Jeong, Sumin In, Sieun Kim, Hannie Shin, Jongheon Jeong, Sang Ho Yoon, Jaewook Chung, Sangpil Kim

机构 * Korea University(韩国大学) KAIST(韩国科学技术院) Samsung Research(三星研究)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

AI总结 FaceShield通过操控扩散模型的注意力机制和面部特征提取器,提出了一种主动防御深度伪造的方案,有效提升对抗性扰动的鲁棒性和不可察觉性。

Comments Accepted to ICCV 2025. Keywords: Deepfake, Adversarial Attack, Diffusion Models, GANs, Face Swap, Proactive Defense

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07410 2025-12-15 cs.CV 57%

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

InterAgent:基于物理的多智能体命令执行通过交互图上的扩散

Bin Li, Ruichi Zhang, Han Liang, Jingyan Zhang, Juze Zhang, Xin Chen, Lan Xu, Jingyi Yu, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) University of Pennsylvania(宾夕法尼亚大学) ByteDance(字节跳动) Stanford University(斯坦福大学) InstAdapt

专题命中 VLA模型 :action model(abstract);分类 cs.CV

AI总结 InterAgent通过交互图上的扩散模型实现了基于物理的多智能体协调控制,能够从文本提示中生成连贯且物理合理的多代理行为。

Comments Project page: https://binlee26.github.io/InterAgent-Page

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06294 2025-12-09 q-bio.MN cs.LG math.PR q-bio.QM stat.ML 57%

Interpretable Neural Approximation of Stochastic Reaction Dynamics with Guaranteed Reliability

可解释的神经近似随机反应动力学:具有保证可靠性的方法

Quentin Badolle, Arthur Theuer, Zhou Fang, Ankit Gupta, Mustafa Khammash

机构 * Department of Biosystems Science and Engineering, ETH Zurich(1 生物系统科学与工程系,苏黎世联邦理工学院)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 DeepSKA通过结合可解释性、可靠性保证和高效计算,为随机反应动力学提供了一种新的神经近似方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06290 2025-12-09 cs.CV 57%

StrokeNet: Unveiling How to Learn Fine-Grained Interactions in Online Handwritten Stroke Classification

StrokeNet: 解析如何在在线手写体识别中学习细粒度交互

Yiheng Huang, Shuang She, Zewei Wei, Jianmin Lin, Ming Yang, Wenyin Liu

机构 * College of Computer Science and Technology(计算机科学与技术学院) Guangdong University of Technology(广东技术大学) CVTE Research(CVTE研究院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

AI总结 StrokeNet通过参考点对表示和动态选择参考点,有效捕捉手写体中细粒度交互,提升识别准确率。

Comments 17 pages, 5 figures

Journal ref ICDAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00080 2025-12-08 cs.SI cs.AI 57%

SoREX: Towards Self-Explainable Social Recommendation with Relevant Ego-Path Extraction

SoREX:面向具有相关自我路径提取的自解释社交推荐

Hanze Guo, Yijun Ma, Xiao Zhou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallagher人工智能学院) Beijing Key Laboratory of Research on Large Models(北京大模型研究关键实验室) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索工程研究中心)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 SoREX通过引入自解释的GNN框架,利用相关自我路径提取提升社交推荐的预测准确性与解释能力。

Comments ACM Transactions on Information Systems (TOIS), 2025. Online AM: 17 Nov 2025. DOI: 10.1145/3777374. Code: https://github.com/antman9914/SoREX

Journal ref ACM Transactions on Information Systems (TOIS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01234 2025-12-03 cs.HC cs.AI 57%

Proactive Agentic Whiteboards: Enhancing Diagrammatic Learning

主动代理白板:增强图示学习

Suveen Ellawela, Sashenka Gamage, Dinithi Dissanayake

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 DrawDash通过实时语音驱动的视觉辅助,主动完成和优化教育图表,以减少教师认知负担并提升图示教学效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00379 2025-12-02 q-bio.BM cs.LG 57%

EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants

EnzyCLIP:一种基于对比学习的跨注意力双编码框架,用于预测酶动力学常数

Anas Aziz Khan, Md Shah Fahad, Priyanka, Ramesh Chandra, Guransh Singh

机构 * SCOPE Vellore Institute of Technology(维洛雷理工学院) BIT Department of Computer Science(计算机科学系) Department of Bioengineering and Biotechnology(生物工程与生物技术系)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 EnzyCLIP通过对比学习和跨注意力机制,结合蛋白质序列和底物分子结构预测酶动力学参数,提升Kcat和Km预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04966 2025-12-01 cs.SD cs.AI eess.AS 57%

LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning

LAPS-Diff: 一种基于扩散的歌唱语音合成框架,具有语言感知的语调风格引导学习

Sandipan Dhar, Mayank Gupta, Preeti Rao

机构 * Department of Electrical Engineering, Indian Institute of Technology, Bombay.(电子工程系,印度理工学院,博亚姆)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 LAPS-Diff通过语言感知嵌入和语音风格引导学习,提升低资源环境下印地语歌唱语音合成的质量。

Comments 10 pages, 5 figures, 3 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19952 2025-11-26 cs.LG 57%

Hierarchical Spatio-Temporal Attention Network with Adaptive Risk-Aware Decision for Forward Collision Warning in Complex Scenarios

具有自适应风险感知决策的分层时空注意力网络用于复杂场景的前方碰撞预警

Haoran Hu, Junren Shi, Shuo Jiang, Kun Cheng, Xia Yang, Changhao Piao

机构 * organization= School of Automation \& School of Industrial Internet, Chongqing University of Posts organization= Platform Technology Development Department, AVATR Technology Co. LTD. , city= Chongqing , postcode= 400000 , country= China organization= School of Vehicle Mobility, Tsinghua University , city= Beijing , postcode= 100084 , country= China

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 本文提出一种结合分层时空注意力网络和动态风险阈值调整算法的前方碰撞预警框架,通过高效模型和自适应机制提升复杂场景下的预警精度与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17443 2025-11-25 cs.HC cs.AI cs.GR 57%

GRAPHIC--Guidelines for Reviewing Algorithmic Practices in Human-centred Design and Interaction for Creativity

GRAPHIC--指导人本设计与交互中算法实践的审查指南

Joana Rovira Martins, Pedro Martins, Ana Boavida

专题命中 VLA模型 :action model(abstract);分类 cs.AI

AI总结 GRAPHIC框架旨在分析应用于图形设计的计算系统,通过三大维度揭示人机协作中的研究空白,促进基于核心设计原则的创新。

Comments 20 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13032 2025-11-18 cs.CV 57%

Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts

Sheng Liu, Yuanzhi Liang, Jiepeng Wang, Sidan Du, Chi Zhang, Xuelong Li

机构 * Nanjing University(南京大学) Institute of Artificial Intelligence, China Telecom (TeleAI)(人工智能研究所)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23203 2025-10-28 cs.CV 57%

DecoDINO: 3D Human-Scene Contact Prediction with Semantic Classification

Lukas Bierling, Davide Pasero, Fleur Dolmans, Helia Ghasemi, Angelo Broere

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22538 2025-10-28 cs.LG 57%

Iteratively Refined Early Interaction Alignment for Subgraph Matching based Graph Retrieval

Ashwin Ramachandran, Vaibhav Raj, Indrayumna Roy, Soumen Chakrabarti, Abir De

机构 * UC San Diego(加州大学圣地亚哥分校) IIT Bombay(印度理工学院班加罗尔)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Journal ref Neurips 2024 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19478 2025-10-23 cs.CV 57%

Mitigating representation bias caused by missing pixels in methane plume detection

Julia Wąsala, Joannes D. Maasakkers, Ilse Aben, Rochelle Schneider, Holger Hoos, Mitra Baratchi

机构 * Leiden Institute for Advanced Computer Science (LIACS)(莱顿高级计算机科学研究所) SRON Space Research Organization Netherlands(荷兰空间研究组织SRON) Department of Earth Sciences(地球科学系) Φ \Phi -lab, ESA-ESRIN(Φ实验室,ESA-ESRIN) Chair for AI Methodology (AIM) at RWTH Aachen(RWTH亚琛大学人工智能方法学教授职位)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Accepted at the MACLEAN workshop at ECML-PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06679 2025-10-09 cs.CV 57%

DreamOmni2: Multimodal Instruction-based Editing and Generation

Bin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang, Jiyang Liu, Jingyao Li, Haoru Tan, Sitong Wu, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) HKUST(香港科技大学) HKU(香港大学) ByteDance Inc(字节跳动公司)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06504 2025-10-09 cs.CV 57%

Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation

Qingxuan Wu, Zhiyang Dou, Chuan Guo, Yiming Huang, Qiao Feng, Bing Zhou, Jian Wang, Lingjie Liu

机构 * University of Pennsylvania(宾夕法尼亚大学) The University of Hong Kong(香港大学) Snap Inc(Snap公司)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05171 2025-10-08 cs.LG cs.CY 57%

Carbon Emission Prediction in China Considering New Quality Productive Forces Using a Deep & Corss Learning Modeling Framework

Haijin Xie, Gongquan Zhang

专题命中 VLA模型 :action model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25085 2025-10-07 cs.CL cs.AI cs.IR 57%

jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking

Feng Wang, Yuqing Li, Han Xiao

机构 * Jina AI GmbH(Jina AI公司) University of Pittsburgh(匹兹堡大学)

专题命中 VLA模型 :action model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏