CommentsAuthor Accepted Manuscript. Accepted for publication in the Proceedings of the 34th ACM International Conference on Multimedia (ACM MM '26). This author-created manuscript is not the ACM Version of Record
Legible and Intuitive Multi-modal Robot State and Intent Communication Validated in Online and Real-world Studies
可读且直观的多模态机器人状态与意图通信:在线和真实世界研究验证
Tim Schreiter, Jens V. Rüppel, Andrey Rudenko, Martin Magnusson, Achim J. Lilienthal
机构
*
Chair of Perception for Intelligent Systems, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(慕尼黑工业大学慕尼黑机器人与机器智能研究所智能系统感知教席)
;
Centre for Applied Autonomous Sensor Systems (AASS), Örebro University(厄勒布鲁大学应用自主传感器系统中心)
;
Robotics Institute Germany (RIG)(德国机器人研究所)
MoCA: Multi-modal Cross-masked Autoencoder for Time Series in Digital Health
MoCA:用于数字健康测量的多模态交叉掩码自编码器
Howon Ryu, Yuliang Chen, Yacun Wang, Andrea Z. LaCroix, Chongzhi Di, Loki Natarajan, Yu Wang, Jingjing Zou
机构
*
Herbert Wertheim School of Public Health and Human Longevity Science, University of California, San Diego, San Diego, CA, USA(赫伯特·韦瑟姆公共卫生与人类长寿科学学院,加州大学圣地亚哥分校)
;
Halıcıoğlu Data Science Institute, University of California San Diego, San Diego, CA, USA(哈利奇奥卢数据科学研究所,加州大学圣地亚哥分校)
;
Department of Computer Science and Engineering, University of California San Diego, San Diego, CA, USA(计算机科学与工程系,加州大学圣地亚哥分校)
;
Division of Public Health Sciences, Fred Hutchinson Cancer Center, Seattle, WA, USA(公共卫生科学部,Fred Hutchinson癌症中心)
MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding
MAC 2026:推动微动作分析迈向细粒度理解
Kun Li, Dan Guo, Jihao Gu, Pengyu Liu, Xiaobai Li, Haoyu Chen, Yanbin Hao, Guoying Zhao, Meng Wang
机构
*
United Arab Emirates University(阿联酋大学)
;
Hefei University of Technology(合肥工业大学)
;
University College London(伦敦大学学院)
;
Zhejiang University(浙江大学)
;
University of Oulu(奥卢大学)
;
CMVS, University of Oulu(奥卢大学计算机视觉与媒体研究中心)
机构
*
State Key Laboratory of Internet of Things for Smart City, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系智慧城市物联网国家重点实验室)
;
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
专题命中
多模态生成
:MLLM(summary_cn,abstract);分类 cs.CV
AI总结
提出SEED系统,结合相似性引导数据增强、单一ViT联合检测与定位、以及基于MLLM的演化报告生成,在ACM MM 2026文本伪造挑战赛中获得第三名。
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
MMaDA-VLA: 基于统一多模态指令与生成的大型扩散视觉-语言-动作模型
Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang
机构
*
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
East China University of Science and Technology(华东理工大学)
;
Huawei Celia Team(华为Celia团队)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
OpenHelix Robotics
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
RealityBridge: 连接可编辑3D高斯泼溅驾驶模拟与现实世界视频
Zhenhua Wu, Yun Pang, Mingkun Chang, Yuwei Ning, Liangzhi Wang, Yi Xiao, Guanbin Li
机构
*
Sun Yat-sen University(中山大学)
;
Guangdong Key Laboratory of Information Security Technology(广东省信息安全技术重点实验室)
;
Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education(教育部机器智能与先进计算重点实验室)
机构
*
South China University of Technology(南方科技大学)
;
StepFun
;
Institute of Automation Chinese Academy of Sciences(中国科学院自动化研究所)
;
Nanyang Technological University(南洋理工大学)
;
The Chinese University of Hong Kong(香港中文大学)
RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care
RESPClinBench:呼吸专科医疗中的多模态临床决策与纵向疾病管理基准测试
Mouxiao Bian, Zhi Chen, Ruiyao Chen, Lu Lu, Hengrui Liang, Chaoyi Huang, Yiluo Lin, Jingru Ding, Yun Zhong, Yueming Su, Jie Xu
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Macau University of Science and Technology(澳门科技大学)
;
First Affiliated Hospital of Guangzhou Medical University(广州医科大学附属第一医院)
;
Guangzhou Institute of Respiratory Health(广州呼吸健康研究院)
Space2Ground 2.0: A Multi-Source Dataset and Framework for Agricultural Monitoring through Fusion of Street-Level and Satellite Imagery
Space2Ground 2.0:结合街景与卫星影像的农业监测多源数据集及框架
Iason Tsardanidis, Alkiviadis Koukos, George Choumos, Vasileios Sitokonstantinou, Charalampos Kontoes
机构
*
Operational Unit BEYOND Centre, IAASARS, National Observatory of Athens(雅典国家天文台IAASARS研究所BEYOND中心业务单元)
;
DHI Water & Environment(DHI水环境公司)
;
Artificial Intelligence Group, Wageningen University & Research(瓦赫宁根大学及研究中心人工智能组)
CommentsThis paper has been accepted for presentation at the 45th EARSeL Symposium, Athens, Greece. The Space2Ground 2.0 dataset, is publicly available through Zenodo at: https://doi.org/10.5281/zenodo.21219542
CommentsFollowing an internal verification, the authors identified an inconsistency in the experimental data pipeline that requires the reported results to be recomputed. We are withdrawing this version while we re-run the analysis
机构
*
East China Normal University Shanghai China
;
Centre for Artificial Intelligence
;
Robotics, Hong Kong Institute of Science \& Innovation, Chinese Academy of Sciences Hong Kong China
;
Department of Computing, The Hong Kong Polytechnic University Hong Kong China
;
East China Normal University
;
Robotics, Hong Kong Institute of Science \& Innovation, Chinese Academy of Sciences
;
Department of Computing, The Hong Kong Polytechnic University