FOGMACHINE -- Leveraging Discrete-Event Simulation and Scene Graphs for Modeling Hierarchical, Interconnected Environments under Partial Observations from Mobile Agents
Lars Ohnemus, Nils Hantke, Max Weißer, Kai Furmans
机构
*
Institute for Material Handling and Logistics, Karlsruhe Institute of Technology(材料搬运与物流研究所,卡尔斯鲁厄技术大学)
专题命中
视觉空间推理
:planning(abstract)
Commentssubmitted to the IEEE for possible publication; 8 pages, 3 figures, 1 table
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
Yan Shu, Hangui Lin, Yexin Liu, Yan Zhang, Gangyan Zeng, Yan Li, Yu Zhou, Ser-Nam Lim, Harry Yang, Nicu Sebe
机构
*
University of Trento(特伦托大学)
;
Hong Kong University of Science and Technology(香港科学与技术大学)
;
University of International Relations(国际关系大学)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Nanjing University of Science and Technology(南京理工大学)
;
VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院)
;
University of Central Florida(佛罗里达中央大学)
ReLI: A Language-Agnostic Approach to Human-Robot Interaction
Linus Nwankwo, Bjoern Ellensohn, Ozan Özdenizci, Elmar Rueckert
机构
*
Chair of Cyber-Physical Systems, Technical University of Leoben(技术大学莱博恩首席物理系统 Chair)
;
Institute of Machine Learning and Neural Computation, Graz University of Technology(格拉茨技术大学机器学习与神经计算研究所)
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
Xinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia, Zhuohang Dang, Basura Fernando, Jun Liu, Mike Zheng Shou
机构
*
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
;
Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室)
;
Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, China(陕西省大数据知识工程重点实验室)
;
IHPC, Agency for Science, Technology and Research, Singapore(新加坡科技研究局IHPC)
;
Show Lab, National University of Singapore(新加坡国立大学Show实验室)
;
College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院)
Building Information Models to Robot-Ready Site Digital Twins (BIM2RDT): An Agentic AI Safety-First Framework
Reza Akhavian, Mani Amani, Johannes Mootz, Robert Ashe, Behrad Beheshti
机构
*
Department of Civil, Construction, and Environmental Engineering, San Diego State University, San Diego, CA, United States(土木、建设与环境工程系,圣地亚哥州立大学)
;
Department of Electrical and Computer Engineering, University of California, San Diego, San Diego, CA, United States(电气与计算机工程系,加州大学圣地亚哥分校)
;
Department of Mechanical and Aerospace Engineering, University of California, San Diego, San Diego, CA, United States(机械与航空航天工程系,加州大学圣地亚哥分校)
;
Department of Computer Science, San Diego State University, San Diego, CA, United States(计算机科学系,圣地亚哥州立大学)
Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery
Ana Davila, Jacinto Colan, Yasuhisa Hasegawa
机构
*
Nagoya University, Japan(名古屋大学)
专题命中
视觉空间推理
:reasoning(abstract)
CommentsTo be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)
DECAMP: Towards Scene-Consistent Multi-Agent Motion Prediction with Disentangled Context-Aware Pre-Training
Jianxin Shi, Zengqi Peng, Xiaolong Chen, Tianyu Wo, Jun Ma
机构
*
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (GZ)(香港科技大学(广州)机器人与自主系统方向)
;
Division of Emerging Interdisciplinary Areas, The Hong Kong University of Science and Technology(香港科技大学新兴交叉领域研究所)
PlaneRecTR++: Unified Query Learning for Joint 3D Planar Reconstruction and Pose Estimation
Jingjia Shi, Shuaifeng Zhi, Kai Xu
机构
*
National University of Defense Technology(国防科技大学)
专题命中
视觉空间推理
:reasoning(abstract)
CommentsTo be published in IEEE T-PAMI 2025. This is the journal extension of our ICCV 2023 paper "PlaneRecTR", which expands from single view reconstruction to simultaneous multi-view reconstruction and camera pose estimation. Note that the ICCV2023 PlaneRecTR paper could be found in the previous arxiv version [v2](arXiv:2307.13756v2)