arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Computer Vision and Pattern Recognition · 会议 · Computer Vision

2026-06-30 至 2026-06-30 共收录 7
2606.30101 2026-06-30 cs.RO cs.CV

SIR: Structured Image Representations for Explainable Robot Learning

SIR: 用于可解释机器人学习的结构化图像表示

Paul Mattes, Jan Schwab, Jens Bosch, Nils Blank, Maximilian Xiling Li, Minh-Trung Tang, Moritz Haberland, Rudolf Lioutikov

机构 * Intuitive Robots Lab, Karlsruhe Institute of Technology, Germany(直观机器人实验室,卡尔斯鲁厄理工学院,德国) Robotics Institute Germany(德国机器人研究所)

AI总结 提出结构化图像表示(SIR),利用场景图作为中间表示,通过端到端稀疏化图结构学习任务相关子图,实现内在可解释性,在RoboCasa上平均成功率19.5%优于基线14.81%,并揭示数据集偏差。

Comments Published at CVPR 2026

Journal ref In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2026. S. 42484-42493

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29267 2026-06-30 cs.CV

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

增强任意开源多模态大语言模型的部件级点定位能力

Jin-Cheng Jhang, Fu-En Wang, Xin Yang, Nan Qiao, Lu Xia, Min Sun, Cheng-Hao Kuo

机构 * National Tsing Hua University(国立清华大学) Amazon(亚马逊)

AI总结 提出一种通用方法,通过冻结原模型参数并引入Q-Synth模块和注意力到点解码器,为任意开源MLLM赋予精确的2D部件级点定位能力,显著提升部件级定位精度。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29097 2026-06-30 cs.CV

TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

TrafficAlign: 对齐大语言模型用于交通场景生成

Zhi Tu, Liangkun Niu, Tianyi Zhang

AI总结 提出TrafficAlign框架,利用真实驾驶视频合成交通场景并对齐大语言模型,生成场景使自动驾驶模型碰撞率提升10.8%,微调后碰撞率降低36.1%。

Comments Accepted to CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2026, pp. 39744-39754

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28604 2026-06-30 cs.CV

IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

IMU-HOI:通过接触感知惯性融合实现连贯的人-物交互与运动捕捉的共生框架

Lizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun, Jiarui Yang, Ling Pei

机构 * Shanghai Jiao Tong University(上海交通大学)

AI总结 提出IMU-HOI框架,利用稀疏IMU同时恢复全身人体姿态和物体6-DoF轨迹,通过推断手-物接触并融合运动学与惯性推理,实现无漂移的人-物交互运动捕捉。

Comments 10 pages, 5 figures. Accepted by CVPR 2026

Journal ref Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026, pp. 42901-42910

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28561 2026-06-30 cs.RO cs.AI

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

为小型无人机系统合作战术解冲突进行大型语言模型微调

Iman Sharifi, Alex Zongo, Peng Wei

机构 * George Washington University(乔治华盛顿大学)

AI总结 本文研究了利用微调策略使大型语言模型在合作多智能体战术解冲突中做出决策,通过模拟生成数据提升决策准确性与一致性。

Comments 15 pages, 6 figures, to be published in CVPR 2026 Workshop Proceedings

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 1067-1076

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12175 2026-06-30 cs.CV

Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment

重新定义质量标准与距离感知评分建模用于图像编辑评估

Xinjie Zhang, Qiang Li, Xiaowen Ma, Axi Niu, Li Yan, Qingsen Yan

机构 * Northwestern Polytechnical University(西北工业大学) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)

AI总结 本文提出DS-IEQA框架,通过反馈驱动指标优化和令牌解耦距离回归损失,改进图像编辑质量评估的指标定义与评分连续性建模。

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 2812-2820

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05689 2026-06-30 cs.CV cs.AI

CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration

CRFT:用于跨模态图像配准的一致-递归特征流变换器

Xuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li, Xichao Teng

机构 * Northeastern University, China(东北大学(中国)) National University of Defense Technology, China(国防科技大学(中国))

AI总结 CRFT提出了一种基于特征流学习的统一粗到细框架,通过特征对齐和流估计实现鲁棒的跨模态图像配准,通过多尺度特征相关性和层次特征融合提升精度和鲁棒性。

Comments Accepted to CVPR 2026

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 34784-34794

详情

展开后加载摘要…

URL PDF HTML 收藏