arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

European Conference on Computer Vision · 会议 · Computer Vision

2026-08-20 至 2026-08-20 共收录 13
2608.18166 2026-08-20 eess.IV cs.CV 新提交

TractoGraphVLM: A Unified Vision-Language Framework for White Matter Tractography

TractoGraphVLM:用于白质纤维束成像的统一视觉-语言框架

Gurucharan Marthi Krishna Kumar, Janine Dale Mendola, Amir Shmuel

AI总结 TractoGraphVLM是统一视觉-语言框架,可完成白质纤维束的分类、检索、描述、问答四项任务,在HCP数据集上表现良好,具跨年龄迁移鲁棒性,仅从语言学习神经解剖学知识。

Comments Accepted as a Spotlight at the ECCV 2026 Workshop on Artificial Intelligence for Medical 3D Vision (AI4M3D). Our codebase, including all training and evaluation pipelines, is publicly available at this https URL (https://github.com/AS-Lab/Marthi-et-al-2026-TractoGraphVLM-Unified-Vision-Language-White-Matter-Tractography)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19000 2026-08-20 cs.CV 新提交

Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

场景布置(Mise-en-Scène):用于人机协同设计共创的扩散Transformer中隐式布局的涌现

Zipeng Xu, Ryan Murdock, Umberto Michieli

机构 * Canva Research(Canva研究院)

AI总结 本文提出Mise-en-Scène框架,通过微调扩散Transformer实现隐式布局涌现,结合匹配放置步骤保证素材保真度,在PrismLayersPlus基准上生成的设计感知质量显著优于现有方法。

Comments Best Paper Award at ECCV Human-AI Co-Creation Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18734 2026-08-20 cs.CV 新提交

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

CL4D:用于动态场景视觉-语言推理的对比语言-4D预训练

Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella, Dulanga Weerakoon, Vigneshwaran Subbaraju, Ranga Rodrigo

机构 * University of Moratuwa(莫拉图瓦大学) Singapore-MIT Alliance for Research & Technology (SMART) Centre(新加坡-麻省理工研究与技术联盟(SMART)中心) Agency for Science, Technology and Research (A*STAR)(新加坡科学、技术与研究局(A*STAR))

AI总结 本研究提出CL4D(首个4D视觉编码器)及基于其的4DVLM,在自建DynAction4D数据集上训练,CL4D性能较现有方法提升约16.75%,4DVLM优于Gemini、GPT-5等前沿视频VLM。

Comments Accepted at the 19th European Conference on Computer Vision (ECCV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18711 2026-08-20 cs.CV eess.SP 新提交

EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment

EgoHRV:用于自主神经反应与技能评估的自我中心式系统的连续心率变异性估计

Berken Utku Demirel, Christian Holz

机构 * ETH Zürich(苏黎世联邦理工学院)

AI总结 本文提出EgoHRV方法,利用自我中心头戴设备的凝视摄像头结合3D骨干网络与低-高分解模块等,实现从凝视视频中连续估计HRV与HR,在HR/HRV估计中达最优准确率,集成后使EgoExo4D熟练度估计器准确率提升17.8%

Comments Accepted to the European Conference on Computer Vision (ECCV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18602 2026-08-20 cs.CV 新提交

Teach a Molmo2Fish: Towards interactive fish tracking with natural language guidance

Teach a Molmo2Fish:面向自然语言引导的交互式鱼类追踪

Kai Van Brunt (1), Justin Kay (1), Sara Beery (1) ((1) Massachusetts Institute of Technology)

机构 * Massachusetts Institute of Technology(麻省理工学院)

AI总结 本研究针对声呐鱼类追踪数据集定制了多模态大语言模型工具Molmo2Fish,通过交互式预测修正工作流开展实验,发现其在鱼类追踪和轨迹修正上性能良好,但自然语言引导的融入仍需提升。

Comments 29 pages, 6 figures, to be published in Third Workshop on Computer Vision for Ecology at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18523 2026-08-20 cs.CV cs.AI 新提交

Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection

用于通用AI生成图像检测的先验条件高斯判别模型

Shashank Kotyan, Makoto Shing, Yuki Imajuku, Rujikorn Charakorn, Tarin Clanuwat

机构 * Sakana AI

AI总结 该研究针对AI生成图像检测在多因素变化下失效的问题,提出先验条件高斯判别梯方法,在Percept-Lens数据集上验证其性能,推动相关报告与基线的优化。

Comments Accepted in ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18317 2026-08-20 cs.CV cs.RO 新提交

Reproducible Multimodal Affordance Prediction

可复现的多模态可供性预测

Tommaso Apicella, Alessio Xompero, Andrea Cavallaro

机构 * Istituto Italiano di Tecnologia(意大利技术研究院) EPFL(洛桑联邦理工学院)

AI总结 针对可供性预测方法评估难的问题,本文提出Affordance Sheet以规范任务表述等信息,实现可供性模型的可复现基准测试与现实场景可靠评估。

Comments Paper accepted to Workshop on Human-Centered Multimodal Intelligence in the Wild (HCMIW) in European Conference on Computer Vision (ECCV) 2026; 18 pages, 3 figures, 7 tables. Project webpage at this https URL (https://apicis.github.io/aff-sheet)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18246 2026-08-20 cs.CV cs.AI cs.LG 新提交

Visual-Prompt Guided Wildlife Instance-Level Recognition

视觉提示引导的野生动物实例级识别

Mufhumudzi Muthivhi, Jiahao Huo, Terence van Zyl, Fredrik Gustafsson

机构 * University of Johannesburg(约翰内斯堡大学) Linköping University(林雪平大学)

AI总结 针对细粒度野生动物重识别挑战,提出单阶段端到端检测与重识别模型,采用DINOv2、MegaDescriptor及提示增强技术,在mAP指标上取得与两阶段方法相近的竞争力表现。

Comments Accepetd in ECCV Instance-Level Recognition and Generation Workshop 2026, Malmö Sweden

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18191 2026-08-20 cs.LG cs.SD eess.AS 新提交

ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy

ChiroEcho:将自动化蝙蝠叫声分类扩展至学习分类体系之外

Burooj Ghani, Welmoed Eversteijn, Milan van Hirtum, Juan Sebastián Cañas, Vincent J. Kalkman, Dan Stowell, A. Leonie Baier

机构 * Naturalis Biodiversity Center(自然生物多样性中心) Department of Cognitive Science and Artificial Intelligence, Tilburg University(蒂尔堡大学认知科学与人工智能系) People and Nature Lab, University College London(伦敦大学学院人与自然实验室) Leiden Institute of Advanced Computer Science, Leiden University(莱顿大学莱顿高级计算机科学研究所)

AI总结 ChiroEcho框架联合预测蝙蝠的属与物种,结合地理分布将欧洲蝙蝠自动分类覆盖从73%提升至85%,为解决未见过的细粒度类别问题提供了原理验证。

Comments 24 pages, 3 figures. Accepted at the CV4E workshop, ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17935 2026-08-20 cs.CV 版本更新

Beyond Instrument Motion: Recognizing Tissue Tension Toward Surgical Skill Assessment

超越器械运动:面向手术技能评估的组织张力识别

Marko Haralović, Zhiqi Miao, Alexander Machiel Bont, Jiapan Guo, Frans van Workum, Estefanía Talavera

机构 * University of Zagreb(萨格勒布大学) University of Twente(特文特大学) University of Groningen(格罗宁根大学) Radboud University Medical Center(拉德堡德大学医学中心) Canisius-Wilhelmina Hospital(卡尼修斯-威廉明娜医院)

AI总结 针对现有手术视频理解方法未捕捉组织张力的问题,本文构建SurgTension数据集,提出TensionTRAC框架,实现了腹腔镜及机器人辅助直肠癌手术的组织张力识别,为手术技能评估提供客观基准。

Comments The paper is accepted by ECCV 2026 Workshop On Medical Video Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17803 2026-08-20 cs.CV 版本更新

Scale Matters: Adaptive Granularity Selection for Cross-Species 3D Plant Organ Segmentation

尺度重要:跨物种三维植物器官分割的自适应粒度选择

Carla Salazar, Lazaros Nalpantidis

机构 * Technical University of Denmark (DTU)(丹麦技术大学(DTU)) Pioneer Centre for Artificial Intelligence(人工智能先锋中心)

AI总结 针对跨物种三维植物器官分割中固定空间粒度泛化能力差的问题,提出结合冻结Utonia基础模型与自适应粒度选择的小样本方法AGS-PlantSeg,在三个数据集上实现88.9%平均mIoU,性能优于固定粒度基线且与全监督架构相当。

Comments Accepted at the Computer Vision in Plant Phenotyping and Agriculture (CVPPA) Workshop at the European Conference on Computer Vision (ECCV) 2026. Project page: this https URL (https://dtu-pas.github.io/ags-plantseg/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12442 2026-08-20 cs.CV 版本更新

MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis

MV2:用于新视角合成的多视角多车辆驾驶数据集

Sanjay Bhargav Dharavath, Hanvitha Saraswathi Mukkamala, Faizan Farooq Khan, Ioannis Kakogeorgiou, Aditya Arun, C V Jawahar, Zakaria Laskar

机构 * International Institute of Information Technology, Hyderabad(海得拉巴国际信息技术学院) King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) IIT, National Centre for Scientific Research “Demokritos”(德谟克利特国家科学研究中心 IIT) Adobe MDSR, India(奥多比印度MDSR部门) Indian Institute of Science Education and Research, Thiruvananthapuram(特里凡得琅印度科学教育与研究学院)

AI总结 针对驾驶场景新视角合成的难题,提出含50个场景12000张图像的MV2多车辆数据集,经严格配准验证,基准测试显示视角差异增大NVS性能下降,前馈位姿估计器弱于优化方法。

Comments 38 pages, 25 figures, ECCV 2026 accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21079 2026-08-20 cs.CV 版本更新

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models

聚焦推理:面向视觉语言模型的有状态、基于动作的视觉聚焦

Juhong Min, Lazar Valkov, Vitali Petsiuk, Hossein Souri, Deen Dayal Mohan

机构 * AI Center -- Mountain View, Samsung Electronics(三星电子山景城人工智能中心)

AI总结 本文提出Foveated Reasoner,通过整合聚焦与推理机制,提升视觉语言模型在高分辨率图像下的效率与准确性,实验证明其在多种基准测试中表现优异。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏