arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4882 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4882 篇

2505.16791 2026-05-08 cs.LG cs.AI 57%

Cohort-Based Active Modality Acquisition

基于队列的主动模态获取

Tillmann Rheude, Roland Eils, Benjamin Wild

机构 * Berlin Institute of Health, Charité - Universitätsmedizin Berlin(柏林健康研究所,柏林查理大学) Intelligent Medicine Institute, Fudan University(智能医学研究院,复旦大学) Department of Mathematics and Computer Science, Freie Universität Berlin(数学与计算机科学系,柏林自由大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出基于队列的主动模态获取方法,通过填补缺失模态的预期效用来指导额外模态的获取,实验证明其在资源受限环境下更有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00187 2026-05-05 cs.HC cs.AI cs.ET 57%

Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era

可解释AI用于盲人和低视力用户:在智能体时代导航信任、模态和可解释性

Abu Noman Md Sakib, Protik Dey, Zijie Zhang, Taslima Akter

机构 * University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨了盲人和低视力用户在智能体时代对可解释AI的独特需求,指出需多模态界面和责任意识设计以提升可解释性与信任度。

Comments Proceedings of the CHI 2026 Workshop on Human-Centered Explainable AI (HCXAI), April 13-17, 2026, Barcelona, Spain

Journal ref Proceedings of the CHI 2026 Workshop on Human-Centered Explainable AI (HCXAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13486 2026-05-01 cs.CV 57%

Uncertainty Quantification Framework for Aerial and UAV Photogrammetry through Error Propagation

通过误差传播的航空和无人机摄影测量不确定性量化框架

Debao Huang, Rongjun Qin

机构 * Geospatial Data Analytics Laboratory, The Ohio State University(地理空间数据分析实验室,俄亥俄州立大学) Department of Civil, Environmental and Geodetic Engineering, The Ohio State University(土木、环境与大地测量工程系,俄亥俄州立大学) Department of Electrical and Computer Engineering, The Ohio State University(电气与计算机工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究院,俄亥俄州立大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出通过误差传播的不确定性量化框架,解决多视立体阶段的不确定性估计问题,利用自校准方法提升摄影测量点云的鲁棒性和可验证性。

Comments 27 pages, 12 figures, this manuscript has been accepted to ISPRS Journal of Photogrammetry and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24062 2026-04-28 cs.AI 57%

Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer

在泛化之前建立基础:人工智能如何与人类在因果迁移中不同

Liangru Xiang, Yuxi Ma, Zhihao Cao, Yixin Zhu, Song-Chun Zhu

机构 * Department of Automation(清华大学自动化系) Institute for Artificial Intelligence(北京大学人工智能研究院) School of Psychological and Cognitive Sciences(北京大学心理与认知科学系) State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室) Beijing Key Laboratory of Behavior and Mental Health(北京行为与心理健康重点实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究探讨了人工智能与人类在因果迁移中的差异,发现AI模型在缺乏环境基础映射时难以高效迁移,而人类能利用先验结构知识。文本条件中AI表现优异,但视觉信息反而降低性能,揭示AI依赖符号处理而非多模态推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20018 2026-04-28 cs.DB cs.AI 57%

MINT: Multi-Vector Search Index Tuning

MINT: 多向量搜索索引调优

Jiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein, Vivek Narasayya, Surajit Chaudhuri

机构 * University of California, San Diego(加州大学圣地亚哥分校) Microsoft Research(微软研究院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出MINT框架,针对多向量搜索场景,通过算法优化实现低延迟和存储召回约束下的索引调优,相比基线提升2.1至8.3倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16794 2026-04-21 cs.CV 57%

Improving Radio Interferometry Imaging by Explicitly Modeling Cross-Domain Consistency in Reconstruction

通过显式建模跨域一致性来改进射电干涉成像

Kai Cheng, Ruoqi Wang, Qiong Luo

机构 * The Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出CDCRec方法,通过显式建模跨域一致性提升射电干涉成像质量,采用多任务多阶段框架增强域间交互,实验表明其在干涉域转换中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15341 2026-04-20 cs.HC cs.AI 57%

MRGEN: A Conceptual Framework for LLM-Powered Mixed Reality Authoring Tools for Education

MRGEN:一种基于大语言模型的混合现实教育作者工具的概念框架

Mohammed Oussama Seddini, Mohamed Ez-Zaouia, Ngoc Luyen Le, Iza Marfisi

机构 * LIUM, Le Mans Université(里摩斯大学) IRISA, Université de Rennes(里莫斯大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出MRGEN框架,利用大语言模型帮助教师创建适用于移动设备的混合现实学习活动,实验显示AI辅助显著缩短任务时间,90%以上参与者认为AI支持有助于构思和对齐学习目标。

Journal ref The Mobile Learning 2026 International Conference, Mar 2026, Zagreb, Croatia

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14069 2026-04-16 cs.CV 57%

Towards Unconstrained Human-Object Interaction

迈向无约束的人-物交互

Francesco Tonini, Alessandro Conti, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci

机构 * Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系) Fondazione Bruno Kessler(布鲁诺·克塞勒基金会) Department of Computer Science, University of Verona(威尼斯大学计算机科学系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文通过多模态大语言模型探索无约束人-物交互检测,提出无需预定义交互词汇的新任务,并展示MLLMs在该任务中的潜力。

Comments Accepted to the 20th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05957 2026-04-16 cs.DC cs.AI 57%

Domain-Adaptive Model Merging Across Disconnected Modes

跨断开模式的领域适应性模型合并

Junming Liu, Yusen Zhang, Rongchao Zhang, Wenkai Zhu, Tian Wu

机构 * Tongji University(同济大学) Peking University(北京大学) Southeast University(东南大学) Nanchang University(南昌大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出DMM框架,通过独立训练领域模型、合并相似模型及生成伪数据精炼知识,实现无需数据的模型合并,提升性能。

Comments 5 pages, 1 figure, 3 tables; Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13017 2026-04-15 cs.AI cs.HC 57%

PAL: Personal Adaptive Learner

PAL:个性化学习者

Megha Chakraborty, Darssan L. Eswaramoorthi, Madhur Thareja, Het Riteshkumar Shah, Finlay Palmer, Aryaman Bahl, Michelle A Ihetu, Amit Sheth

机构 * Artificial Intelligence Institute, University of South Carolina(人工智能研究院,南卡罗来纳大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 PAL通过实时分析多模态内容和动态调整问题难度,为学习者提供个性化学习体验,提升教育响应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12508 2026-04-15 cs.CV 57%

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception

从衰减到注意:用于细粒度视觉感知的变分信息流操控

Jilong Zhu, Yang Feng

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(智能信息处理重点实验室,计算技术研究所,中国科学院(ICT/CAS)) State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院(ICT/CAS)) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出变分信息流框架,通过建模问题-答案对相关的视觉显著性作为潜在分布,解决多模态大语言模型在细粒度感知任务中的信息衰减问题,提升模型表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12180 2026-04-15 cs.LG cs.AI 57%

CycloneMAE: A Scalable Multi-Task Learning Model for Global Tropical Cyclone Probabilistic Forecasting

CycloneMAE:一种用于全球热带气旋概率预报的可扩展多任务学习模型

Renlong Hang, Zihao Xu, Jiuwei Zhao, Runling Yu, Leye Cheng, Qingshan Liu

机构 * School of Computer Science, School of Software(计算机科学学院、软件学院) Nanjing University of Information Science and Technology(南京信息工程大学) School of Atmospheric Sciences(大气科学学院) Shanghai Typhoon Institute(上海台风研究所) China Meteorological Administration(中国气象局) Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 CycloneMAE通过多模态数据学习可迁移的热带气旋表示,结合离散概率格机制与预训练/微调范式,实现确定性预报和概率分布输出,优于现有NWP系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10547 2026-04-15 q-bio.NC cs.AI cs.LG 57%

Pursuit of biomarkers of brain diseases: Beyond cohort comparisons

寻找脑部疾病的生物标志物:超越队列比较

Pascal Helson, Arvind Kumar

机构 * Science For Life Laboratory(生命科学实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出通过多模态和纵向脑数据指导分组,以定义多维生物标志物,而非依赖单一数据类型的队列比较。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11809 2026-04-14 cs.CV 57%

Who Handles Orientation? Investigating Invariance in Feature Matching

谁负责方向?在特征匹配中研究不变性

David Nordström, Johan Edstedt, Fredrik Kahl, Georg Bökman

机构 * Chalmers University of Technology(挑战者技术大学) Linköping University(林奈大学) University of Amsterdam(阿姆斯特丹大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 研究在现代稀疏匹配流程中,特征描述符和匹配器中引入旋转不变性对性能的影响,发现早期在描述符中学习旋转不变性可提升匹配效率并提高泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11306 2026-04-14 cs.RO cs.AI 57%

Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment

学习遗忘 -- 分层事件记忆用于终身机器人部署

Leonard Bärmann, Joana Plewnia, Alex Waibel, Tamim Asfour

机构 * Institute for Anthropomatics and Robotics, Karlsruhe Institute for Technology(卡尔斯鲁厄理工学院人形机器人与机器人研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出H$^2$-EMV框架,通过用户交互帮助机器人学习选择性遗忘,以应对长期多模态感知存储限制,实验证明其在减少内存和查询时间的同时保持回答准确性,并随时间适应用户需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10912 2026-04-14 cs.CV 57%

TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation

TAMISeg: 基于文本的多尺度医学图像分割与语义编码蒸馏

Qiang Gao, Yi Wang, Yong Zhang, Yong Li, Yongbing Deng, Lan Du, Cunjian Chen

机构 * Department of Data Science and AI, Monash University(蒙纳士大学数据科学与人工智能系) College of Computer Science, Chongqing University(重庆大学计算机学院) Chongqing University Central Hospital, School of Medicine(重庆大学附属中心医院医学院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 TAMISeg通过整合临床语言提示和语义蒸馏增强视觉理解,减少对像素级标注的依赖,优于现有单模态和多模态方法。

Comments Accepted by IEEE International Conference on Multimedia and Expo (ICME), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03197 2026-04-14 cs.CV 57%

Specificity-aware reinforcement learning for fine-grained open-world classification

具有特异性的强化学习用于细粒度开放世界分类

Samuele Angheben, Davide Berasi, Alessandro Conti, Elisa Ricci, Yiming Wang

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出SpeciaRL框架,通过强化学习提升细粒度开放世界分类的准确性和特异性,优于现有方法。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01881 2026-04-10 cs.LG cs.AI 57%

Tractable Uncertainty-Aware Meta-Learning

可 tractable 的不确定性意识元学习

Young-Jin Park, Cesar Almecija, Apoorva Sharma, Navid Azizan

机构 * MIT(麻省理工学院) NVIDIA(英伟达)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出 LUMA,一种用于回归的元学习方法,能够高效地进行分布内任务的概率预测,检测分布外上下文数据,并有效处理异质多模态任务分布。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27492 2026-04-07 cs.RO cs.AI cs.HC cs.LG 57%

Copilot-Assisted Second-Thought Framework for Brain-to-Robot Hand Motion Decoding

辅助第二思考框架用于脑机臂运动解码

Yizhe Li, Shixiao Wang, Jian K. Liu

机构 * University of Birmingham(伯明翰大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出混合CNN-注意力模型解码EEG手部运动轨迹,在单受试者实验中取得高PCC值,进一步扩展至EEG-EMG多模态解码,并引入copilot框架提升轨迹精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03334 2026-04-07 cs.CV 57%

Bridging the Dimensionality Gap: A Taxonomy and Survey of 2D Vision Model Adaptation for 3D Analysis

弥合维度差距:2D视觉模型适应3D分析的分类与综述

Akshat Pandya, Bhavuk Jain

机构 * Independent Researcher(独立研究员)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文综述了将2D视觉模型适应3D分析的策略,分类为数据导向、架构导向和混合方法,探讨了计算复杂度、预训练依赖性和几何归纳偏置的权衡。

Comments VISAPP 2026

Journal ref Proceedings of the 21st International Conference on Computer Vision Theory and Applications - Volume 3: VISAPP 2026; ISBN 978-989-758-804-4; ISSN 2184-4321, SciTePress, pages 353-364

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14351 2026-04-07 cs.LG cs.AI 57%

WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control

WIMLE:具有IMLE的不确定性感知世界模型用于高效连续控制

Mehran Aghabozorgi, Alireza Moazeni, Yanshu Zhang, Ke Li

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 WIMLE通过扩展IMLE至模型基于强化学习框架,学习随机多模世界模型并估计预测不确定性,提升了连续控制任务的样本效率和稳定性。

Comments Accepted at ICLR 2026. Website: https://mehranagh20.github.io/wimle/ Code: https://github.com/mehranagh20/wimle

Journal ref In Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14336 2026-04-07 cs.CV 57%

ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental Understanding

ArchMap:用于牙齿计数和结构牙科理解的拱形扁平化和知识引导的视觉语言模型

Bohan Zhang, Yiyi Miao, Taoyu Wu, Tong Chen, Ji Jiang, Zhuoxiao Li, Zhe Tang, Limin Yu, Jionglong Su

机构 * Xi'an Jiaotong-Liverpool University(西交利物浦大学) University of Liverpool(利物浦大学) Zhejiang University of Technology(浙江工业大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出ArchMap,一种无需训练的知识引导框架,通过几何归一化和本体引导的多模态推理,提升3D口腔扫描的结构化分析能力,实现牙齿计数、解剖分区、牙列阶段分类及临床状况识别。

Journal ref In Proceedings of the 2025 IEEE International Conference on Big Data (BigData), pp. 7529-7538, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00363 2026-04-02 cs.RO cs.CV 57%

A Dual-Stream Transformer Architecture for Illumination-Invariant TIR-LiDAR Person Tracking

一种用于光照不变TIR-LiDAR人体跟踪的双流Transformer架构

Yuki Minase, Kanji Tanaka

机构 * Graduate School of Engineering, University of Fukui(福井大学工学研究科)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种双流Transformer架构,利用LiDAR和TIR摄像头实现全天候鲁棒的人体跟踪,通过知识迁移策略提升性能,实验显示在AO和SR指标上优于传统方法。

Comments 6 pages, 4 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24956 2026-04-01 cs.RO cs.AI cs.LG 57%

MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation

MSG:多流生成策略用于样本高效的机器人操作

Jan Ole von Hartz, Lukas Schweizer, Joschka Boedecker, Abhinav Valada

机构 * University of Freiburg(弗莱堡大学)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出MSG多流生成策略,通过结合多个对象中心策略提升泛化能力和样本效率,实现在五次演示下学习高质量生成策略,性能提升89%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28120 2026-03-31 cs.CV 57%

MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding

MedLoc-R1: 为基于GRPO的医疗视觉接地任务的性能感知课程奖励调度

Guangjing Yang, Ziyuan Qin, Chaoran Zhang, Chenlin Du, Jinlin Wang, Wanran Sun, Zhenyu Zhang, Bing Ji, Qicheng Lao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Emory University(埃默里大学) Peking University(北京大学) Shandong University(山东大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出MedLoc-R1框架,通过性能感知奖励调度提升医疗视觉接地任务的定位准确性和训练稳定性,无需引入辅助网络。

Comments 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27969 2026-03-31 cs.CV 57%

Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs

Hg-I2P:通过异构图实现通用的图像到点云配准

Pei An, Junfeng Ding, Jiaqi Yang, Yulong Wang, Jie Ma, Liangliang Nan

机构 * Huazhong University of Science and Technology(华中科技大学) Northwestern Polytechnical University(西北工业大学) Huazhong Agricultural University(华中农业大学) Delft University of Technology(代尔夫特理工大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出Hg-I2P方法,通过异构图融合多模态特征,提升图像与点云配准的泛化能力和精度。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27536 2026-03-31 cs.AI 57%

Dual-Stage LLM Framework for Scenario-Centric Semantic Interpretation in Driving Assistance

双阶段LLM框架用于驾驶辅助中的场景中心语义解释

Jean Douglas Carvalho, Hugo Taciro Kenji, Ahmad Mohammad Saber, Glaucia Melo, Max Mauro Dias Santos, Deepa Kundur

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出双阶段LLM框架,用于评估城市驾驶场景中基于LLM的风险推理可复现性。通过多模态驾驶数据构建确定性场景窗口,并在固定提示约束和闭合数字风险方案下评估,揭示模型间在风险等级分配、高风险升级、证据使用和因果归因上的系统性差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22339 2026-03-31 cs.LG cs.CL stat.ML 57%

Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits

Chinchilla方法2的问题2:等效FLOP拟合中的系统性偏差

Eric Czech, Zhiwei Xu, Yael Elmatad, Yixin Wang, William Held

机构 * Open Athena AI Foundation(开放雅典娜AI基金会) Department of Statistics, University of Michigan, Ann Arbor(密歇根大学安娜堡分校统计系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文研究了Chinchilla方法2在拟合神经网络扩展定律时的系统性偏差,指出其在无噪声合成数据中仍存在偏差,影响参数分配,并提出通过变量投影方法改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02703 2026-03-27 cs.CV 57%

Structure Causal Models and LLMs Integration in Medical Visual Question Answering

结构因果模型与大语言模型在医学视觉问答中的整合

Zibo Xu, Qiang Li, Weizhi Nie, Weijie Wang, Anan Liu

机构 * School of Microelectronics, Tianjin University(天津大学微电子学院) School of Electrical and Information Engineering, Tianjin University(天津大学电气与信息工程学院) Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种因果推理框架,通过消除图像与问题间的混杂效应,提升医学视觉问答的准确性,并引入因果图结构和多变量重采样方法以增强模型对复杂医疗数据的理解能力。

Comments Accepted by IEEE TMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19371 2026-03-23 cs.CV 57%

Factored Levenberg-Marquardt for Diffeomorphic Image Registration: An efficient optimizer for FireANTs

因子化Levenberg-马夸尔特法用于可微同构图像配准:一种高效的优化器用于FireANTs

Rohit Jena, Pratik Chaudhari, James C. Gee

机构 * Department of Computer and Information Science(计算机与信息科学系) University of Pennsylvania(宾夕法尼亚大学) Department of Electrical and Systems Engineering(电气与系统工程系) Department of Radiology & Penn Image Computing and Science Laboratory (PICSL)(放射科及宾夕法尼亚大学图像计算与科学实验室(PICSL)) Perelman School of Medicine, University of Pennsylvania(佩尔曼医学学院,宾夕法尼亚大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种改进的Levenberg-马夸尔特优化器,通过信任区域方法自适应调整单个标量阻尼参数,减少内存消耗并保持性能,适用于大体积图像和跨模态配准。

详情

展开后加载摘要…

URL PDF HTML 收藏