arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Harvard University(哈佛大学)

共收录 1302
2603.26117 2026-03-30 eess.IV cs.CV

FINDER: Zero-Shot Field-Integrated Network for Distortion-free EPI Reconstruction in Diffusion MRI

FINDER:用于扩散磁共振成像中无失真EPI重建的零样本场积分网络

Namgyu Han, Seong Dae Yun, Chaeeun Lim, Sunghyun Seok, Sunju Kim, Yoonhwan Kim, Yohan Jun, Tae Hyung Kim, Berkin Bilgic, Jaejin Cho

机构 * Sejong University(世宗大学) Forschungszentrum Jülich(于利希研究中心) Athinoula A. Martinos Center for Biomedical Imaging(Athinoula A. Martinos生物医学成像中心) Harvard Medical School(哈佛医学院) Hongik University(弘益大学)

AI总结 本文提出FINDER网络,通过联合优化图像和B0场图,解决EPI重建中的几何失真问题,提升扩散成像质量。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24999 2026-03-30 stat.AP cs.AI

Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

高效检测劣质基准项目:基于新颖可扩展系数的方法

Michael Hardy, Joshua Gilbert, Benjamin Domingue

机构 * Stanford University, CA, United States(斯坦福大学,美国加利福尼亚州) Harvard University, MA, United States(哈佛大学,美国马萨诸塞州)

AI总结 本文提出基于交叉项目等比回归的非参数可扩展系数,用于高效检测全球劣质项目,通过Kendall's τ保持关联方向,实现无假设线性或参数模型的项目评分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26008 2026-03-30 cs.CV cs.AI

FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

FairLLaVA: 大规模视觉-语言助手中的公平性感知参数高效微调

Mahesh Bhosale, Abdul Wasi, Shantam Srivastava, Shifa Latif, Tianyu Luan, Mingchen Gao, David Doermann, Xuan Gong

机构 * University at Buffalo(布法罗大学) University of Kashmir(克什米尔大学) Accenture(埃森哲) Harvard Medical School(哈佛医学院)

AI总结 FairLLaVA通过减少目标属性间的互信息,实现视觉指令微调中的公平性改进,提升医疗影像生成的公平性和自然语言生成质量。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20964 2026-03-30 cs.CV cs.AI

Evidence-based diagnostic reasoning with multi-agent copilot for human pathology

基于多智能体助手的证据驱动诊断推理

Luca L. Weishaupt, Chengkuan Chen, Drew F. K. Williamson, Richard J. Chen, Guillaume Jaume, Tong Ding, Bowen Chen, Anurag Vaidya, Long Phi Le, Guillaume Jaume, Ming Y. Lu, Faisal Mahmood

机构 * Health Sciences and Technology, Harvard-MIT(哈佛-MIT健康科学与技术) Department of Pathology, Massachusetts General Hospital, Harvard Medical School(麻省总医院病理科,哈佛医学院) Cancer Program, Broad Institute of Harvard and MIT(哈佛-MIT博德研究所癌症项目) Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保尔森工程与应用科学学院) Electrical Engineering and Computer Science, Massachusetts Institute of Technology (MIT)(麻省理工学院电气工程与计算机科学) Harvard Data Science Initiative, Harvard University(哈佛大学数据科学计划)

AI总结 本文提出PathChat+,一种专为人类病理设计的多模态大语言模型,通过大量病理特定指令样本训练,显著优于现有模型,在多图像理解与自主诊断推理方面表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04810 2026-03-27 cs.NE cs.IT cs.LG math.IT q-bio.NC

Correlative Information Maximization: A Biologically Plausible Approach to Supervised Deep Neural Networks without Weight Symmetry

相关信息最大化:一种生物合理的方法用于无权重对称性的监督深度神经网络

Bariscan Bozkurt, Cengiz Pehlevan, Alper T Erdogan

机构 * Gatsby Computational Neuroscience Unit, UCL(伦敦大学学院盖茨比计算神经科学单元) KUIS AI Center, Koc University(科奇大学KUIS人工智能中心) EEE Department, Koc University(科奇大学电气与电子工程系) John A. Paulson School of Engineering & Applied Sciences and Center for Brain Science, Harvard University(哈佛大学约翰·A·保尔森工程与应用科学学院及脑科学中心) Kempner Institute for the Study of Natural and Artificial Intelligence(肯普纳自然与人工智能研究所)

AI总结 本文提出相关信息最大化作为生物神经网络信号传播的替代规范方法,解决了传统神经网络和反向传播算法的生物合理性问题,并提供了一种无权重对称性的解决方案。

Comments Neurips published version

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24676 2026-03-27 cs.AI cond-mat.dis-nn cond-mat.stat-mech physics.bio-ph physics.soc-ph

When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs

集体智慧何时是一场彩票?LLM中膜etic漂移的多智能体缩放定律

Hidenori Tanaka

机构 * CBS–NTT Program in Physics of Intelligence(物理智能中心-NTT项目) Center for Brain Science(脑科学中心) Harvard University(哈佛大学) Physics of Artificial Intelligence Group(人工智能物理研究组) NTT Research, Inc.(NTT研究公司)

AI总结 研究通过QSG模型揭示LLM多智能体系统中集体共识的形成机制,发现共识可能由膜etic漂移驱动,提出基于群体规模、通信带宽等因素的缩放定律。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21152 2026-03-27 physics.geo-ph cs.AI

TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology

TRACE:一种用于地震学自主物理推理的多智能体系统

Feng Liu, Jian Xu, Xin Cui, Xinghao Wang, Zijie Guo, Jiong Wang, S. Mostafa Mousavi, Xinyu Gu, Hao Chen, Ben Fei, Lihua Fang, Fenghua Ling, Zefeng Li, Lei Bai

机构 * School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Earth and Space Sciences, University of Science and Technology of China(中国科学技术大学地球和空间科学学院) Department of Earth and Planetary Sciences, Harvard University(哈佛大学地球与行星科学系) Institute of Earthquake Forecasting, China Earthquake Administration(中国地震局地震预报研究所)

AI总结 TRACE通过结合大语言模型规划与正式地震学约束,从原始观测中推导出可审计的物理机制,解决了地震序列物理机制推断的挑战,推动了地球科学从专家依赖分析向知识引导的自主发现发展。

Comments 25 pages for main text and 164 pages for appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.09534 2026-03-27 cs.IT cs.LG eess.SP math.IT

Discriminative reconstruction via simultaneous dense and sparse coding

通过同时密集和稀疏编码进行判别重建

Abiy Tasissa, Emmanouil Theodosis, Bahareh Tolooshams, Demba Ba

机构 * Department of Mathematics, Tufts University(塔夫茨大学数学系) School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院)

AI总结 本文提出结合表示能力和判别特征的新型密集稀疏编码模型,通过几何条件和凸优化方法恢复密集和稀疏向量,并验证了密集稀疏自编码器在图像去噪和频率内容捕捉方面的有效性。

Comments 27 pages. Made changes to improve the clarity and presentation of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13767 2026-03-26 cs.CV

VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI

VocSegMRI:多模态学习在实时MRI中的精确声带段分割

Daiqi Liu, Johannes Enk, Maureen Stone, Fangxu Xing, Tomás Arias-Vergara, Jerry L. Prince, Jana Hutter, Jonghye Woo, Andreas Maier, Paula Andrea Pérez-Toro

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃朗根-纽伦堡大学) University of Maryland School of Dentistry(马里兰大学牙科学院) Harvard Medical School/Massachusetts General Hospital(哈佛医学院/麻省总医院) Universidad de Antioquia UdeA(安蒂奥基亚大学 UdeA) Johns Hopkins University(约翰霍普金斯大学) Leibniz University Hannover(汉诺威莱布尼茨大学)

AI总结 本文提出VocSegMRI,通过跨注意力融合和对比学习目标,整合视频、音频和语音学输入,实现实时MRI中声带结构的精确分割,实验表明其优于单模态和多模态基线方法。

Comments Preprint submitted to MIDL short paper 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23522 2026-03-26 cs.CL cs.AI

Qworld: Question-Specific Evaluation Criteria for LLMs

Qworld: 为LLMs设计的问题特定评估标准

Shanghua Gao, Yuchang Su, Pengwei Sui, Curtis Ginder, Marinka Zitnik

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Medicine, Brigham and Women’s Hospital(布里洛妇产科医院医学部) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究 institute) Broad Institute of MIT and Harvard(MIT 和哈佛大学Broad研究所) Harvard Data Science Initiative(哈佛大学数据科学计划)

AI总结 Qworld通过递归扩展树生成问题特定评估标准,覆盖89%专家标准并产生79%新标准,揭示LLM在长期影响、公平性等维度的能力差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23322 2026-03-25 stat.AP cs.AI cs.CY physics.geo-ph

Leveraging LLMs and Social Media to Understand User Perception of Smartphone-Based Earthquake Early Warnings

利用大型语言模型和社会媒体理解智能手机地震预警的用户感知

Hanjing Wang, S. Mostafa Mousavi, Patrick Robertson, Richard M. Allen, Alexie Barski, Robert Bosch, Nivetha Thiruverahan, Youngmin Cho, Tajinder Gadh, Steve Malkos, Boone Spooner, Greg Wimpey, Marc Stogaitis

机构 * Department of Earth and Planetary Sciences, Harvard University(哈佛大学地球与行星科学系) Google LLC(谷歌公司) Seismological Laboratory, University of California, Berkeley(加州大学伯克利分校地震实验室) Google Germany GmbH(谷歌德国公司)

AI总结 本文通过分析社交媒体数据,探讨用户对智能手机地震预警系统的感知,发现用户信任与预警及时性密切相关,揭示了系统准确性的用户定义差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22231 2026-03-24 cs.IR cs.AI cs.GT cs.LG

One Model, Two Markets: Bid-Aware Generative Recommendation

一个模型,两个市场:基于竞价的生成推荐

Yanchen Jiang, Zhe Feng, Christopher P. Mah, Aranyak Mehta, Di Wang

机构 * Harvard University(哈佛大学) Google Research(谷歌研究)

AI总结 本文提出GEM-Rec框架,整合商业相关性和 monetization 目标,通过控制令牌和竞价感知解码机制实现动态优化,确保高价值物品生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22198 2026-03-24 cs.CV

Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance Learning

混合小专家:克服多实例学习中的线性层瓶颈

Daniel Shao, Joel Runevic, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, Andrew H. Song, Faisal Mahmood

机构 * Massachusetts Institute of Technology(麻省理工学院) Harvard University(哈佛大学) Emory University(埃默里大学) MD Anderson Cancer Center(MD安德森癌症中心)

AI总结 本文提出MAMMOTH模块,通过低秩变换提升多实例学习中任务特定特征提取,实验显示其在130种配置中提升性能,平均提升3.8%。

Comments Published in ICLR 2026 (37 pages, 16 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21566 2026-03-24 cs.CV cs.AI cs.DB cs.LG cs.RO

CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation

白内障手术分割与可扩展地面真实标注的领域适应模型CataractSAM-2

Mohammad Eslami, Dhanvinkumar Ganeshkumar, Saber Kazeminasab, Michael G. Morley, Michael V. Boland, Michael M. Lin, John B. Miller, David S. Friedman, Nazlee Zebardast, Lucia Sobrin, Tobias Elze

机构 * Thomas Jefferson High School for Science and Technology(托马斯·杰弗逊高中科学与技术学校) Mass Eye and Ear, Harvard Medical School(麻省眼耳研究所,哈佛医学院) Harvard Ophthalmology AI Lab, Schepens Eye Research Institute of Mass Eye and Ear, Harvard Medical School(哈佛眼科学人工智能实验室,麻省眼耳研究所的谢普恩斯眼研究学院,哈佛医学院)

AI总结 本文提出CataractSAM-2,一种针对白内障眼科手术视频的实时语义分割模型,结合稀疏提示与视频掩码传播框架,提升标注效率并推动高质量地面真实标注的可扩展生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19217 2026-03-24 cs.CL

Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+

模态匹配至关重要:在URIEL+中校准语言距离以实现跨语言迁移

York Hay Ng, Aditya Khan, Xiang Lu, Matteo Salloum, Michael Zhou, Phuong H. Hoang, A. Seza Doğruöz, En-Shiun Annie Lee

机构 * University of Toronto, Canada(多伦多大学) University of Michigan, USA(密歇根大学) Harvard University, USA(哈佛大学) Carnegie Mellon University, USA(卡内基梅隆大学) LT3, IDLab, Universiteit Gent, Belgium(IDLab,根特大学) Ontario Tech University, Canada(安大略技术大学)

AI总结 本文提出一种类型匹配的语言距离框架,通过结构感知的表示方法提升跨语言迁移性能,特别是在任务相关距离类型下表现更优。

Comments Accepted to EACL 2026 SRW

Journal ref In Proceedings of EACL 2026 (Volume 4: Student Research Workshop), pages 110 to 130. ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21398 2026-03-24 cs.AI cs.GT

Persona Vectors in Games: Measuring and Steering Strategies via Activation Vectors

游戏中的角色向量:通过激活向量测量和引导策略

Johnathan Sun, Andrew Zhang

机构 * Harvard University(哈佛大学)

AI总结 本文通过对比激活添加方法,在博弈论场景中构建了代表利他、宽恕和他人期望的角色向量,发现激活引导能系统性地改变策略选择和自然语言解释,但同时也观察到言辞与策略可能产生分歧。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21206 2026-03-24 cs.CV

Boundary-Aware Instance Segmentation in Microscopy Imaging

显微成像中的边界感知实例分割

Thomas Mendelson, Joshua Francois, Galit Lahav, Tammy Riklin-Raviv

机构 * The School of Electrical and Computer Engineering, Ben-Gurion University of the Negev(巴以大学电气与计算机工程学院) Department of Systems Biology, Harvard Medical School(哈佛医学院系统生物学系)

AI总结 本文提出一种无需提示的边界感知实例分割框架,通过预测符号距离函数实现平滑且几何一致的细胞轮廓建模,采用改进的豪斯多夫距离损失提升边界准确性和实例分离性能。

Comments Accepted for publication in IEEE International Symposium on Biomedical Imaging (ISBI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09313 2026-03-24 cs.AI

Curveball Steering: The Right Direction To Steer Isn't Always Linear

曲球操控:引导的正确方向并不总是线性

Shivam Raval, Hae Jin Song, Linlin Wu, Abir Harrasse, Jeff M. Phillips, Fazl Barez, Amirali Abdullah

机构 * Harvard Berkman Klein Center(哈佛伯克曼克莱因中心) Harvard University(哈佛大学) University of Utah(犹他大学) University of Oxford(牛津大学)

AI总结 本文提出曲球操控方法,通过多项式核PCA在特征空间中进行干预,更尊重学习到的激活几何,优于线性PCA方法,尤其在几何扭曲强的场景中表现更佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01641 2026-03-24 cs.CV

FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring

FideDiff:高效的高保真图像运动去模糊扩散模型

Xiaoyang Liu, Zhengyan Zhou, Zihang Xu, Jiezhang Cao, Zheng Chen, Yulun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Harvard University(哈佛大学)

AI总结 本文提出FideDiff,一种高效的单步扩散模型,用于高保真图像运动去模糊。通过将运动去模糊转化为扩散过程,结合Kernel ControlNet和自适应时间步预测,提升了去模糊性能。

Comments Accepted to ICLR 2026. Code is available at https://github.com/xyLiu339/FideDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20396 2026-03-24 cs.AI math.LO

Compression is all you need: Modeling Mathematics

压缩才是关键:数学建模

Vitaly Aksenov, Eve Bodnia, Michael H. Freedman, Michael Mulligan

机构 * Center of Mathematical Sciences and Applications(数学科学与应用中心) Harvard University(哈佛大学) Department of Physics and Astronomy(物理与天文学系) University of California(加州大学) Riverside, CA 92521, USA(河滨分校,CA 92521,美国)

AI总结 本文探讨了人类数学通过层次化定义压缩的特点,通过单调群模型展示数学表达的扩展性,并通过MathLib测试验证了人类数学在形式数学中的多项式增长特性。

Comments 28 pages, 5 figures, 1 appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20320 2026-03-24 cs.SE cs.AI cs.LG

The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents

工具可及性对LLM代理安全对齐的因果影响

Shasha Yu, Fiona Carroll, Barry L. Bentley

机构 * Cardiff Metropolitan University(卡地夫 Metropolitan 大学) Harvard Medical School(哈佛医学院) Clark University(克拉克大学) Harvard University(哈佛大学)

AI总结 研究通过对比文本聊天机器人与具备工具访问权限的代理行为,发现工具可及性显著增加安全违规率,表明文本评估不足以评估代理系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20229 2026-03-24 cs.CY cs.AI

Characterizing the ability of LLMs to recapitulate Americans' distributional responses to public opinion polling questions across political issues

刻画LLMs在不同政治议题上再现美国人对公共意见调查问题分布性反应的能力

Eric Gong, Nathan E. Sanders, Bruce Schneier

机构 * Harvard College(哈佛大学学院) Berkman-Klein Center, Harvard University(伯克曼-克莱因中心,哈佛大学) Harvard Kennedy School(哈佛肯尼迪学院)

AI总结 本文提出一种框架,通过直接提示LLM预测多项选择政治问题的响应分布,相比传统方法更准确且成本更低,且在不同人口统计学和问题上表现更系统可预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20101 2026-03-23 cs.AI

Pitfalls in Evaluating Interpretability Agents

评估可解释性代理中的陷阱

Tal Haklay, Nikhil Prakash, Sana Pandey, Antonio Torralba, Aaron Mueller, Jacob Andreas, Tamar Rott Shaham, Yonatan Belinkov

机构 * Kempner Institute at Harvard University(哈佛大学 Kempner 机构) Northeastern University(东北大学) Boston University(波士顿大学) MIT(麻省理工学院)

AI总结 本文研究了自动化电路分析中可解释性代理的评估挑战,提出基于模型组件功能互换性的无监督内在评估方法,揭示了复制评估的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19607 2026-03-23 cs.CV

Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning

Physion-Eval:通过人类推理评估生成视频的物理真实性

Qin Zhang, Peiyu Jing, Hong-Xing Yu, Fangqiang Ding, Fan Nie, Weimin Wang, Yilun Du, James Zou, Jiajun Wu, Bing Shuai

机构 * Physion Labs(Physion实验室) Stanford University(斯坦福大学) MIT(麻省理工学院) Harvard University(哈佛大学) Character AI

AI总结 本文提出Physion-Eval基准,通过专家人类推理评估生成视频的物理真实性,揭示当前视频生成模型在物理关键场景中存在显著缺陷,83.3%的外向视角和93.5%的内向视角视频存在可识别的物理错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02142 2026-03-23 cs.RO cs.CV

FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

FD-VLA:力感知的视觉-语言-动作模型用于接触密集的 manipulation

Ruiteng Zhao, Wenshuo Wang, Yicheng Ma, Xiaocong Li, Francis E. H. Tay, Marcelo H. Ang, Haiyue Zhu

机构 * Advanced Robotics Centre, National University of Singapore(新加坡国立大学先进机器人中心) School of Electrical & Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) College of Information Science and Technology, Eastern Institute of Technology(东部技术学院信息科学与技术学院) John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院) Advanced Robotics Centre at National University of Singapore(新加坡国立大学先进机器人中心) Singapore Institute of Manufacturing Technology, Agency for Science, Technology and Research (A*STAR)(新加坡制造技术研究所,科技研究局(A*STAR))

AI总结 本文提出FD-VLA模型,通过力蒸馏模块在不依赖物理力传感器的情况下实现接触密集任务的力感知,提升机器人视觉-语言-动作的鲁棒性与实用性。

Comments ICRA 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15872 2026-03-23 cs.LG

Hidden Breakthroughs in Language Model Training

语言模型训练中的隐藏突破

Sara Kangaslahti, Elan Rosenfeld, Naomi Saphra

机构 * Harvard University(哈佛大学) Google Research(谷歌研究)

AI总结 本文提出POLCA方法,通过分解损失变化来揭示训练过程中的隐藏转折点,通过合成任务验证其在模型能力解读中的潜力。

Comments ICLR 2026 Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19517 2026-03-23 cs.CV cs.LG

ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding

ReXInTheWild:医疗照片理解的统一基准

Oishi Banerjee, Sung Eun Kim, Alexandra N. Willauer, Julius M. Kernbach, Abeer Rihan Alomaish, Reema Abdulwahab S. Alghamdi, Hassan Rayhan Alomaish, Mohammed Baharoon, Xiaoman Zhang, Julian Nicolas Acosta, Christine Zhou, Pranav Rajpurkar

机构 * Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) National Strategic Technology Research Institute, Seoul National University Hospital(首尔国立大学医院国家战略技术研究所) Department of Medicine, Division of Gastroenterology, Massachusetts General Hospital(麻省总医院内科与消化内科部门) Department of Neuroradiology, Heidelberg University Hospital(海德堡大学医院神经放射科部门) King Abdulaziz Medical City, National Guard Health Affairs(国王阿卜杜勒-阿齐兹医疗城,国民守卫健康事务局) Division of Pulmonary, Critical Care, and Sleep Medicine, University of Cincinnati(辛辛那提大学肺科、重症医学及睡眠医学部门)

AI总结 本文提出ReXInTheWild基准,通过484张医学文献来源的照片和7个临床主题的955个多项选择问题,评估视觉-语言模型对医疗照片内容的理解能力,揭示不同模型在细粒度图像理解和医学推理上的性能差异。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13219 2026-03-23 cs.CV

Prompt-based Adaptation in Large-scale Vision Models: A Survey

基于提示的大规模视觉模型适应:综述

Xi Xiao, Yunbei Zhang, Lin Zhao, Yiyang Liu, Xiaoying Liao, Zheda Mai, Xingjian Li, Xiao Wang, Hao Xu, Jihun Hamm, Xue Lin, Min Xu, Qifan Wang, Tianyang Wang, Cheng Han

机构 * University of Alabama at Birmingham, USA(美国阿拉巴马大学伯明翰分校) Tulane University, USA(美国路易斯安那大学) Northeastern University, USA(美国东北大学) University of Missouri-Kansas City, USA(美国密苏里大学堪萨斯城分校) Johns Hopkins University, USA(约翰霍普金斯大学) Ohio State University, USA(俄亥俄州立大学) Carnegie Mellon University, USA(卡内基梅隆大学) Oak Ridge National Laboratory, USA(橡树岭国家实验室) Harvard University, USA(哈佛大学) Mohamed bin Zayed University of Artificial Intelligence, UAE(阿联酋Mohamed bin Zayed人工智能大学) Meta AI, USA(Meta AI)

AI总结 本文综述了基于提示的视觉模型适应方法,探讨了VP和VPT的核心机制及在不同领域的应用,揭示了其在测试时适应和可信AI中的作用,并总结了当前基准和未来研究方向。

Comments Accepted by TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18042 2026-03-20 eess.IV cs.LG

A Novel Framework using Intuitionistic Fuzzy Logic with U-Net and U-Net++ Architecture: A case Study of MRI Bain Image Segmentation

一种利用直觉模糊逻辑与U-Net和U-Net++架构的新型框架:MRI脑图像分割的案例研究

Hanuman Verma, Kiho Im, Akshansh Gupta, M. Tanveer

机构 * Department of Mathematics, Bareilly College, Bareilly (MJP Rohilkhand University), Uttar Pradesh, 243005, India(巴里利学院数学系,巴里利(MJP罗希兰德大学),乌塔尔普拉德什,243005,印度) Fetal Neonatal Neuroimaging and Developmental Science Center, Boston Children’s Hospital, Harvard Medical School, Boston, MA 02115, USA and Division of Newborn Medicine, Boston Children’s Hospital, Harvard Medical School, Boston, MA 02115, USA also with Department of Pediatrics, Harvard Medical School, Boston, MA, USA(波士顿儿童医院胎儿和新生儿神经影像与发育科学中心,哈佛医学院,波士顿,马萨诸塞州02115,美国;波士顿儿童医院新生儿医学科,哈佛医学院,波士顿,马萨诸塞州02115,美国;也与哈佛医学院儿科系,波士顿,马萨诸塞州,美国) National Institute of Science Communication and Policy Research, New Delhi, 110012, India(国家科学传播与政策研究所,新德里,110012,印度) Department of Mathematics, Indian Institute of Technology Indore, Indore (M.P.)-453552, India(印度理工学院印度尔数学系,印度尔(马哈拉施特拉邦)-453552,印度)

AI总结 本文提出一种结合直觉模糊逻辑的U-Net和U-Net++框架,用于提高MRI脑图像分割的准确性,通过处理图像中的不确定性提升分割性能。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03182 2026-03-20 cs.RO cs.AI cs.CL cs.SC

Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning

模拟到规则:一个双VLM框架用于正式视觉规划

Yilun Hao, Yongchao Chen, Chuchu Fan, Yang Zhang

机构 * MIT(麻省理工学院) Harvard University(哈佛大学) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

AI总结 本文提出VLMFP框架,结合SimVLM和GenVLM实现自动生成PDDL问题和领域文件,提升视觉规划的精确性和泛化能力。

Comments 40 pages, 6 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏