arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-03-19 至 2026-03-19 共收录 77 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 77 篇

2603.17067 2026-03-19 cs.CL cs.AI 90%

Evaluating Ill-Defined Tasks in Large Language Models

评估大型语言模型中的模糊任务

Yi Zhou, Basel Shbita

机构 * IBM Research San Jose, CA(IBM研发San Jose分校)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文分析了现有评估基准在处理模糊任务时的不足,通过复杂指令跟随和自然语言到Mermaid序列图的案例研究,揭示了评估方法的不稳定性与非诊断性,推动更稳健的评估设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13526 2026-03-19 cs.CL 89%

MATA: Mindful Assessment of the Telugu Abilities of Large Language Models

MATA:对大型语言模型泰卢格语能力的有意识评估

Chalamalasetti Kranti, Sowmya Vajjala

机构 * Department of Linguistics, University of Potsdam(波恩大学语言学系) National Research Council(加拿大国家研究委员会)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文介绍MATA数据集,用于评估大型语言模型在泰卢格语中的能力,包含729道精心编写的多选和开放式问题,评估11种开源和闭源LLM,并分析其表现,揭示LLM在多选题中依赖表面启发法的现象,比较LLM作为评判者与人类评估在低资源语言中的可靠性。

Comments Accepted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17123 2026-03-19 cs.CR cs.AI 89%

Security Assessment and Mitigation Strategies for Large Language Models: A Comprehensive Defensive Framework

大型语言模型的安全性评估与缓解策略:一种全面的防御框架

Taiwo Onitiju, Iman Vakilinia

机构 * School of Computing University of North Florida(北弗罗里达大学计算学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文评估了大型语言模型在关键基础设施中的安全风险,提出了一种标准化的评估框架和多层次防御系统,通过测试五种主流模型对1万多个对抗性提示的响应,揭示了安全差异,并开发了高准确率的防御框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16934 2026-03-19 cs.CV cs.AI 88%

AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding

AgriChat:一种用于农业图像理解的多模态大语言模型

Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed

机构 * Department of Computer Science, Khalifa University of Science and Technology(卡利法科技大学计算机科学系)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出AgriChat,一种基于多模态大语言模型的农业图像理解系统,通过整合视觉描述与网络增强的科学检索,生成农业基准数据集,提升农业领域的模型鲁棒性和可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05591 2026-03-19 cs.AI 88%

MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models

MLlm-DR: 向多模态大语言模型的可解释性抑郁识别迈进

Wei Zhang, Juan Chen, En Zhu, Wenhong Cheng, YunPeng Li, Yanbo J. Wang

机构 * National University of Defense Technology(国防科技大学) University of Chinese Academy of Sciences(中国科学院大学) Shanghai Mental Health Center, Shanghai Jiao Tong University School of Medicine(上海精神卫生中心,上海交通大学医学院) Nanjing Industria Tenebris Information Technology Co., Ltd.(南京Industria Tenebris信息技术有限公司) National University of Uzbekistan named after Mirzo Ulugbek(乌兹别克斯坦Mirzo Ulugbek命名的国立大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出MLlm-DR模型,通过多模态大语言模型实现可解释的抑郁识别,结合小型LLM和轻量级查询模块,提升诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25620 2026-03-19 cs.CV 88%

LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology

LMOD+: 一个全面的多模态数据集和基准,用于开发和评估眼科中的多模态大语言模型

Zhenyue Qin, Yang Liu, Yu Yin, Jinyu Ding, Haoran Zhang, Anran Li, Dylan Campbell, Xuansheng Wu, Ke Zou, Tiarnan D. L. Keenan, Emily Y. Chew, Zhiyong Lu, Yih Chung Tham, Ninghao Liu, Xiuzhen Zhang, Qingyu Chen

机构 * School of Medicine, Yale University(耶鲁大学医学院) School of Computing, Australian National University(澳大利亚国立大学计算机学院) School of Engineering, Imperial College London(伦敦帝国理工学院工程学院) School of Computing, University of Georgia(佐治亚大学计算机学院) Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨秀隆医学学院) National Eye Institute, National Institutes of Health(美国国立卫生研究院眼科研究所) National Library of Medicine, National Institutes of Health(美国国立卫生研究院国家医学图书馆) School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算机技术学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract)

AI总结 本文提出一个包含32633个实例的多模态眼科基准数据集,涵盖12种常见眼科疾病和5种成像模态,通过扩展数据集、扩展任务覆盖和系统评估24种最先进的MLLMs,推动眼科AI应用发展。

Comments ACM Transactions on Computing for Healthcare

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17094 2026-03-19 cs.CL 87%

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction

评估LLM模拟的对话在建模不一致和不合作行为中的表现

Ryo Kamoi, Ameya Godbole, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou

机构 * Microsoft Corporation(微软公司) Penn State University(宾夕法尼亚州立大学) University of Southern California(南加州大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出CoCoEval框架,通过检测10种不一致和不合作行为评估LLM生成对话与人类对话的差异,发现LLM生成的对话在常规提示下较少出现不一致行为,但通过提示工程难以有效控制,监督微调可能导致过度生成特定行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17217 2026-03-19 cs.CL cs.AI cs.LG 87%

Anonymous-by-Construction: An LLM-Driven Framework for Privacy-Preserving Text

匿名由设计:一个基于LLM的隐私保护文本框架

Federico Albanese, Pablo Ronco, Nicolás D'Ippolito

机构 * Veritran University of Buenos Aires(布宜诺斯艾利斯大学) University of San Andrés(圣安德烈斯大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于LLM的匿名化框架,通过本地LLM实现文本隐私保护,同时保持语义和任务相关性,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17683 2026-03-19 cs.AI cs.LG 86%

Sensi: Learn One Thing at a Time -- Curriculum-Based Test-Time Learning for LLM Game Agents

Sensi:一次学一项——基于课程的学习式测试时间学习用于LLM游戏代理

Mohsen Arjmandi

机构 * Independent Researcher (CTO, evolutionID)(独立研究者(CTO, evolutionID))

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 Sensi通过结构化测试时间学习机制,提升LLM游戏代理在未知环境中的学习效率,其核心方法包括双玩家架构、课程学习系统和数据库作为控制平面,显著提高了样本效率。

Comments Preprint. 18 pages, 5 figures, 2 tables. Independent research. Code and Colab demo coming soon on GitHub

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24384 2026-03-19 cs.CL cs.AI 86%

HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment

HarmMetric Eval: 对LLM有害性评估的度量和裁判进行基准测试

Langqi Yang, Tianhang Zheng, Yixuan Chen, Kedong Xiu, Hao Zhou, Wangze Ni, Lei Chen, Zhan Qin, Kui Ren

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Zhejiang University(浙江大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院) Pretraining Team, xAI(预训练团队,xAI) Hong Kong University of Science and Technology (GuangZhou)(香港科技大学(广州))

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出HarmMetric Eval,用于评估有害性度量和裁判的质量,发现传统参考型指标在细粒度评估中表现优于LLM裁判,改进的裁判通过结合细粒度标准和参考型指标提升了效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17680 2026-03-19 cs.CV cs.AI 85%

WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models

WeatherReasonSeg:一种用于视觉语言模型中天气感知推理分割的基准测试

Wanjun Du, Zifeng Yuan, Tingting Chen, Fucai Ke, Beibei Lin, Shunli Zhang

机构 * Beijing Jiaotong University(北京交通大学) National University of Singapore(新加坡国立大学) Monash University(墨尔本大学)

专题命中 评测与基准 :language model(title,abstract);LLM(abstract);prompting(abstract);分类 cs.AI

AI总结 本文提出WeatherReasonSeg基准测试,评估视觉语言模型在恶劣天气下的推理分割能力,通过合成和真实数据集分析天气对模型性能的影响,发现天气严重程度与模型性能呈单调下降关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17172 2026-03-19 cs.LG 85%

Noise-Response Calibration: A Causal Intervention Protocol for LLM-Judges

噪声响应校准:一种用于LLM判官的因果干预协议

Maxim Khomiakov, Jes Frellsen

机构 * Technical University of Denmark(丹麦技术大学) Normal Computing Corporation(Normal Computing公司) Pioneer Centre for Artificial Intelligence(人工智能先锋中心)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出基于受控输入干预的校准协议,通过信号噪声比扰动和词法扰动评估LLM判官在不同模态下的表现,揭示了文本与表格数据在噪声下的差异行为。

Comments Published as a conference paper at CAO Workshop at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06396 2026-03-19 cs.AI cs.CR 85%

Efficient LLM Safety Evaluation through Multi-Agent Debate

通过多智能体辩论实现高效的LLM安全评估

Dachuan Lin, Guobin Shen, Zihao Yang, Tianrong Liu, Dongcheng Zhao, Yi Zeng

机构 * Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) Beijing Key Laboratory of Safe AI and Super Alignment(北京安全人工智能与超对齐重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) CSE, The Chinese University of Hong Kong(香港中文大学电子工程系) University of Chinese Academy of Sciences(中国科学院大学) Department of Mathematics, The Chinese University of Hong Kong(香港中文大学数学系) Long-term AI(长期人工智能)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出多智能体辩论框架,利用HAJailBench基准测试,提升LLM安全评估的可靠性与经济性,验证了少量辩论轮次即可获得显著效果。

Comments 15 pages, 5 figures, 10 tables. Updated abstract to fix an incconsistency issue with the main paper: HAJailBench size (12,000 -> 11,100)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10204 2026-03-19 cs.SE cs.LG 85%

Code Roulette: How Prompt Variability Affects LLM Code Generation

代码掷骰子:提示变化如何影响LLM代码生成

Andrei Paleyes, Radzim Sendyka, Diana Robinson, Christian Cabrera, Neil D. Lawrence

机构 * University of Cambridge(剑桥大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 研究探讨提示变化对LLM代码生成质量的影响,提出评估流程以量化模型对输入变化的敏感性,通过实验验证方法有效性。

Comments Extended version of the paper accepted to LLM4Code @ ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04893 2026-03-19 cs.CV cs.AI 85%

SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models

SCAM:多模态基础模型的现实世界字形鲁棒性评估

Justus Westerhoff, Erblina Purelku, Jakob Hackstein, Jonas Loos, Leo Pinetzki, Erik Rodner, Lorenz Hufe

机构 * BLISS e.V.(BLISS协会) Berliner Hochschule für Technik (BHT)(柏林技术大学) Technische Universität Berlin(柏林技术大学) KI Werkstatt, Hochschule für Technik und Wirtschaft Berlin (HTW)(柏林技术经济学院) Merantix Momentum(Merantix Momentum公司) Fraunhofer Heinrich-Hertz-Institut, Berlin, Germany(柏林弗劳恩霍夫 Heinrich-Hertz 研究所)

专题命中 评测与基准 :foundation model(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出SCAM数据集,用于评估多模态基础模型对字形攻击的鲁棒性,发现训练数据和模型架构影响攻击敏感性,且大语言模型基座可降低攻击影响。

Comments Accepted at CVPR 2025 Workshop EVAL-FoMo-2

Journal ref Journal of Data-centric Machine Learning Research (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21679 2026-03-19 cs.CL 84%

ReviewScore: Misinformed Peer Review Detection with Large Language Models

ReviewScore:利用大语言模型进行误信同行评审检测

Hyun Ryu, Doohyuk Jang, Hyemin S. Lee, Joonhyun Jeong, Gyeongman Kim, Donghyeon Cho, Gyouk Chu, Minyeong Hwang, Hyeongwon Jang, Changhun Kim, Haechan Kim, Jina Kim, Joowon Kim, Yoonjeon Kim, Kwanhyung Lee, Chanjae Park, Heecheol Yun, Gregor Betz, Eunho Yang

机构 * KAIST(韩国科学技术院) Carnegie Mellon University(卡内基梅隆大学) MIT(麻省理工学院) KRAFTON(KRAFTON公司) AITRICS(AITRICS研究院) KIT(卡尔斯鲁厄理工学院)

专题命中 评测与基准 :large language model(title);language model(title);分类 cs.CL

AI总结 本文提出ReviewScore方法,通过检测同行评审中的错误前提和可回答问题来识别低质量评审。利用大语言模型评估评审点的误信程度,实验显示模型在评估一致性上表现中等,但仍有改进空间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17017 2026-03-19 cs.CL cs.AI 84%

LLM NL2SQL Robustness: Surface Noise vs. Linguistic Variation in Traditional and Agentic Settings

LLM NL2SQL鲁棒性:传统与智能设置中表面噪声与语言变异的比较

Lifu Tu, Rongguang Wang, Tao Sheng, Sujjith Ravi, Dan Roth

机构 * Oracle AI

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文评估了NL2SQL系统在表面噪声和语言变异下的鲁棒性,发现传统和智能设置中模型表现各异,揭示了实现鲁棒NL2SQL系统的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13417 2026-03-19 cs.AI cs.CL 84%

Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discourse

通过气候 discourse 中的隐含因果链发现评估 LLM 推理

Liesbeth Allein, Nataly Pineda-Castañeda, Andrea Rocci, Marie-Francine Moens

机构 * Department of Computer Science, KU Leuven(卢森堡大学计算机科学系) Institute of Argumentation, Linguistics, and Semiotics (IALS)(论证、语言学与象征学研究所)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文通过隐含因果链发现任务评估 LLM 的因果推理能力,发现其生成的因果步骤数量和粒度存在差异,但主要依赖关联模式匹配而非真实因果推理。

Comments LREC 2026, appendix to be found in ACL anthology

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17522 2026-03-19 cs.CL cs.AI 83%

Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions

检测机器:跨架构、领域和对抗条件的AI生成文本检测器全面基准

Madhav S. Baidya, S. S. Baidya, Chirag Chawla

机构 * Indian Institute of Technology (BHU)(印度理工学院 (BHU)) Indian Institute of Technology Guwahati(印度理工学院古瓦哈蒂)

专题命中 评测与基准 :LLM(abstract,comments);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出跨HC3和ELI5两个语料库的全面基准,评估不同检测方法在领域转移和对抗鲁棒性上的表现,发现Transformer模型在分布内表现优异但易受领域偏移影响,XGBoost模型在可解释性上表现良好。

Comments ~30 pages, 10+ figures. Code available at: https://github.com/MadsDoodle/Human-and-LLM-Generated-Text-Detectability-under-Adversarial-Humanization

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23562 2026-03-19 cs.LG cs.AI cs.CL 82%

VL-RouterBench: A Benchmark for Vision-Language Model Routing

VL-RouterBench:一种用于视觉-语言模型路由的基准测试

Zhehao Huang, Baijiong Lin, Jingyuan Zhang, Jingying Wang, Yuhang Liu, Ning Lu, Tao Li, Xiaolin Huang

机构 * Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(1 图像处理与模式识别研究所,上海交通大学) The Hong Kong University of Science and Technology (Guangzhou)(2 香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(3 香港科学与技术大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出VL-RouterBench,用于系统评估视觉-语言模型路由系统的能力,通过构建质量与成本矩阵,评估10种路由方法并发现显著的路由改进空间,揭示了路由架构改进的潜力。

Comments CVPR 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17952 2026-03-19 cs.CL 81%

Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures

机器翻译中的性别歧义化解:解码器-only架构的诊断评估

Chiara Manna, Hosein Mohebbi, Afra Alishahi, Frédéric Blain, Eva Vanmassenhove

机构 * Research Center for Cognitive Science and Artificial Intelligence(认知科学与人工智能研究中心) Tilburg School of Humanities and Digital Sciences(蒂尔堡人文与数字科学学院) Tilburg University(蒂尔堡大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);instruction tuning(abstract);post-training(abstract)

AI总结 本文针对机器翻译中性别偏见问题,提出新的评估框架,揭示解码器-only模型在性别特定指标上的表现及后训练对偏见的缓解作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17832 2026-03-19 cs.CL cs.AI cs.LG 80%

Text-to-Stage: Spatial Layouts from Long-form Narratives

文本到舞台:从长篇叙述中生成空间布局

Jefferson Hernandez, Swarnadeep Saha, Chenxi Whitehouse, Sanjeel Parekh, Calvin Murdock, Yuliang Li, W. Owen Brimijoin, Vamsi Krishna Ithapu, Ishwarya Ananthabhotla

机构 * Rice University(里士大学) Meta FAIR Meta Reality Labs Research(Meta现实实验室)

专题命中 评测与基准 :LLM(abstract);language model(abstract);SFT(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究如何通过语言模型从无结构文本推断戏剧舞台布局,提出确定性评估体系和结合拒绝SFT与RL的训练方法,在多个指标上优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17231 2026-03-19 cs.CL eess.AS 79%

Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models

语音生成大音频-语言模型中的神经层面情感控制

Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn, Björn Schuller, Berrak Sisman

机构 * 1 Center for Language Speech Processing (CLSP), Johns Hopkins University, USA 2 Group on Language, Audio \& Music (GLAM), Imperial College London, UK

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL

AI总结 本文首次在语音生成大音频-语言模型中研究神经层面的情感控制,通过识别紧凑的情感敏感神经元(ESNs)实现无需训练的情感引导,实验表明其在不同模型上均能提升情感表现并支持自动与人工评估。

Comments 11 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17111 2026-03-19 cs.CV cs.AI 79%

Hidden Clones: Exposing and Fixing Family Bias in Vision-Language Model Ensembles

隐藏克隆:揭示并修复视觉-语言模型集成中的家族偏见

Zacharie Bugaud

机构 * Astera Institute(Astera研究院)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

AI总结 本文研究了视觉-语言模型集成中同家族模型的关联误差问题,提出三种家族感知方法,通过层级家族投票、QualRCCV和学得候选评分,提升了集成性能并验证了方法的泛化能力。

Comments 15 pages, 6 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17048 2026-03-19 cs.LG cs.CV 79%

SCE-LITE-HQ: Smooth visual counterfactual explanations with generative foundation models

SCE-LITE-HQ:基于生成基础模型的平滑视觉反事实解释

Ahmed Zeid, Sidney Bender

机构 * BIFOLD – Berlin Institute for the Foundations of Learning(柏林学习与数据基础研究所)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 本文提出SCE-LITE-HQ框架,利用预训练生成模型在潜在空间中生成稳定且多样的反事实解释,避免训练专用生成模型的开销,适用于高分辨率数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17813 2026-03-19 cs.CV 78%

M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking

M2P:通过Mask-to-Point弱监督学习改进视觉基础模型以实现密集点跟踪

Qiangqiang Wu, Tianyu Yang, Bo Fang, Jia Wan, Matias Di Martino, Guillermo Sapiro, Antoni B. Chan

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系) Meituan, Shenzhen, China(美团(深圳,中国)) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Department of Electrical and Computer Engineering, Duke University(杜克大学电气与计算机工程系)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 本文提出M2P学习方法,利用视频对象分割掩码注解改进视觉基础模型,通过局部结构一致性损失、掩码标签一致性损失和掩码边界约束,提升密集点跟踪性能,实验表明其在TAP-Vid-DAVIS基准上优于DINOv2和DINOv3。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17675 2026-03-19 cs.CV 78%

DeepCORO-CLIP: A Multi-View Foundation Model for Comprehensive Coronary Angiography Video-Text Analysis and External Validation

DeepCORO-CLIP:一种多视图基础模型,用于综合冠状动脉造影视频-文本分析和外部验证

Sarra Harrabi, Yichen Wu, Geoffrey H. Tison, Minhaj Ansari, Milos Vukadinovic, David Ouyang, Joshua P. Barrios, Jacques Delfrate, Robert Avram

机构 * Division of Cardiology, Department of Medicine, Montreal Heart Institute(蒙特利尔心脏研究所心血管科,医学部) McGill, Faculty of Engineering(麦吉尔大学工程学院) Division of Cardiology, Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学部心血管科) Cardiovascular Research Institute, University of California, San Francisco(加州大学旧金山分校心血管研究所) Bakar Computational Health Sciences Institute, University of California, San Francisco(加州大学旧金山分校Bakar计算健康科学研究所) Department of Cardiology, Smidt Heart Institute, Cedars-Sinai Medical Center, Los Angeles, CA(Cedars-Sinai 医疗中心洛杉矶分院心血管科,Smidt心脏研究所) Department of Bioengineering, University of California Los Angeles(加州大学洛杉矶分校生物工程系) Department of Medicine, Kaiser Permanente, San Francisco, USA(Kaiser Permanente 旧金山分院医学部) Department of Medicine, Université de Montréal(蒙特利尔大学医学部)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 DeepCORO-CLIP通过多视图视频-文本对比学习,实现了对冠状动脉造影的全面分析与外部验证,显著提升了斑块检测和疾病预测的准确性。

Comments 69 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15967 2026-03-19 cs.CV 78%

A Comprehensive Benchmark of Histopathology Foundation Models for Kidney Digital Pathology Images

一种针对肾数字病理图像的病理基础模型综合基准测试

Harishwar Reddy Kasireddy, Patricio S. La Rosa, Akshita Gupta, Anindya S. Paul, Jamie L. Fermin, William L. Clapp, Meryl A. Waldman, Tarek M. El-Ashkar, Sanjay Jain, Luis Rodrigues, Kuang Yu Jen, Avi Z. Rosenberg, Michael T. Eadon, Jeffrey B. Hodgin, Pinaki Sarder

机构 * Department of Electrical and Computer Engineering, University of Florida(佛罗里达大学电气与计算机工程系) Seed Production Innovation, Crop Science Division, Bayer Company(拜耳公司种子生产创新部) Division of Medicine – Quantitative Health, University of Florida(佛罗里达大学医学部-定量健康分部) Department of Health Outcomes and Biomedical Informatics, University of Florida(佛罗里达大学健康结果与生物医学信息学系) Department of Pathology, Immunology and Laboratory Medicine, University of Florida College of Medicine(佛罗里达大学医学院病理学、免疫学与实验室医学系) Kidney Disease Branch, National Institute of Diabetes and Digestive and Kidney Diseases, National Institutes of Health(美国国立卫生研究院糖尿病、消化系统与肾病研究所肾病分支) Indiana University School of Medicine(印第安纳大学医学院) Departments of Medicine, Washington University School of Medicine(华盛顿大学医学院医学部) Universidade de Coimbra(科英布拉大学) Department of Pathology and Laboratory Medicine, University of California at Davis School of Medicine(加州大学戴维斯分校医学院病理学与实验室医学系) Department of Pathology, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院病理学系)

专题命中 评测与基准 :foundation model(title,abstract)

AI总结 本文评估了11种公开的病理基础模型在11个肾脏特定下游任务中的表现,揭示了其在肾病诊断中的适用性及改进方向。

Comments 31 Pages, 14 Tables, 12 figures, Co-correspondence to jhodgin@med.umich.edu and pinaki.sarder@ufl.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09262 2026-03-19 cs.CV 78%

MultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models

MultiMedEval:用于评估医学视觉-语言模型的基准和工具包

Corentin Royer, Bjoern Menze, Anjany Sekuboyina

机构 * University of Zürich(苏黎世大学)

专题命中 评测与基准 :language model(title,abstract)

AI总结 MultiMedEval通过23个数据集和11个医学领域,评估六种多模态任务,提供开源工具简化VLM评估,促进公平统一的基准测试。

Comments Accepted at MIDL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17737 2026-03-19 cs.LG 77%

Embedding World Knowledge into Tabular Models: Towards Best Practices for Embedding Pipeline Design

将世界知识嵌入表格模型:迈向嵌入管道设计的最佳实践

Oksana Kolomenko, Ricardo Knauer, Erik Rodner

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文通过系统评估256种管道配置,探讨了如何设计有效的LLM嵌入管道,发现嵌入方式和模型选择对预测性能有显著影响,梯度提升树表现优异。

Comments Computational Intelligence 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏