arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

共收录 410
2604.14656 2026-04-17 cs.AI cs.CL cs.CV

Rethinking Patient Education as Multi-turn Multi-modal Interaction

重新思考患者教育作为多轮多模态交互

Zonghai Yao, Zhipeng Tang, Chengtao Lin, Xiong Luo, Benlu Wang, Juncheng Huang, Chin Siang Ong, Hong Yu

机构 * VA Bedford Health Care(VA贝德福德医疗中心) UMass Amherst(马萨诸塞大学阿默斯特分校) UMass Lowell(马萨诸塞大学洛厄尔分校) Yale University(耶鲁大学) National University of Singapore(新加坡国立大学) Yale School of Medicine(耶鲁医学院)

AI总结 本文提出MedImageEdu基准,通过多轮多模态交互提升患者教育效果,评估咨询过程和最终响应质量,发现多模态模型在视觉 grounding、安全性和情绪互动方面存在不足。

Comments Equal contribution for the first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14256 2026-04-17 cs.IR cs.AI

Evaluation of Agents under Simulated AI Marketplace Dynamics

对模拟AI市场动态下代理的评估

To Eun Kim, Alireza Salemi, Hamed Zamani, Fernando Diaz

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出Marketplace Evaluation框架,通过模拟市场动态评估信息访问系统,补充传统准确度指标,探讨市场留存和份额等指标。

Comments SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13814 2026-04-16 cs.HC cs.AI

Cognitive Offloading in Agile Teams: How Artificial Intelligence Reshapes Risk Assessment and Planning Quality

敏捷团队中的认知卸载:人工智能如何重塑风险评估与规划质量

Adriana Caraeni, Alexander Shick, Andrew Lan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

AI总结 本文通过实验比较AI、人类和混合规划模型,发现AI虽高效但风险评估差,人类虽灵活但成本高,提出混合框架提升规划质量。

Comments 7 pages, 5 Tables, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13453 2026-04-16 cs.LG

FAST: A Synergistic Framework of Attention and State-space Models for Spatiotemporal Traffic Prediction

FAST:一种结合注意机制和状态空间模型的协同框架用于时空交通预测

Xinjin Li, Jinghan Cao, Mengyue Wang, Yue Wu, Longxiang Yan, Yeyang Zhou, Ziqi Sha, Yu Ma

机构 * Columbia University(哥伦比亚大学) San Francisco State University(旧金山州立大学) University of California, Berkeley(加州大学伯克利分校) New York University(纽约大学) University of Pennsylvania(宾夕法尼亚大学) University of California, San Diego(加州大学圣地亚哥分校) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) Carnegie Mellon University(卡内基梅隆大学)

AI总结 FAST结合注意机制和状态空间模型,提出了一种可扩展的时空交通预测框架,通过时空-时空架构和Mamba基的空模块,有效捕捉短期和长期时间模式及长距离传感器依赖,实验表明其在精度、可扩展性和泛化性上均优于现有方法。

Comments Accepted by ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06168 2026-04-16 cs.CV cs.RO

Action Images: End-to-End Policy Learning via Multiview Video Generation

动作图像:通过多视图视频生成实现端到端策略学习

Haoyu Zhen, Zixian Gao, Qiao Sun, Yilin Zhao, Yuncong Yang, Yilun Du, Pengsheng Guo, Tsun-Hsuan Wang, Yi-Ling Qiao, Chuang Gan

机构 * UMass Amherst(马萨诸塞大学阿默斯特分校) NVIDIA(英伟达) Harvard University(哈佛大学) Genesis AI

AI总结 本文提出Action Images,通过多视图视频生成实现策略学习,利用像素化的动作表示,使视频模型本身成为零样本策略,提升视频-动作联合生成质量。

Comments Project Page: https://actionimages.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13115 2026-04-16 cs.CL cs.IR

Agentic Conversational Search with Contextualized Reasoning via Reinforcement Learning

基于强化学习的上下文化推理对话搜索代理

Fengran Mo, Yifan Gao, Sha Li, Hansi Zeng, Xin Liu, Zhaoxuan Tan, Xian Li, Jianshu Chen, Dakuo Wang, Meng Jiang

机构 * University of Montreal(蒙特利尔大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University of Notre Dame(诺特丹大学) Northeastern University(东北大学)

AI总结 本文提出一种基于强化学习的对话搜索代理,通过在对话轮次中交替搜索与推理,实现探索性与适应性行为,优于现有基准方法。

Comments Accepted by ACL 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05206 2026-04-15 cs.CR cs.AI cs.CL cs.CV

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

大规模安全:大型模型和智能体安全的全面综述

Xingjun Ma, Yifeng Gao, Yixu Wang, Ruofan Wang, Xin Wang, Ye Sun, Yifan Ding, Hengyuan Xu, Yunhao Chen, Yunhan Zhao, Hanxun Huang, Yige Li, Yutao Wu, Jiaming Zhang, Xiang Zheng, Yang Bai, Zuxuan Wu, Xipeng Qiu, Jingfeng Zhang, Yiming Li, Xudong Han, Haonan Li, Jun Sun, Cong Wang, Jindong Gu, Baoyuan Wu, Siheng Chen, Tianwei Zhang, Yang Liu, Mingming Gong, Tongliang Liu, Shirui Pan, Cihang Xie, Tianyu Pang, Yinpeng Dong, Ruoxi Jia, Yang Zhang, Shiqing Ma, Xiangyu Zhang, Neil Gong, Chaowei Xiao, Sarah Erfani, Tim Baldwin, Bo Li, Masashi Sugiyama, Dacheng Tao, James Bailey, Yu-Gang Jiang

机构 * Fudan University(复旦大学) The University of Melbourne(墨尔本大学) Singapore Management University(新加坡国立大学) Deakin University(德肯大学) Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) University of Oxford(牛津大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) The University of Sydney(悉尼大学) Griffith University(格里菲斯大学) University of California, Santa Cruz(加州大学圣克鲁兹分校) Sea AI Lab(Sea AI实验室) Tsinghua University(清华大学) Virginia Tech(弗吉尼亚理工大学) CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全部) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) Purdue University(普渡大学) Duke University(杜克大学) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) RIKEN(理化学研究所) The University of Tokyo(东京大学)

AI总结 本文综述了大型模型和智能体的安全性,分析了各类攻击威胁及防御策略,指出安全评估、防御机制和数据实践的重要性,强调研究社区和国际合作的必要性。

Comments 706 papers, 60 pages, 3 figures, 14 tables; GitHub: https://github.com/xingjunm/Awesome-Large-Model-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12066 2026-04-15 cs.AI cs.CY

Mathematics Teachers Interactions with a Multi-Agent System for Personalized Problem Generation

数学教师与多智能体系统在个性化问题生成中的互动

Candace Walkington, Theodora Beauchamp, Fareya Ikram, Merve Koçyiğit Gürbüz, Fangli Xia, Margan Lee, Andrew Lan

机构 * Southern Methodist University(南方 Methodist 大学) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) Worchester Polytechnic Institute(沃斯顿理工学院)

AI总结 研究探讨了多智能体教师闭环系统在中学数学问题个性化中的应用,通过教师输入基础问题和主题,由LLM生成问题,再由四个AI代理评估问题的数学准确性、真实性、可读性和现实性。

Comments Paper accepted to AIED 2026 - South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15876 2026-04-15 cs.CY cs.AI

Should There be a Teacher In-the-Loop? A Study of Generative AI Personalized Tasks Middle School

应该有教师在循环中吗?生成式AI个性化任务中学中学的探讨

Candace Walkington, Mingyu Feng, Itffini Pruitt-Britton, Theodora Beauchamp, Andrew Lan

机构 * Department of Teaching and Learning(教学与学习系) Southern Methodist University(南方 Methodist 大学) WestEd(西斯特教育) American Institutes for Research(美国研究机构) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

AI总结 研究探讨生成式AI在中学个性化任务中的应用,分析教师与AI协作的效率及学生对不同粒度个性化任务的偏好。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05902 2026-04-14 cs.CR cs.CL

Defending against Backdoor Attacks via Module Switching

通过模块切换防御后门攻击

Weijun Li, Ansh Arora, Xuanli He, Mark Dras, Qiongkai Xu

机构 * Macquarie University(麦考瑞大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University College London(伦敦大学学院)

AI总结 本文提出模块切换防御方法,通过验证理论依据和实验效果,展示其在减少后门攻击有效性方面优于现有方法,适用于深层网络和多种架构,实验表明其在较少模型情况下具有更强的防御能力。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11036 2026-04-14 cs.CL cs.AI

Uncertainty-Aware Web-Conditioned Scientific Fact-Checking

具有不确定性的网络条件科学事实核查

Ashwin Vinod, Katrin Erk

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出一种基于原子谓词-论元分解和校准的不确定性和网络验证方法,通过嵌入对齐局部片段,使用紧凑的证据基础检查器验证事实,并仅在不确定支持的事实触发领域受限的网络搜索。系统支持二元和三元分类,优于现有基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13987 2026-04-14 cs.CL

RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine

RiTeK:一个用于大型语言模型在医学领域文本知识图谱上复杂推理的数据集

Jiatan Huang, Mingchen Li, Zonghai Yao, Dawei Li, Yuxin Zhang, Zhichao Yang, Yongkang Xiao, Feiyun Ouyang, Xiaohan Li, Shuo Han, Hong Yu

机构 * University of Connecticut(康涅狄格大学) University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) School of Computing, and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院) UMass Chan Medical School(马萨诸塞大学陈医学院) University of Minnesota(明尼苏达大学) Rollins School of Public Health, Emory University(埃默里大学罗林斯公共卫生学院) Optum AI University of Massachusetts, Lowell(马萨诸塞大学洛厄尔分校)

AI总结 本文提出RiTeK数据集,用于评估大型语言模型在医学领域文本知识图谱上的复杂推理能力,通过合成高质量查询和专家评估,揭示现有检索方法的不足。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09494 2026-04-13 cs.CL cs.AI cs.IR cs.LG

RecaLLM: Addressing the Lost-in-Thought Phenomenon with Explicit In-Context Retrieval

RecaLLM:通过显式上下文检索解决‘思维迷失’现象

Kyle Whitecross, Negin Rahimi

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 RecaLLM通过交替推理与显式上下文检索解决推理过程中因长上下文导致的检索性能下降问题,显著提升了RULER和HELMET基准测试表现。

Comments Code, data, and models available at https://github.com/kswhitecross/RecaLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07659 2026-04-10 cs.CL

Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction

高效且有效的内部内存检索用于基于大语言模型的医疗预测

Mingchen Li, Jiatan Huang, Zonghai Yao, Hong yu

机构 * University of Connecticut(康涅狄格大学) University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) University of Massachusetts, Lowell(马萨诸塞大学洛厄尔分校) UMass Chan Medical School(马萨诸塞大学陈医学院)

AI总结 本文提出K2K框架,通过内部键基知识访问提升LLM在医疗预测中的效率与准确性,实验显示在四个基准数据集上达到最佳性能。

Comments ACL 2026 (Findings), reviewer score: 3.5,3.5,4

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07350 2026-04-09 cs.CV cs.GR cs.LG

Fast Spatial Memory with Elastic Test-Time Training

快速空间记忆与弹性测试时间训练

Ziqiao Ma, Xueyang Yu, Haoyu Zhen, Yuncong Yang, Joyce Chai, Chuang Gan

机构 * MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) University of Michigan(密歇根大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出弹性测试时间训练方法,用于改进长上下文3D重建中的灾难性遗忘问题,通过引入弹性权重巩固,提升模型在处理任意长序列时的稳定性和适应性。

Comments Project Page: https://fast-spatial-memory.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05227 2026-04-08 cs.CV

Active Measurement of Two-Point Correlations

两点相关函数的主动测量

Max Hamilton, Daniel Sheldon, Subhransu Maji

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出一种人机协同框架,通过预训练分类器指导采样,高效估计目标源的两点相关函数,降低方差并减少标注工作量。

Comments AIStats 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00408 2026-04-07 cs.CV cs.AI

Low-Bitrate Video Compression through Semantic-Conditioned Diffusion

基于语义的低比特率视频压缩

Lingdong Wang, Guan-Ming Su, Divya Kothandaraman, Tsung-Wei Huang, Mohammad Hajiesmaili, Ramesh K. Sitaraman

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Dolby Laboratories(杜比实验室)

AI总结 本文提出DiSCo框架,通过语义条件扩散模型在低比特率下实现高质量视频压缩,采用文本描述、空间时间退化视频和可选草图等模态,提升感知质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03387 2026-04-07 cs.AI

Hume's Representational Conditions for Causal Judgment: What Bayesian Formalization Abstracted Away

休谟因果判断的表征条件:贝叶斯形式化抽象了什么

Yiling Wu

机构 * BridgeM, Inc.(BridgeM公司) Department of Philosophy, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校哲学系)

AI总结 本文探讨休谟因果判断的三个表征条件,并分析其在从休谟到贝叶斯主义的正式化过程中如何被保留和抽象,指出大语言模型展示了统计更新而不满足这些条件的现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02892 2026-04-07 cs.LG stat.ME

Improving Generative Methods for Causal Evaluation via Simulation-Based Inference

通过基于模拟的推断改进因果评估的生成方法

Pracheta Amaranath, Vinitra Muralikrishnan, Amit Sharma, David Jensen

机构 * College of Information and Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校信息与计算机科学学院) Microsoft Research(微软研究院)

AI总结 本文提出SBICE框架,通过推断生成方法及其参数的后验分布,提高因果估计器评估的可靠性。

Comments 13 pages main text, 68 pages total

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00855 2026-04-07 cs.AI cs.LG cs.MA

A Multi-Agent Reinforcement Learning Framework for Public Health Decision Analysis

面向公共卫生决策分析的多智能体强化学习框架

Dinesh Sharma, Ankit Shah, Chaitra Gopalappa

机构 * University of South Florida(南佛罗里达大学) Indiana University(印第安纳大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出多智能体强化学习框架,用于优化公共卫生干预策略,通过考虑跨行政区的流行病学互动,提高资源分配效率,实验表明其在减少新感染方面优于传统单智能体方法。

Comments Updated to the accepted version published in Healthcare Analytics (November 2025)

Journal ref Healthcare Analytics, 8 (2025) 100436

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03141 2026-04-06 cs.CL

Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation

超越精确度:面向长文本LLM生成的事实性评估中的重要性感知召回

Nazanin Jafari, James Allan, Mohit Iyyer

机构 * UMass Amherst(马萨诸塞大学阿默斯特分校) University of Maryland(马里兰大学)

AI总结 本文提出一种综合衡量精确度与召回的框架,通过外部知识源构建参考事实,并引入基于相关性和显著性的权重方案,揭示当前LLM在长文本生成中事实性不完整的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02686 2026-04-06 cs.LG cs.AI

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models

超越语义操控:奖励模型中的标记空间攻击

Yuheng Zhang, Mingyue Huo, Minghao Zhu, Mengxue Zhang, Nan Jiang

机构 * UIUC(伊利诺伊大学厄巴纳-香槟分校) Independent Researcher(独立研究员) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出Token Mapping Perturbation Attack(TOMPA)框架,通过在标记空间中进行对抗优化,发现非语言标记模式以提升奖励模型性能,揭示了当前RLHF流程的关键漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02382 2026-04-06 cs.SE cs.AI

Ambig-IaC: Multi-level Disambiguation for Interactive Cloud Infrastructure-as-Code Synthesis

Ambig-IaC:交互式云基础设施即代码合成的多级消歧

Zhenning Yang, Kaden Gruizenga, Tongyuan Miao, Patrick Tser Jern Kon, Hui Guan, Ang Chen

机构 * University of Michigan(密歇根大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出Ambig-IaC框架,通过多级消歧方法提升IaC配置生成的准确性,针对IaC配置的不可逆性,设计了基于结构分歧的评估框架,在结构和属性评估上分别提升18.4%和25.4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00290 2026-04-06 cs.AI cs.MA

ClinicalReTrial: Clinical Trial Redesign with Self-Evolving Agents

临床试验重设计:具有自进化代理的临床试验重设计

Sixue Xing, Kerui Wu, Xuanye Xia, Meng Jiang, Jintai Chen, Tianfan Fu

机构 * University of Notre Dame(圣母大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Georgia Institute of Technology(佐治亚理工学院) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanjing University(南京大学)

AI总结 本文提出ClinicalReTrial系统,通过自进化代理对临床试验协议进行迭代重设计,提升成功率并降低成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21745 2026-04-03 cs.AI cs.CL

The Presupposition Problem in Representation Genesis

表征生成中的预设问题

Yiling Wu

机构 * BridgeM, Inc.(BridgeM公司) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文探讨了大语言模型在未经历表征生成过渡时的认知能力问题,指出现有哲学框架存在预设结构导致解释延迟,并提出表征回归概念。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21736 2026-04-03 cs.AI cs.CL

The Reasoning Error About Reasoning: Why Different Types of Reasoning Require Different Representational Structures

推理中的推理错误:为何不同类型的推理需要不同的表示结构

Yiling Wu

机构 * BridgeM, Inc.(BridgeM公司) Department of Philosophy, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校哲学系)

AI总结 本文探讨不同推理类型对表示系统结构需求的差异,提出四个结构性质,并指出结构保证对推理类型的重要性,支持了结构重组的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00207 2026-04-02 cs.LG

Lead Zirconate Titanate Reservoir Computing for Classification of Written and Spoken Digits

铅锆钛酸铅共振计算用于手写和语音数字分类

Thomas Buckley, Leslie Schumm, Manor Askenazi, Edward Rietman

机构 * Manning College of Information and Computer Science, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校曼宁信息与计算机科学学院) Independent Researcher, Chicago, IL(独立研究员,芝加哥,伊利诺伊州) Biomedical Hosting, Arlington, MA(Biomedical Hosting,阿灵顿,马萨诸塞州)

AI总结 本文利用物理共振计算处理手写和语音数字数据,展示PZT在MNIST上达到89%准确率,优于线性方法,但AudioMNIST表现与基线相当,表明共振计算在中等难度任务中更有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04485 2026-04-02 cs.CV

Not All Birds Look The Same: Identity-Preserving Generation For Birds

并非所有鸟都看起来一样:用于鸟类的身份保持生成

Aaron Sun, Oindrila Saha, Subhransu Maji

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校)

AI总结 本文提出NABirds Look-Alikes数据集,用于评估鸟类身份保持生成,显示基于物种、年龄和性别的训练提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29852 2026-04-01 cs.GR cs.AI cs.CV

VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing

VectorGym:一个多任务基准用于SVG代码生成、草图和编辑

Juan Rodriguez, Haotian Zhang, Abhay Puri, Tianyang Zhang, Rishav Pramanik, Meng Lin, Xiaoqing Xie, Marco Terral, Darsh Kaushik, Aly Shariff, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vazquez, Christopher Pal, Marco Pedersoli

机构 * ServiceNow Research(ServiceNow 研究院) Mila, Quebec AI Institute(Mila 魁北克人工智能研究所) Columbia University(哥伦比亚大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Stony Brook University(石溪大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Minzu University of China(中央民族大学) University of Waterloo(滑铁卢大学) Canada CIFAR AI Chair(加拿大 CIFAR 人工智能讲席) Computer Vision Center(计算机视觉中心)

AI总结 VectorGym提出一个涵盖SVG生成、复杂编辑和视觉理解的多任务基准,通过专家标注解决现有基准的不足,采用多任务强化学习方法,训练出性能领先的Qwen3-VL模型,公开可用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27033 2026-03-31 cs.CV

RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs

RealBirdID:在MLLM时代鸟类物种识别的基准测试

Logan Lawrence, Mustafa Chasmai, Rangel Daroya, Wuao Liu, Seoyun Jeong, Aaron Sun, Max Hamilton, Fabien Delattre, Oindrila Saha, Subhransu Maji, Grant Van Horn

机构 * Computer Vision Lab, UMass Amherst(马萨诸塞大学阿默斯特分校计算机视觉实验室)

AI总结 研究针对野外细粒度鸟类识别难题,提出RealBirdID基准测试,要求系统在无法识别时提供明确依据。发现开源和专有模型在可回答案例上准确率低,且大模型在无法回答时难以给出正确理由。

Comments Accepted to CVPR26. 23 pages, 23 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏