arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-03-10 至 2026-03-10 共收录 443 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 96 篇

2603.07401 2026-03-10 cs.CV 67%

VIVECaption: A Split Approach to Caption Quality Improvement

VIVECaption:一种改进标题质量的分步方法

Varun Ananth, Baqiao Liu, Haoran Cai

机构 * Adobe Inc.(Adobe公司)

专题命中 评测与基准 :language model(abstract);SFT(abstract)

AI总结 VIVECaption通过双面方法改进标题质量,利用分层抽样和模型对齐策略提升图像-标题对齐效果,提供高质量训练数据解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05587 2026-03-10 cs.AI cs.CL cs.DB cs.LG 67%

MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark

MMTU: 一个大规模多任务表格理解和推理基准

Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Lingjiao Chen, Dongmei Zhang, Surajit Chaudhuri, H. V. Jagadish

机构 * University of Michigan(密歇根大学) Microsoft Corporation(微软公司)

专题命中 评测与基准 :foundation model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MMTU是一个大规模多任务表格理解和推理基准,旨在评估模型在专家级别处理真实表格的能力,揭示了当前模型在表格理解、推理和编码方面的挑战。

Comments Full version of a paper accepted at NeurIPS 2025; Code and data available at https://github.com/MMTU-Benchmark/MMTU and https://huggingface.co/datasets/MMTU-benchmark/MMTU

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06700 2026-03-10 cs.CV 67%

SIQA: Toward Reliable Scientific Image Quality Assessment

SIQA:迈向可靠的科学图像质量评估

Wenzhe Li, Liang Chen, Junying Wang, Yijing Guo, Ye Shen, Farong Wen, Chunyi Li, Zicheng Zhang, Guangtao Zhai

机构 * TongJi University(同济大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 SIQA提出了一种多维框架,通过SIQA-U和SIQA-S评估科学图像的科学正确性和感知清晰度,揭示模型在评分一致性与科学理解上的差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16160 2026-03-10 cs.CV 67%

Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning

Video2Layout: 重建用于空间推理的度量基础认知图

Yibin Huang, Wang Xu, Wanyue Zhang, Helu Zhi, Jingjing Huang, Yangbin Xu, Yangang Sun, Conghui Zhu, Tiejun Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所)

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 Video2Layout通过重建度量基础的空间布局,提升多模态大语言模型在空间推理中的性能,其方法结合监督微调与强化微调,实现更精确的空间认知图构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13172 2026-03-10 cs.CV 67%

WHU-STree: A Multi-modal Benchmark Dataset for Street Tree Inventory

WHU-STree: 一个用于街道树盘点的多模态基准数据集

Ruifei Ding, Zhe Chen, Wen Fan, Chen Long, Huijuan Xiao, Yelu Zeng, Zhen Dong, Bisheng Yang

专题命中 评测与基准 :large language model(abstract);language model(abstract)

AI总结 WHU-STree是一个多模态的街道树盘点数据集,包含丰富的标注和多任务支持,用于提升城市街道树管理的自动化和智能化水平。

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing, 2026, 233: 519-542

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07394 2026-03-10 cs.CV cs.AI cs.CL 62%

AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions

AQuA:迈向具有模糊性视觉问题的策略性响应生成

Jihyoung Jang, Hyounghun Kim

机构 * Graduate School of Artificial Intelligence, POSTECH(人工智能研究生院,POSTECH) Department of Computer Science and Engineering, POSTECH(计算机科学与工程系,POSTECH)

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.AI

AI总结 AQuA通过细粒度数据集和策略微调,提升VLMs在模糊视觉问答中的策略性响应生成能力。

Comments ICLR 2026 (28 pages); Project website: https://aqua-iclr2026.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07365 2026-03-10 cs.SD cs.AI cs.CL cs.MM eess.AS 62%

Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning

多领域音频问答基准:面向声音内容推理

Chao-Han Huck Yang, Sreyan Ghosh, Qing Wang, Jaeyeon Kim, Hengyi Hong, Sonal Kumar, Guirui Zhong, Zhifeng Kong, S Sakshi, Vaibhavi Lokegaonkar, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha, Gunhee Kim, Jun Du, Rafael Valle, Bryan Catanzaro

机构 * NVIDIA University of Maryland, College Park(马里兰大学 College Park 分校) University of Science and Technology of China(中国科学技术大学) Seoul National University(首尔国立大学) Adobe

专题命中 评测与基准 :language model(abstract);分类 cs.CL、cs.AI

AI总结 DCASE 2025挑战赛提出多领域音频问答基准,通过生物声学、时间声音景观和复杂问答子集测试音频-语言模型在多样声音场景中的交互式问答能力,推动音频理解和推理能力发展。

Comments Dataset: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official DCASE Task-5 challenge: dcase.community/challenge2025/task-audio-question-answering. Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06976 2026-03-10 cs.CL cs.AI 62%

A Systematic Investigation of Document Chunking Strategies and Embedding Sensitivity

文档分块策略与嵌入敏感性系统性研究

Muhammad Arslan Shaukat, Muntasir Adnan, Carlos C. N. Kuhn

机构 * Open Source Institute(开源研究所) Faculty of Science and Technology(科学与技术学院) University of Canberra(堪培拉大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文系统研究了文档分块策略与嵌入敏感性,发现内容感知分块显著提升检索效果,段落分组分块在多个领域表现最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06677 2026-03-10 cs.CV cs.AI cs.LG 62%

Chart Deep Research in LVLMs via Parallel Relative Policy Optimization

通过并行相对策略优化进行LVLMs中的图表深度研究

Jiajin Tang, Gaoyang, Wenjie Wang, Sibei Yang, Xing Chen

机构 * ByteDance(字节跳动) ShanghaiTech University(上海科技大学) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 评测与基准 :post-training(abstract);分类 cs.AI、cs.LG

AI总结 本文提出PRPO和MCDR-Bench,通过并行优化和可控错误注入提升图表深度研究的训练与评估能力。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08450 2026-03-10 cs.CL 57%

A Dataset for Probing Translationese Preferences in English-to-Swedish Translation

用于探测英语到瑞典语翻译中翻译倾向的语料库

Jenny Kunz, Anja Jarochenko, Marcel Bollmann

专题命中 评测与基准 :language model(abstract);分类 cs.CL

AI总结 本文提出首个英语到瑞典语翻译倾向语料库,用于评估语言模型对翻译倾向与习语的偏好,发现模型倾向于翻译倾向表达,即使无上下文也如此。

Comments To appear at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00086 2026-03-10 q-fin.ST cs.AI cs.CE 57%

Impact of LLMs news Sentiment Analysis on Stock Price Movement Prediction

大型语言模型新闻情感分析对股价变动预测的影响

Walid Siala, Ahmed Khanfir, Mike Papadakis

机构 * SnT, University of Luxembourg(卢森堡大学SnT学院) RIADI, ENSI, University of Manouba(突尼斯曼努巴大学RIADI学院)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 本文研究了基于大型语言模型的新闻情感分析对股价预测的影响,发现DeBERTa模型在预测准确度上表现最佳,集成模型进一步提升了预测性能。

Journal ref ICLR 2026 Workshop on Advances in Financial AI (AFA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07405 2026-03-10 cs.CL cs.CY 57%

SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations

SPOT:一个标注的法语语料库和基准,用于检测在线对话中的关键干预

Manon Berriche, Célia Nouri, Chloée Clavel, Jean-Philippe Cointet

专题命中 评测与基准 :prompting(abstract);分类 cs.CL

AI总结 SPOT通过标注法语语料库和基准,检测在线对话中的关键干预,验证了监督学习在非英语社交媒体任务中的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07899 2026-03-10 cs.LG stat.ML 57%

Bayesian Transformer for Probabilistic Load Forecasting in Smart Grids

用于智能电网的概率负荷预测的贝叶斯变换器

Sajib Debnath, Md. Uzzal Mia

机构 * O&M Analytics, AES Clean Energ(O&M分析部,AES清洁能源) The AES Corporation(AES公司) Department of Information and Communication Engineering(信息与通信工程系) Pabna University of Science and Technology(帕纳大学科学与技术)

专题命中 评测与基准 :post-training(abstract);分类 cs.LG

AI总结 本研究提出了一种贝叶斯变换器框架,通过整合三种不确定性机制提升智能电网中概率负荷预测的准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07612 2026-03-10 cs.CL 57%

KohakuRAG: A simple RAG framework with hierarchical document indexing

KohakuRAG:一种具有层次文档索引的简单RAG框架

Shih-Ying Yeh, Yueh-Feng Ku, Ko-Wei Huang, Buu-Khang Tu

机构 * National Tsing Hua University Comfy Org Research Kohaku-Lab

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 KohakuRAG通过层次文档索引和集合推理技术,提升了检索增强生成系统在高精度引用任务中的性能,取得挑战赛第一名。

Comments 38pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01396 2026-03-10 cs.AI cs.CE q-bio.QM 57%

HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts

HarmonyCell: 在语义和分布偏移下自动化单细胞扰动建模

Wenxuan Huang, Mingyu Tsoi, Yanhao Huang, Xinjie Mao, Xue Xia, Hao Wu, Jiaqi Wei, Yuejin Yang, Lang Yu, Cheng Tan, Xiang Zhang, Zhangyang Gao, Siqi Sun

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology(香港科学与技术大学) Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) Shanghai Innovation Institute(上海创新研究院) University of British Columbia(不列颠哥伦比亚大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.AI

AI总结 HarmonyCell通过语义统一和自适应搜索机制,在语义和分布偏移下实现单细胞扰动建模的自动化,有效执行率达95%并超越专家基线。

Comments 18 pages total (8 pages main text + appendix), 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07286 2026-03-10 cs.CL 57%

Taiwan Safety Benchmark and Breeze Guard: Toward Trustworthy AI for Taiwanese Mandarin

台湾安全基准与Breeze Guard:迈向可信的台湾台语AI

Po-Chun Hsu, Meng-Hsi Chen, Tsu Ling Chao, Chia Tien Han, Da-shan Shiu

机构 * MediaTek Research(联发科研究实验室) National Taiwan University(国立台湾大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 本文提出TS-Bench基准和Breeze Guard模型,通过文化基础预训练提升台湾台语安全检测性能,显著优于现有模型。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07169 2026-03-10 cs.LG stat.ML 57%

Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts

使LLMs像专家一样优化多场景CUDA内核

Yuxuan Han, Meng-Hao Guo, Zhengning Liu, Wenguang Chen, Shi-Min Hu

机构 * Tsinghua University(清华大学)

专题命中 评测与基准 :LLM(abstract);分类 cs.LG

AI总结 本文提出CUDAMaster系统,通过多代理和硬件感知方法,实现对多场景CUDA内核的高效优化,显著提升性能并超越现有闭源库。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06987 2026-03-10 cs.RO cs.AI 57%

Foundational World Models Accurately Detect Bimanual Manipulator Failures

基础世界模型准确检测双臂机械臂故障

Isaac R. Ward, Michelle Ho, Houjun Liu, Aaron Feldman, Joseph Vincent, Liam Kruse, Sean Cheong, Duncan Eddy, Mykel J. Kochenderfer, Mac Schwager

机构 * Stanford University(斯坦福大学) Watney Robotics(Watney机器人公司)

专题命中 评测与基准 :foundation model(abstract);分类 cs.AI

AI总结 本文提出基于视觉基础模型的双臂机械臂故障检测方法,通过压缩潜在空间中的世界模型提升检测精度,相比传统方法在参数效率和故障检测率上均表现更优。

Comments 8 pages, 5 figures, accepted at the 2026 IEEE International Conference on Robotics and Automation

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06942 2026-03-10 cs.CL 57%

Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks

深度研究,浅层评估:一个针对长形式问答基准的元评估案例研究

Jena D. Hwang, Varsha Kishore, Amanpreet Singh, Dany Haddad, Aakanksha Naik, Malachi Hamada, Jonathan Bragg, Mike D'Arcy, Daniel S. Weld, Lucy Lu Wang, Doug Downey, Sergey Feldman

机构 * Allen Institute for AI(Allen人工智能研究所)

专题命中 评测与基准 :LLM(abstract);分类 cs.CL

AI总结 本文研究了长形式问答基准的元评估方法,指出成对偏好排名适用于系统级评估,而指标级注释和专家评估对可靠评估至关重要,同时揭示了主观性挑战。

Comments 11 pages (including Limitations), 10 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06693 2026-03-10 cs.CV cs.LG 57%

Soft Equivariance Regularization for Invariant Self-Supervised Learning

用于不变自监督学习的软等价性正则化

Joohyung Lee, Changhun Kim, Hyunsu Kim, Kwanhyung Lee, Juho Lee

机构 * AITRICS Voronoi Inc. Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 评测与基准 :pretraining(abstract);分类 cs.LG

AI总结 本文提出软等价性正则化,通过解耦不变性和等价性强制的层,提升自监督学习在ImageNet-1k和ImageNet-C/P等任务中的性能。

Comments 14th International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06680 2026-03-10 cs.CV cs.AI 57%

VB: Visibility Benchmark for Visibility and Perspective Reasoning in Images

VB:用于图像中可见性与视角推理的可见性基准

Neil Tripathi

机构 * New York University(纽约大学)

专题命中 评测与基准 :language model(abstract);分类 cs.AI

AI总结 VB基准测试视觉-语言模型在图像可见性推理中的能力,通过可见性声明和置信度评分评估模型性能。

Comments 18 pages, 1 figure, 3 tables. Code and data: https://github.com/neilt93/Paper-with-Davis

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06583 2026-03-10 cs.HC cs.CY cs.LG 57%

XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being

XInsight:集成化阶段一致的心理咨询支持代理用于数字福祉

Fei Wang, Jiangnan Yang, Junjie Chen, Yuxin Liu, Kun Li, Yanyan Wei, Dan Guo, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) Anhui University(安徽大学) United Arab Emirates University(阿联酋大学) Intelligent Interconnected Systems Laboratory of Anhui Province (HFUT)(安徽省智能互联系统实验室(HFUT))

专题命中 评测与基准 :LLM(abstract);分类 cs.LG

AI总结 XInsight是一种集成化多代理框架,通过阶段一致的工作流程和统一循环,提升网络应用中数字福祉的心理支持效果。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05143 2026-03-10 cs.CV cs.CL 57%

A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering

一种用于可解释作物疾病视觉问答的双阶段多任务视觉-语言框架

Md. Zahid Hossain, Most. Sharmin Sultana Samu, Md. Rakibul Islam, Md. Siam Ansary

专题命中 评测与基准 :pretraining(abstract);分类 cs.CL

AI总结 本文提出了一种轻量且可解释的双阶段多任务视觉-语言框架,用于作物疾病视觉问答,实现了高准确率和良好的可解释性。

Comments Preprint, manuscript is under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08347 2026-03-10 cs.CV 50%

Local-Global Prompt Learning via Sparse Optimal Transport

通过稀疏最优传输实现局部-全局提示学习

Deniz Kizaroğlu, Ülku Tuncer Küçüktas, Emre Çakmakyurdu, Alptekin Temizel

专题命中 评测与基准 :language model(abstract)

AI总结 SOT-GLP通过稀疏最优传输实现局部-全局提示学习,提升少样本分类和分布外检测性能。

Comments 9 pages, 3 figures, 4 tables. Code available at GitHub

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08260 2026-03-10 cs.RO 50%

Seed2Scale: A Self-Evolving Data Engine for Embodied AI via Small to Large Model Synergy and Multimodal Evaluation

Seed2Scale: 一种通过小到大模型协同和多模态评估的自我进化数据引擎用于具身AI

Cong Tai, Zhaoyu Zheng, Haixu Long, Hansheng Wu, Zhengbin Long, Haodong Xiang, Rong Shi, Zhuo Cui, Shizhuang Zhang, Gang Qiu, He Wang, Ruifeng Li, Biao Liu, Zhenzhe Sun, Tao Shen

机构 * ZTE Corporation(中兴通讯公司)

专题命中 评测与基准 :language model(abstract)

AI总结 Seed2Scale通过小到大模型协同和多模态评估,实现具身AI的自我进化数据引擎,显著提升性能并提供可扩展的开发路径

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07652 2026-03-10 cs.CV 50%

GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence

GLASS: 图与视觉-语言辅助的语义形状对应

Qinfeng Xiao, Guofeng Mei, Qilong Liu, Chenyuan Yi, Fabio Poiesi, Jian Zhang, Bo Yang, Yick Kit-lun

机构 * Hong Kong Polytechnic University, HK SAR(香港理工大学) Fondazione Bruno Kessler, Italy(布鲁诺·凯斯勒基金会) Laboratory for Artificial Intelligence in Design, HK SAR(设计中的人工智能实验室) University of Technology Sydney, Australia(悉尼科技大学)

专题命中 评测与基准 :foundation model(abstract)

AI总结 GLASS通过整合几何频谱分析与视觉-语言模型的语义先验,实现了在非等距变形和跨类设置中鲁棒的3D形状语义对应学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07552 2026-03-10 cs.CV cs.RO 50%

ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction

ReconDrive: 用于自动驾驶场景重建的快速前馈4D高斯点扩散

Haibao Yu, Kuntao Xiao, Jiahang Wang, Ruiyang Hao, Yuxin Huang, Guoran Hu, Haifang Qin, Bowen Jing, Yuntian Bo, Ping Luo

机构 * Tuojing Intelligence The University of Hong Kong(香港大学) King's College London(伦敦国王学院) The University of Sydney(悉尼大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 评测与基准 :foundation model(abstract)

AI总结 ReconDrive通过改进的4D高斯点扩散方法,在自动驾驶场景重建中实现高效高保真生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07446 2026-03-10 cs.HC 50%

GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader Users

GeoVisA11y: 一种基于AI的地理可视化问答系统,用于屏幕阅读用户

Chu Li, Rock Yuren Pang, Arnavi Chheda-Kothary, Ather Sharif, Henok Assalif, Jeffrey Heer, Jon E. Froehlich

专题命中 评测与基准 :LLM(abstract)

AI总结 GeoVisA11y通过自然语言交互为屏幕阅读用户提供可访问的地理可视化系统,并揭示了不同用户群体的交互模式差异。

Comments This manuscript has been accepted at CHI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06982 2026-03-10 cs.CV cs.IR 50%

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

优化多模态模型用于基于图像的形状检索:预对齐和硬对比学习的作用

Paul Julius Kühn, Cedric Spengler, Michael Weinmann, Arjan Kuijper, Saptarshi Neil Sinha

机构 * Fraunhofer IGD(弗劳恩霍夫图像研究中心) Delft University of Technology(代尔夫特理工大学)

专题命中 评测与基准 :pretraining(abstract)

AI总结 本文提出通过预对齐和硬对比学习优化多模态模型,提升基于图像的形状检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02419 2026-03-10 cs.CV 50%

DINOv3 Visual Representations for Blueberry Perception Toward Robotic Harvesting

DINOv3视觉表示在蓝莓感知中的应用:面向机器人采摘

Rui-Feng Wang, Daniel Petti, Yue Chen, Changying Li

机构 * Bio-Sensing, Automation, and Intelligence Laboratory, Department of Agricultural and Biological Engineering, Institute of Food and Agricultural Sciences, University of Florida(生物感知、自动化与智能实验室,农业与生物工程系,食品与农业科学研究院,佛罗里达大学) Department of Biomedical Engineering, Georgia Institute of Technology(生物医学工程系,佐治亚理工学院)

专题命中 评测与基准 :foundation model(abstract)

AI总结 DINOv3作为语义主干网络,在蓝莓机器人采摘中用于果实和损伤分割及检测,但检测受限于目标尺度变化和空间聚合建模的挑战。

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏