arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-21 至 2026-04-21 共收录 45 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 45 篇

2506.05405 2026-04-21 cs.CV 92%

A VLM-based Method for Visual Anomaly Detection in Robotic Scientific Laboratories

基于VLM的方法用于机器人科学实验室中的视觉异常检测

Shiwei Lin, Chenxu Wang, Xiaozhen Ding, Yi Wang, Boyuan Du, Lei Song, Chenggang Wang, Huaping Liu

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Physics and Electronic Information, Yantai University(烟台大学物理与电子信息学院) School of Computer and Big Data, Fuzhou University(福州大学计算机与大数据学院) Department of Automation, Shanghai Jiao Tong University(上海交通大学自动化系)

专题命中 视觉推理 :VLM(title,title_cn);vision-language model(abstract);visual reasoning(abstract);分类 cs.CV

AI总结 本文提出基于VLM的视觉推理方法,通过四个逐步信息提示配置实现多级监督,构建了专用视觉基准以评估其在科学流程异常检测中的有效性,实验表明提供更多上下文信息可提升检测准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16918 2026-04-21 cs.CL cs.LG 91%

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning

具有新鲜度意识的优先经验回放用于LLM/VLM强化学习

Weiyu Ma, Yongcheng Zeng, Yan Song, Xinyu Cui, Jian Zhao, Xuhui Liu, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡布尔大学科学与技术大学) Chinese Academy of Sciences, Institute of Automation(中国科学院自动化研究所) AI Centre, Department of Computer Science, University College London(伦敦大学学院人工智能中心) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究所)

专题命中 视觉推理 :VLM(title,title_cn);vision-language model(abstract);分类 cs.LG

AI总结 本文提出Freshness-Aware PER,通过引入指数衰减机制解决LLM/VLM强化学习中优先级过时问题,显著提升样本效率,在多个任务中取得优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05920 2026-04-21 cs.CV 86%

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

通过多轮基于定位的强化学习实现高分辨率视觉推理

Xinyu Huang, Yuhao Dong, Weiwei Tian, Bo Li, Rui Feng, Ziwei Liu

机构 * Fudan University(复旦大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 视觉推理 :grounding(title,abstract);visual reasoning(title);分类 cs.CV

AI总结 本文提出MGPO框架,通过多轮对话机制使LMMs在无需额外标注的情况下提升视觉定位能力,实验表明其在MME-Realworld和V*Bench任务中表现优异。

Comments Accepted by ACL(Findings) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21278 2026-04-21 cs.CV cs.AI cs.CL cs.LG 86%

GeoRC: A Benchmark for Geolocation Reasoning Chains

GeoRC:地理定位推理链基准测试

Mohit Talreja, Joshua Diao, Jim Thannikary James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉推理 :VLM(summary_cn,abstract);vision language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出GeoRC基准,通过Champion级GeoGuessr专家生成的800条真实推理链,评估VLM生成推理链的准确性,发现大模型在生成可审计推理链上仍逊于人类专家,而小型模型表现更差。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23404 2026-04-21 cs.CV cs.CL 85%

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

通过文本表示引导推理来解锁多模态大语言模型中的空间推理

Jiacheng Hua, Yishu Yin, Yuhang Wu, Tai Wang, Yifei Huang, Miao Liu

机构 * College of AI, Tsinghua University, Beijing, China(清华大学人工智能学院,北京,中国) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国) The University of Tokyo, Tokyo, Japan(东京大学,东京,日本)

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出TRACE方法,通过生成文本表示来提升多模态大语言模型的空间推理能力,实验表明其在多个基准测试中表现优异。

Comments Accepted to ACL 2026. 22 pages, 6 figures, 10 tables. Project page: https://trace-reasoning.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17243 2026-04-21 cs.CV 85%

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

RemoteShield:为地球观测启用鲁棒的多模态大语言模型

Rui Min, Liang Yao, Shiyu Miao, Shengxiang Xu, Yuxuan Liu, Chuanyi Zhang, Shimin Di, Fan Liu

机构 * Hohai University(河海大学) Nanjing University(南京大学) Southeast University(东南大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出RemoteShield,一种针对地球观测的鲁棒多模态大语言模型,通过偏好学习提升在现实输入变化下的稳定性与一致性,实验显示其在多模态扰动下表现优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20879 2026-04-21 cs.CV 85%

DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

DriveAgent-R1: 通过主动感知和混合思维推进基于视觉语言模型的自动驾驶

Weicheng Zheng, Xiaofei Mao, Nanfei Ye, Pengxiang Li, Kun Zhan, Xianpeng Lang, Hang Zhao

机构 * Shanghai Qi Zhi Institute(上海启智研究院) LiAuto Tongji University(同济大学) Tsinghua University(清华大学)

专题命中 视觉推理 :VLM(title);vision-language model(abstract);visual reasoning(abstract);grounding(abstract)

AI总结 DriveAgent-R1通过主动感知和混合思维框架,提升自动驾驶中的视觉推理能力,采用三阶段训练策略,在复杂场景中实现高效决策,展现与顶级模型相当的性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02871 2026-04-21 cs.CL cs.AI 85%

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning

位置:多模态大语言模型可以显著推动科学推理

Yibo Yan, Shen Wang, Jiahao Huo, Jingheng Ye, Zhendong Chu, Xuming Hu, Philip S. Yu, Carla Gomes, Bart Selman, Qingsong Wen

机构 * Squirrel AI HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Tsinghua University(清华大学) University of Illinois at Chicago(伊利诺伊大学香槟分校) Cornell University(康奈尔大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.AI

AI总结 本文探讨多模态大语言模型在科学推理中的应用,提出四阶段研究路线,指出当前模型在跨领域推理中的潜力与挑战,为实现通用人工智能提供新视角。

Comments Accepted by The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026, Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03331 2026-04-21 cs.CV cs.AI cs.LG 85%

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

MMErroR:面向视觉-语言模型错误推理的基准测试

Yang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu, Mingxuan Huang, Jingchao Wang, Zhihong Zhu, Boyan Xu, Zhiqi Huang

机构 * Guangdong University of Technology(广东工业大学) Hong Kong Baptist University(香港 Baptist大学) Sun Yat-sen University(中山大学) Peking University(北京大学)

专题命中 视觉推理 :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出MMErroR基准,通过1997个包含单一推理错误的样本,评估视觉-语言模型检测错误推理的能力,发现即使最佳模型也仅能正确分类66.65%的错误。

Comments Accepted by ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22368 2026-04-21 cs.CV cs.AI 84%

When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations

当视觉不是问题:在误导性数据可视化上评估视觉-语言模型

Harsh Nishant Lalai, Raj Sanjay Shah, Hanspeter Pfister, Sashank Varma, Grace Guo

机构 * Birla Institute of Technology and Science, Pilani(巴尔·印度理工学院和科学学院,皮拉尼) Georgia Institute of Technology(佐治亚理工学院) Harvard University(哈佛大学)

专题命中 视觉推理 :vision-language model(title);vision language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 本文评估视觉-语言模型在识别误导性数据可视化中的能力,发现模型在检测可视化设计错误上更可靠,但常误判非误导性可视化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23941 2026-04-21 cs.LG cs.CV 84%

Vision Language Models are Biased

视觉语言模型存在偏见

An Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Vy Tuong Dang, Anh Totti Nguyen, Daeyoung Kim

机构 * KAIST(韩国科学技术院) College of William and Mary(威廉与玛丽学院) University of Alberta(阿尔伯塔大学) Auburn University(阿伯茨维尔大学)

专题命中 视觉推理 :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG

AI总结 研究探讨了视觉语言模型在计数和识别任务中的偏见问题,发现其在不同领域表现不佳,去除背景可显著提升准确率,揭示上下文视觉线索导致偏见。

Comments Code and qualitative examples are available at: vlmsarebiased.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17830 2026-04-21 cs.RO 83%

SYMBOLIZER: Symbolic Model-free Task Planning with VLMs

SYMBOLIZER:基于VLMs的符号模型无关任务规划

Sami Azirar, Zlatan Ajanovic, Hermann Blum

机构 * RWTH Aachen(亚琛理工大学) University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究院)

专题命中 视觉推理 :VLM(summary_cn,abstract);visual language model(abstract)

AI总结 本文提出SYMBOLIZER框架,利用VLMs将图像转化为符号状态,结合启发式搜索实现跨领域任务规划,优于直接VLM规划并可与基于启发式的传统方法媲美。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17233 2026-04-21 cs.CV cs.AI 82%

Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM

通过基于资料的多模态大语言模型增强零样本个性化图像审美评估

Chun Wang, Chenfeng Wei, Chenyang Liu, Weihong Deng

机构 * Mashang Consumer Finance Co., Ltd.(马莎消费金融有限公司) Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉推理 :MLLM(summary_cn,abstract);分类 cs.CV、cs.AI

AI总结 本文提出P-MLLM,通过用户资料引导的多模态大语言模型实现零样本个性化图像审美评估,实验表明其在无历史数据情况下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17054 2026-04-21 cs.CV cs.AI 82%

mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval

mEOL:无需训练的指令引导多模态嵌入器用于矢量图形和图像检索

Kyeong Seon Kim, Baek Seong-Eun, Lee Jung-Mok, Tae-Hyun Oh

机构 * KAIST(韩国科学技术院) POSTECH

专题命中 视觉推理 :MLLM(abstract,abstract_cn);visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出无需训练的指令引导多模态嵌入框架,通过多模态大语言模型将文本、位图和SVG代码映射到对齐的嵌入空间,利用模态特定指令和结构化SVG提示实现嵌入方向控制,构建首个文本到SVG检索基准,展示出优于传统基线的性能。

Comments Round 1 early acceptance to WACV 2026, Project page: https://scene-the-ella.github.io/meol

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21064 2026-04-21 cs.AI cs.CV 81%

OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

OVOD-Agent:一种用于主动视觉推理和自进化检测的马尔可夫-多臂框架

Chujie Wang, Jianyu Lu, Zhiyuan Luo, Xi Chen, Chu He

机构 * Wuhan University(武汉大学)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出OVOD-Agent框架,通过将文本优化扩展为可解释的视觉CoT,结合弱马尔可夫决策过程和多臂模块,实现主动视觉推理和自进化检测,提升稀有类别检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06803 2026-04-21 cs.CL cs.CV 79%

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning

森林在树之前:用于高效视觉推理的潜在叠加

Yubo Wang, Juntian Zhang, Yichen Wu, Yankai Lin, Nils Lukas, Yuhan Liu

机构 * MBZUAI Fudan University(复旦大学) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Harvard University(哈佛大学)

专题命中 视觉推理 :visual reasoning(title);vision-language model(abstract);分类 cs.CV

AI总结 本文提出Laser方法,通过动态窗口对齐学习实现视觉推理的'森林在树之前'认知层次,提升模型在保持全局特征叠加的同时,实现高效且可解释的视觉推理。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16857 2026-04-21 cs.CV cs.RO 79%

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

BOP-ASK:面向视觉-语言模型的对象交互推理

Vineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich, Prashanth Krishnamurthy, Ramesh Karri, Stan Birchfield, Farshad Khorrami, Jonathan Tremblay

机构 * New York University(纽约大学) NVIDIA(英伟达)

专题命中 视觉推理 :vision-language model(title);vision language model(abstract);分类 cs.CV

AI总结 本文提出BOP-ASK数据集,用于训练和评估视觉-语言模型的对象交互推理能力,包含15万张图像和3300万对问题答案,涵盖六个任务,验证了模型在复杂环境中的空间推理能力。

Comments Accepted at CVPR 2026. Code, Datasets & Benchmark available at https://bop-ask.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04334 2026-04-21 cs.CV 79%

GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models

GeoArena:在大视觉-语言模型中评估开放世界地理推理

Pengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon Li

机构 * Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) Department of Computer Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出GeoArena框架,通过人偏好评估动态图像中的地理推理能力,评估17种前沿LVLM,分析模型行为与人类偏好可靠性。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16617 2026-04-21 cs.CV cs.MM cs.SD 79%

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

AVRT:通过单模态教师模型实现音频-视觉推理迁移

Edson Araujo, Saurabhchand Bhati, M. Jehanzeb Mirza, Brian Kingsbury, Samuel Thomas, Rogerio Feris, James R. Glass, Hilde Kuehne

机构 * University of Tübingen, Germany(图宾根大学) MIT, Cambridge MA, USA(麻省理工学院) IBM Research, USA(IBM研究院) MIT-IBM Watson AI Lab, USA(麻省理工-IBM沃森人工智能实验室) Tuebingen AI Center, Germany(图宾根人工智能中心)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV

AI总结 本文提出AVRT框架,通过单模态教师模型生成高质量音频-视觉推理轨迹,并利用LLM合并模型整合轨迹,从而在多模态设置中实现推理迁移,取得音视频和音频任务的最优效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08539 2026-04-21 cs.CV cs.AI cs.CL 79%

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks

OpenVLThinkerV2: 一个多领域视觉任务的通用多模态推理模型

Wenbo Hu, Xin Chen, Yan Gao-Tian, Yihe Deng, Nanyun Peng, Kai-Wei Chang

机构 * University of California, Los Angeles (UCLA)(加州大学洛杉矶分校)

专题命中 视觉推理 :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出G$^2$RPO训练目标,通过非线性分布匹配解决多视觉任务中的奖励拓扑差异和感知与推理平衡问题,构建了鲁棒的OpenVLThinkerV2模型,在18个基准测试中表现优异。

Comments code at: https://github.com/uclanlp/openvlthinker

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04509 2026-04-21 cs.CL 78%

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

ErrorRadar:通过错误检测基准测试多模态大语言模型的复杂数学推理能力

Yibo Yan, Shen Wang, Jiahao Huo, Hang Li, Boyan Li, Jiamin Su, Xiong Gao, Yi-Fan Zhang, Tianlong Xu, Zhendong Chu, Aoxiao Zhong, Kun Wang, Hui Xiong, Philip S. Yu, Xuming Hu, Qingsong Wen

机构 * Squirrel AI HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) MSU(密歇根州立大学) UCAS(中国科学技术大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 视觉推理 :multimodal large language model(title,abstract)

AI总结 ErrorRadar通过错误检测任务评估多模态大语言模型在复杂数学推理中的能力,包含2500个高质量K-12数学问题,实验显示GPT-4o仍比人类评价差10%。

Comments Accepted by The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026, Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17504 2026-04-21 cs.CV cs.AI 73%

RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding

RS-HyRe-R1: 一种混合奖励机制以克服遥感图像理解中的感知惯性

Gaozhi Zhou, Hu He, Peng Shen, Jipeng Zhang, Liujue Zhang, Linrui Xu, Zeyuan Wang, Ziyu Li, Xuezhi Cui, Wang Guo, Haifeng Li

机构 * School of Mechanical and Electrical Engineering, Central South University(中南大学机械与电气工程学院) School of Geosciences and Info-Physics, Central South University(中南大学地球科学与信息物理学院)

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 本文提出RS-HyRe-R1混合奖励机制,旨在克服遥感图像理解中的感知惯性问题,通过引入空间推理激活奖励、感知正确性奖励和视觉-语义路径演化奖励,提升模型的全面视觉证据挖掘能力,实验显示其在多个任务上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17052 2026-04-21 cs.CV 70%

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

OASIS:按需分层事件记忆用于流媒体视频推理

Zhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang, Haonan Lu, Guanbin Li

机构 * Sun Yat-sen University(中山大学) OPPO AI Center(OPPO人工智能中心) Shenzhen Loop Area Institute(深圳河套学院) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

专题命中 视觉推理 :MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 OASIS通过结构化按需检索解决流媒体视频推理中记忆检索的挑战,提升长周期准确性和组合推理能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20409 2026-04-21 cs.CL cs.AI cs.CY 70%

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

认知链式推理(CoCoT):关于社会情境的结构化多模态推理

Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim, Motahhare Eslami, Maarten Sap

机构 * Seoul National University(首尔国立大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学)

专题命中 视觉推理 :VLM(abstract,abstract_cn);分类 cs.AI

AI总结 本文提出CoCoT框架,通过感知、情境推断和规范应用三个阶段提升多模态社会推理能力,在多个任务中取得显著提升。

Comments Under review; 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12320 2026-04-21 cs.CV cs.AI cs.MM 62%

EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports

EgoEsportsQA:一种用于电子竞技感知与推理的视角视频基准

Jianzhe Ma, Zhonghao Cao, Shangkui Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * Renmin University of China(中国人民大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文提出EgoEsportsQA基准,通过六阶段流程收集1745对高质量问答对,评估视频大语言模型在虚拟环境中的感知与推理能力,揭示模型在战术推理和微操作方面的不足。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07562 2026-04-21 cs.CL cs.AI cs.CY cs.LG 62%

Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs

基于推理的无监督文本聚类细化方法

Tunazzina Islam

机构 * Department of Computer Science(计算机科学系)

专题命中 视觉推理 :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出利用大语言模型作为语义评判者,对无监督聚类结果进行推理细化,通过三个阶段提升聚类的连贯性与可解释性,实验证明其在社交媒体数据中优于传统方法。

Comments Accepted to the Findings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16893 2026-04-21 cs.CV cs.LG 62%

EasyVideoR1: Easier RL for Video Understanding

EasyVideoR1: 更容易的视频理解强化学习

Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学网络安全学院)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 本文提出EasyVideoR1框架,通过高效预处理、统一奖励系统和混合训练策略,提升视频理解任务的训练效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16528 2026-04-21 cs.CV cs.AI 62%

Expert-Annotated Embryo Image Dataset with Natural Language Descriptions for Evidence-Based Patient Communication in IVF

具有自然语言描述的专家标注胚胎图像数据集:用于基于证据的患者沟通的体外受精

Nicklas Neu, Thomas Ebner, Jasmin Primus, Bernhard Schenkenfelder, Raphael Zefferer, Mathias Brunbauer, Florian Kromp

机构 * Software Competence Center Hagenberg(海因斯贝格软件能力中心) Kepler Universitätsklinikum, Kinderwunsch Zentrum(凯普勒大学医院,生育中心) Wunschkind Klinik Dr. Brunbauer(布鲁纳auer生育诊所)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一个包含胚胎图像和自然语言描述的专家标注数据集,用于提升体外受精中胚胎选择的可解释性和透明度,支持基于证据的决策和患者沟通。

Comments 7 pages, 3 figures, in submission to Nature Scienfitic Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18484 2026-04-21 cs.CV cs.MM cs.RO 57%

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

XEmbodied:一种具有增强几何和物理线索的大型规模具身环境基础模型

Kangan Qian, ChuChu Xie, Yang Zhong, Jingrui Pang, Siwen Jiao, Sicong Jiang, Zilin Huang, Yunlong Wang, Kun Jiang, Mengmeng Yang, Hao Ye, Guanghao Zhang, Hangjun Ye, Guang Chen, Long Chen, Diange Yang

机构 * Tsinghua University(清华大学) Automotive and Robotics, Xiaomi Corporation(小鹏汽车与机器人部,小鹏科技) National University of Singapore(新加坡国立大学) McGill University(麦吉尔大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV

AI总结 XEmbodied通过整合3D几何表示和物理线索,提升视觉-语言-动作模型在大规模具身环境中的空间推理和语义理解能力,其在18个公开基准测试中表现出色。

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17584 2026-04-21 cs.AI 57%

DIRCR: Dual-Inference Rule-Contrastive Reasoning for Solving RAVENs

DIRCR:双推理规则对比推理用于解决RAVENs

Jiachen Zhang, Chengtai Li, Jianfeng Ren, Linlin Shen, Zheng Lu, Ruibin Bai

机构 * School of Computer Science, University of Nottingham Ningbo China, China(诺丁汉大学宁波校区计算机科学学院) Nottingham Ningbo China Beacons of Excellence Research and Innovation Institute, China(诺丁汉宁波中国卓越研究与创新研究所) School of Artificial Intelligence, Shenzhen University, China(深圳大学人工智能学院)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI

AI总结 本文提出DIRCR模型,通过双推理推理模块和规则对比学习模块,解决RAVENs中全局与局部关系整合不足的问题,提升推理鲁棒性和泛化能力。

Comments Accepted By ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏