arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-06-08 至 2026-06-08 共收录 275 信号源:cs.CL, cs.AI, cs.LG

1. 其他LLM 19 篇

2606.06853 2026-06-08 cs.CV cs.AI 新提交 79%

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

MotionEnhancer: 利用视频扩散模型增强运动感知的视觉-语言模型

Yifan Xu, Chao Zhang, Ruifei Ma, Fei Gao, Zhifei Yang, Jiaxing Qi, Zhipeng Chen

机构 * School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院) Beijing Digital Native Digital City Research Center(北京数字原生数字城研究中心) School of Computer Science, Peking University(北京大学计算机学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)

专题命中 其他LLM :language model(title,abstract);分类 cs.AI

AI总结 提出MotionEnhancer,通过从视频扩散模型中提取运动先验并利用注意力对齐增强视觉-语言模型的运动理解能力,无需额外参数或架构修改,在运动级视频理解基准上取得一致提升。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22278 2026-06-08 cs.CV cs.LG 版本更新 79%

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

视觉-语言模型中空间变量绑定的双重机制

Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Northeastern University(东北大学) Sony Playstation(索尼PlayStation)

专题命中 其他LLM :language model(title,abstract);分类 cs.LG

AI总结 本文揭示视觉-语言模型通过语言骨干中的内容无关空间关系编码和视觉编码器中的全局布局表示两种机制实现空间变量绑定,其中视觉编码器起主导作用。

Comments 37 pages, 53 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26714 2026-06-08 cs.LG cs.AI 版本更新 73%

On the importance of multiple training seeds for evaluating machine unlearning

关于多个训练种子在评估机器遗忘中的重要性

Jamie Lanyon, Axel Finke, Petros Andreou, Georgina Cosma

机构 * Department of Computer Science(计算机科学系) School of Mathematics(数学学院) School of Science(科学学院) Statistics and Physics(统计学与物理学) Loughborough University(洛桑大学) Newcastle University(新castle大学)

专题命中 其他LLM :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文指出评估机器遗忘算法时仅使用单个训练种子可能导致结果不具代表性,并通过图像分类、联邦学习排序和大语言模型实验验证了问题普遍性,最后给出选择训练和遗忘种子数量的指导。

Comments mini paper, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13092 2026-06-08 cs.LG cs.AR 版本更新 57%

Breaking the Tuning Barrier: Zero-Hyperparameters Yield Multi-Corner Analysis Via Learned Priors

打破调优壁垒:通过先验学习实现零超参数的多角分析

Wei W. Xing, Kaiqi Huang, Jiazhan Liu, Hong Qiu, Shan Shen

机构 * School of Mathematical and Physical Science, University of Sheffield(谢菲尔德大学数学与物理科学学院) SZU–UoS Joint Centre for Innovation and Entrepreneurship, College of Mechatronics and Control Engineering, Shenzhen University(深大-乌兹别克斯坦联合创新与创业中心,机电控制工程学院,深圳大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 其他LLM :foundation model(abstract);分类 cs.LG

AI总结 针对电路多角分析中仿真成本高且现有方法需大量调参的问题,提出基于预训练基础模型的上下文学习方法,无需调优即可匹配最先进精度,将验证成本降低10倍以上。

Comments Accepted by DAC2026. Camera-ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07161 2026-06-08 cs.CV 新提交 50%

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance

TraRA: 面向城市监控视频文本识别的轨迹级识别聚合方法

Duc Tri Tran, Trung Thanh Nguyen, Vijay John, Phi Le Nguyen, Yasutomo Kawanishi

机构 * RIKEN(日本理化学研究所) Hanoi University of Science and Technology(河内科学技术大学) Nagoya University(名古屋大学) Lawrence Technological University(劳伦斯技术大学) Ritsumeikan University(立命馆大学)

专题命中 其他LLM :language model(abstract)

AI总结 提出TraRA方法,通过轨迹级文本识别聚合,利用时间与多模态一致性,解决监控视频中运动模糊、遮挡等导致的帧级识别不一致问题,在多个基准上提升跟踪与识别性能。

Comments 22nd IEEE International Conference on Advanced Visual and Signal-Based Systems

详情

展开后加载摘要…

URL PDF HTML 收藏