arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-28 至 2026-04-28 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 11 篇

2512.10362 2026-04-28 cs.CV cs.AI 81%

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models

视觉漏斗:缓解多模态大语言模型中的上下文盲区

Woojun Jung, Jaehoon Go, Mingyu Jeon, Sunjae Yoon, Junyeong Kim

机构 * Department of AI, Chung-Ang University(Chung-Ang 大学人工智能系)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出视觉漏斗方法,通过上下文锚定和熵缩放投资组合缓解多模态大语言模型中的上下文盲区问题,通过动态调整裁剪尺寸和中心点,提升视觉细节与全局上下文的关联性。

Comments Accepted to CVPR 2026(Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07555 2026-04-28 cs.LG cs.IT cs.NI math.IT 78%

Multimodal Remote Inference

多模态远程推断

Keyuan Zhang, Yin Sun, Bo Ji

机构 * Department of Computer Science, Virginia Tech(弗吉尼亚理工大学计算机科学系) Department of Electrical and Computer Engineering, Auburn University(阿伯伯大学电气与计算机工程系)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文研究多模态远程推断系统,通过优化调度减少机器学习模型的推断误差,提出误差感知切换与传输策略(EAST)及低复杂度策略,实验表明EAST在多模态情况下显著降低推断误差。

Comments Submitted to IEEE/ACM Transactions on Networking; a preliminary version appeared in the Proceedings of the 22nd IEEE International Conference on Mobile Ad-Hoc and Smart Systems (MASS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24082 2026-04-28 cs.CR cs.AI cs.CL 62%

Jailbreaking Frontier Foundation Models Through Intention Deception

通过意图欺骗突破前沿基础模型

Xinhe Wang, Katia Sycara, Yaqi Xie

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种多轮欺骗方法,通过模拟良性意图和模型一致性,引导模型生成有害输出,并揭示了未被注意到的para-jailbreaking漏洞。

Comments Accepted at CVPR 2026 Findings Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24062 2026-04-28 cs.AI 57%

Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer

在泛化之前建立基础:人工智能如何与人类在因果迁移中不同

Liangru Xiang, Yuxi Ma, Zhihao Cao, Yixin Zhu, Song-Chun Zhu

机构 * Department of Automation(清华大学自动化系) Institute for Artificial Intelligence(北京大学人工智能研究院) School of Psychological and Cognitive Sciences(北京大学心理与认知科学系) State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室) Beijing Key Laboratory of Behavior and Mental Health(北京行为与心理健康重点实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究探讨了人工智能与人类在因果迁移中的差异,发现AI模型在缺乏环境基础映射时难以高效迁移,而人类能利用先验结构知识。文本条件中AI表现优异,但视觉信息反而降低性能,揭示AI依赖符号处理而非多模态推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20018 2026-04-28 cs.DB cs.AI 57%

MINT: Multi-Vector Search Index Tuning

MINT: 多向量搜索索引调优

Jiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein, Vivek Narasayya, Surajit Chaudhuri

机构 * University of California, San Diego(加州大学圣地亚哥分校) Microsoft Research(微软研究院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出MINT框架,针对多向量搜索场景,通过算法优化实现低延迟和存储召回约束下的索引调优,相比基线提升2.1至8.3倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23649 2026-04-28 cs.CR 50%

Rényi Pufferfish Privacy with Gaussian-based Priors: From Single Gaussian to Mixture Model

基于高斯先验的Rényi Pufferfish隐私:从单高斯到混合模型

Wenjin Yang, Ni Ding, Zijian Zhang, Zhen Li, Jing Sun, Jincheng An, Yong Liu, Liehuang Zhu

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文研究了基于高斯和高斯混合先验的Rényi Pufferfish隐私机制,推导了高斯扰动后的Rényi散度,提出了放松的闭式充分条件,并通过实验展示了其在隐私-效用权衡上的改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22911 2026-04-28 cs.RO 50%

RecoverFormer: End-to-End Contact-Aware Recovery for Humanoid Robots

RecoverFormer: 末端执行器感知的人形机器人恢复控制

Zihui Liu

机构 * Stanford University(斯坦福大学)

专题命中 其他多模态 :multi-modal(abstract)

AI总结 本文提出RecoverFormer,一种端到端的人形机器人恢复策略,通过学习切换恢复行为(如补偿步态、手-环境接触、质心重塑)实现鲁棒恢复,验证了其在多种扰动条件下的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22866 2026-04-28 cs.CR cs.ET 50%

Risk Models as Mediating Artifacts: A Postphenomenological Analysis of the CIIM Framework in Cybersecurity Practice

风险模型作为中介化 artifacts:对网络安全实践中 CIIM 框架的后现象学分析

Rommel Salas-Guerra

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文通过后现象学理论分析网络安全风险管理中的 CIIM 框架,探讨其作为中介化工具对安全实践者认知和行动的影响,提出'崩溃现象学'概念。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22812 2026-04-28 cs.CY cs.LG stat.AP 50%

Cross-Course Generalizability of SRL-Aligned Predictive Models Using Digital Learning Traces

基于数字学习轨迹的SRL对齐预测模型的跨课程泛化能力

Jakob Schwerter, Loreen Sabel, Judith Bose, Matthew L. Bernacki, Di Xu, Marko Schmellenkamp, Thomas Zeume, Philipp Doebler

机构 * Hector Research Institute of Education Sciences and Psychology, University of Tübingen(教育科学与心理学赫克托研究 institute,图宾根大学) TU Dortmund University(多特蒙德技术大学) University of North-Carolina, Chapel Hill(北卡罗来纳大学教堂山分校) University of California, Irvine(加州大学尔湾分校) Ruhr University Bochum(波鸿鲁尔大学)

专题命中 其他多模态 :multimodal(abstract)

AI总结 研究通过分析多模态数字轨迹数据,评估SRL对齐模型在不同课程和机构中的预测性能,发现Elastic Net在跨情境泛化中表现更稳健,但模型泛化需谨慎考虑不同情境下的风险率差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15483 2026-04-28 cs.LG cs.RO 50%

$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

$π_{0.7}$: 一种可定向的通用机器人基础模型与涌现能力

Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, Vedant Choudhary, Foster Collins, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, Maitrayee Dhaka, Jared DiCarlo, Danny Driess, Michael Equi, Adnan Esmail, Yunhao Fang, Chelsea Finn, Catherine Glossop, Thomas Godden, Ivan Goryachev, Lachlan Groom, Haroun Habeeb, Hunter Hancock, Karol Hausman, Gashon Hussein, Victor Hwang, Brian Ichter, Connor Jacobsen, Szymon Jakubczak, Rowan Jen, Tim Jones, Gregg Kammerer, Ben Katz, Liyiming Ke, Mairbek Khadikov, Chandra Kuchi, Marinda Lamb, Devin LeBlanc, Brendon LeCount, Sergey Levine, Xinyu Li, Adrian Li-Bell, Vladislav Lialin, Zhonglin Liang, Wallace Lim, Yao Lu, Enyu Luo, Vishnu Mano, Nandan Marwaha, Aikys Mongush, Liam Murphy, Suraj Nair, Tyler Patterson, Karl Pertsch, Allen Z. Ren, Gavin Schelske, Charvi Sharma, Baifeng Shi, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, Will Stoeckle, Jiaming Tang, Jimmy Tanner, Shalom Tekeste, Marcel Torne, Kyle Vedder, Quan Vuong, Anna Walling, Haohuan Wang, Jason Wang, XuDong Wang, Chris Whalen, Samuel Whitmore, Blake Williams, Charles Xu, Sukwon Yoo, Lili Yu, Wuming Zhang, Zhuoyang Zhang, Ury Zhilinsky

机构 * Physical Intelligence

专题命中 其他多模态 :multimodal(abstract)

AI总结 $π_{0.7}$通过多样化的上下文条件训练,实现跨环境任务执行与零样本泛化,支持多阶段任务和复杂操作,如咖啡机操作,性能媲美专门强化学习微调模型。

Comments Website: https://www.pi.website/blog/pi07

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07825 2026-04-28 stat.ME 50%

Estimation Strategies for Causal Decomposition Analysis with Allowability Specifications

可允许性规范下的因果分解分析估计策略

John W. Jackson, Ting-Hsuan Chang, Aster Meche, Trang Q. Nguyen

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文探讨了因果分解分析在可允许性规范下的估计策略,提出新的估计方法以解决密度建模挑战,并通过模拟研究和实际数据验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏