arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Columbia University(哥伦比亚大学)

2026-04-17 至 2026-04-17 共收录 5
2604.14525 2026-04-17 cs.AI

Quantifying Cross-Query Contradictions in Multi-Query LLM Reasoning

量化多查询LLM推理中的跨查询矛盾

Rohit Kumar Salla, Ramya Manasa Amancherla, Manoj Saravanan

机构 * Virginia Tech(弗吉尼亚理工大学) Columbia University(哥伦比亚大学)

AI总结 本文研究多查询推理中的逻辑一致性,提出包含390个实例的基准测试,引入集级指标,通过提取承诺、验证全局可满足性及反例引导修复,减少跨查询矛盾并保持单查询准确性。

Comments Accepted at the ICLR 2026 Workshop on Logical Reasoning of Large Language Models. 9 pages, 6 tables, code and data at https://huggingface.co/datasets/rohitspider/cross_query_benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14370 2026-04-17 stat.ME cs.LG

Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance

AI辅助干预的部署:容量限制与嘈杂合规性

Carri W. Chan, Yi Han, Hannah Li, Benjamin L. Ranard

机构 * Decision, Risk, and Operations, Columbia Business School(哥伦比亚商学院决策、风险与运营部门) Department of Statistics, Columbia University(哥伦比亚大学统计学系) Division of Pulmonary, Allergy, and Critical Care Medicine, Department of Medicine, Vagelos College of Physicians and Surgeons, Columbia University(哥伦比亚大学瓦格洛斯医学与外科学院肺科、过敏与危重医学分会)

AI总结 本文研究了在服务容量限制和行为响应不确定的情况下,AI辅助干预的最优阈值设置问题,提出Operational AUC作为改进的算法选择指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06296 2026-04-17 cs.LG cs.AI cs.MA cs.SE

AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent

AgentOpt v0.1 技术报告:基于LLM的代理的客户端优化

Wenyue Hua, Sripad Karne, Qian Xie, Armaan Agrawal, Nikos Pagonas, Kostis Kaffes, Tianyi Peng

机构 * Microsoft Research, AI Frontiers(微软研究院,人工智能前沿) Cornell University(康奈尔大学) Columbia University(哥伦比亚大学)

AI总结 本文提出AgentOpt框架,用于解决LLM代理客户端资源分配问题,通过多种搜索算法优化模型选择和资源分配,提升效率与效果。

Comments 24 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05024 2026-04-17 stat.ME cs.AI cs.LG

Model-Free Assessment of Simulator Fidelity via Quantile Curves

通过分位数曲线进行无模型的模拟器保真度评估

Garud Iyengar, Yu-Shiou Willy Lin, Kaizheng Wang

机构 * Department of IEOR and Data Science Institute, Columbia University(工业工程与数据科学学院,哥伦比亚大学)

AI总结 本文提出通过分位数曲线评估模拟器保真度,构建置信区间以估计潜变量参数差异,支持统计推断和风险测量,适用于多种输出空间。

Comments 39 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23468 2026-04-17 cs.RO cs.AI cs.LG

Multi-Modal Manipulation via Multi-Modal Policy Consensus

多模态操控 via 多模态策略共识

Haonan Chen, Jiaming Xu, Hongyu Chen, Kaiwen Hong, Binghao Huang, Chaoqi Liu, Jiayuan Mao, Yunzhu Li, Yilun Du, Katherine Driggs-Campbell

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Columbia University(哥伦比亚大学) Massachusetts Institute of Technology(麻省理工学院) Harvard University(哈佛大学)

AI总结 本文提出通过多模态策略共识实现多模态操控,采用扩散模型分解策略并利用路由网络动态组合模态贡献,提升机器人操控任务中多模态推理能力。

Comments 8 pages, 7 figures. Project website: https://policyconsensus.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏