arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Columbia University(哥伦比亚大学)

共收录 1073
2601.21649 2026-01-30 cs.LG cs.AI cs.CL cs.SE

SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning

SWE-Spot:构建基于仓库的小型专家

Jinjun Peng, Magnus Saebo, Tianjun Zhong, Yi-Jie Cheng, Junfeng Yang, Baishakhi Ray, Simin Chen, Yangruibo Ding

机构 * Department of Computer Science, Columbia University, New York, USA(哥伦比亚大学计算机科学系) Computer Science Department, University of California, Los Angeles, California, USA(加州大学洛杉矶分校计算机科学系)

AI总结 SWE-Spot提出仓库导向学习范式,通过参数知识获取训练小型专家模型,提升代码任务表现并降低推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18710 2026-01-27 cs.ET cs.LG quant-ph

Analyzing Images of Blood Cells with Quantum Machine Learning Methods: Equilibrium Propagation and Variational Quantum Circuits to Detect Acute Myeloid Leukemia

利用量子机器学习方法分析血细胞图像:平衡传播与变分量子电路用于检测急性髓系白血病

A. Bano, L. Liebovitch

机构 * Electrical and Computer Engineering(电气与计算机工程) Rutgers University(罗格斯大学) AC4 in the Climate School(气候学校AC4) Columbia University(哥伦比亚大学)

AI总结 本文利用量子机器学习方法,通过平衡传播和变分量子电路在血细胞图像中检测急性髓系白血病,展示了量子方法在医疗影像中的可行性。

Comments 5 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17690 2026-01-27 cs.SD cs.AI cs.IR cs.LG eess.AS

Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance

片段长度至关重要:对音频指纹识别性能中片段长度的研究

Ziling Gong, Yunyan Ouyang, Iram Kamdar, Melody Ma, Hongjie Chen, Franck Dernoncourt, Ryan A. Rossi, Nesreen K. Ahmed

机构 * Data Science Institute Columbia University(数据科学研究所 哥伦比亚大学) Dolby Laboratories(杜比实验室) Adobe Research(Adobe研究) Cisco Research(思科研究)

AI总结 本文研究了片段长度对音频指纹识别性能的影响,发现短片段长度(0.5秒)表现更优,并评估了大语言模型推荐最佳片段长度的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17577 2026-01-27 cs.HC cs.AI cs.CL

Status Hierarchies in Language Models

语言模型中的地位层级

Emilio Barkett

机构 * Brigham Young University–Hawaii COLUMBIA UNIVERSITY(哥伦比亚大学)

AI总结 语言模型在多智能体环境中会因地位线索形成层级,高地位分配反而降低高能力模型的服从,揭示AI系统中的新兴社会行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17510 2026-01-27 stat.ML cs.AI cs.LG

"Rebuilding" Statistics in the Age of AI: A Town Hall Discussion on Culture, Infrastructure, and Training

在人工智能时代重建统计学:关于文化、基础设施和培训的圆桌讨论

David L. Donoho, Jian Kang, Xihong Lin, Bhramar Mukherjee, Dan Nettleton, Rebecca Nugent, Abel Rodriguez, Eric P. Xing, Tian Zheng, Hongtu Zhu

机构 * Department of Statistics, Stanford University(斯坦福大学统计学系) Department of Biostatistics, University of Michigan, Ann Arbor(密歇根大学安娜堡分校生物统计学系) Harvard T.H. Chan School of Public Health(哈佛大学T.H. Chan公共卫生学院) Department of Statistics, Harvard University(哈佛大学统计学系) Broad Institute(Broad研究所) Yale School of Public Health(耶鲁大学公共卫生学院) Department of Statistics and Data Science, Yale University(耶鲁大学统计学与数据科学系) Department of Statistics, Iowa State University(爱荷华州立大学统计学系) Department of Statistics and Data Science, Carnegie Mellon University(卡内基梅隆大学统计学与数据科学系) Baskin School of Engineering, University of California, Santa Cruz(加州大学圣克鲁兹分校Baskin工程学院) Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院) Department of Statistics, Columbia University(哥伦比亚大学统计学系) Department of Biostatistics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物统计学系)

AI总结 本文记录了2024年JSM圆桌讨论,探讨统计学在人工智能时代的发展,聚焦文化、基础设施和培训等关键问题。

Comments 35 pages, 3 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17231 2026-01-27 cs.RO cs.AR

Real-Time, Energy-Efficient, Sampling-Based Optimal Control via FPGA Acceleration

基于FPGA加速的实时、节能的采样最优控制

Tanmay Desai, Brian Plancher, R. Iris Bahar

机构 * Colorado School of Mines(科罗拉多矿业学院) Barnard College(巴纳德学院) Columbia University(哥伦比亚大学) Dartmouth College(达特茅斯学院)

AI总结 本文提出一种基于FPGA加速的MPPI控制方法,通过深度流水线和并行性优化,实现嵌入式平台上的高效能和低能耗控制。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22860 2026-01-27 cs.CL q-bio.NC

Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement

远离浅层:通过残差解耦实现脑预测推理嵌入

Linyang He, Tianjun Zhong, Richard Antonello, Gavin Mischler, Micah Goldblum, Nima Mesgarani

机构 * Zuckerman Mind Brain Behavior Institute, Columbia University(扎克曼脑行为研究所,哥伦比亚大学) Department of Electrical Engineering, Columbia University(电气工程系,哥伦比亚大学) Department of Computer Science, Columbia University(计算机科学系,哥伦比亚大学)

AI总结 通过残差解耦方法,研究实现了对语言推理过程的神经嵌入解耦,揭示了推理在大脑中的独特预测能力及时间特征。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09882 2026-01-27 cs.CL

iBERT: Interpretable Embeddings via Sense Decomposition

iBERT:通过语义分解实现可解释嵌入

Vishal Anand, Milad Alshomary, Kathleen McKeown

机构 * Microsoft(微软公司) Columbia University(哥伦比亚大学)

AI总结 iBERT通过语义分解生成可解释嵌入,提升风格表示效果并保持作者身份验证性能。

Comments Accepted to the Main Proceedings of EACL 2026. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12371 2026-01-27 cs.LG stat.ML

Path-specific effects for pulse-oximetry guided decisions in critical care

脉氧引导决策中特定路径的偏见影响

Kevin Zhang, Yonghan Jung, Divyat Mahajan, Karthikeyan Shanmugam, Shalmali Joshi

机构 * Columbia University(哥伦比亚大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Mila & Université de Montréal(Mila与蒙特利尔大学) Google DeepMind(谷歌DeepMind) Purdue University(普渡大学)

AI总结 本研究通过因果推理方法,利用路径特定效应分析脉氧测量中的种族偏见对ICU侵入性通气的影响,揭示种族差异对通气持续时间的影响,并强调因果方法在医疗公平性评估中的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00003 2026-01-27 cs.CL

A Review of Incorporating Psychological Theories in LLMs

对将心理理论纳入大语言模型的综述

Zizhou Liu, Ziwei Gong, Lin Ai, Zheng Hui, Run Chen, Colin Wayne Leach, Michelle R. Greene, Julia Hirschberg

机构 * Department of Computer Science, Columbia University(哥伦比亚大学计算机科学系) Department of Psychology, Barnard College(巴纳德学院心理学系) The Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室)

AI总结 本文综述了心理学理论如何指导和增强大语言模型开发,整合了六个心理学子领域,旨在促进心理学更深入地融入NLP研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16922 2026-01-26 cs.LG stat.ML

Group-realizable multi-group learning by minimizing empirical risk

多组学习的组可实现方法通过最小化经验风险

Navid Ardeshir, Samuel Deng, Daniel Hsu, Jingwen Liu

机构 * Columbia University(哥伦比亚大学)

AI总结 本文提出了一种通过最小化经验风险来实现多组学习的方法,该方法在组可实现设置下能提高样本复杂性,并提出了基于不当学习的替代方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17629 2026-01-23 cs.LG

Boundary-Aware Adversarial Filtering for Reliable Diagnosis under Extreme Class Imbalance

边界感知对抗过滤用于极端类别不平衡下的可靠诊断

Yanxuan Yu, Michael S. Hughes, Julien Lee, Jiacheng Zhou, Andrew F. Laine

机构 * Columbia University, USA(哥伦比亚大学) Columbia University Irving Medical Center, USA(哥伦比亚大学伊万斯医学中心)

AI总结 AF-SMOTE通过对抗过滤和边界效用模型提高极端类别不平衡下的召回率和校准,优于现有过采样方法。

Comments Accepted to IEEE ISBI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08620 2026-01-23 cs.LG

CT-PatchTST: Channel-Time Patch Time-Series Transformer for Long-Term Renewable Energy Forecasting

CT-PatchTST:用于长周期可再生能源预测的通道-时间补丁时间序列变压器

Kuan Lu, Menghao Huo, Yuxiao Li, Qiang Zhu, Zhenrui Chen

机构 * School of Electrical and Computer Engineering(电气与计算机工程学院) Cornell University(康奈尔大学) Department of Electrical and Computer Engineering(电气与计算机工程系) Northeastern University(东北大学) Fu Foundation School of Engineering and Applied Science(富兰克林基金会工程与应用科学学院) Columbia University in the City of New York(纽约市哥伦比亚大学) School of Engineering(工程学院) Santa Clara University(圣克拉拉大学) Department of Mechanical and Aerospace Engineering(机械与航空航天工程系) University of Houston(休斯顿大学)

AI总结 CT-PatchTST通过捕捉时间依赖性和跨通道相关性,提升风能和太阳能的长周期预测精度,优化能源存储调度,增强电网稳定性与响应性。

Comments Published in: 2025 10th International Conference on Computer and Information Processing Technology (ISCIPT)

Journal ref 2025 10th International Conference on Computer and Information Processing Technology (ISCIPT), pp. 86-95

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14518 2026-01-22 cs.CL

Business Logic-Driven Text-to-SQL Data Synthesis for Business Intelligence

基于业务逻辑的文本到SQL数据合成用于商务智能

Jinhui Liu, Ximeng Zhang, Yanbo Ai, Zhou Yu

机构 * Department of Computer Science, Columbia University(哥伦比亚大学计算机科学系)

AI总结 本文提出基于业务逻辑的数据合成方法,生成高现实性的SQL数据,显著提升文本到SQL模型的性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13456 2026-01-21 cs.LG cs.DC

Federated Learning Under Temporal Drift -- Mitigating Catastrophic Forgetting via Experience Replay

联邦学习中的时间漂移——通过经验回放缓解灾难性遗忘

Sahasra Kokkula, Daniel David, Aaditya Baruah

机构 * Columbia University(哥伦比亚大学)

AI总结 本研究通过客户端侧经验回放方法缓解联邦学习中的时间漂移问题,实验表明该方法能有效恢复模型性能并防止灾难性遗忘。

Comments 8 pages, 5 figures. Course project for Neural Networks & Deep Learning COMSW4776 course at Columbia University

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07772 2026-01-21 cs.AI

An approach for systematic decomposition of complex llm tasks

一种复杂大语言模型任务的系统分解方法

Tianle Zhou, Jiakai Xu, Guanhong Liu, Jiaxiang Liu, Haonan Wang, Eugene Wu

机构 * Columbia University(哥伦比亚大学)

AI总结 本文提出ACONIC框架,通过形式复杂度度量系统分解任务,提升大语言模型在复杂任务中的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12227 2026-01-21 cs.LG

Learning Longitudinal Health Representations from EHR and Wearable Data

从电子健康记录和可穿戴数据中学习纵向健康表示

Yuanyun Zhang, Han Zhou, Li Feng, Yilin Hong, Shi Li

机构 * University of the Chinese Academy of Sciences, Columbia University(中国科学院大学,哥伦比亚大学)

AI总结 本文提出一种多模态基础模型,通过联合表示电子健康记录和可穿戴数据,提升纵向健康预测的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11920 2026-01-21 cs.CL cs.AI

Enhancing LLM-Based Data Annotation with Error Decomposition

通过错误分解增强基于LLM的数据标注

Zhen Xu, Vedant Khatri, Yijun Dai, Xiner Liu, Siyan Li, Xuanming Zhang, Renzhe Yu

机构 * Columbia University(哥伦比亚大学) University of California, Irvine(加州大学尔湾分校) University of Pennsylvania(宾夕法尼亚大学)

AI总结 通过错误分解增强基于LLM的数据标注,提出诊断评估范式以区分任务固有模糊性与模型不准确,提升标注质量评估的准确性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11768 2026-01-21 eess.AS cs.AI cs.LG cs.SD eess.SP

Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music

轻量级自监督检测基频及单音音乐中发声概率的准确估计

Venkat Suprabath Bitra, Homayoon Beigi

机构 * Dept. of Computer Science, Columbia University, New York, US(计算机科学系,哥伦比亚大学,纽约,美国) Dept. of Mechanical Engineering and Dept. of Electrical Engineering, Columbia University(机械工程系和电气工程系,哥伦比亚大学) Recognition Technologies, Inc., New York, US(识别技术公司,纽约,美国)

AI总结 本文提出了一种轻量级自监督框架,用于单音音乐中基频和发声概率的准确估计,通过转置等价学习和EM风格的迭代加权方案实现无需标注的高效训练。

Comments 12 pages, 6 figures, 3 tables, and an appendix, Accepted for publication at ICPRAM 2026 in Marbella, Spain, on March 2, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09841 2026-01-21 cs.LG cs.AI

A pipeline for enabling path-specific causal fairness in observational health data

一种实现路径特定因果公平性的观察性健康数据管道

Aparajita Kashyap, Sara Matijevic, Noémie Elhadad, Steven A. Kushner, Shalmali Joshi

机构 * Department of Biomedical Informatics(生物医学信息学系) Columbia University(哥伦比亚大学) Big Data Institute(大数据研究所) University of Oxford(牛津大学)

AI总结 本文提出了一种通用管道,用于训练能够解决直接和间接医疗偏见的因果公平机器学习模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24555 2026-01-21 cs.LG

From Perception to Punchline: Empowering VLM with the Art of In-the-wild Meme

从感知到 punchline:通过野生表情包艺术赋能 VLM

Xueyan Li, Yingyi Xue, Mengjie Jiang, Qingzi Zhu, Yazhe Niu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Software Engineering(软件工程学院) Columbia Engineering(哥伦比亚工程学院) Columbia University(哥伦比亚大学) The Chinese University of Hong Kong MMLab(香港中文大学MMLab)

AI总结 HUMOR 通过分层推理和群体偏好对齐,提升 VLM 在多模态生成中的推理多样性与幽默质量。

Comments 46 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19945 2026-01-16 stat.ML cs.LG math.PR

Data-Driven Dynamic Factor Modeling via Manifold Learning

通过流形学习的数据驱动动态因子建模

Graeme Baker, Agostino Capponi, J. Antonio Sidaoui

机构 * Columbia University(哥伦比亚大学)

AI总结 本文提出了一种基于流形学习的数据驱动动态因子建模方法,通过各向异性扩散图学习低维嵌入,以提高高维协变量和响应的联合演变预测精度,并在投资组合压力测试中实现了显著的误差改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08342 2026-01-14 cs.CL

Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue

通过合成多说话对话检测心理操控

Run Chen, Wen Liang, Ziwei Gong, Lin Ai, Julia Hirschberg

机构 * Columbia University(哥伦比亚大学) Red Hat(红帽公司)

AI总结 本文提出首个通过合成多说话对话检测心理操控的研究,揭示语音与文本在检测准确性上的差异,强调多模态对话系统中模态意识的重要性。

Comments Accepted to IWSDS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07854 2026-01-14 q-bio.OT cs.AI

Immunological Density Shapes Recovery Trajectories in Long COVID

免疫密度塑造长期新冠恢复轨迹

Jing Wang, Tong Zhang, Xing Niu, Jie Shen, Yiming Luo, Qiaomin Xie, Amar Sra, Zorina Galis, Jeremy Weiss

机构 * National Library of Medicine(国家医学图书馆) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stevens Institute of Technology(史蒂文斯理工学院) Columbia University(哥伦比亚大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) George Washington University(乔治华盛顿大学) National Heart, Lung, and Blood Institute(国家心肺血液研究所)

AI总结 研究发现长期新冠症状严重程度由免疫密度决定,恢复主要通过重复接种疫苗实现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20031 2026-01-13 cs.RO cs.CV

MG-SLAM: Structure Gaussian Splatting SLAM with Manhattan World Hypothesis

MG-SLAM:基于曼哈顿世界假设的结构高斯点撒SLAM

Shuhong Liu, Tianchen Deng, Heng Zhou, Liuzhuozheng Li, Hongyu Wang, Danwei Wang, Mingrui Li

机构 * Department of Information Science and Technology and Department of Complexity Science and Engineering, The University of Tokyo(信息科学与技术系和复杂科学与工程系,东京大学) Institute of Medical Robotics and Department of Automation, Shanghai Jiao Tong University(医疗机器人研究所和自动化系,上海交通大学) Department of Mechanical Engineering, Columbia University(机械工程系,哥伦比亚大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电子与电气工程学院,南洋理工大学) Department of Computer Science, Dalian University of Technology(计算机科学系,大连理工大学)

AI总结 MG-SLAM基于曼哈顿世界假设,通过融合线段和平面假设提升室内场景重建的几何精度和完整性,实现高斯SLAM的先进性能。

Comments IEEE Transactions on Automation Science and Engineering

Journal ref IEEE Transactions on Automation Science and Engineering 22 (2025) 17034-17049

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18411 2026-01-13 cs.RO cs.CV

SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles

SeePerSea:用于自主水面车辆的水下物体多模态感知数据集

Mingi Jeong, Arihant Chadda, Ziang Ren, Luyang Zhao, Haowen Liu, Monika Roznere, Aiwei Zhang, Yitao Jiang, Sabriel Achong, Samuel Lensgraf, Alberto Quattrini Li

机构 * Department of Computer Science, Dartmouth College(达特茅斯学院计算机科学系) IQT Labs(IQT实验室) Department of Computer Science, Columbia University(哥伦比亚大学计算机科学系) Department of Computer Science, University of Maryland College Park(马里兰大学计算机科学系) The Institute for Human and Machine Cognition and The University of West Florida(人机认知研究所与西佛罗里达大学) School of Computing, Binghamton University(宾夕法尼亚州立大学计算学院)

AI总结 SeePerSea数据集为自主水面车辆提供多模态水下物体感知数据,通过训练测试现有深度学习算法,推动海洋自主技术发展。

Comments Topic: Special Issue on ICRA 2024 Workshop on Field Robotics

Journal ref IEEE Transactions on Field Robotics 2 (2025) - 737-752

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06216 2026-01-13 cs.CY cs.AI

LLM Agents in Law: Taxonomy, Applications, and Challenges

法律中的LLM代理:分类、应用与挑战

Shuang Liu, Ruijia Zhang, Ruoyun Ma, Yujia Deng, Lanyi Zhu, Jiayu Li, Zelong Li, Zhibin Shen, Mengnan Du

机构 * Carnegie Mellon University(卡内基梅隆大学) National University of Singapore(国立新加坡大学) Stanford University(斯坦福大学) University of Washington(华盛顿大学) The University of Chicago(芝加哥大学) Rutgers University(罗格斯大学) Columbia University(哥伦比亚大学) New Jersey Institute of Technology(新泽西理工学院)

AI总结 本文探讨了法律领域中LLM代理的分类、应用及挑战,分析了技术转变、应用分类、评估方法及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02212 2026-01-13 cs.LG cs.AI

DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning

DiFFPO:通过强化学习训练扩散大语言模型以快速且高效地推理

Hanyang Zhao, Dawen Liang, Wenpin Tang, David Yao, Nathan Kallus

机构 * Columbia University(哥伦比亚大学) Netflix

AI总结 DiFFPO通过强化学习训练扩散大语言模型,使其在推理速度和准确性上取得平衡,提升推理效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11720 2026-01-13 cs.LG cs.AI

RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences

RPO:通过丰富的视觉-语言偏好进行视觉生成模型微调

Hanyang Zhao, Haoxian Chen, Yucheng Guo, Genta Indra Winata, Tingting Ou, Ziyu Huang, David D. Yao, Wenpin Tang

机构 * Columbia University(哥伦比亚大学) Amazon(亚马逊) Princeton University(普林斯顿大学) Capital One

AI总结 RPO通过利用视觉语言模型的丰富反馈信号,改进偏好对的编纂,从而提升视觉生成模型的微调效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05420 2026-01-13 cs.RO

RoboPanoptes: The All-seeing Robot with Whole-body Dexterity

RoboPanoptes:具备全身灵活性的全方位机器人

Xiaomeng Xu, Dominik Bauer, Shuran Song

机构 * Stanford University(斯坦福大学) Columbia University(哥伦比亚大学)

AI总结 RoboPanoptes通过全身视觉和运动策略实现全身灵活性,能够高效操作和适应复杂环境。

Comments Project website: https://robopanoptes.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏