arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM SIGIR Conference on Research and Development in Information Retrieval · 会议 · Information Retrieval

共收录 1420
2608.09934 2026-08-12 cs.CL cs.AI 新提交

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

LLM智能体工厂:领域特定大语言模型智能体的检索

Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy

机构 * Sber AI Moscow Institute of Physics and Technology(莫斯科物理技术学院) National University of Science and Technology MISIS(莫斯科国立科学技术大学MISIS) Skolkovo Institute of Science and Technology(斯科尔科沃科学技术研究所) Artificial Intelligence Research Institute(人工智能研究所)

AI总结 本研究提出LLM Agents Factory框架,通过检索2万余个预设智能体配置文件按需构建领域特定智能体,在多基准测试中实现高准确率与低推理成本,为工业应用提供动态智能体生成的替代方案。

Comments 7 pages, 1 figure, SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08467 2026-08-11 cs.AI cs.CL cs.IR 新提交

LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

MCP中的大语言模型很重要:测量由大语言模型驱动的低效资源利用

Minhan Cho, Soyoung Park, Kihyeon Jeong, Byeongkyu Jeon, Daejin Choi, Jinyoung Han

机构 * Sungkyunkwan University(成均馆大学) National Assembly Research Service(国会研究服务处) AlphaBridge(阿尔法桥公司) Ewha Womans University(梨花女子大学)

AI总结 该研究针对24个LLM开展54000次试验,发现MCP中服务器嵌入的参考数据会因LLM偏好导致低效资源利用,提出需将服务器指令置于客户端LLM工具选择之前的改进方向。

Comments 4 pages, 1 table. Accepted at the AgentSearch Workshop at SIGIR 2026, Melbourne, Australia (non-archival). Code and data: https://github.com/rabqatab/llm-in-mcp-matters

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04302 2026-08-06 cs.CV cs.IR cs.MM 新提交

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

CLIP-CC-Bench:评估视频语言模型中的段落级视频描述

Mukhtiar Ali, Harsh Dubey, Sugam Mishra, Chulwoo Pack

机构 * South Dakota State University(南达科他州立大学)

AI总结 CLIP-CC-Bench是针对视频语言模型的长篇段落级视频描述评估套件,采用多LLM嵌入模型集成与粗细粒度语义匹配方法,评估17种模型填补了现有短片段基准的空白。

Comments Accepted and presented at EvalMG 2026, the Second Workshop on Evaluation for Multimodal Generation, co-located with ACM SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28971 2026-08-03 cs.IR cs.LG 新提交

Don't Contrast the Impossible: Region-Constrained Batching for Contrastive User Modeling on a Local Community Platform

不要对比不可能的:面向本地社区平台对比式用户建模的区域约束批处理

Seungho Han, Byeongchang Kim, Jin Yu

AI总结 本研究针对本地社区平台用户建模中对比学习的不可能负样本问题,提出区域约束批采样方法,可提升用户表示质量并优化推荐、广告排序,相关嵌入已部署生产。

Comments Accepted at SIGIR 2026 (Industry Track)

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), pages 4639-4643, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27959 2026-07-31 cs.CV cs.IR 新提交

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

FiRE:利用细粒度上下文学习增强多模态大语言模型(MLLMs)以实现复杂图像检索

Bohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen, Meng Liu, Yupeng Hu, Xiangyu Zhao

机构 * Shandong University(山东大学) City University of Hong Kong(香港城市大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Shandong Jianzhu University(山东建筑大学)

AI总结 本研究针对MLLMs在复杂图像检索任务中细粒度建模不足的问题,提出自动化细粒度多模态五元组数据集构建流程与两阶段微调策略,在零样本设置下于五个数据集上取得优于现有方法的性能。

Journal ref Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27654 2026-07-31 cs.CL cs.AI 新提交

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

从单文档到跨文档:大语言模型多粒度事件分析的基准测试

Tao Wen, Shuai Shao, Pei Ke, Xu Han, Jie Zou, Guannan Li, Tao Tian, Jinjie Qiu, Lan Wang, Ke Qin

机构 * University of Electronic Science and Technology of China(电子科技大学) Tsinghua University(清华大学)

AI总结 本文推出MiGUE-Bench基准及MiGUE-Pipeline框架,通过四项核心任务评估LLMs多粒度事件分析能力,明确其能力边界与缺陷,为该领域改进提供方向。

Comments 9 pages. Published in the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), pp. 3464-3472, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23561 2026-07-28 cs.IR 新提交

Towards a Relevance Posterior in Neural Information Access

迈向神经信息访问中的相关性后验

Andrew Parry, Emmanouil Georgios Lionis, Debasis Ganguly, Sean MacAvaney

AI总结 研究现代信息检索系统中相关性操作的局限,提出将其理解为近似后验推理,扩展经典概率检索形式,通过显式似然 - 先验分解转移计算,纳入查询独立文档效用,经实验证明能提升检索效果并给出研究方向。

Comments SIGIR 2026 Perspectives Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14914 2026-07-28 cs.LG cs.IR 版本更新

Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation

加性控制变量化身自归一化在非策略评估中的主导地位

Olivier Jeunen, Shashank Gupta

机构 * Microsoft(微软)

AI总结 本文证明加性控制变量化身自归一化在非策略评估中具有更优的均方误差性能,理论支持了从自归一化到最优基线修正的转变。

Comments Published at SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05204 2026-07-27 cs.IR 版本更新

Entities as Retrieval Signals: A Systematic Study of Coverage, Supervision, and Evaluation in Entity-Oriented Ranking

实体作为检索信号:面向实体导向排序的系统研究

Shubham Chatterjee

AI总结 研究探讨了实体导向排序中覆盖、监督与评估的系统性问题,发现实体通道限制导致覆盖与区分度难以兼得,强调需改进评估方法以区分条件与开放世界场景。

Comments v2: Corrects RelCov@20 in Table 6 (previously approximated from entity document frequencies; now computed exactly at document level). Reframes the evaluation axis as leaked vs. clean entity supervision rather than document-pool restriction. Adds discussion of Boudens et al. (SIGIR 2026), linking density statistics, and an OER-oracle diagnostic

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19793 2026-07-23 cs.AI cs.CV 新提交

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

多模态智能体搜索中的无声故障:一种诊断分类法与跨评判者评估

Zhengxian Wu, Junjie Gao, Kai Yang

机构 * Ant Group(蚂蚁集团)

AI总结 研究多模态智能体搜索中无声故障这一隐藏可靠性问题,引入六类分类法,构建轨迹级诊断管道,通过实验表明表面准确性高估真实正确性,且无声故障与能力相关、常转移。

Journal ref SIGIR 2026 SynthIR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18712 2026-07-22 cs.IR 新提交

An Epistemic Position-Based Click Model: From Interactions to Epistemic Distributions of Relevance and Bias

基于认知位置的点击模型:从交互到相关性和偏差的认知分布

Oscar Rolando Ramirez Milian, Harrie Oosterhuis

AI总结 研究用户与排名交互中点击概率建模问题,提出基于证据深度学习的方法,输出贝塔分布捕捉认知不确定性,经实验验证该方法有效,是将贝叶斯不确定性纳入点击建模的重要进展。

Comments Published at SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18626 2026-07-22 cs.IR cs.CL 新提交

PLAID-PRF: Pseudo-Relevance Feedback with Centroid-like Tokens in PLAID

PLAID-PRF:在PLAID中使用类质心令牌的伪相关反馈

Xiao Wang, Sean MacAvaney, Craig Macdonald

机构 * University of International Business and Economics(国际经济贸易大学) University of Glasgow(格拉斯哥大学)

AI总结 研究在PLAID基础上提出PLAID-PRF方法,通过对顶部检索结果执行伪相关反馈来重新制定查询向量,利用质心向量降低计算成本。实验表明该方法能有效提升检索效果,相比PLAID有显著改进,且计算开销小,实现高效反馈感知后期交互检索。

Comments SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18421 2026-07-22 cs.CL cs.AI 版本更新

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

TReB:评估大语言模型表格推理能力的综合基准

Ce Li, Xiaofan Liu, Zhiyan Song, Ce Chi, Boshen Shi, Chen Zhao, Guanguang Chang, Zhendong Wang, Kexin Yang, Xing Wang, Chao Deng, Junlan Feng

机构 * JIUTIAN Research(天研机构)

AI总结 针对大语言模型处理表格数据推理能力评估基准缺失的问题,提出TReB基准,用分类法涵盖26个子任务,经数据处理构建高质量数据集,设计含三种推理模式的评估框架,揭示现有模型在表格任务上有改进空间,且数据和框架均公开。

Comments published by SIGIR 2026

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026, 3267-3275

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17457 2026-07-21 cs.IR 新提交

The Matryoshka Hypencoder

套娃式假设编码器

Majd Alkawaas, Sean MacAvaney

AI总结 研究基于套娃表示学习扩展假设编码器,支持多种大小Q-Nets以权衡有效性和效率。该“套娃式假设编码器”在域内参数大幅减少,吞吐量提升,为假设编码器实际部署奠定基础。

Comments SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17347 2026-07-21 cs.IR 新提交

Adapting Embedding Models for Agent Capability Retrieval

使嵌入模型适应智能体能力检索

Tingwei Chen, Yunxiao Shi, Zhengdong Chu, Qingsong Wen, Min Xu

AI总结 研究如何让一般文本检索的现成模型适应智能体能力检索,通过微调三个模型在特定数据集上训练,测试其在未训练目录上的迁移能力,结果显示适应对两个目录均有帮助。

Comments Accepted for oral presentation at the AgentSearch Workshop, SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11254 2026-07-21 cs.IR

MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval

MIRA:一个基于大语言模型的多类别集成检索评估基准

Mehmet Deniz Türkmen, Suchana Datta, Dwaipayan Roy, Daniel Hienert, Philipp Mayr, Derek Greene

AI总结 本文提出MIRA基准,通过真实用户查询构建,支持多类别跨领域检索评估,利用大语言模型生成主题描述和相关性评估,降低测试集生成成本。

Comments Accepted to SIGIR 2026. Resource Paper. 8 pages, 2 figures. DOI:10.1145/3805712.3808614

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), 2026, pp. 3426-3433

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27599 2026-07-20 cs.IR cs.LG

One Pass, Any Order: Position-Invariant Listwise Reranking for LLM-Based Recommendation

一次通过,任意顺序:基于LLM的推荐系统位置不变列表重排序

Ethan Bito, Yongli Ren, Estrid He

机构 * RMIT University(皇家墨尔本理工大学)

AI总结 本文提出InvariRank框架,通过结构化注意力掩码和RoPE共享位置框架,实现位置不变的列表重排序,提升推荐系统稳定性与效率。

Comments Accepted at SIGIR 2026

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026, pp. 3625-3629

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12714 2026-07-15 cs.IR 新提交

Learning to Forget: Satiation-Aware Long-Sequence Transducers for Mitigating Post-Purchase Redundancy

学习遗忘:用于减轻购买后冗余的饱腹感感知长序列变换器

Yipin Dai, Ruocong Tang, Xing Fang, Yang Huang, Jing Wang, Zhentao Song, He Guo

AI总结 针对电商场景中购买行为常意味兴趣终止,现有模型存在行动-意图不对称致购买后冗余的问题,提出饱腹感感知机制SAM,含双路径交叉注意力等三个关键组件,实验表明其显著降低购买后重复率超60%。

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), Industry Track, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12578 2026-07-15 cs.IR 新提交

Cheaper is Better: A Discount-Aware Network for Conversion Rate Prediction in E-commerce Recommendation System

更便宜更好:电子商务推荐系统中用于转化率预测的折扣感知网络

Ruocong Tang, Yang Huang, Xing Fang, Chenyi Yan, Chuike Sun, Jing Wang

AI总结 电商推荐系统中点击后转化率预测面临诸多挑战,本文提出折扣感知网络DANet,通过时频变换、分布去偏和监督回归辅助任务建模商品折扣率与CVR关系,实验证明其能提升预测性能,已成功部署。

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), Industry Track, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12098 2026-07-15 cs.IR 新提交

Explaining When PRF Fails: Participatory Auditing for Selective Query Expansion

解释伪相关反馈何时失败:用于选择性查询扩展的参与式审计

Zeyan Liang, Graham McDonald, Iadh Ounis

AI总结 研究PRF失败问题,提出两阶段审计然后自动化框架,第一阶段通过对用户审计揭示PRF利弊,第二阶段利用基于语言模型的重排器作预测器,解释PRF危害及决策过程,使检索组件可审计且以用户为基础。

Comments Accepted at WExIR @ SIGIR 2026, the 2nd Workshop on Explainability in Information Retrieval, Melbourne (Naarm), Australia, 24 July 2026. 3 pages, 1 figure. Extended abstract building on the SIGIR 2026 short paper "Auditing Query Drift: Do Users Actually Benefit from Pseudo-Relevance Feedback?" (doi:10.1145/3805712.3809916)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12783 2026-07-15 cs.IR cs.AI 版本更新

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

SQuTR:一种在语音噪声下 spoken query 到文本检索的鲁棒性基准

Yuejie Li, Ke Yang, Yueying Hua, Berlin Chen, Jianhao Nie, Yueping He, Caixin Kang

机构 * Huazhong University of Science and Technology(华中科技大学) The University of Hong Kong(香港大学) Soochow University(苏州大学) University of Science and Technology of China(中国科学技术大学) Wuhan University(武汉大学) Tsinghua University(清华大学) The University of Tokyo(东京大学)

AI总结 SQuTR通过大规模数据集和统一评估协议,评估语音检索系统在复杂噪声环境下的鲁棒性,揭示了极端噪声下检索性能显著下降的问题。

Comments Accepted by SIGIR 2026

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), July 20--24, 2026, Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10910 2026-07-14 cs.IR cs.LG 新提交

ZoRRO: A Zero-Weight Personalized Recommender System for Scalable News Recommendation

ZoRRO:用于可扩展新闻推荐的零权重个性化推荐系统

Johannes Kruse, Ryotaro Shimizu, Kasper Lindskow, Jon Tofteskov, Michael Riis Andersen, Julian McAuley, Jes Frellsen

机构 * Technical University of Denmark(丹麦技术大学) ZOZO Research(ZOZO研究公司) University of California San Diego(加州大学圣地亚哥分校) Pioneer Centre for Artificial Intelligence(先锋人工智能中心)

AI总结 研究针对可扩展新闻推荐提出零权重、无需训练的ZoRRO框架,在离线评估中优于神经基线,在线测试中点击率性能与先进模型相近且速度快600多倍,并揭示相关性能差距,凸显该框架对大规模新闻推荐的实用性及多指标评估的重要性。

Comments 6 pages, 2 figures. Accepted at the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), Melbourne, Australia, July 20-24, 2026. Code available at https://github.com/johanneskruse/zorro

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09908 2026-07-14 cs.CL cs.IR 新提交

RouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation

RouteRec:推荐代理选择与聚合的严格评估

Kaiji Zhou, Vladimir Kalmykov, Yue Feng

机构 * University of Birmingham(伯明翰大学)

AI总结 研究在推荐系统异构代理选择中,RouteRec框架在成本约束下比较请求级硬选与项目级学习聚合,发现在MovieLens-1M数据集上,请求级选择粗糙,项目级聚合更有前景,不同聚合方式有不同效果。

Comments 8 pages, 7 figures. Accepted at AgentSearch 2026 (The First Workshop on Indexing, Retrieval, and Ranking of AI Agents), co-located with SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03615 2026-07-13 cs.IR 新提交

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise

无作用的噪声:实验设置如何塑造报告的噪声能力

Michał Mazuryk, Fleur Dolmans, Louis Gehringer, Ina Klaric, Jia-Huei Ju, Mohammad Aliannejadi

AI总结 研究重现Cuconasu等人关于噪声对检索增强生成系统问答性能影响的发现,在扩展实验设置下评估其稳健性,发现该效应受推理配置影响大,强调审视推理设计的重要性。

Comments SIGIR 26 Repro

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08540 2026-07-10 cs.IR cs.CL 新提交

Improving Ad-hoc Search Effectiveness for Conversational Information Retrieval via Model Merging

通过模型合并提高对话信息检索的即席搜索效果

Ahmed Rayane Kebir, Jose G. Moreno, Lynda Tamine

机构 * University of Toulouse, IRIT(图卢兹大学,IRIT)

AI总结 针对对话信息检索难题,以往方法有缺陷。本文引入模型合并策略,通过线性和非线性参数合并,设计单一检索模型,无需额外微调,在多数据集实验,显著增强即席搜索能力,提高泛化性,零样本下NDCG@3提升15%。

Comments Accepted to SIGIR 2026. 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06765 2026-07-09 cs.IR 新提交

When and How to Ask: Dynamic Preference Elicitation Strategies for Conversational Recommendation

何时以及如何询问:对话推荐的动态偏好引出策略

Feng Xia, Shuo Zhang, Xi Wang

AI总结 研究对话推荐中偏好引出策略,发现最优策略具阶段依赖性与上下文敏感性。引入InPE数据集,提出COPE架构。离线评估表明上下文感知策略有益,分析预测策略揭示对话阶段趋势,为交互模式提供证据。

Comments Accepted at SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05970 2026-07-08 cs.IR cs.AI 新提交

Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search

忠实还是可查找?评估用于RDF数据集搜索的大语言模型生成的元数据

Riccardo Terrenzi, Serkan Ayvaz

机构 * University of Southern Denmark(南部丹麦大学)

AI总结 研究RDF数据集六种元数据生成设置,评估其检索有效性与忠实性。发现无约束重写检索增益强但忠实性低,更有根据的设置提高了忠实性,基于配置文件的重写权衡最佳,将合成元数据定位为需综合评估多方面的系统级信息检索问题。

Comments 5 pages, 1 figure, accepted at SynthIR @ SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03886 2026-07-07 cs.IR cs.AI 新提交

Enhancement of E-commerce Sponsored Search Relevancy with LLM

利用大语言模型提升电子商务赞助搜索的相关性

Md Omar Faruk Rokon, Andrei Simion, Weizhi Du, Musen Wen, Hong Yao, Kuang-chih Lee

机构 * Walmart AdTech(沃尔玛广告技术)

AI总结 研究利用预训练大语言模型在电子商务赞助搜索框架中开发先进广告相关性模型,通过LoRA改编LLAMA2 7B模型,引入新型分类器,经训练显著提升搜索精度和效率。

Comments eCom 24: ACM SIGIR Workshop on eCommerce, July 18, 2024, Washington, DC, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03880 2026-07-07 cs.IR cs.AI 新提交

Next-Gen Sponsored Search: Crafting the Perfect Query with Inventory-Aware RAG (InvAwr-RAG) Based GenAI

下一代赞助搜索:基于库存感知检索增强生成式人工智能(InvAwr-RAG)打造完美查询

Md Omar Faruk Rokon, Weizhi Du, Zhaodong Wang, Musen Wen

机构 * Walmart AdTech(沃尔玛广告科技)

AI总结 研究电商赞助搜索中为给定查询识别相关关键词的挑战,引入基于库存感知检索增强生成式人工智能模型,结合多种查询提升相关性和用户参与度,初步结果显示填充率等显著提升。

Comments Published in eCom@SIGIR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06774 2026-07-07 cs.SE 版本更新

OpenCoderRank: Personalized Technical Assessments with Generative AI

OpenCoderRank: 基于生成式AI的个性化技术评估

Hridoy Sankar Dutta, Sana Ansari, Swati Kumari, Shounak Ravi Bhalerao

AI总结 OpenCoderRank是一种轻量级自托管平台,通过模拟真实世界限时技术评估,为资源受限环境提供定制化评估解决方案,结合BERTScore和LLM评估方法验证其有效性。

Comments Accepted to SynthIR Workshop (SIGIR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏