AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows
AgentCo-op: 基于检索的互操作多智能体工作流合成
Shuaike Shen, Wenduo Cheng, Shike Wang, Mingqian Ma, Jian Ma
机构
*
Ray and Stephanie Lane Computational Biology Department, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院雷和斯蒂芬妮·兰德计算生物学系)
;
Machine Learning Department, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院机器学习系)
机构
*
Northwestern University in Qatar(卡塔尔西北大学)
;
Independent Researcher(独立研究员)
;
Hamad Bin Khalifa University(哈马德·本·卡伊夫大学)
;
University of Tübingen(图宾根大学)
专题命中
视觉定位与Grounding
:vision language model(abstract);grounding(abstract)
机构
*
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, China(合肥工业大学计算机科学与信息工程学院)
;
Wuhan University, Wuhan, China(武汉大学)
;
Lab for Intelligence and visiON (LION)(智能与视觉实验室)
;
Xi'an Jiaotong University(西安交通大学)
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
IndusAgent: 通过智能工具增强开放词汇工业异常检测
Rongbin Tan, Fangfang Lin, Zhenlong Yuan, Min Qiu, Kejin Cui, Mengmeng Wang, Yi Wang, Zijian Song, Zhiyuan Wang, Jiyuan Wang, Yue Wang, Shuhan Song§, Huawei Cao
机构
*
State Key Lab of Processors, Institute of Computing Technology, CAS(处理器国家重点实验室,计算技术研究所,中国科学院)
;
Santa Clara University(圣克拉拉大学)
;
LongCat Team(LongCat团队)
;
Independent Researcher(独立研究者)
;
New York University(纽约大学)
;
Sun Yat-sen University(孙中山大学)
;
Nanyang Technological University(南洋理工大学)
;
Stanford University(斯坦福大学)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.CV
DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos
DeformMaster: 一个用于从视频中生成变形物体交互物理-神经世界模型
Can Li, Zhoujian Li, Ren Li, Jie Gu, Lei Lei, Jingmin Chen, Lei Sun
机构
*
Nankai University(南开大学)
;
Zhejiang University(浙江大学)
;
Southern University of Science and Technology(南方科技大学)
;
Rightly Robotics, A4X(Rightly Robotics,A4X)
;
University of Science and Technology of China(中国科学技术大学)
机构
*
Dream-X Team(Dream-X团队)
;
Stanford University(斯坦福大学)
;
National University of Singapore(新加坡国立大学)
;
Independent Researcher(独立研究者)
;
Case Western Reserve University(凯斯西储大学)
;
UC Santa Cruz(加州大学圣克鲁兹分校)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.CV
CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization
CosFly-Track: 一个大规模多模态数据集,用于通过多约束轨迹优化的无人机视觉跟踪
Xiangyue Wang, Hanxuan Chen, Songsheng Cheng, Ruilong Ren, Jie Zheng, Shuai Yuan, Tianle Zeng, Hanzhong Guo, Kangli Wang, Ji Pei
机构
*
Autel Robotics(Autel机器人公司)
;
Nanjing University(南京大学)
;
Peking University(北京大学)
;
Southern University of Science and Technology(南方科技大学)
;
University of Hong Kong(香港大学)
Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models
使用视觉-语言模型进行意大利议会演讲的转录与识别
Luigi Curini, Alfio Ferrara, Giovanni Pagano, Sergio Picascia
机构
*
Università degli Studi di Milano(米兰大学)
;
Department of Social and Political Sciences(社会科学系)
;
Department of Literary Studies, Philology and Linguistics(文学研究、语言学与语言学系)
;
Department of Computer Science(计算机科学系)
Commentsto be published in: ParlaCLARIN V: Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora, organized within the 15th Language Resource and Evaluation Conference (2026)
机构
*
Microsoft Research Asia(微软亚洲研究院)
;
Nanjing University(南京大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Wuhan University(武汉大学)
;
University of Technology Sydney(悉尼科技大学)
;
Tsinghua University(清华大学)
How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
如何移动决定了将来的行动:轨迹条件的自身视角预测
Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara, Lorenzo Torresani
机构
*
Khoury College of Computer Sciences, Northeastern University, Boston(北德文斯克学院,东北大学,波士顿)
;
Korea Advanced Institute of Science and Technology, Daejeon(韩国科学技术院,大田)
;
Department of Mathematics and Computer Science, University of Catania, Italy(卡塔尼亚大学数学与计算机科学系,意大利)
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
SurgOnAir: 基于层次感知的实时手术视频评论
Jingyi He, Yue Zhou, Long Bai, Kun Yuan, Nassir Navab, Yuan Bi
机构
*
Computer Aided Medical Procedures (CAMP), TU Munich, Germany Munich Center for Machine Learning (MCML), Munich, Germany University of Strasbourg, France The Chinese University of Hong Kong, Hong Kong
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
HalluCXR: 评估和缓解医疗视觉-语言模型在胸部X光解读中的幻觉
Haoyu Wang, Zitong Li
机构
*
Department of Biostatistics & Health Informatics, Institute of Psychiatry, Psychology & Neuroscience, King’s College London(生物统计学与健康信息学系,精神病学、心理学与神经科学研究所,伦敦国王学院)
Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow Matching
通过黎曼流匹配对预训练视觉语言模型进行知识不确定性量化
Li Ju, Mayank Nautiyal, Andreas Hellander, Ekta Vats, Prashant Singh
机构
*
Department of Information Technology, Uppsala University, Uppsala, Sweden(瑞典乌普萨拉大学信息科技系)
;
Science for Life Laboratory, Uppsala University, Uppsala, Sweden(瑞典乌普萨拉大学生命科学实验室)