Paper2Video: Automatic Video Generation from Scientific Papers
机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments Project Page: https://showlab.github.io/Paper2Video/
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments Project Page: https://showlab.github.io/Paper2Video/
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments Accepted to PMLR 298, 10th Machine Learning for Healthcare Conference (MLHC)
机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) ; University of Surrey(Surrey大学) ; School of Computer Science, Wuhan University(武汉大学计算机学院) ; School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) ; College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) ; Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables
机构 * seas.upenn.edu(宾夕法尼亚大学塞as学院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) ; Beijing Institute of Technology(北京理工大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Hunan University(湖南大学) ; Shanghai AI Lab(上海人工智能实验室) ; HEBUST
专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV、cs.LG
Comments Accepted to NeurIPS 2025. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF
机构 * Niantic Spatial ; KAUST(科威特科学与技术研究中心) ; UCL(伦敦大学学院)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments ICCV 2025. Project page: https://nianticlabs.github.io/placeit3d/
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments 52 pages
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments Project Page: https://deeptracereward.github.io/
机构 * Aalto University(阿alto大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
机构 * Korea University(韩国大学) ; University of Seoul(首尔大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
Comments 23 pages, 17 figures
机构 * Fractal AI Research(Fractal AI研究)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
机构 * KAIST(韩国科学技术院) ; Seoul National University(首尔国立大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments Accepted to EMNLP 2025 (Main Conference)
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) ; Beijing Electronic Science and Technology Institute(北京电子科学技术研究所) ; Wangxuan Institute of Computer Technology(北京大学计算机技术研究院) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * Nanjing University of Science and Technology(南京理工大学) ; Beijing Normal University(北京师范大学)
专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.AI
Comments 10 pages, 5 figures
机构 * Amazon(亚马逊公司) ; University of California, Los Angeles (UCLA)(加州大学洛杉矶分校) ; Duke University(杜克大学)
专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.AI
Comments Submitted to CVPR 2025 and Published at CVPR 2025 AI for Content Creation workshop
机构 * University of Southern California(南加州大学) ; DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
Comments Accepted at NeurIPS 2025, 32 pages, 5 figures
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.LG
机构 * La Sapienza University of Rome(拉维亚大学罗马分校) ; Peoples' Friendship University of Russia (RUDN University)(俄罗斯人民友谊大学) ; Joint Institute for Nuclear Research(联合核子研究所) ; Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) ; University of Skövde(斯德哥尔摩大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) ; Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院) ; Peter Munk Cardiac Centre, University Health Network(皮特·蒙克心脏中心,大学健康网络) ; Medical Biophysics, University of Toronto(医学生物物理系,多伦多大学) ; Department of Computer Science, University of Toronto(计算机科学系,多伦多大学) ; Department of Laboratory Medicine and Pathobiology, University of Toronto(实验室医学与病理学系,多伦多大学) ; AI Hub, University Health Network(人工智能中心,大学健康网络)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.LG
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments 18 pages
机构 * KAUST(科威特科学与技术王国(KAUST)) ; CIIRC CTU(布拉格技术大学智能信息与计算研究中心(CIIRC CTU)) ; Adobe Research(Adobe研究)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ; ISTI-CNR(意大利国家研究委员会ISTI) ; University of Pisa(比萨大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
Comments ICCV 2025
机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) ; Department of Civil, Construction and Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
机构 * Sharif University of Technology(谢里夫理工大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments Accepted to TMLR 2025
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * Accenture Labs(埃森哲实验室)
专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV、cs.LG
机构 * TOELT LLC AI lab(TOELT LLC人工智能实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Ho Chi Minh University of Science(胡志明市科学大学) ; Deakin University(德肯大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments Published at ACMMM2025 (Dataset track)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments 9 pages,7 figures,conference