VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.MM
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.MM
机构 * New York University(纽约大学) ; Yale University(耶鲁大学) ; Stanford University(斯坦福大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Project page: https://vision-x-nyu.github.io/thinking-in-space.github.io/
机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) ; Department of Software and Microelectronics, Peking University(软件与微电子系,北京大学) ; School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) ; State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University(虚拟现实技术与系统国家重点实验室,北航软件学院) ; Peng Cheng Laboratory(鹏城实验室) ; Baidu Inc(百度公司)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments ICML 2025 Spotlight
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Under review for ACM SIGSPATIAL 2025
机构 * Department of Civil, Construction and Environmental Engineering, North Dakota State University(土木、建筑与环境工程系,北达科塔州立大学)
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV
机构 * Shanghai Jiao Tong University(上海交通大学) ; Michigan State University(密歇根州立大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments 10 pages, 9 figures, published in NeurIPS 2024
机构 * AI Thrust, Information Hub, HKUST(GZ)(香港科技大学(广州)人工智能 thrust 与信息中心) ; College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) ; School of Medicine, Stanford University(斯坦福大学医学院) ; Computer Science Department, UIUC(伊利诺伊大学厄巴纳-香槟分校计算机科学系) ; Medical Big Data Center, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University(广东省人民医院(广东省医学科学院)医学大数据中心,南方医科大学) ; Innovation Institute for Artificial Intelligence in Medicine of Zhejiang University, College of Pharmaceutical Sciences, Zhejiang University(浙江大学人工智能医学创新研究院,浙江大学药学院) ; The Second Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院附属第二医院) ; GE HealthCare(通用电气医疗) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; IQVIA ; Computer Science Department, Stanford University(斯坦福大学计算机科学系) ; Informatics, Harvard Medical School, Harvard University(哈佛医学院,哈佛大学informatics部门) ; State Key Laboratory for Novel Software Technology at Nanjing University, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室,南京大学计算机科学学院)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
Comments accepted by Nature Scientific Data
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments 17 pages (main paper), 7 pages appendix. ICLR 2025 conference paper
机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) ; PCA-Lab, School of Computer Science and Engineering, Nanjing University of Science and Technology, China(PCA实验室,计算机科学与工程学院,南京理工大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Journal ref Journal of the American Medical Informatics Association, 2025;, ocaf074
机构 * Full author list in Contributions(完整作者列表)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments Astra Technical Report
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted in 2025 IEEE International Conference on Image Processing (ICIP)
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; National Key Laboratory of Space Integrated Information System(空间信息集成系统国家重点实验室) ; Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) ; Department of Computer Science and Technology(计算机科学与技术系)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments This work has been submitted to the IEEE for possible publication
机构 * Sber AI Lab(Sber AI实验室)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
机构 * School of Economics and Management, North China Electric Power University(华北电力大学经济管理学院) ; Department of Electrical Engineering, Tsinghua University(清华大学电气工程系) ; School of Control and Computer Engineering, North China Electric Power University(华北电力大学控制与计算机工程学院) ; School of Electrical Engineering, Southeast University(东南大学电气工程学院)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
机构 * Department of Informatics, Graduate School of Informatics and Engineering, The University of Electro-Communications(信息学院、信息与工程研究生院、电通通信大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Accepted at the International Conference on Advanced Concepts for Intelligent Vision Systems (ACIVS 2025)
机构 * ARC Lab, Tencent PCG(腾讯PCG广告实验室) ; City University of Hong Kong(香港城市大学)
专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV
Comments Homepage: https://github.com/TencentARC/Video-Holmes
机构 * CUHK(香港中文大学) ; HKUST(香港理工大学) ; RUC(中国人民大学)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
机构 * The University of Tokyo, Graduate School of Arts and Sciences(东京大学艺术与科学研究生院) ; Center for Information and Neural Networks (CiNet), National Institute of Information and Communications Technology(信息与神经网络中心(CiNet),信息与通信技术国家研究所) ; The University of Osaka, Graduate School of Frontier Biosciences(大阪大学前沿生命科学研究生院) ; The University of Tokyo, Faculty of Engineering(东京大学工学部) ; The University of Osaka, Graduate School of Engineering Science(大阪大学工学研究院) ; The University of Osaka, Graduate School of Medicine(大阪大学医学研究院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments 25 pages, 7 figures
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Peng Cheng Laboratory(鹏城实验室) ; National University of Singapore(新加坡国立大学) ; Meituan Inc.(美团公司)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Journal ref in Proceedings of 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2025)
机构 * Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(斯克尔科沃科学与技术研究所智能空间机器人实验室,数字工程中心) ; Institute of Automation, Qilu University of Technology (Shandong Academy of Sciences)(齐鲁科技大学自动化研究所,山东科学院)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted by ICRA
机构 * University of Genoa(热力学与系统工程系,基因瓦大学) ; The University of Danang–University of Science and Technology(丹那大学–科学与技术大学) ; Japan Advanced Institute of Science and Technology(日本先进科学研究院)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
机构 * Department of Civil and Environmental Engineering(土木及环境工程系) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI
机构 * Simon Fraser University(西蒙弗雷泽大学) ; McGill University(麦吉尔大学) ; The University of Hong Kong(香港大学) ; Wild Salmon Center(野生鲑鱼中心) ; Pacific Salmon Foundation(太平洋鲑鱼基金会) ; Haida Fisheries Program(海达渔业计划)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI
Comments 10 pages, accepted by IJCAI 2025, AI and Social Good Track
机构 * University of California, Irvine(加州大学尔湾分校)
专题命中 视频多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Kuaishou Technology(快手科技)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments CVPR 2025. Homepage: https://zhuomanliu.github.io/PhysFlow/
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Georgia Institute of Technology(佐治亚理工学院)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments 8 pages, 6 figures