PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
PRISM:通过基于图像的自我奖励机制进行文本到图像生成的提示优化
Guo Tang, HongJie Luo, Tianxu Wang, Ying Zhang, Hao Wang
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Sun Yat-sen University(中山大学)
;
South China University of Technology(华南理工大学)
;
Guangdong University of Technology(广东工业大学)
ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
ARMOR++:用于对深度伪造检测器进行可转移攻击的多域原语集的智能编排
Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis
机构
*
Department of Informatics, Aristotle University of Thessaloniki(塞萨洛尼基亚里士多德大学信息学系)
;
Infocomm Technology Cluster, Singapore Institute of Technology(新加坡科技学院信息通信技术集群)
;
Department of Applied Physics and Applied Mathematics, Columbia University(哥伦比亚大学应用物理与应用数学系)
;
Division of Natural and Applied Sciences, Duke Kunshan University(昆山杜克大学自然科学与应用科学部)
;
Department of Electrical and Computer Engineering, University of Toronto(多伦多大学电气与计算机工程系)
A Good Initialization is All You Need for Faithful Visual Attribution
忠实视觉归因只需一个良好的初始化
Zihan Gu, Jiayu Wang, Hua Zhang, Yue Hu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Southern California(南加州大学)
;
DeerLab LLC(DeerLab有限责任公司)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Clemson University(克莱姆森大学)
;
Google(谷歌)
;
San Jose State University(圣何塞州立大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
The University of Hong Kong(香港大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Carnegie Mellon University(卡内基梅隆大学)
;
LIGHTSPEED Shenzhen(LIGHTSPEED深圳)
;
LIGHTSPEED Los Angeles(LIGHTSPEED洛杉矶)
ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving
ASSCG:自动驾驶中快慢LLM规划的恰到好处门控
Sining Ang, Yuan Chen, Liu Haiyan, Xuanyao Mao, Jason Bao, Xuliang, Bingchuan Sun, Yan Wang
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Department of Automation, University of Science and Technology of China(中国科学技术大学自动化系)
;
Beijing University of Aeronautics and Astronautics(北京航空航天大学)
;
Lenovo Group Limited(联想集团)
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
P-MTP: 通过渐进深度缩放的多令牌预测实现高效文档解析
Le Xiang, Chenxi Zhai, Shu Wei, Jingjing Wu, Qunyi Xie, Xiao Tan, Kunbin Chen, Wei He
机构
*
Department of Computer Vision Technology (VIS) Baidu Inc China(百度计算机视觉技术部(VIS))
;
Tsinghua University Shenzhen International Graduate School China(清华大学深圳国际研究生院)
GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation
GTA-Net:合作博弈论在胸部X光报告生成中的视觉-语言对齐
Saif ur Rehman Khan, Imad Ahmed Waqar, Sebastian Vollmer, Andreas Dengel, Muhammad Nabeel Asim
机构
*
Department of Computer Science, Rhineland-Palatinate Technical University of Kaiserslautern-Landau(莱茵兰-普法尔茨凯泽斯劳滕-兰道工业大学计算机科学系)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
IntelligentX GmbH
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think
微调视觉-语言-动作模型所需的层数比你想象的少
Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha, Khoa Vo, Philip Lund Møller, Quang T. Nguyen, Long Dinh, Tung M. Luu, Tuan Dam, Vu Duong, Trung Le, Nghi D. Q. Bui, Minh Vu, Tran Nguyen Le, An Thai Le, Ngan Le, Daniel Sonntag, James Zou, Jan Peters, Duy M. H. Nguyen, Ngo Anh Vien
机构
*
Center for AI Research, VinUniversity(VinUniversity人工智能研究中心)
;
VinRobotics
;
University of Arkansas(阿肯色大学)
;
Technical University of Denmark(丹麦技术大学)
;
Hanoi University of Science and Technology(河内科技大学)
;
KAIST(韩国科学技术院)
;
Monash University(莫纳什大学)
;
Oldenburg University(奥尔登堡大学)
;
DFKI(德国人工智能研究中心)
;
University of Stuttgart(斯图加特大学)
;
IMPRS-IS(国际马克斯·普朗克智能系统研究学院)
;
Stanford University(斯坦福大学)
;
Technische Universität Darmstadt(达姆施塔特工业大学)
机构
*
Cho Chun Shik Graduate School of Mobility, Korea Advanced Institute of Science and Technology(韩国科学技术院赵春植移动研究生院)
;
Department of Mechanical Engineering, Hanyang University(汉阳大学机械工程系)
;
Narnia Labs(纳尼亚实验室)