Diffusion-CAM: Faithful Visual Explanations for dMLLMs
扩散-CAM:面向dMLLMs的可信视觉解释
Haomin Zuo, Yidi Li, Luoxiao Yang, Xiaofeng Zhang
机构
*
Department of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知系)
;
Sun Yat-sen University(中山大学)
;
Northwestern University(西北大学)
;
Technion - Israel Institute of Technology(以色列理工学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.AI
Wei Chen, Qibin Zhao, John Paisley, Junmei Yang, Delu Zeng
机构
*
School of Mathematics, South China University of Technology(华南理工大学数学学院)
;
Tensor Learning Team, RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心张量学习团队)
;
Department of Electrical Engineering, Columbia University(哥伦比亚大学电气工程系)
;
School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息工程学院)
机构
*
Shanghai Artificial Intelligence Laboratory, OpenDataLab(上海人工智能实验室,OpenDataLab)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
CASIVISION(中科视语)
;
THU(清华大学)
;
HNU(湖南大学)
;
HUST(华中科技大学)
Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
通过可验证奖励的强化学习指导大语言模型进行谈判
Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Lei Lei, Yuheng Zhang, Yisong Yue, David Simchi-Levi
机构
*
Purdue University(普渡大学)
;
California Institute of Technology(加州理工学院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Massachusetts Institute of Technology(麻省理工学院)
机构
*
Soochow University(苏州大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
FiT, Tencent(腾讯FiT)
;
Alibaba Group(阿里巴巴集团)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
School of Education, Zhejiang Normal University(浙江师范大学教育学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.CV
Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning
通过跨模态语义不一致性学习检测企业AI洗钱
Zhanjie Wen, Jingqiao Guo
机构
*
School of Economics and Trade, Guangdong University of Finance(广东金融学院经济贸易学院)
;
Department of Computer Science, Faculty of Science, Hong Kong Baptist University(香港浸会大学理学院计算机科学系)
PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation via Multi-Agent LLMs
PosterGen:基于多智能体LLM的美观化论文到海报生成
Zhilin Zhang, Xiang Zhang, Jiaqi Wei, Yiwei Xu, Chenyu You
机构
*
Stony Brook University(石溪大学)
;
New York University(纽约大学)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Zhejiang University(浙江大学)
;
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构
*
Shandong University(山东大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
;
New York University(纽约大学)
;
Xi'an Jiaotong University(西安交通大学)
;
Australian National University(澳大利亚国立大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.AI