Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching
两全其美:通过统一离散流匹配实现多模态推理与生成
Onkar Susladkar, Tushar Prakash, Gayatri Deshmukh, Kiet A. Nguyen, Jiaxun Zhang, Adheesh Juvekar, Tianshu Bao, Lin Chai, Sparsh Mittal, Inderjit S Dhillon, Ismini Lourentzou
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Google Research(谷歌研究)
;
Sony Research, India(索尼印度研究)
;
Indian Institute of Technology, Roorkee(印度理工学院罗尔基分校)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents
先感知后推理:一种用于高效可靠主动移动代理的预推理感知框架
Zhijie Ding, Weinan Hong, Zicheng Zhu, Lei Li, Dezhi Kong, Hao Wang, Peng Zhou, Xuchu Jiang, Jiaming Xu
机构
*
HyperAI Team, Xiaomi Corporation(HyperAI团队,小米公司)
;
Zhongnan University of Economics and Law(中南财经政法大学)
;
Jilin University(吉林大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
UNISON: 通过深度LLM融合的统一声音生成与编辑框架
Zhaoqing Li, Haoning Xu, Jingran Su, Yaofang Liu, Zhefan Rao, Huimeng Wang, Jiajun Deng, Tianzi Wang, Zengrui Jin, Rui Liu, Haoxuan Che, Xunying Liu
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
City University of Hong Kong(香港城市大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Tsinghua University(清华大学)
;
Huawei Research Hong Kong(华为香港研究)
CommentsAccepted to ICML 2026. We have also released an updated version of the benchmark, WISE_Verified. Please refer to https://github.com/PKU-YuanGroup/WISE for the latest version
Reversible Superdense Ordering of Tetragonal Lithium in a Layered Material
层状材料中四方锂的可逆超密有序
Natalie L. Williams, Stephen D. Funni, Sihun Lee, Drake Niedzielski, Mingjia Fang, Josh Leeman, Ratnadwip Singha, Saif Siddique, Tyler Hendee, Shiyu Xu, Jeffrey Kaaret, Giovanni Sartorello, Nicole A. Benedek, Leslie M. Schoop, Lynden A. Archer, Tomás A. Arias, Judy J. Cha
GLENS: Global Search via Learning from Solver Iterates with Diffusion Models
GLENS: 通过扩散模型从求解器迭代中学习进行全局搜索
Anjian Li, Bartolomeo Stellato, Ryne Beeson
机构
*
Department of Electrical and Computer Engineering, Princeton University(电气工程与计算机科学系,普林斯顿大学)
;
Department of Operations Research and Financial Engineering, Princeton University(运筹学与金融工程系,普林斯顿大学)
;
Department of Mechanical and Aerospace Engineering, Princeton University(机械与航空航天工程系,普林斯顿大学)
机构
*
Tsinghua University(清华大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
East China Normal University(华东师范大学)
;
Tongji University(同济大学)
;
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments
FRED:面向洪水道路环境的多模态自动驾驶数据集
Connor Malone, Sebastien Demmel, Sebastien Glaser
机构
*
Queensland University of Technology(昆士兰理工大学)
;
ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3)(农村和偏远地区自动化车辆培训中心(AVR3))
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
回到柏拉图的洞穴:大规模检验跨模态表示收敛性
A. Sophia Koepke, Daniil Zverev, Shiry Ginosar, Alexei A. Efros
机构
*
UC Berkeley(伯克利大学)
;
Technical University Munich, MCML(慕尼黑技术大学)
;
University of Tübingen, Tübingen AI Center(图宾根大学)
;
Toyota Technical Institute at Chicago(芝加哥丰田技术研究所)
The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset
自动驾驶的未来之路:KITScenes多模态数据集
Richard Schwarzkopf, Fabian Immel, Alexander Blumberg, Jonas Merkert, Nils Rack, Kaiwen Wang, Fabian Konstantinidis, Julian Truetsch, Carlos Fernandez, Annika Bätz, Kevin Rösch, Marlon Steiner, Willi Poh, Yinzhe Shen, Royden Wagner, Felix Hauser, Dominik Strutz, Jaime Villa, Gleb Stepanov, Holger Caesar, Ömer Şahin Taş, Frank Bieder, Jan-Hendrik Pauls, Christoph Stiller
机构
*
FZI Research Center for Information Technology(弗劳恩霍夫信息技术研究中心)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
University Charles III of Madrid(马德里第三大学)
;
Delft University of Technology(代尔夫特理工大学)
Multi-Modal Machine Learning for Breast Cancer Recurrence Prediction
多模态机器学习用于乳腺癌复发预测
Jiahao Shao, Xudong Wang, Anam Nawaz Khan, Christopher Brett, Xueping Li, Bing Yao
机构
*
Department of Industrial & Systems Engineering, The University of Tennessee, Knoxville, TN 37996, USA(工业与系统工程系,田纳西大学,诺克斯维尔,TN 37996,美国)
;
The University of Tennessee Medical Center, Knoxville, TN 37920, USA(田纳西大学医学中心,诺克斯维尔,TN 37920,美国)
A Fast Screening Approach for High-dimensional Outcomes and High-dimensional Predictors
高维结果与高维预测变量的快速筛选方法
Hongju Park, Zhenyao Ye, Shuo Chen
机构
*
School of Pharmacy, University of North Carolina, Chapel Hill(北卡罗来纳大学药学院)
;
Maryland Psychiatric Research Center, School of Medicine, University of Maryland(马里兰大学医学院精神病研究中心)
SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series
SagaQA:面向电视剧长篇叙事理解的多跳推理基准
Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen
机构
*
IRIT, University of Toulouse, France(法国图卢兹大学IRIT中心)
;
Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局)
;
CNRS, IRIT, France(法国CNRS与IRIT)