Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
面向视觉原生多模态深度搜索智能体的在策略数据演化
Shijue Huang, Hangyu Guo, Guanting Dong, Chenxin Li, Junting Lu, Xinyu Geng, Zhaochen Su, Zhenyu Li, Shuang Chen, Hongru Wang, Yi R. Fung
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Renmin University of China(中国人民大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
;
Tsinghua University(清华大学)
;
University of Edinburgh(爱丁堡大学)
Compact Task-Aligned Imitation Learning for Laboratory Automation
紧凑的任务对齐模仿学习用于实验室自动化
Kanata Suzuki, Hanon Nakamura, Kana Miyamoto, Tetsuya Ogata
机构
*
Spatial Robotics Research Center, Fujitsu Limited.(富士通株式会社空间机器人研究中心)
;
Faculty of Science and Engineering, Waseda University(早稻田大学理工学部)
;
National Institute of Advanced Industrial Science and Technology(国家工业科学与技术研究院)
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
通过特定任务自我报告揭示隐藏模型行为
Taras Kutsyk, Bartosz Zieliński
机构
*
Jagiellonian University, Faculty of Mathematics and Computer Science(雅盖隆大学数学与计算机科学系)
;
Jagiellonian University, Doctoral School of Exact and Natural Sciences(雅盖隆大学精确与自然科学研究博士学院)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
OmniAD:基于多模态推理的工业异常检测与理解
Shifang Zhao, Yiheng Lin, Lu Han, Yao Zhao, Yunchao Wei
机构
*
Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院)
;
Visual Intelligence + X International Joint Laboratory of the Ministry of Education(教育部视觉智能+X国际合作实验室)
;
Key Laboratory of Noise and Vibration Research, Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所噪声与振动重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA
Q-VGM: 基于Q引导的值梯度匹配的流匹配VLA策略
Ziqian Wang, Yitian Liu, Xingjian Mao, Minqian Wang, Yao Mu
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
University of Michigan, Ann Arbor(密歇根大学安娜堡分校)
;
University of Electronic Science and Technology of China(电子科技大学)
CommentsWithdrawn by the authors to allow for substantial revision in light of valuable reviewer feedback. A thoroughly revised version may be submitted in the future
机构
*
Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
;
State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学)
;
Department of Automation, Tsinghua University(自动化系,清华大学)
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
超越狄拉克 delta:缓解强化微调中的多样性崩溃以实现多功能图像生成
Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
机构
*
Nanjing University(南京大学)
;
Australia National University(澳大利亚国立大学)
;
Wuhan University(武汉大学)
;
National University of Singapore(新加坡国立大学)
;
University of Technology Sydney(技术科技大学)
Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
Flick:在多任务低资源语言中使用K感知中间学习进行少标签文本分类
Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi, Imran Razzak
机构
*
University of New South Wales Australia(新南威尔士大学(澳大利亚))
;
University of Technology Sydney Australia(悉尼科技大学(澳大利亚))
;
University of Southampton United Kingdom(南安普顿大学(英国))
;
Macquarie University Australia(麦考瑞大学(澳大利亚))
;
MBZUAI UAE(阿联酋人工智能研究所)