A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
A-ProS:通过多模型反馈实现可靠的自主编程
Anika Tabassum, Md Sifat Hossain, Md. Fahim Arefin, Tariqul Islam, Tarannum Shaila Zaman
机构
*
Dept. of Computer Science and Engineering, University of Dhaka(达卡大学计算机科学与工程系)
;
Dept. of Information Systems, University of Maryland, Baltimore County(马里兰大学巴尔的摩县信息学院)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县)
机构
*
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学)
;
Nankai University(南开大学)
;
The Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州))
;
School of Software Engineering, State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(软件学院,人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学)
机构
*
School of AI for Science, Peking University(科学人工智能学院,北京大学)
;
School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学)
;
School of Computer Science, Peking University(计算机科学学院,北京大学)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
改进方法而非提示:针对大语言模型的进化式 jailbreak 攻击合成
Yunhao Chen, Xin Wang, Juncheng Li, Yixu Wang, Jie Li, Yan Teng, Yingchun Wang, Xingjun Ma
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
机构
*
Institute for AI, Data Analysis and Systems (AIDAS) School of Engineering, Computing and Mathematics, Oxford Brookes University, UK(人工智能、数据分析和系统研究所(AIDAS)工程、计算与数学学院,英国奥克斯福德布鲁克斯大学)
Train the Trainers -- An Agentic AI Framework for Peer-Based Mental Health Support in Battlefield Environments
训练培训者——一种基于同伴的AI框架,用于战场环境中的同伴心理支持
Atmaram Yarlagadda, Eranga Bandara, Ross Gore, Anita H. Clayton, Preston Samuel, Christopher K. Rhea, Sachin Shetty, Ravi Mukkamala, Xueping Liang, Amin Hass, Abdul Rahman
机构
*
McDonald Army Health Center(麦克唐纳陆军健康中心)
;
Old Dominion University(老 Dominion 大学)
;
Blanchfield Army Community Hospital(布莱恩菲尔德陆军社区医院)
;
Department of Psychiatry and Neurobehavioral Sciences(精神病学与神经行为科学系)
;
University of Virginia School of Medicine(弗吉尼亚大学医学院)
;
Florida International University(佛罗里达国际大学)
;
Accenture Technology Labs(埃森哲技术实验室)
Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning
超越策略优化:一种数据整理飞轮用于稀疏奖励长周期规划
Yutong Wang, Pengliang Ji, Kaixin Li, Baolong Bi, Tao Feng, Guillaume Sartoretti
机构
*
Department of Mechanical Engineering, National University of Singapore(新加坡国立大学机械工程系)
;
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
;
School of Computing, National University of Singapore(新加坡国立大学计算机科学学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval
OPERA: 一种增强强化学习的协调规划-执行架构用于面向推理的多跳检索
Yu Liu, Yanbing Liu, Fangfang Yuan, Cong Cao, Youbang Sun, Kun Peng, Weizhuo Chen, Jianjun Li, Zhiyuan Ma
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
;
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
CommentsV2 of the article: - Added AdaLN-zero - Added table comparing JEPA-WMs with baselines with std translating per-seed variability only, no variability across epochs - Reordered figures in main body of the paper V3: added data scaling experiments, theoretical appendix section on autoregressive rollout, acceptance at TMLR