Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal
共识在策略上是不充分的:推理轨迹分歧作为知识表示信号
Michał Wawer, Jarosław A. Chudziak
机构
*
Laboratory of The New Ethos(新伦理实验室)
;
Warsaw University of Technology(华沙理工大学)
;
Institute of Computer Science(计算机科学研究所)
;
Faculty of Electronics and Information Technology(电子与信息技术学院)
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents
重新思考持续经验内化以实现自我进化的大语言模型智能体
Jingwen Chen, Wenkai Yang, Shengda Fan, Wenbo Nie, Chenxing Sun, Shaodong Zheng, Yangen Hu, Lu Pan, Ke Zeng, Yankai Lin
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院 Gallagher 学院)
;
School of Software, Beihang University(北航软件学院)
;
Meituan(美团)
机构
*
School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院)
LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment
LaVIDE: 通过地图-图像对齐的语言提示卫星变化检测
Shuguo Jiang, Fang Xu, Chuandong Liu, Hong Tan, Shengyang Li, Lei Yu, Wen Yang, Sen Jia, Gui-Song Xia
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
;
Technology and Engineering Center for Space Utilization and the Key Laboratory of Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与重点实验室)
;
School of Aeronautics and Astronautics, University of Chinese Academy of Sciences(中国科学院大学航空宇航学院)
;
School of Electronic Information, Wuhan University(武汉大学电子信息学院)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
机构
*
Centre for Artificial Intelligence Research(人工智能研究中心)
;
Department of Information and Communication Technology(信息与通信技术系)
;
University of Agder, Norway(阿格德大学,挪威)
Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling
通过非参数时空子目标建模实现逆强化学习
Adrian Šošić, Elmar Rueckert, Jan Peters, Abdelhak M. Zoubir, Heinz Koeppl
机构
*
Signal Processing Group(信号处理组)
;
Institute for Robotics and Cognitive Systems(机器人与认知系统研究所)
;
Autonomous Systems Labs(自主系统实验室)
;
Bioinspired Communication Systems(生物启发通信系统)
Sample-Efficient Policy Learning based on Completely Behavior Cloning
基于完全行为克隆的高效策略学习
Qiming Zou, Ling Wang, Ke Lu, Yu Li
机构
*
Department of Computer Science and Technology, Harbin Institute of Technology, China(计算机科学与技术系,哈尔滨工业大学,中国)
;
Department of Management Science and Engineering, Anhui University of Technology, China(管理科学与工程系,安徽理工大学,中国)
Edward Balaban, Stephen B. Johnson, Mykel J. Kochenderfer
机构
*
Intelligent Systems Division, NASA Ames Research Center(美国国家航空航天局阿姆斯研究中心智能系统部门)
;
Dependable System Technologies, LLC(可靠系统技术有限公司)
;
Jacobs ESSCA Group at NASA Marshall Space Flight Center(美国国家航空航天局马歇尔太空飞行中心Jacobs ESSCA小组)
;
Department of Aeronautics and Astronautics, Stanford University(斯坦福大学航空与航天系)