首页 | 官方网站   微博 | 高级检索  
     

场景图谱驱动目标搜索的多智能体强化学习
引用本文:陆升阳,赵怀林,刘华平.场景图谱驱动目标搜索的多智能体强化学习[J].智能系统学报,2023,18(1):207-215.
作者姓名:陆升阳  赵怀林  刘华平
作者单位:1. 上海应用技术大学 电气与电子工程学院,上海 201418;2. 清华大学 计算机科学与技术系,北京 100084
摘    要:针对强化学习在视觉语义导航任务中准确率低,导航效率不高,容错率太差,且部分只适用于单智能体等问题,提出一种基于场景先验的多智能体目标搜索算法。该算法利用强化学习,将单智能体系统拓展到多智能体系统上将场景图谱作为先验知识辅助智能体团队进行视觉探索,利用集中式训练分布式探索的多智能体强化学习的方法以大幅度提升智能体团队的准确率和工作效率。通过在AI2THOR中进行训练测试,并与其他算法进行对比证明此方法无论在目标搜索的准确率还是效率上都优先于其他算法。

关 键 词:多智能体  强化学习  视觉语义导航  场景图谱  先验知识  分布式探索  集中式训练  目标搜索

Multi-agent reinforcement learning for scene graph-driven target search
LU Shengyang,ZHAO Huailin,LIU Huaping.Multi-agent reinforcement learning for scene graph-driven target search[J].CAAL Transactions on Intelligent Systems,2023,18(1):207-215.
Authors:LU Shengyang  ZHAO Huailin  LIU Huaping
Affiliation:1. School of electrical and Electronic Engineering, Shanghai Institute of Technology, Shanghai 201418, China;2. Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China
Abstract:To solve the problems of reinforcement learning in the visual semantic navigation task, such as low accuracy, low navigation efficiency, poor fault tolerance rate, and the suitability of only some problems for a single agent, we propose a multi-agent target search algorithm based on scene prior. This algorithm extends the single-agent system to a multi-agent system through reinforcement learning. It mainly includes two aspects: first, a scene atlas is used as prior knowledge to assist the agent team in visual exploration; second, the multi-agent reinforcement learning method of centralized training and distributed exploration is used to greatly improve the accuracy and work efficiency of the agent team. Training tests in AI2THOR and comparison with other algorithms prove that this method is superior to other algorithms in target search accuracy and efficiency.
Keywords:multi-agent  reinforcement learning  visual semantic navigation  scene graph  prior knowledge  distributed exploration  centralized training  target search
点击此处可从《智能系统学报》浏览原始摘要信息
点击此处可从《智能系统学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司    京ICP备09084417号-23

京公网安备 11010802026262号