场景图谱驱动目标搜索的多智能体强化学习 Multi-agent reinforcement learning for scene graph-driven target search期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

场景图谱驱动目标搜索的多智能体强化学习

引用本文：	陆升阳,赵怀林,刘华平.场景图谱驱动目标搜索的多智能体强化学习[J].智能系统学报,2023,18(1):207-215.

作者姓名：	陆升阳赵怀林刘华平

作者单位：	1. 上海应用技术大学电气与电子工程学院，上海 201418;2. 清华大学计算机科学与技术系，北京 100084

摘要：	针对强化学习在视觉语义导航任务中准确率低，导航效率不高，容错率太差，且部分只适用于单智能体等问题，提出一种基于场景先验的多智能体目标搜索算法。该算法利用强化学习，将单智能体系统拓展到多智能体系统上将场景图谱作为先验知识辅助智能体团队进行视觉探索，利用集中式训练分布式探索的多智能体强化学习的方法以大幅度提升智能体团队的准确率和工作效率。通过在AI2THOR中进行训练测试，并与其他算法进行对比证明此方法无论在目标搜索的准确率还是效率上都优先于其他算法。
关键词：	多智能体强化学习视觉语义导航场景图谱先验知识分布式探索集中式训练目标搜索
Multi-agent reinforcement learning for scene graph-driven target search

LU Shengyang,ZHAO Huailin,LIU Huaping.Multi-agent reinforcement learning for scene graph-driven target search[J].CAAL Transactions on Intelligent Systems,2023,18(1):207-215.

Authors:	LU Shengyang ZHAO Huailin LIU Huaping

Affiliation:	1. School of electrical and Electronic Engineering, Shanghai Institute of Technology, Shanghai 201418, China;2. Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China

Abstract:	To solve the problems of reinforcement learning in the visual semantic navigation task, such as low accuracy, low navigation efficiency, poor fault tolerance rate, and the suitability of only some problems for a single agent, we propose a multi-agent target search algorithm based on scene prior. This algorithm extends the single-agent system to a multi-agent system through reinforcement learning. It mainly includes two aspects: first, a scene atlas is used as prior knowledge to assist the agent team in visual exploration; second, the multi-agent reinforcement learning method of centralized training and distributed exploration is used to greatly improve the accuracy and work efficiency of the agent team. Training tests in AI2THOR and comparison with other algorithms prove that this method is superior to other algorithms in target search accuracy and efficiency.

Keywords:	multi-agent reinforcement learning visual semantic navigation scene graph prior knowledge distributed exploration centralized training target search

	点击此处可从《智能系统学报》浏览原始摘要信息
	点击此处可从《智能系统学报》下载全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏