APPLIED RESEARCH

Development and investigation of an agent based on Deep Q-Network algorithm for control tasks in dynamic environments

Authors

  • Nikita S. Kurtash Peter the Great St. Petersburg Polytechnic University, 195251, St. Petersburg, Polytekhnicheskaya str., 29
  • Sofia D. Bryutova Peter the Great St. Petersburg Polytechnic University, 195251, St. Petersburg, Polytekhnicheskaya str., 29

How to cite

GOST Kurtash N. S., Bryutova S. D. Development and investigation of an agent based on Deep Q-Network algorithm for control tasks in dynamic environments // STROITEL'NYE I DOROZHNYE MASHINY. 2025. Vol. 69. No. 12. P. 136-144.
APA Kurtash, N. S. & Bryutova, S. D. (2025). Development and investigation of an agent based on Deep Q-Network algorithm for control tasks in dynamic environments. STROITEL'NYE I DOROZHNYE MASHINY, 69(12), 136-144.

Abstract

Deep reinforcement learning (Deep RL) is one of the most promising areas of machine learning. Deep reinforcement learning is an important area of machine learning that finds application in autonomous control, robotics, and gaming systems. This paper presents the implementation of the Deep Q-Network (DQN) algorithm for training an agent to control the Snake game. The methodology includes the formation of a compact vector representation of the environment state from 11 binary features, the use of a fully connected neural network to approximate the Q-function, the use of an experience replay mechanism with a buffer of 100,000 records, and an ε-greedy strategy to ensure a balance between exploring the environment and exploiting the knowledge gained. During the study, the agent was trained over 500 game episodes with different environment configurations. The results showed a steady increase in the average score from 0 to 23.7 during training and an average score of 25.5 during testing, which is 7.6% higher than the final training score. The maximum score achieved was 74 points. The results confirm the applicability of Deep RL for solving control problems in stochastic environments and demonstrate the ability of the DQN agent to generalize the learned strategy.

Keywords

reinforcement learning Deep Q-Network neural networks Q-learning game agents machine learning

References

Chen C., Ying V., Laird D. Глубокое Q-обучение с рекуррентными нейронными сетями: отчет о проекте // CS229 Final Project Report. Stanford University. 2016. С. 1–6. URL: http://cs229.stanford.edu/proj2016/report/ChenYingLairdDeepQLearningWithRecurrentNeuralNetwords-report.pdf (дата обращения: 26.12.2025).

Mnih V., Kavukcuoglu K., Silver D. и др. Обучение игре Atari с использованием глубокого обучения с подкреплением // arXiv.org. 2013. С. 1–9. URL: https://arxiv.org/abs/1312.5602 (дата обращения: 26.12.2025).

Mnih V., Kavukcuoglu K., Silver D. и др. Управление на уровне человека с использованием глубокого обучения с подкреплением // Nature. 2015. Т. 518. № 7540. С. 529–533. DOI: 10.1038/nature14236.

Osband I., Blundell C., Pritzel A., Van Roy B. Глубокое исследование среды с использованием Bootstrapped DQN // Advances in Neural Information Processing Systems (NeurIPS). 2016. Т. 29. С. 4026–4034. URL: https://arxiv.org/abs/1602.04621 (дата обращения: 26.12.2025).

Schaul T., Quan J., Antonoglou I., Silver D. Приоритетное воспроизведение опыта // arXiv.org. 2015. С. 1–21. URL: https://arxiv.org/abs/1511.05952 (дата обращения: 26.12.2025).

Schulman J., Wolski F., Dhariwal P. и др. Алгоритмы проксимальной оптимизации политики // arXiv.org. 2017. С. 1–12. URL: https://arxiv.org/abs/1707.06347 (дата обращения: 26.12.2025).

Sutton R.S., Barto A.G. Обучение с подкреплением: введение. 2-е изд. Кембридж: MIT Press, 2018. 548 с. URL: http://incompleteideas.net/book/the-book-2nd.html (дата обращения: 26.12.2025).

Van Hasselt H., Guez A., Silver D. Глубокое обучение с подкреплением с использованием Double Q-learning // Proceedings of the AAAI Conference on Artificial Intelligence. 2016. Т. 30. № 1. С. 2094–2100. URL: https://arxiv.org/abs/1509.06461 (дата обращения: 26.12.2025).

Watkins C.J.C.H., Dayan P. Q-learning // Machine Learning. 1992. Т. 8. № 3–4. С. 279–292. DOI: 10.1007/BF00992698.

Yuwono F., Yen G.P., Christopher J. Гоночные автомобили с автопилотом: применение глубокого обучения с подкреплением // arXiv.org. 2024. С. 1–8. URL: https://arxiv.org/abs/2410.22766 (дата обращения: 26.12.2025).

Metrics

272 views
0 downloads
Want to publish with us?
Submit an article

Machine-readable metadata

Similar Articles

1 2 3 > >> 

You may also start an advanced similarity search for this article.