Supervisory control of partially observed discrete event systems based on a reinforcement learning
Toshimitsu Ushio, T. Yamasaki
Abstract
Toshimitsu Ushio, T. Yamasaki
Abstract
In discrete event systems, the supervisor controls events to satisfy the control specifications given by formal languages. However a precise description of the specifications and the discrete event systems is required for constructing the supervisor. So, this paper proposes a method to construct a supervisor based on a reinforcement learning for partially observed discrete event systems. In the proposed method, specifications are given by rewards, and an optimal supervisor is derived by considering rewards for the occurrence of events and disabling events. Moreover learning speed is accelerated by updating plural Q values. It is done by utilizing characteristics of a supervisory control. An efficiency of the proposed method is examined by computer simulation. The proposed method shows a new approach for applying a supervisory control in the case of implicit specifications and uncertain environment.
OpenAlex reports 14 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In discrete event systems, the supervisor controls events to satisfy the control specifications given by formal languages. However a precise description of the specifications and the discrete event systems is required for constructing the supervisor. So, this paper proposes a method to construct a supervisor based on a reinforcement learning for partially observed discrete event systems. In the proposed method, specifications are given by rewards, and an optimal supervisor is derived by considering rewards for the occurrence of events and disabling events. Moreover learning speed is accelerated by updating plural Q values. It is done by utilizing characteristics of a supervisory control. An efficiency of the proposed method is examined by computer simulation. The proposed method shows a new approach for applying a supervisory control in the case of implicit specifications and uncertain environment.
Key concepts: Supervisor, Supervisory control, Event (particle physics), Computer science, Plural, Control (management), Construct (python library), Supervisory control theory