We address the problem of real-time remote tracking of a finite-state symmetric Markov source in an energy harvesting status update system. Different from most existing works that consider perfect feedback channels, we consider that both forward and feedback channels are error-prone. Our goal is to find a transmission policy that minimizes the time average of a (generic) distortion function subject to an energy limitation constraint. Since the feedback channel is error-prone, the transmitter only has partial knowledge about the estimate of the source state at the sink. Therefore, we model the problem as a partially observable Markov decision process (POMDP) and then cast it as a belief-MDP problem. To solve the problem, we develop a learning-based transmission policy by integrating R-learning into the deep Q-learning framework. Simulation results show the effectiveness of the proposed policy compared to two benchmark policies.
Funding Agencies|Infotech Oulu; Research Council of Finland [323698]; 6G Flagship program [346208]; Swedish Research Council [2022-03664]; HPY Research Foundation; Nokia Foundation