ALEXANDRIA, Va., Sept. 15 -- United States Patent no. 12,737,693, issued on Sept. 15, was assigned to MITSUBISHI HEAVY INDUSTRIES LTD. (Tokyo).
"Learning device that performs reinforcement learning of a policy of an agent, learning method, and computer-readable storage medium" was invented by Sotaro Karakama (Tokyo) and Natsuki Matsunami (Tokyo).
According to the abstract* released by the U.S. Patent & Trademark Office: "A learning device is configured to perform reinforcement learning of a policy of an agent by self-play under a multi-agent environment. The multi-agent environment is an asymmetric environment in which at least one of a type of an action performed by the agent, a type of a state acquired by the agent, and a definition of ...