ALEXANDRIA, Va., July 16 -- United States Patent no. 12,670,406, issued on June 30, was assigned to Intuit Inc. (Mountain View, Calif.).
"Confidence-based reward for group relative policy optimization in language models" was invented by Sagiv Antebi (Tel Aviv, Israel), Matan Vetzler (Tel Aviv, Israel), Shai Ardazi (Tel Aviv, Israel) and Ofir Ben Shoham (Tel Aviv, Israel).
According to the abstract* released by the U.S. Patent & Trademark Office: "Certain aspects of the disclosure provide a method for training a language model (LM) including: generating, using an LM, one or more outputs; computing a confidence score of an output of the one or more outputs based on a perplexity value of the output; determining, by a group relative policy op...