ALEXANDRIA, Va., July 14 -- United States Patent no. 12,682,211, issued on July 14, was assigned to Character Technologies Inc. (Palo Alto, Calif.).
"Optimizing key value cache for large language model inference" was invented by Bowen Liang (Sunnyvale, Calif.), Noam Mordechai Shazeer (Palo Alto, Calif.) and Myle Ott (New York).
According to the abstract* released by the U.S. Patent & Trademark Office: "An input sequence is received from a client device. Large language model inference is performed by processing the input sequence through a series of transformer layers to generate one or more tokens including by performing hybrid attention, multi-query attention, and cross-layer key value sharing. The one or more generated tokens are provid...