Dataflow in Generative LLMs
This is a rough sketch of the inference loop in an LLM, as it predicts one token at a time. I find new research hard to follow unless it's clear where the change plugs in. Code helps, even pseudo-code that doesn't run, since it makes the primitives and the control-flow explicit. It's also been helpful to ground explanations from AI assistants, in the style I prefer.
The snippets below only convey the shape, and aren't real implementations. For more background on why tranformers are the way they are, there are a lot of good sources. My favorite piece of exposition is on Cosma Shalizi's blog, which you should go read!