I have run AWS infrastructure for years. When I started looking at how large language models actually get served, I hit a wall on something basic. Someone said "the model sits on the GPU.
Source: [Dev.to](https://dev.to/aj_aws_sa_sg/a-gpu-is-two-things-and-only-one-of-them-holds-your-model-4bbh)