It's not about the model
There is a particular kind of satisfaction in watching a model load. The layers march across the terminal one at a time as llama.cpp offloads them to the GPU, the numbers climb, and then comes the line that tells you how much VRAM is left over. I remember sitting in front of that line the first evening the second card was in the machine. Two Radeon AI PRO R9700, sixty four gigabytes of VRAM between them, and llama.cpp reporting [XX.X] GB for the weights of a thirty billion parameter coder model at four bit quantization. Room to spare, I thought. The hard part is over. Everything from here is just usage.