On Memory, Inference, and Overcapacity in AI

AI LLM inference edge-computing

Large language models and transformers are very, very powerful.

However, my claim is that a large portion of what LLM inferences need by companies and corporations, as well as end consumers in their day-to-day lives, is something that is constrained by memory bandwidth and memory capacities.123