Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020.
If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute
Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.
In the not-so-distant future, you’ll be able to book a trip from your front door to any other front door in the world in 1 ticket/payment.
Imagine not even having to switch vessels with the vehicle morphing from freeway driving to personal flight even to in-town or scenic driving (or flight) with different routes, handling and displays.
I’d like to float around in a bubble, like Glinda the good witch
If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.