AI has moved fast, faster than most of us expected, and the cost of keeping up is starting to show. What began as inexpensive, wide-open access to powerful cloud models has shifted into a landscape of rising token prices, stricter quotas, and unpredictable availability. More teams are starting to ask a question that would have sounded unrealistic a year ago: should we start hosting our AI models locally?
In this post, I’ll walk through why token costs and model access are becoming harder to rely on, why open models and local hardware are closing the gap, and what actually happened when I ran a full agentic workflow on my own machine, starting with the cost problem that kicked this off.





