Ollama.com just started a 'transparency' campaign on token expenditure. Two months ago I left Claude for GLM+K+Deepseek clankers and have never once gotten close to blow the same budget on open-source paid inferences. Bare metal is the way, otherwise inference providers is the only comprimise. You can't measure quality/cost otherwise.