[HN] Ask HN: Burning $100K/week on LLM tokens – what are you doing to cut costs?

Summary

A company running AI coding agents for trading infrastructure is spending over $100K weekly on LLM tokens, primarily due to context window bloat from long tool traces and repeated system prompts. Despite efforts like prompt caching, using smaller models for routine tasks, and aggressive conversation truncation, they are still seeking more effective cost-cutting strategies. This scenario underscores the significant operational cost challenges associated with deploying large language models at scale.

Continue Reading

Explore related coverage about community news and adjacent AI developments: [r/ML] [D] MYTHOS-INVERSION STRUCTURAL AUDIT, [r/LocalLLaMA] karpathy / autoresearch, [r/ML] [R] Agentic AI and Occupational Displacement: A Multi-Regional Task Exposure Analysis (236 occupations, 5 US metros), [r/ML] Building behavioural response models of public figures using Brain scan data (Predict their next move using psychological modelling) [P].

[HN] Ask HN: Burning $100K/week on LLM tokens – what are you doing to cut costs?

Summary

Continue Reading

Related Articles

Comments