Optimize AI Agent Token Costs: Caching, Compression & Model Tiering
Optimize AI Agent Token Costs: Caching, Compression & Model Tiering: Designed as a zero-dependency, open-source TypeScript architecture under the MIT License with native Model Context Protocol (MCP) support and deterministic phase state machines.
- Prompt caching reduces repeated system prompt costs by up to 90%
- Context compaction keeps sessions economical over long tasks
- Route simple tasks to smaller models, complex to large models
- Tool output truncation prevents context bloat from large file reads
Anthropic Prompt Caching
Claude's prompt caching lets you cache the system prompt and large static context blocks. Subsequent requests that hit the cache cost 90% less for cached tokens. Enable it by structuring your system prompt as a large static block and using Anthropic's cache control headers.
For a typical coding agent with a 4,000-token system prompt running 100 turns per session, prompt caching saves approximately $0.36 per session at Claude Sonnet rates.
Put rarely-changing content first in your system prompt
Anthropic caches from the beginning of the context. Keep dynamic content (current task, file snippets) at the end to maximize cache hits on the static system instructions.
Dynamic Model Tiering
Not every agent task needs a frontier model. Build a tiered routing strategy:
- Simple tasks (rename a variable, fix a typo, add a comment): Use
claude-3-haikuorgpt-4o-miniat 10-20x lower cost. - Standard tasks (write a function, refactor a module): Use
claude-3-5-sonnetorgpt-4o. - Complex tasks (architect a system, debug a subtle race condition): Use
claude-3-7-sonnetoro3.
Smoke Monkey\'s provider config can be set dynamically per agent instance, letting you implement this tiering logic in your orchestrator.
Frequently Asked Questions
Q:How much do context compaction savings add up in practice?
For a 10-turn agent session that accumulates 30,000 tokens of context, compaction can trim it to 8,000 tokens before the next turn, saving roughly 70% on input token costs for the remainder of the session.
Q:Can I set a token budget limit to prevent runaway costs?
Yes. Subscribe to the model.response event, accumulate total token counts, and call agent.stop() when you hit your budget threshold. The agent will complete its current tool call and halt gracefully.
Related Alternatives & Comparisons
Build with Smoke Monkey Harness
Zero dependencies. 24 built-in tools. Human-in-the-loop safety. 100% open source under the MIT License.