Cost Optimization
8 min readUpdated: October 2026

Optimize AI Agent Token Costs: Caching, Compression & Model Tiering

Technical Review: Smoke Monkey Core Architecture Team
Tested on Node.js 18+ & BunTypeScript 5.x
Quick Answer & Executive Definition

Optimize AI Agent Token Costs: Caching, Compression & Model Tiering: Designed as a zero-dependency, open-source TypeScript architecture under the MIT License with native Model Context Protocol (MCP) support and deterministic phase state machines.

Key Architectural Takeaways

Anthropic Prompt Caching

Claude's prompt caching lets you cache the system prompt and large static context blocks. Subsequent requests that hit the cache cost 90% less for cached tokens. Enable it by structuring your system prompt as a large static block and using Anthropic's cache control headers.

For a typical coding agent with a 4,000-token system prompt running 100 turns per session, prompt caching saves approximately $0.36 per session at Claude Sonnet rates.

Put rarely-changing content first in your system prompt

Anthropic caches from the beginning of the context. Keep dynamic content (current task, file snippets) at the end to maximize cache hits on the static system instructions.

Dynamic Model Tiering

Not every agent task needs a frontier model. Build a tiered routing strategy:

  • Simple tasks (rename a variable, fix a typo, add a comment): Use claude-3-haiku or gpt-4o-mini at 10-20x lower cost.
  • Standard tasks (write a function, refactor a module): Use claude-3-5-sonnet or gpt-4o.
  • Complex tasks (architect a system, debug a subtle race condition): Use claude-3-7-sonnet or o3.

Smoke Monkey\'s provider config can be set dynamically per agent instance, letting you implement this tiering logic in your orchestrator.

Google Search Questions & Answers

Frequently Asked Questions

Q:How much do context compaction savings add up in practice?

For a 10-turn agent session that accumulates 30,000 tokens of context, compaction can trim it to 8,000 tokens before the next turn, saving roughly 70% on input token costs for the remainder of the session.

Q:Can I set a token budget limit to prevent runaway costs?

Yes. Subscribe to the model.response event, accumulate total token counts, and call agent.stop() when you hit your budget threshold. The agent will complete its current tool call and halt gracefully.

Related Alternatives & Comparisons

Build with Smoke Monkey Harness

Zero dependencies. 24 built-in tools. Human-in-the-loop safety. 100% open source under the MIT License.

npm install smoke-monkey-harness