How an AI Coding Agent Burned 301 Million Tokens in 24 Hours: Postmortem & Pre-Action Fix

โ€ข
2 min read
Igor Ganapolsky
Founderยท
How an AI Coding Agent Burned 301 Million Tokens in 24 Hours: Postmortem & Pre-Action Fix
Article

Autonomous AI coding agents promise 10x developer leverage. But without runtime execution guardrails, a single failing tool call can send an agent into an exponential token-burning compaction spiral.

The Incident: 301,424,414 Tokens Burned in 24 Hours

Using the HALO (Hierarchical Agent Loop Optimizer) diagnostic engine across our local development workstation, we captured OpenInference telemetry across 8 agent sessions over a 24-hour cycle. The ground truth was alarming:

  • Total Tokens Burned: 301,424,414 input tokens.
  • Runaway Session Culprit: Session a3de7a8b burned 240,463,806 tokens alone.
  • Loop Profile: 1,727 Bash tool executions and 168 context compactions.

Why Agents Loop

When an agent runs a tool that fails (such as attempting to read a file that does not exist, getting an EPERM permission error on a directory, or encountering a git merge conflict), standard agent harnesses dump the raw stack trace into the prompt context.

The model attempts the action again with trivial alterations. Context size rapidly blows past 100k tokens, triggering expensive compaction passes while repeating the identical failure over and over.

The Top Recurring Failure Taxonomies Harvested

  1. tool:Read (6 trace failures): Target file does not exist in working directory.
  2. tool:Write (6 trace failures): EPERM: operation not permitted on sandboxed directories.
  3. tool:Bash (4 trace failures): Repeated git merge conflicts retried in a tight loop.
  4. tool:mcp__screenpipe (2 trace failures): Service offline / port 3030 connection refused.
  5. tool:mcp__browseros-neo__act (2 trace failures): Schema parameter mismatch (unknown field action).

The Fix: Pre-Action Interdiction at the Tool Boundary

Instead of paying a frontier LLM zsh.05 per tool call to inspect actions post-hoc, ThumbGate (thumbgate on npm) compiles recurring trace failure clusters directly into sub-millisecond, CPU-local PreToolUse pre-validations:

  1. Pre-Action File Existence Check: Before sending a Read call to the LLM, ThumbGate checks file existence locally. Missing paths fail-fast in 0.05ms with 0 tokens burned.
  2. Sandbox Permission Guard: Writes targeting non-writable directories are interdicted before OS EPERM exceptions occur.
  3. Fail-Fast Retry Cap: Hard-caps repeated identical failures at 2 attempts, cutting off runaway compaction loops before they hit frontier models.
// ThumbGate PreToolUse: Interdict before remote inference tax
const { evaluateHaloGuardrails } = require('thumbgate/scripts/halo-harness-optimizer');

// Fails fast at 0 tokens and <1ms latency
const verdict = evaluateHaloGuardrails('Read', { file_path: '/missing/file.ts' });
// { ok: false, code: 'HALO_FILE_NOT_FOUND', reason: 'Target file does not exist...' }

Stop the Leak

If your background agents are eating your API budget and CI runners:

Comments (0)

Join the discussion by logging into your account.

No comments yet. Be the first to comment!

Igor Ganapolsky

Builder of AI agent governance and safety systems - such as ThumbGate.ai and ThumbGate.app

Subscribe to Igor Ganapolsky's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.