ZYVOPMulti-Platform Sync
SeriesAI NewsWhy ZyVOPJoin Discord
LoginGet Started
ZYVOPMulti-Platform Sync

The Developer Publishing Hub. Write once, publish everywhere, and make your work citation-ready with built-in SEO, AEO, and GEO discovery support. Zero reader paywalls.

Content

  • Categories
  • Tags
  • Badges
  • Leaderboard
  • Write Article
  • Newsletter

Company

  • About Us
  • Why ZyVOP
  • Changelog
  • Compare Platforms
  • Hashnode vs ZyVOP
  • DEV vs ZyVOP
  • Developer API & CLI
  • Author Handbook
  • Contact

Connect

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • DMCA Policy
  • Code of Conduct

© 2026 ZyVOP. Developer Publishing Hub.

Zero paywalls · Full content ownership
All systems operational
HomeArchitectureThe Harness is the Moat: Building a Deterministic Agent Runtime with Context Pruning
Architecture

The Harness is the Moat: Building a Deterministic Agent Runtime with Context Pruning

Lê Đức Minh
Lê Đức Minh
AI Engineer
September 16, 2026
6 min read
The Harness is the Moat: Building a Deterministic Agent Runtime with Context Pruning
#AI
👍1

Every agent looks brilliant in a five-turn demo. You give it a neat prompt, wire up two mock tools, ask it to look up a customer ID and draft a confirmation email, and it flies through without missing a step. You feel like you have automated half your workday before lunch.

Then you deploy it into a live staging environment with thirty tools and an actual database, and watch it fall apart by turn twenty-eight.

The failure is rarely a profound reasoning lapse. The model does not suddenly forget English or lose its grasp of basic logic. What actually happens is dirtier: twenty-seven tool calls have dumped eighty kilobytes of raw JSON payloads, terminal outputs, and HTTP error traces into the message history. The original system prompt has been squeezed out of the attention spotlight. The model hallucinates a parameter that never existed in your API schema, chokes on a database timeout, and enters an infinite retry loop that burns twelve dollars of API credits before your rate limiter kicks in.

Autonomous LLM agents fail in multi-turn production environments primarily due to context poisoning and unconstrained state divergence, not underlying model reasoning deficiencies. A production-grade agent harness mitigates this by replacing raw message arrays with deterministic state machines, selective semantic context pruning, strict output schemas, and transactional checkpoint rollbacks.

The core lesson from two years of shipping production agents is simple: the model is a commodity, but the harness is the moat. If you want durability over forty turns, you have to stop treating an agent loop as a growing conversation and start treating it as a managed operating system process.


Why Naive Message Arrays Poison the Attention Window

Most starter agent implementations use an append-only array. Step one runs, you append the user prompt. Step two calls a tool, you append the tool call. Step three receives tool output, you append the raw tool payload.

After fifteen turns, your message list looks like a junk drawer:

Message Index

Message Type

Content Size

Operational Value at Step 25

0

System Prompt

1.8 KB

Critical (invariant constraints)

1

User Request

0.2 KB

Critical (original user objective)

2-14

Step 1-6 Tool Calls

8.4 KB

Zero (completed intermediate actions)

15

Raw SQL Output

42.1 KB

Zero (only 2 rows mattered)

16-24

Step 7-12 Scratchpad

14.2 KB

Low (noise and abandoned paths)

25

Current Step

0.5 KB

Active working context

By step twenty-five, over 80% of the active context is dead weight: raw database dumps, obsolete API errors from retries, and verbose scratchpad reasoning from tasks resolved twenty minutes ago.

This causes two catastrophic failures:

  1. Recency Bias Dilution: The model pays closer attention to the massive SQL dump twenty tokens away than to the safety constraint defined in the system prompt sixty thousand tokens back.

  2. Context-Window Bloat: Every single turn bills you for re-reading those eighty kilobytes. As I discussed in Why your coding agent's bill grows faster than the chat, the token bill scales quadratically with turn count on naive message lists.


The Expedition Model: Separating Pack Weight from Base Camps

When you climb a mountain, you do not carry every empty water bottle, wrapper, and broken crampon from camp one up to the summit. You carry essentials, leave cache points behind you, and only pack what is needed for the current pitch.

In an agent runtime, this means splitting execution state into three distinct tiers:

Tier 1: Invariant State

This never changes during the run. It contains your core operational constraints, approved tool signatures, and the original user instruction. In our architecture, this sits at the very beginning of the prompt to maximize KV cache reuse across turns.

Tier 2: The Persistent Task Ledger

This is a structured summary table maintained outside the conversation array. When an agent runs a database query to find a customer record, the raw result array (forty kilobytes) is parsed immediately. The ledger records: Customer ID: 89412, Status: Active, Plan: Enterprise. The raw database output is dropped from the prompt entirely.

Tier 3: Active Working Memory

Only the current step and the immediate previous turn retain full raw tool outputs. Once step N completes and is validated, its tool payload is compressed into a one-line fact for the ledger, and the raw payload is purged from the next inference prompt.


Implementing Semantic Compaction in Python

Here is the exact pattern we use to prune tool outputs without losing state. Instead of handing raw responses back to the model, the harness forces compaction:

from dataclasses import dataclass, field
from typing import Any, Dict, List


@dataclass
class AgentState:
    original_objective: str
    invariant_rules: List[str]
    facts_ledger: Dict[str, Any] = field(default_factory=dict)
    active_history: List[Dict[str, str]] = field(default_factory=list)

    def record_fact(self, key: str, value: Any) -> None:
        """Store verified operational fact in persistent ledger."""
        self.facts_ledger[key] = value

    def append_turn(self, role: str, content: str) -> None:
        self.active_history.append({"role": role, "content": content})
        # Keep only the last 4 messages in raw working memory
        if len(self.active_history) > 4:
            self.active_history = self.active_history[-4:]

    def assemble_prompt_messages(self) -> List[Dict[str, str]]:
        """Assembles prompt with invariant rules, ledger, and recent turns."""
        ledger_lines = [
            f"- {k}: {v}" for k, v in self.facts_ledger.items()
        ]
        ledger_block = (
            "CURRENT VERIFIED FACTS:\n" + "\n".join(ledger_lines)
            if ledger_lines
            else "No verified facts recorded yet."
        )

        system_content = (
            f"OBJECTIVE: {self.original_objective}\n\n"
            f"RULES:\n" + "\n".join(f"- {r}" for r in self.invariant_rules) + "\n\n"
            f"{ledger_block}"
        )

        messages = [{"role": "system", "content": system_content}]
        messages.extend(self.active_history)
        return messages

Notice what happens here: the prompt length remains bounded even if the agent runs for a hundred steps. The system prompt remains pinned to message index zero, the ledger grows slowly by a few dozen tokens per milestone, and the raw churn of tool calls never exceeds four messages.


Transactional Tool Calls and Checkpoint Rollbacks

What happens when an agent calls a bash script that errors out, or attempts to write an invalid file path?

In naive agent loops, the runtime appends the error traceback to the chat: FileNotFoundError: [Errno 2] No such file or directory. The model sees the failure, tries to explain itself, apologizes, tries another broken path, and fills twenty turns arguing with its own error log.

In a deterministic harness, tool executions are transactional. If a tool call fails validation:

  1. The failed action and its error log are not appended to the primary active history.

  2. The harness checks an internal state machine. If the failure is recoverable, it injects a concise single-line correction: Error: path /var/data/out.csv not found. Valid directories: /var/data/raw, /var/data/processed.

  3. If the failure indicates loop thrashing (three consecutive identical failures), the harness triggers a checkpoint rollback: it reverts the working memory to the state prior to the first failed tool call and forces a different execution branch.

As detailed in What DeepSeek's Open-Source Agent Harness Gets Right, treating the harness as a state machine with hard transition guards is what separates reproducible systems from probabilistic toys.


Share your thoughts in the comments — I'd love to hear how this technology is impacting your industry.

Follow me: LinkedIn | GitHub


Three Hard Heuristics for Production Agent Runtimes

When architecting agent systems for multi-step tasks, enforce these operational rules:

  1. Raw Payloads Never Enter Context Unfiltered: If a tool returns a JSON array with twenty keys, run a schema filter or extractor first. Pass only the keys requested by the current step.

  2. Decouple Storage from Context: Store files, full documents, and database rows in an external store (SQLite, Redis, or disk). Pass identifiers and excerpt ranges, not blob contents.

  3. Hard Cap on Loop Depth: Every agent execution must have a hard boundary on both clock time and turn count. If an agent cannot reach a verified milestone in fifteen steps, a human escalation or deterministic abort is strictly superior to allowing twenty more hallucinated steps.

Building reliable agents does not require waiting for next year's model release. It requires building the engineering scaffolding around today's models that keeps their attention focused, their state verified, and their memory clean.


References

  • For the foundational four-pillar agent anatomy, read Prompt, Context, Harness, Loop: An Agent's Anatomy.

  • For structuring tool constraints with rigid ontologies instead of soft prompts, see AI Agent Guardrails: Why Ontologies Beat Prompting.

  • For measuring the exact cost mechanics of context bloat, refer to Why your coding agent's bill grows faster than the chat.


FAQ

  1. Why do autonomous LLM agents fail after multiple turns?

Autonomous agents fail primarily due to context poisoning and state divergence. As multiple tool executions dump raw payloads and error logs into the message history, the model loses sight of initial invariant constraints and begins hallucinating parameters or entering infinite retry loops

  1. What is semantic context compaction in an agent harness?

Semantic context compaction is an architectural pattern that splits agent state into invariant system rules, a persistent factual ledger, and transient working memory. Once a tool execution completes, its raw payload is pruned and reduced to verified facts, preventing context-window bloat

Comments (0)

Login to post a comment.

Lê Đức Minh
Lê Đức Minh

AI Engineer

Just Another Dev Guy

Subscribe to Lê Đức Minh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.

More from Lê Đức Minh

View profile

What DeepSeek's Open-Source Agent Harness Gets Right

A few weeks ago, my dad and I spent an afternoon at the kitchen table trying to figure out what an "AI agent" actually is.

8 minSep 16

I Tested 3 LLMs on Vietnamese Idioms. They Got It Backwards.

State‑of‑the‑art LLMs, even Vietnamese‑specialized ones, score below 50 % on the VIVID benchmark of 1,636 authentic Vietnamese idioms and fail completely on a hand‑picked set of six complex idioms. Th

6 minSep 5

AI Reads Vietnamese Slang as Angry. 3 Models Don't.

Benchmark reveals Llama-3.1-8B mislabels Vietnamese slang as angry, losing 20 F1 points, while three other LLMs correctly interpret social media comments.

6 minSep 5

The Real Vietnamese LLM Tokenizer Cost (It's Not 4.5x)

Discover why Vietnamese tokenizers cost only 1.05‑2.14× English tokens, not 4.5×, and how updating tokenizers can cut LLM expenses by a third.

7 minSep 5

A Simple Look at TokPress: A Compressor That Uses an LLM's Tokenizer

TokPress compresses tiny JSON log lines by tokenizing with OpenAI's o200k_base tokenizer, then applying LZ77 and rANS for smaller files.

2 minSep 1