Pi 1.0 Is Here: What Changed in the Minimal Coding Agent (and What Pi Durable Is For)

Earendil's minimal coding agent hits 1.0 with Codemode, built-in MCP, virtual models and a fullscreen TUI, plus Pi Durable for long-running agents.

•
10 min read
Ankit Singh
Software Engineer
Pi 1.0 Is Here: What Changed in the Minimal Coding Agent (and What Pi Durable Is For)

TL;DR

  • Earendil shipped Pi 1.0 on October 1, 2026. It's MIT licensed and installs with a one-line script.

  • The headline list: Codemode with built-in MCP support, virtual models, deferred tool loading, prompt cache warming, mid-conversation system messages, a new theme, and a fullscreen terminal UI by default.

  • The same day they released Pi Durable, an experimental framework for long-running agents that survive crashes. It doesn't replace the Pi you run in your terminal.

  • Most of those features actually landed over the last two months of 0.x releases. 1.0 is the "we're calling this stable" moment.

📝 How I put this together: I read the 1.0 announcement, the Pi Durable post, the MCP post that came out two days earlier, and the full changelog. I haven't lived with 1.0 for weeks, so there are no benchmarks or long-term impressions here. This is what's documented, plus my read on it.


So what is Pi, exactly?

Pi is a terminal-based coding agent. Earendil calls it an agent harness: the loop and tooling around a model that lets it read files, run commands, edit code and call tools.

Its whole personality is minimalism. The team says it waits for an idea to prove itself before adding it, and weighs every feature against the complexity it brings. The flip side is that Pi is built to be bent to your will: extensions, packages, custom providers, your own tools.

Earendil also says hundreds of thousands of people use Pi every week. That's their number, not something I could check, but it explains why they wanted a stable label on it.

Install it:

# macOS / Linux
curl -fsSL https://pi.dev/install.sh | sh
# Windows
powershell -c "irm https://pi.dev/install.ps1 | iex"

1.0 didn't arrive all at once

The announcement lists the new features as part of 1.0. If you read the changelog, though, most of them showed up earlier:

Release

Date

What landed

0.84.0

Aug 6

Fullscreen TUI mode (opt-in at the time), plus Mermaid and LaTeX rendering in transcripts

0.86.0

Sep 19

Prompt cache warming, transcript-aware prompt and tool changes

0.99.0

Sep 29

Codemode, MCP, tool search, experimental virtual models, classifier models, system theme

0.99.2

Sep 30

MCP servers stop slowing down the first prompt

1.0.0

Oct 1

Fullscreen by default, a leaner codemode, image generation from codemode scripts, "Sign in with Radius" in /login, MCP OAuth hardening

Even the version numbers tell a story: Pi went from 0.87.1 on September 22 to 0.99.0 a week later.

So 1.0.0 itself is a smaller diff than the feature list suggests. I don't mean that as a knock. For a 1.0, boring and stable is the point, and it's a clear signal to anyone who was waiting before building on it.


Codemode, and the MCP u-turn

This is the biggest philosophical shift. Pi's website used to say, proudly, that Pi doesn't support MCP. Mario Zechner wrote a whole post about why you might not need it. Now MCP is part of the core, and Earendil published a post two days before 1.0 titled "You Said No MCP!" to explain the change of heart.

Their short version: MCP in 2026 isn't the MCP of a year ago, and what Pi needed (a sandbox to orchestrate tools) happened to be what MCP needed too.

What Codemode actually is

Normally an agent calls one tool, gets the result back into its context, then calls the next. Codemode lets the model write a JavaScript script that calls many tools, loops over results and combines them. The script runs in a QuickJS sandbox on the harness side (where the agent loop runs), not where bash runs. Its state is kept in the session transcript rather than on the file system.

you ── prompt ──▶ model
                              │  writes a script
                              ▼
                ┌──────────────────────────┐
                │  codemode (QuickJS box)  │  runs where the agent loop runs
                └───┬──────────┬─────────┬─┘
                    ▼          ▼         ▼
              built-in tools   MCP    models.classify()
                              servers  models.generateImages()
                    └──────────┴─────────┘
                              │  compact result only
                              ▼
                            model

Their demo is a good illustration. The agent pulls open issues from Linear over MCP, then runs a small classifier model over each issue thread to rate how frustrated the commenters sound, four at a time. Hundreds of tool calls happen inside the script, and the model only sees a short summary at the end. The intermediate noise never touches the context window. (Per the docs, Pi runs classifier calls at most four at a time per script.)

Codemode is also useful without MCP at all. The docs list the obvious wins: run several tool calls in parallel, and filter large output before it ever reaches the model.

💡 Who's Jev? It's a classifier model from TypeSafe, not a chat LLM. Pi can run classifiers like this from codemode, and it can turn any llama.cpp chat model into a classifier too.

Turning it on

  • MCP servers come from mcp.json (global, or per project once trusted), or you manage them with /mcp and pi mcp add|remove|list|login|logout.

  • Codemode loads automatically when MCP is configured.

  • You can enable it without MCP by adding it to your default tools: "defaultTools": ["+codemode"]. That +name form adds a tool on top of the defaults (read, bash, edit, write) instead of replacing them.

  • For a one-off run, list every tool you want, because --tools replaces the selection: pi --tools read,bash,edit,write,codemode.

  • MCP tools default to codemode exposure: scripts can call them, but they aren't declared to the model. Scripts find them with searchTools(). If you want a server's tools visible to the model directly, set its exposure to direct.

What 1.0 changed about it

  • Roughly 40% fewer prompt tokens. With default tools and codemode active, the changelog shows one GPT-5.6 request shrinking from about 5,300 to 3,300 tokens.

  • Errors that tell the model how to recover. Typo a tool name and the error suggests the close match.

  • Image generation. Scripts can call models.generateImages() using the session's credentials, and the cost counts toward the session.


Virtual models: one "model" that routes to many

Extensions can now register a virtual model with pi.registerVirtualModel(). For every request, the extension decides which physical model and thinking level to use. The footer shows which one was picked, and /session breaks cost down per physical model.

The announcement's demo has Pi write itself an extension: a router/auto model that plans with Claude Opus and implements with GPT, with the Jev classifier deciding when to switch. There's also an example router in the repo (examples/extensions/jev-router.ts).

A few details from the docs:

  • A virtual model shows up in the model picker like any other, and it can be listed under any provider.

  • A virtual model can't route to another virtual model.

  • Resuming a session restores the virtual selection. If that model is no longer registered, Pi falls back to the physical model that answered last.

One honest note: the 0.99.0 changelog calls virtual models experimental, so expect the API to move.

This is the feature I'm most curious about. Planning with one model and implementing with another is easy to describe and fiddly to do by hand. Having it as a first-class concept, with per-model cost reporting, is a nice idea.


The plumbing trio

Three features in the list won't make a flashy demo, but they're the kind of thing that matters once you use an agent all day.

Feature

What it does

Why it matters

Deferred tool loading

Tools aren't all declared to the model up front. A tool_search tool finds and declares them on demand.

The prompt stays small even as your tool and MCP server count grows.

Cache warming

Keeps eligible prompt caches warm during active runs and, optionally, between runs. Controlled by the global cacheWarming setting: "off", "streaming" (the default) or "idle".

A long test run shouldn't mean re-paying for your whole context afterwards.

Mid-conversation system messages

Changes to the system prompt or tools are recorded in the transcript, so they survive resume and branch navigation, while preserving cached prefixes on supported models.

Extensions can adjust instructions or tools mid-session without losing the change on resume.

🔎 A detail on cache warming: the 1.0 announcement lists it as a feature for Anthropic models, but the settings docs describe it more generally. It only runs when the model declares a cache lifetime and Pi estimates at least $0.05 in avoided cache-miss cost. Refreshes count toward session totals, but they don't enter the model's context, and /session shows the next decision.


Fullscreen by default (and how to opt out)

The TUI now runs fullscreen: the editor and footer stay pinned, the transcript scrolls on its own, there's a draggable scrollbar, and you can search the transcript with Ctrl+Shift+F.

If you'd rather keep your terminal's normal scrollback, set tuiMode to "regular" or launch with --tui-mode regular.

A couple of smaller touches:

  • The system theme (default since 0.99.0) derives its colors from your terminal's own palette.

  • quietStartup: "header" keeps the version and key hints but hides the model scope line and the loaded-resource listing.


Pi Durable: the other half of the launch

Here's the part that's easy to miss if you only skim the announcement.

The Pi you know is built for one person in one terminal. If the process dies, you look at what happened and tell it to continue. That's by design and it's not changing.

But Earendil wants Pi's ideas to work outside the terminal too: in a Slack bot, a background job, or something several people steer at once. That needs a harness that can run anywhere, survive failures and handle very long conversations. So they built Pi Durable.

⚠️ It's experimental. Earendil says the API might still change. It doesn't replace the Pi coding agent. It's a framework for building agentic apps, coding agents included.

The mental model

A harness is storage plus the machinery to run many model conversations in parallel, along with the tools those models call and the environments the tools run in. Everything the harness does, from calling the model to running a tool, is a task. Storage options shipped today are memory, SQLite and JSONL, and the interfaces are small enough to put your own backend behind.

Surviving crashes

Every step of a run is a task that saves a checkpoint before moving on. If the process dies, a new process opens the same storage, finds unfinished tasks and continues from the last checkpoint.

submit(job, requestId) ─▶ task checkpoint ─▶ tool call … 💥 process dies

new process opens same storage ─▶ finds unfinished tasks ─▶ resumes

The details are what I like:

  • A model request that got cut off is sent again, and the partial answer stays in the transcript marked as aborted.

  • A tool call that got cut off reruns only if it's declared safe to. Otherwise the model is told it was interrupted and decides what to do.

  • A requestId makes a submission exactly-once, so a client retrying after a crash gets the original back instead of asking twice.

Here's the shape of that safe-versus-unsafe distinction, trimmed from their post:

const searchIssues = defineTool({
  name: "search_issues",
  description: "Search the issue tracker",
  parameters: Type.Object({ query: Type.String() }),
  replay: "safe", // only reads, so rerunning after a crash is fine
  execute: async (args, api) => { /* ... */ },
});

const deploy = defineTool({
  name: "deploy",
  description: "Deploy a version to production",
  // no `replay`: after a crash the model is told the call was
  // interrupted, and the deploy is never silently repeated
  // ...
});

Forcing you to say "this one is safe to repeat" is the kind of small decision that saves you from a double deploy.

The rest of the toolbox

  • Many conversations at once, with forking. Think a Slack channel as one conversation and each thread as a fork of it. Both run concurrently.

  • Extensions bundle system prompt sections, tools, hooks and tasks, and can be hot-swapped while conversations run.

  • Background compaction, so long conversations keep going instead of stopping to summarize. Older messages stay in storage.

  • Documents, typed JSON for app state (a todo list, a ticket) committed atomically with the transcript.

  • Multiplayer: any number of clients can attach to a conversation, watch it and steer it.

The whole source, without tests, is about 15,000 lines. Earendil points out that's small enough for your agent to read when building on it. They also say that in the coming weeks they'll show the small tools they build with it for their own work, giving a Slack bot and a GitHub triage bot as examples, though they're keeping the details quiet for now.

Try it:

npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord

Upgrading: what to check

The 1.0.0 release itself has no "Breaking Changes" section in the changelog. The items below are the things most likely to surprise you, including a few from the earlier releases you'd pick up if you're coming from an older version.

Area

What to check

Fullscreen

Don't like it? tuiMode: "regular" or --tui-mode regular.

Codemode scripts

Scripts that probed with typeof tools.name must now use "name" in tools.

MCP OAuth

Credentials are now stored per server name and URL. Ones previously stored by URL alone move to the first server that uses them.

--provider flag

--provider without --model used to be silently ignored. It now fails with an error.

Extensions

If you maintain any, skim the breaking-changes sections of 0.84, 0.86 and 0.87, where session, event and provider APIs changed. From 0.99.0, built-in extensions are named builtin:<name>, and --no-extensions now disables them too.

Startup noise

quietStartup: "header" if the banner is too chatty.


My take

A few things stand out to me.

The MCP u-turn is a good sign. Changing your mind in public, with a written explanation, is rarer than it should be. They also didn't just bolt MCP on. 0.99.0 reworked how tools are exposed (direct, model-only, codemode, deferred or hidden) alongside it, so MCP fits the same design as everything else.

The 1.0 label is as much a signal as a feature set. The diff between 0.99.2 and 1.0.0 is modest. What changed is the promise: Earendil is telling people and businesses they can depend on it.

For a "minimal" tool, there's a lot of new surface area. Codemode, virtual models, classifiers, image generation, and cache warming are all in now. Their argument is that these are primitives that stuck after a long filtering process, and they say the list of ideas that didn't make it is much longer. I'd keep an eye on whether that restraint holds.

Pi Durable is the bet to watch. Splitting it into its own package means they can experiment without risking the coding agent, and lessons that prove out are meant to flow back. Just don't build anything you can't afford to rewrite on it yet. It's explicitly experimental.

If you're already on 0.99.x, the jump to 1.0 looks small, and the one change you'll notice right away is fullscreen by default. If you're curious about Pi, install it and try codemode with one MCP server. That's probably the quickest way to see what the fuss is about.


Sources

Comments (0)

Join the discussion by logging into your account.

No comments yet. Be the first to comment!

Ankit Singh
Ankit Singh

Software Engineer

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Ankit Singh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.