A podcast conversation with Mario Zechner (creator of Pi, now at Earendil) and Armin Ronacher (MCPO at Earendil) about the new Pi Durable release. The discussion covers why Pi was subsumed into their company Earendil, how AI has changed their development workflow (what they call “slop operations” — with clear limits), and a deep dive into the architecture of Pi Durable: a small, generic primitive for making long-running agents survive process crashes without losing state. They also compare agent harnesses (AMP’s Orbs, Claude Code, OpenCode, Deep Sea Carnus), debate why LLMs remain bad at architecture despite excelling at binary-verifiable tasks, and share what’s next: a self-modifiable Pi with a plugin system, mobile-first agent experiences, and building products beyond coding agents.
Armin introduces himself as MCPO (“chief mopping officer”) at Earendil and “resident intern”. His friend founded the company with him and they subsumed Pi in February, in the wake of the “OpenClaw craziness”, after a long attempt to get Mario on board. Mario self-deprecatingly describes his open source history: once a project gets users, he feels obligated to maintain it unpaid, and the negativity can wear you down. His goal has been to spread maintenance so the project’s truck factor no longer depends on him — a goal he feels has largely been achieved. Armin adds a parallel anecdote: he started a meetup in Poland that exploded into the country’s largest on meetup.com and now runs itself without him organizing.
Mario credits most of a year of issue triage and fixing for delaying Pi Durable from February to July. He doesn’t trust agents to triage which issues matter, and while agents can fix issues, they almost always add code while doing so — “it’s almost impossible that it will fix a problem without adding quote”. For design or architecture concerns, a human is still needed because agents fix locally and worsen the whole. Armin confesses to being “a victim to my own clanker use”, sending 200-file, 20,000-line PRs while insisting that’s not how things should be.
They embrace “slop operations” at the right layers: Pi’s TUI components (markdown, text, editor) were generated by agents against a hand-written API — Mario has never read a line of that code and doesn’t care, as long as tests verify rendering. He only steps in when performance degrades, because agents identify symptoms well but not the architectural cause of slowness. Armin notes that fully automated codebases naturally grow with every bug fix.
On model evolution, Armin feels new models are getting worse at architecture in practice — because he trusts them with more control, the net outcome disappoints. Both complain about models’ increasingly thesaurus-heavy communication style; both settled on models that “talk like normal people”. Mario highlights a genuine leap in security/reverse-engineering work: a model replaying a crack of his protected JVM software by reversing native memory structures — “no reverser on Earth would have gone that far”. Armin’s USB CarPlay adapter similarly had its firmware protection figured out autonomously. Mario also had DeepSeek build a Nintendo DS emulator from scratch to run a ROM hack. The unifying rule: tasks with a binary oracle (tests, fuzzing, emulation correctness) are won; fuzzing plus an LLM that reasons about uncovered cases is “rocket fuel for testing”.
Pushing back on DHH’s claim that architecture skills will also be automated, Mario argues models are decent at discussing architecture but the likely-completion trajectory in the weights points elsewhere: there is no training data for the process of designing something. Everything else in programming is textualized; this isn’t. The group jokes about selling meeting recordings to model labs (“two million dollars to OpenAI”) and discusses how one would classify high-value traces at scale — a classification problem best solved by mapping conversation traces to what actually shipped in repositories. Armin hears through the grapevine that labs aren’t doing anything sophisticated with traces yet; Mario expects self-improvement pressure (RSI) to force the issue, citing the latest Opus making daily breakthroughs on hardware-to-software navigation. Armin counters that such hill-climbing skill doesn’t transfer to building apps.
Mario defines durability in five sentences: an agent runs in a process, calls tools, the process dies; you want to restart and have it continue exactly where it left off, with no lost tool results or assistant streams, and with application state durably persisted and resumable — no typing “continue”. Two drivers: long-running agents (anything Claw-like, Slack bots living for years with enormous histories) and Pi’s historical limitation of loading the entire conversation tree into memory. It’s also deployable nearly anywhere: Vercel, Cloudflare Durable Objects, E2B, a laptop, even a phone.
The host’s guess (event queue with committed actions) is corrected: Pi Durable deliberately does not do event sourcing, unlike everyone else, because of its many problems. Instead, the core problem is that execution can crash before, during, or after an effect. The solution is a three-part sandwich per task: persist the intent (prepared inputs — e.g. the full LLM request with context window, thinking level, model), execute the effect, persist the result. For external systems that support idempotency (Armin’s favorite word), like scheduling a GitHub Actions run with a caller-chosen ID, a crash mid-effect is recoverable; for non-idempotent effects (deleting a file, charging a credit card) you must handle things carefully. Every operation in the harness — generating an LLM answer, executing one tool call — is such a task, and tasks can spawn and await other tasks. “That’s the whole magic.”
For application state they built a tiny document store where documents are associated with transcript positions. Each version can be recalled either as-of a point in time (the old to-do list when the conversation was at that point) or as the latest state regardless of position. Example: a canvas document storing a list of strokes that evolves as the agent and user interact; rewinding gives the agent the old drawing. They first considered JSON Patch (the RFC for expressing JSON mutations) but found it lacking, so they built an agentic-optimized format. Mutations are captured via JavaScript proxies — you mutate a plain-looking object and each change is recorded as an operation for playback. Armin reflects, after twenty years of web development, that expressing structural updates to nested data (the problem React is “all about”) still hasn’t converged on a standard — Immer mostly gets it right, Automerge’s proxy approach is “a huge hack”.
Existing agents (Hermes, OpenClaw, AMP with Orbs, OpenCode) each solved durability internally, but none is a generalized primitive you can take and build the next product or internal swarm on. Pi Durable is deliberately small (~12,000 lines) and generic, aiming to do everything an agent harness needs out of the box, including compaction. Armin frames the motivation: agents are already unreliable organisms; the orchestration layer should at least be understandable and sane, and keep working when no human is babysitting.
On other harnesses: they only read OpenAI’s Codex repo to decode backend quirks. Mario likes AMP’s team for committing fully to their vision (“willing to bet their house”) and finds Orbs works well; he praises OpenCode’s B2 durability battle from a distance. Armin used AMP last year but isn’t sticky to web interfaces; he prefers mobile apps (Claude Mobile, Codex Mobile) and isn’t enamored with terminals. Deep Sea Carnus impresses them both — Mario read its source, finds its plugin architecture smart (if over-promising), and cites it as conceptual inspiration for Pi’s evolution, as is Claude Code’s new self-modification (“the mods — how could we not have Claude modify Claude Code itself?”), though Claude Code still leaks sessions into each other. MCP support in Pi landed after a year of “excruciating painful deliberation”, only because code mode (their Jeff-classifier integration) needed it — and some community members were upset that he changed his mind at all.
Building an extensible, multi-user web interface for agents is security-hard: not the server sandboxing (“almost trivial”, Docker-solvable) but the shared web UI itself, where the agent-rendered interface can become literal XSS to exfiltrate cookies, and the main mitigations involve UX-hostile iframes. Trust gradients complicate connecting to someone else’s session — you don’t want to fully trust the session owner. Mario’s own anecdote: an agent recently served his entire filesystem to the public internet for a day. IPFS comes up and is dismissed as not addressing exfiltration.
What’s next: dogfooding internally; evolving Pi from “one clanker” toward “many clankers on many machines talking to many humans” with a plugin system, in whatever human/agent collaboration configuration you want. Mario’s phone prototype “PIM” already replaces Claude for Android — everything runs on the phone except the LLM, and it can ssh into his Hetzner box to do sysadmin with the brain on the phone. Earendil is also building products that aren’t coding agents at all, inspired by their earlier email-based personal agent Lefos — while wary of the personal-agent security “catastrophe” brewing (a fresh example: someone’s Grok bot posting their personal finances to the company Slack).
Armin is okay with AI as it is and wants the UX around it to improve; societally he hopes the hype dies down and the technology delivers something truly valuable (room-temperature superconductors, cancer medicine), and that communities form around building great software rather than “slop forks of Photoshop”. Mario celebrates PS5 emulation being solved by agents as game preservation. Both see the current “AI working on AI” phase as a crypto-like market mania that will pass. The thing they most want: onboarding non-users — Mario’s wife did linguistic research with Claude, his four-year-old helped write a birthday game — and they’re frustrated that everyone ships “a chat box and a sidebar” instead of real onboarding. The final exchange praises Droid’s minimal, well-engineered approach (one of the few “S tier” takes, even if dunked on from OpenAI), because cheap code generation otherwise leaves you “sitting with so much slop that is thirty percent finished” — the minimal route sidesteps the failure modes models are naturally prone to.