JournalE.04September 30, 202612 min read

Let it crash: what Elixir supervisors teach AI agents

Supervision trees as a mental model for fault-tolerant, multi-agent systems.

d

Dipesh Chaulagain

AI-Native Fullstack Developer

ElixirArchitecture

The try/catch that ate the agent

Every agent codebase I've reviewed has the same function. It wraps a model call, a tool call and a memory update in one enormous try, and the catch logs a warning and carries on. It feels responsible. It is how most agents end up in states nobody designed.

Consider what “carries on” means. The tool returned an HTML error page instead of JSON, the parse failed, the catch swallowed it, and the half-updated context went into the next model call. Now the model is reasoning about a 502 page. It calls the same tool again. And again. No exception, no alert, a growing bill, and a user watching a spinner.

Telecom engineers hit this wall in the 1980s, with switches that had to stay up for years while running code with bugs in it. Their answer, built into Erlang and inherited by Elixir, sounds reckless and turns out to be the opposite: let it crash. This essay is about what that means precisely, and how much of it you can carry into a TypeScript agent stack today.

What “let it crash” means

It doesn't mean ignore errors. It means separate the code that does the work from the code that decides what to do when the work fails, and make the second kind small, generic and boring.

A worker handles the cases it understands and asserts everything else. If the model's reply doesn't parse, the worker doesn't invent a recovery. It dies, and its state dies with it. A supervisor, whose only job is watching workers, notices and starts a fresh one from a known-good state. The corrupted context is gone because the process holding it is gone.

The insight is statistical. Most production failures are transient: a rate limit, a dropped connection, a malformed reply, a bad interleaving. Retrying the exact same state often reproduces the problem. Restarting from clean state usually doesn't. Joe Armstrong's thesis on Erlang called this building reliable systems in the presence of software errors. Not by removing them, which nobody can, but by containing them.

Agents are processes

The mapping from OTP to agents is almost suspiciously clean. An Erlang process is a lightweight, isolated unit with its own state and a mailbox. It shares no memory with anyone. That's what an agent should be: a planner, a researcher and a tool runner, each owning its own context and talking through messages.

Isolation is what makes crashing safe. If the researcher shares a mutable context object with the writer, killing the researcher leaves the writer holding whatever half-written state it had. If each agent owns its context and communicates by message, a crash takes out exactly one agent's state and nothing else. Most multi-agent frameworks get this wrong by default, with one big shared “memory” that every agent mutates.

Strategies encode dependencies

Supervisors form a tree. Each one watches its children and decides, when one dies, who else has to restart with it. That decision is the strategy, and it's really a declaration of which components depend on each other's state.

Fig. 01A supervision tree for agents
Click crash. Click a strategy.
appdown
max 3 in 5s
session storedown
conversation memory
researchdown
share one plan
plannerdown
restarts 0
researcherdown
restarts 0
writerdown
restarts 0
toolsdown
independent
searchdown
restarts 0
browserdown
restarts 0
code sandboxdown
restarts 0

iex> Supervisor.which_children(App.Supervisor)

Crash search: only search restarts. Tools are independent, so the supervisor uses :one_for_one. Crash the planner: all three research agents restart, because the researcher and writer are executing a plan that no longer exists in anyone's memory. That's :one_for_all. Click the app strategy until it reads :rest_for_one and crash the session store: everything started after it restarts, because they all read from it, while nothing before it is touched.

In Elixir, the whole tree is a few lines of configuration:

lib/agents/application.ex
defmodule Agents.Application do
  use Application

  @impl true
  def start(_type, _args) do
    children = [
      Agents.SessionStore,
      Agents.ResearchSupervisor,
      Agents.ToolSupervisor
    ]

    Supervisor.start_link(children,
      strategy: :one_for_one,
      max_restarts: 3,
      max_seconds: 5,
      name: Agents.Supervisor
    )
  end
end

defmodule Agents.ResearchSupervisor do
  use Supervisor

  def start_link(opts), do: Supervisor.start_link(__MODULE__, opts, name: __MODULE__)

  @impl true
  def init(_opts) do
    children = [Agents.Planner, Agents.Researcher, Agents.Writer]
    Supervisor.init(children, strategy: :one_for_all)
  end
end

The code sandbox is :transient: restarted if it crashes, left alone if it finishes normally. The other options are :permanent, always restart, and :temporary, never restart, which suits one-shot jobs whose failure the caller handles.

A supervisor in TypeScript

None of this requires the BEAM to be useful. The core of a supervisor is small enough to read in one sitting. This is the implementation behind every figure on this page:

lib/supervise/supervisor.ts
export type Restart = "permanent" | "transient" | "temporary"
export type Strategy = "one_for_one" | "one_for_all" | "rest_for_one"

export type ChildSpec = {
  id: string
  restart?: Restart
  start: (signal: AbortSignal) => Promise<void>
}

export type SupervisorEvent =
  | { type: "started"; id: string; pid: number }
  | { type: "exited"; id: string; pid: number }
  | { type: "crashed"; id: string; pid: number; reason: string }
  | { type: "terminated"; id: string; pid: number }
  | { type: "gave_up"; id: string; restarts: number }

export type SupervisorOptions = {
  id: string
  strategy: Strategy
  maxRestarts?: number
  withinMs?: number
  restart?: Restart
  children: ReadonlyArray<ChildSpec>
  onEvent?: (event: SupervisorEvent) => void
  now?: () => number
}

export class Escalation extends Error {
  constructor(
    readonly supervisor: string,
    readonly restarts: number
  ) {
    super(`${supervisor}: ${restarts} restarts exceeded intensity`)
  }
}

let nextPid = 100

const reasonOf = (e: unknown) => (e instanceof Error ? e.message : String(e))

export const supervisor = ({
  id,
  strategy,
  maxRestarts = 3,
  withinMs = 5000,
  restart = "permanent",
  children,
  onEvent = () => {},
  now = Date.now,
}: SupervisorOptions): ChildSpec => ({
  id,
  restart,
  start: (signal) =>
    new Promise<void>((resolve, reject) => {
      const running = new Map<
        string,
        { controller: AbortController; pid: number }
      >()
      const history: Array<number> = []
      let stopped = false

      const launch = (spec: ChildSpec) => {
        const controller = new AbortController()
        const pid = nextPid++
        running.set(spec.id, { controller, pid })
        onEvent({ type: "started", id: spec.id, pid })
        spec.start(controller.signal).then(
          () => exit(spec, pid, null),
          (e: unknown) => exit(spec, pid, e ?? new Error("crashed"))
        )
      }

      const terminate = (childId: string) => {
        const child = running.get(childId)
        if (!child) return
        running.delete(childId)
        child.controller.abort()
        onEvent({ type: "terminated", id: childId, pid: child.pid })
      }

      const shutdown = () => {
        stopped = true
        for (const spec of [...children].reverse()) terminate(spec.id)
      }

      const exit = (spec: ChildSpec, pid: number, error: unknown) => {
        if (stopped || running.get(spec.id)?.pid !== pid) return
        running.delete(spec.id)
        onEvent(
          error === null
            ? { type: "exited", id: spec.id, pid }
            : { type: "crashed", id: spec.id, pid, reason: reasonOf(error) }
        )

        const policy = spec.restart ?? "permanent"
        if (
          policy === "temporary" ||
          (policy === "transient" && error === null)
        )
          return

        const t = now()
        history.push(t)
        while (history.length && history[0] <= t - withinMs) history.shift()
        if (history.length > maxRestarts) {
          onEvent({ type: "gave_up", id, restarts: history.length })
          shutdown()
          reject(new Escalation(id, history.length))
          return
        }

        const at = children.indexOf(spec)
        const affected =
          strategy === "one_for_one"
            ? [spec]
            : strategy === "one_for_all"
              ? [...children]
              : children.slice(at)

        for (const c of [...affected].reverse()) if (c !== spec) terminate(c.id)
        for (const c of affected) {
          if (c === spec || c.restart !== "temporary") launch(c)
        }
      }

      signal.addEventListener("abort", () => {
        shutdown()
        resolve()
      })
      for (const spec of children) launch(spec)
    }),
})

export const run = (root: ChildSpec) => {
  const controller = new AbortController()
  const done = root.start(controller.signal)
  done.catch(() => {})
  return { stop: () => controller.abort(), done }
}

A child is anything with an id and a start(signal) that returns a promise: resolve for a normal exit, reject for a crash. A supervisor is itself a child spec, so trees nest for free, and a supervisor that gives up simply rejects, which its parent sees as one more crash. Building the tree from the figure looks like the Elixir, minus the syntax:

tree.ts
import { run, supervisor } from "@/lib/supervise/supervisor"
import type { ChildSpec } from "@/lib/supervise/supervisor"

declare const sessionStore: ChildSpec
declare const planner: ChildSpec
declare const researcher: ChildSpec
declare const writer: ChildSpec
declare const search: ChildSpec
declare const browser: ChildSpec
declare const sandbox: ChildSpec
declare const page: (e: unknown) => void

const app = run(
  supervisor({
    id: "app",
    strategy: "one_for_one",
    maxRestarts: 3,
    withinMs: 5000,
    children: [
      sessionStore,
      supervisor({
        id: "research",
        strategy: "one_for_all",
        children: [planner, researcher, writer],
      }),
      supervisor({
        id: "tools",
        strategy: "one_for_one",
        children: [search, browser, { ...sandbox, restart: "transient" }],
      }),
    ],
  })
)

app.done.catch(page)

Knowing when to give up

A supervisor that restarts forever is just a very fast infinite loop. OTP bounds it with restart intensity: at most max_restarts within max_seconds, by default three in five. Exceed it and the supervisor kills all its children, then dies, handing the problem to its own parent.

Fig. 02When to stop restarting
Real supervisor, real time
3
5s
search
tools
app
0s8s

Run it to watch the supervisor decide.

The two failure modes look identical in a log line and behave nothing alike. Flaky is a rate limit: every so often a request fails, and a restart genuinely fixes it. The supervisor absorbs every crash and nobody notices. Broken is an expired API key: every restart crashes in 120 milliseconds. Retrying can't help, so the supervisor stops trying, escalates, and within about a second the whole application is down.

That's the correct outcome. The restart hierarchy tries progressively bigger resets, first the worker, then its subtree, then the application, and when none of them work, it fails loudly instead of burning tokens in a loop. Turn max_restarts up to ten in broken mode and watch how little it changes: a bug isn't fixed by patience.

Crash on purpose

Here's the part that matters most for agents. Traditional programs crash on their own when something's wrong: a null dereference, a failed match. Agents mostly don't. A confused model doesn't throw; it keeps producing plausible tool calls. So the most important failures never become crashes, and supervision can't help with a failure it never sees.

The fix is tripwires: cheap invariants that turn silent misbehaviour into a crash. Same tool call three times in a row. Step budget exhausted. Token budget exceeded. Output fails its schema. Below, a pricing tool returns an HTML error page exactly once, and the agent's context is poisoned by it.

Fig. 03Crash on purpose
An agent with a poisoned context

The pricing tool will return an HTML error page once, at step 3.

Run the agent.

Without the tripwire, the agent loops on compare prices until the demo cuts it off, and not a single error is raised. With it, the third identical call crashes the agent, the supervisor restarts it from the last checkpoint with a clean context, and it finishes. The poisoned context didn't need to be repaired. It needed to be thrown away.

With the AI SDK, tripwires fit in prepareStep, which runs before every step with the steps so far. Throwing there rejects the whole generateText call:

tripwires.ts
export class Tripwire extends Error {}

type Step = {
  toolCalls: ReadonlyArray<{ toolName: string; input: unknown }>
  usage: { totalTokens: number | undefined }
}

export const tripwires =
  ({ maxSteps = 20, maxTokens = 60_000, maxRepeats = 3 } = {}) =>
  ({ steps }: { steps: ReadonlyArray<Step> }) => {
    if (steps.length >= maxSteps) {
      throw new Tripwire(`step budget of ${maxSteps} exhausted`)
    }

    const tokens = steps.reduce((n, s) => n + (s.usage.totalTokens ?? 0), 0)
    if (tokens > maxTokens) {
      throw new Tripwire(`token budget exceeded: ${tokens}`)
    }

    const calls = steps.flatMap((s) =>
      s.toolCalls.map((c) => `${c.toolName}(${JSON.stringify(c.input)})`)
    )
    const recent = calls.slice(-maxRepeats)
    if (recent.length === maxRepeats && new Set(recent).size === 1) {
      throw new Tripwire(`loop: ${recent[0]} ×${maxRepeats}`)
    }
    return undefined
  }

In Elixir the same instinct is idiomatic. Pattern matches are assertions: {:ok, reply} = crashes on anything but success, and Jason.decode! crashes on malformed JSON. There's no recovery code to write, because the supervisor is the recovery code.

lib/agents/researcher.ex
defmodule Agents.Researcher do
  use GenServer

  def start_link(opts), do: GenServer.start_link(__MODULE__, opts, name: __MODULE__)

  @impl true
  def init(_opts) do
    {:ok, Agents.Checkpoints.load(__MODULE__)}
  end

  @impl true
  def handle_call({:step, input}, _from, state) do
    {:ok, reply} = Agents.LLM.complete(state.context ++ [input])
    %{"findings" => findings} = Jason.decode!(reply)

    if state.steps >= 20, do: raise("step budget exhausted")

    state = %{state | context: state.context ++ [input, reply], steps: state.steps + 1}
    :ok = Agents.Checkpoints.save(__MODULE__, state)
    {:reply, findings, state}
  end
end

Restart to where?

A restart wipes memory. That's the feature, and also the catch: anything the agent must not lose has to live outside it. Toggle checkpoints off in Fig. 03 and the restarted agent searches flights and hotels all over again, paying for both twice.

The rule is to checkpoint the last known-good state, not the latest state. Save after a step succeeds and its output validates, never mid-step. Then init is just “load the checkpoint,” and it has to be fast and idempotent, because it will run far more often than you expect:

researcher.ts
import type { ChildSpec } from "@/lib/supervise/supervisor"

type Job = { id: string; prompt: string }
type Checkpoint = { jobId: string; step: number; context: Array<string> }

declare const queue: { next: (signal: AbortSignal) => Promise<Job> }
declare const checkpoints: {
  load: (agent: string) => Promise<Checkpoint | null>
  save: (agent: string, c: Checkpoint) => Promise<void>
}
declare const step: (
  c: Checkpoint,
  signal: AbortSignal
) => Promise<Checkpoint | "done">

export const researcher: ChildSpec = {
  id: "researcher",
  restart: "permanent",
  start: async (signal) => {
    let current = await checkpoints.load("researcher")

    while (!signal.aborted) {
      current ??= { jobId: (await queue.next(signal)).id, step: 0, context: [] }
      const next = await step(current, signal)
      if (next === "done") {
        current = null
        continue
      }
      await checkpoints.save("researcher", next)
      current = next
    }
  },
}

That's the same step journal as E.01, seen from the other side. Durable steps make restarts cheap; supervisors make restarts happen. Either one alone leaves a gap.

What Node can't do

Honesty matters here, because the port is not the platform. The BEAM schedules processes preemptively, so one runaway process can't starve the rest. Every process has its own heap and garbage collector. And a supervisor can kill any process at any moment, unconditionally.

JavaScript can't do the last one. The TypeScript supervisor stops a child by aborting its signal, which is cooperative: a child that ignores the signal, or blocks the event loop in a synchronous loop, can't be stopped from inside the process. For agents that's mostly fine, because they spend their lives awaiting network calls, and every await is a checkpoint for the signal. For untrusted or CPU-heavy work, such as a code sandbox, put it in a worker thread or a separate process that can be terminated for real.

If your agents are the core of the product, with many long-lived sessions, many concurrent tool calls and a requirement to stay up through anything, that's an argument for running them on the BEAM itself. If they're one feature in a TypeScript app, the ideas travel fine: isolation, a supervisor, strategies, intensity, tripwires, checkpoints.

The checklist

  • Workers assert and crash; they don't improvise recovery.
  • Each agent owns its context. Nothing mutable is shared.
  • A supervisor restarts workers; its strategy mirrors real dependencies.
  • Restart intensity is bounded, and giving up escalates loudly.
  • Tripwires turn loops, runaway steps and bad output into crashes.
  • Checkpoints hold the last known-good state; init just loads one.
  • Anything that can hang for real runs where it can be killed for real.

The model will make mistakes you can't predict, on inputs you can't enumerate. You can't prevent that with defensive code. You can make sure each mistake costs one restart instead of one customer. Let it crash, and design for what happens next.