JournalE.01September 30, 202614 min read

Durable agents with Effect TS

Typed failures, retries, timeouts, fallbacks and crash-proof steps for agents that never silently die.

d

Dipesh Chaulagain

AI-Native Fullstack Developer

Effect v4Agents

The demo worked

It always does. You wire a model to three tools, ask it a question, and it plans, searches, drafts and posts a tidy summary to Slack. Then you ship it, and the agent meets the real world: rate limits at 9 a.m., a provider incident at lunch, a deploy that kills the pod halfway through a twelve-minute run, and a model that decides today is the day it returns JSON with a trailing comma.

None of these are bugs in your prompt. They're distributed systems problems wearing an AI costume. An agent is a long-running, multi-step program whose every step talks to something slow, expensive and unreliable. We've known how to build those for decades. We just keep forgetting when the word agent is involved.

This essay builds a durable agent with Effect v4 from the ground up: typed failures, retry policies that respect the provider, timeouts that actually cancel work, fallbacks, and a step journal that lets a run survive a crash without re-billing a single token. Every figure runs real Effect code in your browser.

Failure is the weather

Start with arithmetic. If each call in a run succeeds 99% of the time, a single call looks fine. But a research agent doesn't make a single call. It plans, calls tools, reads results, re-plans, drafts, critiques and revises. Forty calls is a modest run.

Fig. 01The compounding problem
Drag the sliders
99.0%
40 calls

Runs that finish cleanly

66.9%

33 of every 100 runs hit at least one failure.

At 99% per call and forty calls, a third of your runs fail somewhere. A better model doesn't fix this, because the failures live in the plumbing. What fixes it is making each failure recoverable, so a transient blip costs a few hundred milliseconds instead of the whole run.

In production, the weather comes in five kinds:

  • Rate limits. A 429 with a retry-after header you should respect.
  • Hangs. A stream that stops sending bytes but never closes. No error, just silence.
  • Outages. The provider's status page turns orange, and every retry adds to the pile.
  • Bad output. Valid HTTP, invalid shape: truncated JSON, a missing field, a hallucinated enum.
  • Crashes. Your own process dies from a deploy, an OOM kill or a reclaimed spot instance, and the run's in-memory state dies with it.

The first four are about a single call. The fifth is about the whole run, and it's the one most agent frameworks quietly ignore.

Name every failure

The trouble with try/catch is that catch (e) hands you unknown. You can't write a sensible recovery policy for an error you can't name, so most code does the only safe thing: log it and rethrow.

Effect puts failures in the type. An Effect<A, E, R> succeeds with an A, can fail with an E, and needs services R. Model each failure as a tagged class and the compiler knows every way a call can go wrong.

errors.ts
import { Data, Effect, Schema } from "effect"

export class RateLimited extends Data.TaggedError("RateLimited")<{
  readonly retryAfterMs: number
}> {}

export class ProviderDown extends Data.TaggedError("ProviderDown")<{
  readonly provider: string
}> {}

export class MalformedOutput extends Data.TaggedError("MalformedOutput")<{
  readonly raw: unknown
}> {}

const Completion = Schema.Struct({
  text: Schema.String,
  tokens: Schema.Number,
})

export const makeProvider = (provider: string, url: string) =>
  Effect.fn("llm.complete")(function* (prompt: string) {
    const res = yield* Effect.tryPromise({
      try: (signal) =>
        fetch(url, {
          method: "POST",
          body: JSON.stringify({ prompt }),
          signal,
        }),
      catch: () => new ProviderDown({ provider }),
    })

    if (res.status === 429) {
      const seconds = Number(res.headers.get("retry-after") ?? 1)
      return yield* new RateLimited({ retryAfterMs: seconds * 1000 })
    }
    if (!res.ok) return yield* new ProviderDown({ provider })

    const raw = yield* Effect.tryPromise({
      try: () => res.json(),
      catch: () => new MalformedOutput({ raw: null }),
    })
    return yield* Schema.decodeUnknownEffect(Completion)(raw).pipe(
      Effect.mapError(() => new MalformedOutput({ raw }))
    )
  })

A few things worth noticing:

  • Effect.fn("llm.complete") gives the function a name, a tracing span and a stack frame that points at your code, not the runtime's.
  • tryPromise hands you an AbortSignal. When the effect is interrupted, by a timeout or a crash, the signal fires and fetch actually cancels. Hold onto that; it matters in a minute.
  • The body is decoded with Schema. A 200 with the wrong shape becomes a MalformedOutput right here, not a TypeError three functions later.
  • yield* new RateLimited(...) fails the effect. Tagged errors are yieldable, so failing reads like returning.

Hover makeProvider in your editor and the type tells the whole story: Effect<Completion, ProviderDown | RateLimited | MalformedOutput>. Now we can write a policy.

Retry with intent

Retrying is easy. Retrying well means answering three questions: which failures, how long to wait, and when to give up.

Which failures. Rate limits, timeouts and malformed output are transient: the next attempt has a real chance. An outage usually isn't. A provider in the middle of an incident rarely recovers in the next two seconds, and your retries slow its recovery. So ProviderDown isn't retried at all; it goes straight to the fallback.

How long. Exponential backoff, doubling from 250 ms and capped at 8 seconds. When the provider sends retry-after, wait at least that long. Schedule.modifyDelay sees the error that triggered each retry, so the policy can read the header straight off it.

When to give up. Schedule.max combines the backoff with Schedule.recurs(4) and continues only while both want to. Four retries, then the error moves on.

retry.ts
import { Duration, Effect, Schedule } from "effect"
import type { Cause } from "effect"
import type { MalformedOutput, ProviderDown, RateLimited } from "./errors"

type CallError =
  RateLimited | ProviderDown | MalformedOutput | Cause.TimeoutError

export const isRetryable = (error: CallError) => error._tag !== "ProviderDown"

export const retryPolicy = Schedule.exponential("250 millis").pipe(
  Schedule.jittered,
  Schedule.setInputType<CallError>(),
  Schedule.modifyDelay(({ input, duration }) =>
    Effect.succeed(
      input._tag === "RateLimited"
        ? Duration.max(duration, Duration.millis(input.retryAfterMs))
        : Duration.min(duration, Duration.seconds(8))
    )
  ),
  (backoff) => Schedule.max([backoff, Schedule.recurs(4)])
)

Then there's jitter. It looks like a detail until you watch what happens without it. When a provider blips, every client fails in the same instant. With pure exponential backoff, every client also retries in the same instant, and again, and again: a synchronized stampede that lands on the provider exactly as it's trying to recover. Turn jitter off below and watch the load bars.

Fig. 02Eight clients, one bad second
Toggle jitter
250ms
5×
8s
client 1
client 2
client 3
client 4
client 5
client 6
client 7
client 8
0requests hitting the provider · peak 8 at once8.8s
Schedule.exponential("250 millis").pipe(
  Schedule.jittered,
  Schedule.modifyDelay(({ duration }) =>
    Effect.succeed(Duration.min(duration, Duration.seconds(8)))),
  (backoff) => Schedule.max([backoff, Schedule.recurs(5)]),
)

Schedule.jittered scales each delay by a random factor between 0.8 and 1.2. It's a small nudge, but it turns a wall of simultaneous requests into a spread-out trickle.

Time is a budget

A hang is the worst failure because it isn't one. There's no error to catch and no log line, just a promise that never settles. The agent holds a connection, a worker slot and the user's patience, forever.

Every call needs a deadline. Effect.timeout fails with a TimeoutError when a call runs long, but the important part is what happens to the call itself: Effect interrupts it. Interruption runs finalizers and fires the AbortSignal we passed to fetch, so the socket closes and the provider stops generating tokens you'd be billed for. Promise.race can't do that. It stops waiting, but the request keeps running in the background.

Always have a plan B

Here's the whole call, composed:

resilient.ts
import { Effect } from "effect"
import { makeProvider } from "./errors"
import { isRetryable, retryPolicy } from "./retry"

const primary = makeProvider("anthropic", "https://llm.internal/anthropic")
const fallback = makeProvider("openai", "https://llm.internal/openai")

export const complete = (prompt: string) =>
  primary(prompt).pipe(
    Effect.timeout("20 seconds"),
    Effect.retry({ schedule: retryPolicy, while: isRetryable }),
    Effect.catch(() => fallback(prompt).pipe(Effect.timeout("30 seconds"))),
    Effect.withSpan("llm.resilient", {
      attributes: { "prompt.chars": prompt.length },
    })
  )

Read it top to bottom as a policy: try the primary with a deadline, retry what's retryable, and if it still fails, for any reason, ask the fallback. The span wraps it all, so every attempt shows up in one trace.

Now break it. The figure below runs this exact pipeline against two fake providers whose failures you choose. Time is compressed so you don't sit through twenty-second timeouts: the deadline is 1.2 seconds here.

Fig. 03Failure lab
Real Effect v4, running in your browser
primary(prompt).pipe(
  Effect.timeout("1.2 seconds"),
  Effect.retry({ schedule: retryPolicy, while: isRetryable }),
  Effect.catch(() => fallback(prompt)),
)
anthropic
openai
0mst = 0ms3600ms

Pick a failure, pick an implementation, press run.

Try the hang with await fetch. Nothing happens, and nothing ever will, because the naive version has no concept of too slow. Switch to the Effect pipeline and the attempt is cut at the deadline, interrupted and retried. The outage skips retries entirely and goes straight to the fallback.

Survive the crash

Everything so far protects a single call. But the expensive failure in agent systems isn't a failed call; it's a failed run. Twelve minutes in, forty cents of tokens spent, the pod is evicted for a deploy. A naive agent keeps its progress in memory, so it starts over: it re-plans, re-searches, re-drafts and re-bills.

The fix is the idea behind Temporal, Restate and every workflow engine: journal each step's result to durable storage, and on restart, replay the journal instead of redoing the work. Completed steps return their saved result instantly, and the run picks up exactly where it died.

You don't need a workflow engine to get the core of this. In Effect it's one service and a helper:

journal.ts
import { Context, Effect, Layer, Option } from "effect"

export class Journal extends Context.Service<
  Journal,
  {
    readonly get: (key: string) => Effect.Effect<Option.Option<unknown>>
    readonly put: (key: string, value: unknown) => Effect.Effect<void>
  }
>()("Journal") {}

export const step = <A, E, R>(name: string, run: Effect.Effect<A, E, R>) =>
  Effect.gen(function* () {
    const journal = yield* Journal
    const saved = yield* journal.get(name)
    if (Option.isSome(saved)) return saved.value as A

    const value = yield* run
    yield* journal.put(name, value)
    return value
  }).pipe(Effect.withSpan(`step.${name}`))

export class Kv extends Context.Service<
  Kv,
  {
    readonly get: (key: string) => Effect.Effect<string | undefined>
    readonly set: (key: string, value: string) => Effect.Effect<void>
  }
>()("Kv") {}

export const JournalLive = (runId: string) =>
  Layer.effect(
    Journal,
    Effect.gen(function* () {
      const kv = yield* Kv
      return Journal.of({
        get: (key) =>
          kv
            .get(`${runId}:${key}`)
            .pipe(
              Effect.map((v) =>
                v === undefined ? Option.none() : Option.some(JSON.parse(v))
              )
            ),
        put: (key, value) => kv.set(`${runId}:${key}`, JSON.stringify(value)),
      })
    })
  )

step checks the journal first. If a result exists, it returns it without running anything; otherwise it runs the effect and saves the result. Journal is a service, so storage is swappable: Redis here, Postgres in production, a Map in tests. The agent never knows.

Here's the agent, written as if nothing could go wrong:

agent.ts
import { Effect, Layer } from "effect"
import { complete } from "./resilient"
import { JournalLive, step } from "./journal"
import type { Kv } from "./journal"

declare const search: (query: string) => Effect.Effect<Array<string>>
declare const postToSlack: (message: {
  text: string
  idempotencyKey: string
}) => Effect.Effect<{ ts: string }>
declare const RedisKv: Layer.Layer<Kv>

export const researchAgent = Effect.fn("agent.research")(function* (
  runId: string,
  question: string
) {
  const plan = yield* step("plan", complete(`Plan research for: ${question}`))
  const sources = yield* step("search", search(plan.text))
  const draft = yield* step(
    "draft",
    complete(`Answer "${question}" using:\n${sources.join("\n")}`)
  )
  yield* step(
    "publish",
    postToSlack({ text: draft.text, idempotencyKey: `${runId}:publish` })
  )
  return draft.text
})

export const run = (runId: string, question: string) =>
  researchAgent(runId, question).pipe(
    Effect.provide(JournalLive(runId).pipe(Layer.provide(RedisKv)))
  )

That's the payoff of pushing reliability to the edges. The agent reads like a straight line: plan, search, draft, publish. Retries, timeouts and fallbacks live in complete; durability lives in step. The same runId on restart means the same journal keys, so a restarted run replays instead of redoing.

Fig. 04Kill the pod
Crash it mid-run, then resume
  1. 01 · llm

    step("plan")

    pending

  2. 02 · tool

    step("search")

    pending

  3. 03 · llm

    step("draft")

    pending

  4. 04 · side effect

    step("publish")

    pending

$ waiting for a run…
Tokens billed
0
Slack posts
0
Dupes dropped
0

Tip: crash during “publish” to find the nastiest bug.

Run it, crash it during draft, and resume. plan and search replay in zero milliseconds for zero tokens. Now turn the journal off and do the same thing: every token is billed twice.

Exactly-once is a lie

Now the nasty one. Turn the journal back on, switch off the idempotency key, and crash during publish.

The Slack message goes out twice, journal and all. Look at the order of operations inside step: run the effect, then save the result. If the process dies between the side effect and the save, the journal has no record of it, so on resume the step runs again. No amount of journaling closes that gap. The side effect and the journal write happen in two different systems, and no transaction spans both.

You can't get exactly-once execution. You can get exactly-once effect: at-least-once delivery plus an idempotent receiver. That's what the idempotency key is. It's derived from the run and the step (run_7f3a:publish), so it's identical on every retry and every resume. The receiver remembers keys it has seen and drops duplicates. Payment APIs have worked this way for years; for tools that don't accept a key, put a dedupe table in front of them.

See everything

Durability without visibility is a black box that happens to be reliable. Because Effect.fn, withSpan and step all open spans, every run produces a trace for free: agent.research → step.draft → llm.resilient → llm.complete ×3, with each retry, timeout and fallback visible as its own child span.

Plug in an OpenTelemetry exporter and those spans land in Honeycomb, Grafana or Datadog next to the rest of your stack. When a run takes four minutes instead of one, you'll see that it was the third attempt at draft, against the fallback, after a rate limit. That's its own essay, and it's next.

The checklist

  1. Model every failure as a tagged error. If you can't name it, you can't recover from it.
  2. Retry only transient failures, with capped, jittered, exponential backoff that honors retry-after.
  3. Give every call a deadline, and make sure the timeout cancels the work, not just the wait.
  4. Fall back to a second provider during an outage instead of hammering the first one.
  5. Journal every step so a crashed run resumes instead of restarting.
  6. Keep the code between steps deterministic.
  7. Give every externally visible side effect an idempotency key derived from the run and the step.
  8. Trace all of it, so when something does go wrong you can see which step, which attempt and why.

None of this is exotic. It's the unglamorous engineering that separates a demo from a system, and with Effect it's a few dozen lines at the edges while your agent's logic stays a straight line. Build it once and your agent stops dying silently. It retries, falls back, waits, resumes, and tells you exactly what it did.