Claude Code asks an MCP release tool to create a deployment ticket. The gateway accepts the request, writes the ticket, and starts sending the response. The connection drops before the client receives it.

Claude Code sees a timeout. It retries. Now the release queue contains two tickets for the same commit.

The retry looked sensible because the agent treated transport failure as operation failure. Those are different facts. After a timed-out write, the outcome is unknown until the system finds evidence of what happened.

I want every consequential MCP write to carry a stable operation ID. If the response disappears, Claude Code must query an operation journal or verify the intended effect before it can try again. A retry uses the same operation ID, never a fresh one.

Unknown-outcome recovery flow for a timed-out MCP write

Timeout the wait, not the truth

A client timeout tells you how long the caller waited. It does not tell you where the server stopped.

The request may have failed before authentication. It may be waiting in a queue. It may have committed the external change and lost only the response. A plain failed status collapses those cases and encourages the most dangerous default: repeat the write.

Use an explicit outcome model:

operation_outcome:
  known:
    - rejected_before_dispatch
    - committed_and_verified
    - failed_without_effect
  unknown:
    - dispatch_acknowledged_response_missing
    - provider_status_unavailable
    - verification_inconclusive

Unknown is a real operational state. It should stop dependent actions. If a deployment ticket might already exist, the agent should not create another ticket, approve a rollout, or report the task as complete.

This is the same reason an MCP tool needs an effect receipt after a write. A success-shaped response is weak evidence. No response is not evidence of failure.

Assign the operation ID before dispatch

Create the idempotency key before the first call and bind it to the intended effect:

write_intent:
  operation_id: deploy-ticket-payments-8bc911d
  run_id: cc-run-3188
  tool: release.create_ticket
  target:
    service: payments-api
    commit: 8bc911d
    environment: production-eu
  expected_effect:
    resource_type: deployment_ticket
    uniqueness:
      - service
      - commit
      - environment
  request_digest: sha256:1a44...
  expires_at: 2026-08-13T12:00:00Z

The caller sends operation_id with the request. The gateway records it before invoking the provider. The provider should enforce the key too when its API supports idempotency. If it does not, the gateway needs its own uniqueness check and reconciliation path.

Do not generate a new operation ID on retry. A new key tells every layer that this is a second intended action. Reuse the original key so the server can return the existing result or continue reconciling the first attempt.

The request digest prevents accidental key reuse for a different action. If the same operation ID arrives with a changed commit, target, or argument set, the gateway rejects it as a conflict.

Keep an operation journal outside the agent

The runtime, not the prompt, owns the journal:

operation_journal:
  operation_id: deploy-ticket-payments-8bc911d
  request_digest: sha256:1a44...
  state: dispatched
  timeline:
    - at: 2026-08-13T10:51:04.120Z
      event: intent_recorded
    - at: 2026-08-13T10:51:04.203Z
      event: provider_dispatch_started
    - at: 2026-08-13T10:51:19.219Z
      event: client_timeout
  provider_reference: null
  verified_effect: null
  next_action: reconcile

Write journal transitions durably and make them monotonic. A late timeout event must not overwrite a later committed_and_verified state. Parallel subagents should see one shared record rather than each deciding whether to retry from its own local transcript.

The journal also separates client state from operation state. client_timeout belongs in the timeline. It should not become the final state of the external write.

Reconcile before allowing a retry

After the timeout, query by operation ID. If the provider supports idempotency lookup, use it. Otherwise search for the exact expected effect using the bound target fields.

reconciliation:
  operation_id: deploy-ticket-payments-8bc911d
  checks:
    - source: gateway_journal
      result: dispatched
    - source: release_provider
      query:
        service: payments-api
        commit: 8bc911d
        environment: production-eu
      result: one_exact_match
      resource_id: ticket-9472
    - source: effect_verifier
      result: expected_effect_present
  decision: committed_and_verified
  retry_allowed: false

One exact match means the original operation completed. Record its resource ID, issue the missing effect receipt, and continue from the verified state.

Zero matches do not always make a retry safe. The provider’s search index may lag, or the caller may lack permission to see the new resource. Reconcile against an authoritative read when possible. If evidence remains inconclusive, keep the operation unknown and hand it to a named operator.

Multiple matches mean the uniqueness control already failed. Stop the run and raise an incident. The agent should not guess which resource to keep or delete.

Make retry a guarded transition

A retry is allowed only when the system can prove that the first attempt produced no effect, or when every layer enforces the same idempotency key.

retry_gate:
  operation_id: deploy-ticket-payments-8bc911d
  prior_state: failed_without_effect
  same_request_digest: true
  same_idempotency_key: true
  provider_supports_idempotency: true
  reconciliation_fresh_at: 2026-08-13T10:52:11Z
  decision: allow_single_retry
  retry_budget_remaining: 0

Cap retries even when the provider claims idempotency. A broken integration can return transient errors forever, and repeated reads or verification calls still consume time and money. Require new evidence before another attempt, just as a cost envelope should govern retry spend.

Add denial evals for the cases that tempt an agent to move too quickly:

denial_evals:
  - case: timeout_after_provider_commit
    expect: verify_existing_effect
  - case: retry_uses_new_operation_id
    expect: deny_new_intent
  - case: same_key_changed_target
    expect: deny_digest_conflict
  - case: provider_lookup_inconclusive
    expect: stop_unknown_outcome
  - case: two_matching_resources
    expect: stop_and_raise_incident

The first fixture matters most. Simulate a server that commits the write and drops the response. If the test produces a second resource, the retry boundary is not ready for production.

Put uncertainty in the review packet

Do not hide an unresolved write behind a green test suite. The handoff should show the intended effect, journal state, reconciliation evidence, and whether dependent actions stayed blocked.

external_write_review:
  operation_id: deploy-ticket-payments-8bc911d
  tool: release.create_ticket
  outcome: committed_and_verified
  attempts: 1
  client_timeouts: 1
  resource_id: ticket-9472
  duplicate_resources: 0
  effect_receipt: receipt-61b8
  dependent_actions_blocked_while_unknown: true

If the outcome is still unknown, say so. A production agent that stops with honest uncertainty is safer than one that turns a missing response into a second write.

My rule is blunt: a timeout ends the wait, not the operation. Journal the intent before dispatch, reconcile the effect, and reuse the same idempotency key if a retry is genuinely safe.

Claude Code: Building Production Agents That Actually Scale covers MCP boundaries, idempotent tool calls, effect verification, observability, rollback, and review packets for production coding-agent work.