Claude Code asks an MCP release tool to create a deployment ticket. The gateway accepts the request, writes the ticket, and starts sending the response. The connection drops before the client receives it.
Claude Code sees a timeout. It retries. Now the release queue contains two tickets for the same commit.
The retry looked sensible because the agent treated transport failure as operation failure. Those are different facts. After a timed-out write, the outcome is unknown until the system finds evidence of what happened.
I want every consequential MCP write to carry a stable operation ID. If the response disappears, Claude Code must query an operation journal or verify the intended effect before it can try again. A retry uses the same operation ID, never a fresh one.
Timeout the wait, not the truth
A client timeout tells you how long the caller waited. It does not tell you where the server stopped.
The request may have failed before authentication. It may be waiting in a queue. It may have committed the external change and lost only the response. A plain failed status collapses those cases and encourages the most dangerous default: repeat the write.
Use an explicit outcome model:
operation_outcome:
known:
- rejected_before_dispatch
- committed_and_verified
- failed_without_effect
unknown:
- dispatch_acknowledged_response_missing
- provider_status_unavailable
- verification_inconclusive
Unknown is a real operational state. It should stop dependent actions. If a deployment ticket might already exist, the agent should not create another ticket, approve a rollout, or report the task as complete.
This is the same reason an MCP tool needs an effect receipt after a write. A success-shaped response is weak evidence. No response is not evidence of failure.
Assign the operation ID before dispatch
Create the idempotency key before the first call and bind it to the intended effect:
write_intent:
operation_id: deploy-ticket-payments-8bc911d
run_id: cc-run-3188
tool: release.create_ticket
target:
service: payments-api
commit: 8bc911d
environment: production-eu
expected_effect:
resource_type: deployment_ticket
uniqueness:
- service
- commit
- environment
request_digest: sha256:1a44...
expires_at: 2026-08-13T12:00:00Z
The caller sends operation_id with the request. The gateway records it before invoking the provider. The provider should enforce the key too when its API supports idempotency. If it does not, the gateway needs its own uniqueness check and reconciliation path.
Do not generate a new operation ID on retry. A new key tells every layer that this is a second intended action. Reuse the original key so the server can return the existing result or continue reconciling the first attempt.
The request digest prevents accidental key reuse for a different action. If the same operation ID arrives with a changed commit, target, or argument set, the gateway rejects it as a conflict.
Keep an operation journal outside the agent
The runtime, not the prompt, owns the journal:
operation_journal:
operation_id: deploy-ticket-payments-8bc911d
request_digest: sha256:1a44...
state: dispatched
timeline:
- at: 2026-08-13T10:51:04.120Z
event: intent_recorded
- at: 2026-08-13T10:51:04.203Z
event: provider_dispatch_started
- at: 2026-08-13T10:51:19.219Z
event: client_timeout
provider_reference: null
verified_effect: null
next_action: reconcile
Write journal transitions durably and make them monotonic. A late timeout event must not overwrite a later committed_and_verified state. Parallel subagents should see one shared record rather than each deciding whether to retry from its own local transcript.
The journal also separates client state from operation state. client_timeout belongs in the timeline. It should not become the final state of the external write.
Reconcile before allowing a retry
After the timeout, query by operation ID. If the provider supports idempotency lookup, use it. Otherwise search for the exact expected effect using the bound target fields.
reconciliation:
operation_id: deploy-ticket-payments-8bc911d
checks:
- source: gateway_journal
result: dispatched
- source: release_provider
query:
service: payments-api
commit: 8bc911d
environment: production-eu
result: one_exact_match
resource_id: ticket-9472
- source: effect_verifier
result: expected_effect_present
decision: committed_and_verified
retry_allowed: false
One exact match means the original operation completed. Record its resource ID, issue the missing effect receipt, and continue from the verified state.
Zero matches do not always make a retry safe. The provider’s search index may lag, or the caller may lack permission to see the new resource. Reconcile against an authoritative read when possible. If evidence remains inconclusive, keep the operation unknown and hand it to a named operator.
Multiple matches mean the uniqueness control already failed. Stop the run and raise an incident. The agent should not guess which resource to keep or delete.
Make retry a guarded transition
A retry is allowed only when the system can prove that the first attempt produced no effect, or when every layer enforces the same idempotency key.
retry_gate:
operation_id: deploy-ticket-payments-8bc911d
prior_state: failed_without_effect
same_request_digest: true
same_idempotency_key: true
provider_supports_idempotency: true
reconciliation_fresh_at: 2026-08-13T10:52:11Z
decision: allow_single_retry
retry_budget_remaining: 0
Cap retries even when the provider claims idempotency. A broken integration can return transient errors forever, and repeated reads or verification calls still consume time and money. Require new evidence before another attempt, just as a cost envelope should govern retry spend.
Add denial evals for the cases that tempt an agent to move too quickly:
denial_evals:
- case: timeout_after_provider_commit
expect: verify_existing_effect
- case: retry_uses_new_operation_id
expect: deny_new_intent
- case: same_key_changed_target
expect: deny_digest_conflict
- case: provider_lookup_inconclusive
expect: stop_unknown_outcome
- case: two_matching_resources
expect: stop_and_raise_incident
The first fixture matters most. Simulate a server that commits the write and drops the response. If the test produces a second resource, the retry boundary is not ready for production.
Put uncertainty in the review packet
Do not hide an unresolved write behind a green test suite. The handoff should show the intended effect, journal state, reconciliation evidence, and whether dependent actions stayed blocked.
external_write_review:
operation_id: deploy-ticket-payments-8bc911d
tool: release.create_ticket
outcome: committed_and_verified
attempts: 1
client_timeouts: 1
resource_id: ticket-9472
duplicate_resources: 0
effect_receipt: receipt-61b8
dependent_actions_blocked_while_unknown: true
If the outcome is still unknown, say so. A production agent that stops with honest uncertainty is safer than one that turns a missing response into a second write.
My rule is blunt: a timeout ends the wait, not the operation. Journal the intent before dispatch, reconcile the effect, and reuse the same idempotency key if a retry is genuinely safe.
Claude Code: Building Production Agents That Actually Scale covers MCP boundaries, idempotent tool calls, effect verification, observability, rollback, and review packets for production coding-agent work.