Once an agent can act, you need a dry-run on the live path.
Evals score answers. Firewalls score arguments. Neither one proves the charge, the email, or the ticket. FetchSandbox is the in-flight dry-run for that outbound API action — propose, execute on a twin, keep the receipt.
The short answer
When LangChain, n8n, Claude, or Cursor emits a tool call that would move money or send mail, hold it. Point the same call at a stateful twin. Run the workflow, including the decline and the retry. Attach the receipt. Only then let the live host see the request.
The gap is the action, not the plan
Most writing about agent safety stops one layer too early. Schema validation, policy firewalls, and approval queues decide whether a call is allowed. Code sandboxes decide whether generated shell is isolated. Production still sees the first real Stripe or Resend request.
That is the moment that broke in the wild this month: an agent issued a refund twice because nothing dry-ran the side effect in-flight. The question is not “did the tool JSON look right.” It is “what did the provider actually do.”
Four layers people mix up
| Layer | What it proves | What it misses |
|---|---|---|
| Schema / policy check | Arguments look legal | The side effect never ran |
| Human approval | A person said yes | The live call still fires |
| Code sandbox | Shell and files are isolated | Stripe and Resend are still live |
| Twin dry-run | The action ran on a stateful API twin | Not a policy engine. Not a VM. |
Four actions that skip the dry-run
The refund that fires twice
The agent issued the refund, the provider retried, and nobody held the second delivery. Deduping in code review does not prove the retry. Scenario: webhook_retries, same event id.
The confirm that said paid
Confirm returned a decline and the app still unlocked. The tool call looked fine; the outcome was wrong. Workflow: Stripe accept_payment under payment_declined.
The email that sent to prod
The agent called send. Staging used a real Resend key. There was no dry-run host, so the first test was a customer. Twin: Resend send workflow, no live quota.
The ticket that wrote twice
The tool succeeded, the agent retried, the downstream record duplicated. A mock 200 cannot show that. A twin keeps the created id so the second call is visible.
Propose, dry-run, then execute
FetchSandbox twins are hosts, not live intercepts. Claude or Cursor can quickrun a workflow. An n8n HTTP Request node can change the Stripe URL to the twin. A LangChain tool can wrap the same host. The agent still chooses the action. The twin is where that action runs first.
“Hold this Stripe confirm. Run accept_payment on the twin, then payment_declined. Did the app unlock? Give me the receipt before I hit live.”
“This n8n HTTP node posts to Stripe. Point it at the twin host, replay webhook_retries with the same event id, and tell me if the order fulfilled once or twice.”
The same loop belongs in CI: fail the old handler, pass the new one, exit 0. A green unit test on mocked JSON is not that gate.
The receipt is the hold artifact
Every dry-run produces a shareable URL: requests in order, twin state, webhook events, verdict. That is what you attach to the agent trace or the PR. It is also what a teammate can read when they did not write the tool wrapper.
Proof pattern: Stripe payment_declined on twin a8c3b9fbd9. Confirm returned 402. The old handler still said paid. The new handler did not. Exit 0.
Questions people ask
How do you validate agent actions in production?
Do not start with the live API. Intercept the proposed tool call, execute it against a stateful twin of the provider, and keep a receipt of what the twin did. Schema checks prove the arguments look legal. A twin dry-run proves the side effect — the charge, the email, the ticket — behaved the way you expect, including the retry.
Is this the same as a policy firewall or human approval?
No. A firewall (schema, allow-lists, AEGIS-style pending) decides whether the call is permitted. Human-in-the-loop pauses for a yes. Neither one runs the provider lifecycle. FetchSandbox is the dry-run of the outbound API action after someone has already decided the call is allowed.
Is this an agent execution sandbox like Firecracker or E2B?
No. Those isolate code the agent writes — shell, filesystem, network egress. FetchSandbox isolates the API side effect. The dangerous call here is POST to Stripe or Resend, not a shell command.
Can LangChain, n8n, Claude, or Cursor use this?
Yes. Claude and Cursor talk to FetchSandbox over MCP with quickrun. A LangChain tool can wrap the same twin host. An n8n HTTP Request node can point at the twin instead of the live Stripe or Resend URL. There is no n8n or LangChain spec in the catalogue — the twin is the host your node already calls.
Does FetchSandbox refund on a Stripe twin?
No. The Stripe twin ships accept_payment and payment_declined workflows, plus webhook retry scenarios. It does not ship a refund workflow. For a refund-shaped hold, dry-run the confirm/decline path you actually have, or use a provider that ships capture_and_refund such as Adyen or PayPal.
What does a receipt contain?
Requests in order, the state the twin held, webhook events delivered, and a verdict. HTTP 200 on every step can still come back unproven. A Stripe payment_declined run on twin a8c3b9fbd9 is the pattern: fail the old handler, pass the new one, exit 0.
Start here
Connect MCP once. Put the twin on the path of the next action that can spend or send, before it reaches production.