Skip to main content
How it works

Test what could break. See what happened.

A PR, a database query, a background job, a service. Start with the change your team needs to trust. Agree on the rules, test the failures, and inspect the result.

Start with a change you’re worried about.

We started with API integrations. A successful request looked fine, but a retry, a late event, or a partial failure could still break the result. That is the problem behind FetchSandbox.

The bigger question applies across software: when the code changes, does the behavior you depend on still hold? We’re extending the engine with teams working on services, jobs, queries, and application code.

Bring a specific change and the requirement behind it. Your incident history, tickets, and telemetry help us choose what to test. The verdict comes from observed execution and explicit checks.

Say what must never break.

An invariant is a rule that must stay true. Make it concrete enough to check. “The job works” is too vague. “Retrying the same job leaves one committed result” gives us something to measure.

ChangeRule to checkEvidence to observe
Background jobOne result per job, even after retryJob identity and committed writes
Database queryExisting records keep the correct ownerFixture rows and returned records
Application changeRepeating an operation produces one committed resultOperation identity, execution attempts, and committed state
Connected servicesA repeated event does not repeat its business effectEvent identity and downstream state

These are starting examples for a team pilot. We agree on which rules your environment can expose and which checks we can execute.

Make the test safe and observable.

Use an isolated version of the system with representative test data. Agree on the version under test, the dependencies we can control, and how we will read the result. A process returning 200 is not enough if the database contains the wrong state.

For hosted integration tests, the webhook must be reachable from outside the app workspace. It still needs signature validation. Use the current run’s twin URLs, temporary credentials, signing secret, and required run bindings together.

Keep live billing disabled during twin testing. Remove the temporary settings when finished. For a team system, environment access and failure controls are agreed before the pilot starts.

Test the conditions that could break the rule.

Start with realistic risks: a request times out, a dependency rejects access, a rate limit kicks in, a job restarts, or work overlaps. Pick the scenarios that threaten your requirement.

A declared fault is only the setup. We also need evidence that it happened. A delayed response that finishes before the app’s deadline does not prove timeout recovery. A retry that starts after the first request finishes does not prove overlapping execution.

The available controls depend on the engine, suite, and environment. We agree on that scope rather than assume every failure can be injected into every system.

The receipt should answer what happened.

Read the rule, the scenario, the version tested, and the recorded state together. Use the timeline to connect requests, responses, webhook attempts, and observation checkpoints. Shared evidence is sanitized so credentials and sensitive data stay hidden.

  • Passed: the required evidence was observed and the rule held.
  • Failed: the observed result violated the rule.
  • Not checked: the evidence needed to decide was missing.

If a required check was not measured, the overall result can remain inconclusive. That tells you what still needs work. Results apply to the recorded scenarios and observation windows, not every possible production condition.

Fix what broke. Run the checks again.

Keep the original failure evidence. Test the repaired version against the same rule and relevant failure conditions. Compare the state, not just whether an error disappeared.

Keep the check with your team’s existing regression process. As the software changes, the rule stays. For a team pilot, we agree on how it fits your CI and release gates.

Give the agent a clear job and a clear result.

Connect the free engine through MCP. Ask your building agent to test the integration with FetchSandbox, configure only the isolated environment, inspect every required result, and clean up temporary settings afterward.

Check the tool activity. A statement that the agent tested something is not the proof. You need the recorded run and its evidence. If a check is missing, the agent should inspect the diagnostic and resolve that specific blocker before rerunning.

MCP setup · MCP connection details · Questions and troubleshooting

The integration engine is where we started.

It is available today through MCP and the web app. Provider reference pages document that part of the product. Our team offering takes the same question to your system: what must stay correct as it changes?

Verification guides · Bring your team’s use case