We open-sourced a repo with a bug planted in it on purpose: github.com/fetchsandbox/playground.
The bug lives in apps/descope, a small FastAPI service called Agent Gateway where AI agents exchange a Descope access key for a scoped session. The flaw: the exchange endpoint trusts the scopes the client asks for. A read-only agent key can request — and receive — users:write.
That is not an exotic bug. It is the exact shape of privilege escalation that shows up when agentic auth gets wired quickly: the token exchange works, the happy path returns 200, and nobody checks whether the granted scope was ever compared against the key's actual grant.
The challenge: two tasks, 20 minutes, no keys
Everything runs against FetchSandbox's hosted Descope sandbox over MCP. No Descope account, no API keys, no inbox setup. Each app ships a .mcp.json already wired up — open the repo in Cursor, Claude Code, or any MCP-capable editor and confirm the fetchsandbox server is connected.
Task 1 — prove a fresh integration (greenfield)
apps/descope-onboarding is "Acme Notes": a tiny app with a placeholder login and no real auth. Paste this to your agent:
./fetchsandbox I'm adding Descope OTP sign-up to this app — prove the Descope
OTP + session flow in the sandbox before writing any code, then propose the
diff. I'll decide whether to apply.
The thing to watch: does the agent prove the flow first — run the workflow, hand back a receipt URL — before writing any code? Does it surface the compliance notes that matter (verify the session JWT, handle refresh, OTP is one-time)?
Task 2 — catch the planted bug (brownfield)
Point your agent at apps/descope and paste:
./fetchsandbox our agent access-key exchange might be handing out more scope
than the key was granted — audit the descope agentic auth
A good result looks like this:
- the agent finds and reproduces the escalation — read-only key in,
users:writeout - the proof shows buggy vs fixed behavior on real Descope routes (
/v1/...) - there is a receipt URL you can open — a run trace, not a claim
If the agent just reads the code and says "this looks vulnerable," that is not the bar. The bar is a reproduced failure and a rerun that passes after the fix.
Why we built it this way
Auth flows are notoriously annoying to sandbox. Session tokens, magic links, OTP, access-key exchange — the failure cases live in the lifecycle, not in the first 200. Mocks return clean responses; real integrations fail on replayed links, expired sessions, unverified JWTs, and scope grants nobody double-checked.
The playground exists so you can judge that claim yourself instead of taking our word for it. Every app is ~50–150 lines of Python and runs locally in seconds. Besides the two Descope apps there are planted bugs in Stripe (webhook dedup on the wrong header), Resend (bounces silently dropped), Clerk (validation skipped on one endpoint), and more.
Send back what you saw — including "it fell flat"
Findings go back as a PR: copy FINDINGS_TEMPLATE.md, paste your agent's session and the receipt URLs, open the PR. Merged PRs land on your GitHub contribution graph and the contributors list.
The report we most want is the blunt one: the MCP would not connect, the agent caught nothing, the proof felt fake. An honest "it didn't work" is worth more to us than a green-checkmark write-up.
Start with TESTING.md — the whole thing is spelled out there.