Documentation
Docs
Everything EveryRole does for you, what its results mean, and how to run it after every deploy.
See it in 2 minutes
A real run of EveryRole on rendrio.io, from pasting the address to watching a test’s replay. Waiting is sped up; nothing else is changed.
1 min 50 s, no sound. What the video shows, step by step:
- Paste your app’s address on the home page.
- A free first look: your own pages, as a visitor sees them.
- Sign up with just your email.
- Open the link from your email and answer 6 quick questions.
- Start the first exploration: EveryRole uses your app as a visitor, and every save is blocked.
- It suggests tests in plain English and tries each one twice.
- The suggestions you add become your tests, on the Tests page.
- Run all tests: each one runs in its own browser.
- The run’s report, with a replay of every test.
- Every run stays on the Runs page.
Getting started
- Paste your app’s address. EveryRole explores it signed out first, so you see suggested tests in minutes without giving it any account.
- Add the suggestions you want. Each one is a plain sentence (“A visitor can open the pricing page”) that EveryRole already tried twice for real. Add them one by one (“Add as test”), or select several and press “Add selected as tests”. Reword or dismiss the rest. Nothing runs as a test until you add it.
- Add test accounts for the kinds of user you have, then explore again to get tests for what they can do.
- Turn on daily runs and a deploy hook in the app’s Settings tab.
User types and test accounts
A user type is a kind of user of your app: visitor (signed out), free, paid, admin… Give EveryRole one test account per type, plus the address of your sign-in page. It signs in once per run and per type, and every test runs in its own private browser session.
Use test accounts, not real customers’ accounts. EveryRole can’t get past two-factor codes or CAPTCHAs, so a test account must sign in with email and password.
Starts at tells EveryRole where a type’s own area begins when it isn’t linked from the rest of the site (for example /admin). Never click lists words EveryRole refuses to click or follow, while exploring and while testing: add anything that spends money, deletes data or sends email.
What results mean
- ✓ Passed
- The page showed what the test expects.
- ✕ Broken
- Your app didn’t do what the test expects. This is the only red result, and the only one that can fail a pull request.
- ? Couldn’t check
- A problem on our side, not in your app: our browser stopped answering, the network failed, a machine was restarted. It costs no credits, never alerts you, and the test runs again next time.
- ↻ Passed on 2nd try
- It failed once, then passed. Nothing is broken, but the test may be unstable.
- – Not run
- Something it needs failed first, usually its user type couldn’t sign in.
Pass or fail is decided from checks on the loaded page (text shown, address reached, a control present), never from an AI saying it’s done. Every unit that doesn’t pass is tried once more before you see it.
Daily runs
In an app’s Settings tab, turn on “Run every day” and pick an hour and time zone. All your tests run once a day at that hour. A day with nothing to test is skipped, and a run that is already going isn’t started twice.
Deploy hooks
Each app can have a secret link that starts a run of all its tests. Make it in the app’s Settings tab (on Indie and Pro and Team), and call it when a deploy finishes. Anyone with the link can start a run, so keep it in your CI secrets; you can replace it at any time.
curl -X POST https://staging.everyrole.dev/hooks/deploy/YOUR-SECRET
The body is optional. JSON or form fields:
- base_url
- Test another address than the app’s own, like a preview deploy (Pro and Team). Only its scheme and host are used.
- environment
- production (the default for the app’s own address), preview (the default for another address) or staging. Tests that change data run only in the environments you allowed when you added them.
- commit, ref, url
- Shown with the run so you know which deploy it tested.
The answer is 202 with the queued job. Two deploys in a row while a run is still waiting don’t start two runs: the waiting run tests the newest deploy.
GitHub Actions (works with Vercel and Netlify previews)
Vercel and Netlify report each deploy to GitHub. This workflow runs your tests on every successful one, with the deploy’s own address:
name: EveryRole
on: deployment_status
jobs:
test:
if: github.event.deployment_status.state == 'success'
runs-on: ubuntu-latest
steps:
- run: |
curl -fsS -X POST "${{ secrets.EVERYROLE_DEPLOY_HOOK }}" \
-H "Content-Type: application/json" \
-d "{\"base_url\": \"${{ github.event.deployment_status.environment_url }}\",
\"environment\": \"${{ github.event.deployment_status.environment == 'Production' && 'production' || 'preview' }}\",
\"commit\": \"${{ github.sha }}\"}"
Netlify
In your site’s settings, under deploy notifications, add an outgoing webhook for “Deploy succeeded” with your deploy link as the URL. It tests your app’s own address. For deploy previews, use the GitHub Actions workflow above instead.
Vercel
With Vercel’s Git integration, use the GitHub Actions workflow above: Vercel reports every production and preview deploy to GitHub.
Any other pipeline
Add the curl line as the last step after your deploy succeeds.
Use EveryRole from your coding agent
Your coding agent (Claude Code, Cursor, Codex) can check its own work before it says it’s done: it runs the tests you added to your app, as every kind of user, on a preview deploy, and reads what broke, where, and the replay. It uses your workspace’s apps, tests and credits, with the same plan limits as the app.
1. Make a key
In Settings → API keys, the workspace owner makes a key and copies it (it’s shown once). Give each agent or machine its own key: you see what each one used, and can set a monthly credit cap per key or revoke it at any time.
2. Call the API
Send the key as a bearer token. Answers are JSON.
# your apps, their kinds of user and tests
curl -H "Authorization: Bearer $EVERYROLE_API_KEY" https://staging.everyrole.dev/api/v1/apps
# run an app's tests on a preview deploy (both fields optional) -> {"job": 42, ...}
curl -X POST -H "Authorization: Bearer $EVERYROLE_API_KEY" -H "Content-Type: application/json" -d '{"base_url": "https://my-app-git-branch.vercel.app", "environment": "preview"}' https://staging.everyrole.dev/api/v1/apps/YOUR-APP/runs
# wait for it: status, progress, and the run's id when done
curl -H "Authorization: Bearer $EVERYROLE_API_KEY" https://staging.everyrole.dev/api/v1/jobs/42
# the results: per test its verdict, the failing step and a link to its replay
curl -H "Authorization: Bearer $EVERYROLE_API_KEY" https://staging.everyrole.dev/api/v1/runs/7
POST /api/v1/apps/YOUR-APP/explore explores the app again, so new pages get suggested tests. A preview address must be on the public internet, and testing one is part of Pro and Team; for localhost, use the open-source runner.
- 401
- No key, a wrong one, or a revoked one.
- 402
- Not enough credits, or the key’s monthly cap is reached.
- 403
- Your plan doesn’t include it (a preview address, more explorations this month).
- 409
- A job of that kind is already waiting or running for the app (its id is in the answer), or the app has no tests yet.
- 422 · 429
- A bad address or environment · too many requests from this key in a minute: wait and try again.
3. Or connect it as an MCP server
EveryRole is an MCP server at https://staging.everyrole.dev/mcp: nothing to install, your key as a bearer token. Its tools: setup_guide, list_apps, get_app, create_app, set_test_accounts, set_environments, explore, list_suggestions, add_suggestions, dismiss_suggestion, reword_suggestion, run_tests, job_status, run_results, get_deploy_link and setup_pipeline. In Claude Code:
claude mcp add --transport http everyrole https://staging.everyrole.dev/mcp --header "Authorization: Bearer er_live_..."
In Cursor (.cursor/mcp.json) and other agents that read the same format:
{
"mcpServers": {
"everyrole": {
"url": "https://staging.everyrole.dev/mcp",
"headers": {"Authorization": "Bearer er_live_..."}
}
}
}
An agent that only runs local servers can use the same tools from the package: pip install everyrole, then claude mcp add everyrole -e EVERYROLE_API_URL=https://staging.everyrole.dev -e EVERYROLE_API_KEY=er_live_... -- everyrole mcp.
Keep the key out of your repository: put it in your agent’s settings or environment, not in a file you commit. Each call counts toward the key’s requests a minute and shows in its usage, like the API.
Ask Claude to set it up
With EveryRole connected, say to your agent: “Add EveryRole to this project.” It follows EveryRole’s setup guide (the setup_guide tool, also the set_up_everyrole prompt):
- Add the app from its public address (create_app). A first look finds its sign-in page and other addresses.
- Test accounts: it asks you for one account per kind of user (free, paid, admin…) and saves them (set_test_accounts). They’re encrypted like in the app and never shown again, not even to the agent.
- Explore the app as every kind of user (explore). Each suggested test is run twice for real first.
- Add the suggestions you want (add_suggestions). “All that worked twice” never includes a test that changes data; those are added only by name, and run on preview deploys only. Nothing becomes a test unless it’s added.
- Run the tests (run_tests, then run_results). A broken test says its step, what was expected and what the page showed, so the agent can fix the code without watching the replay.
- Put it in your pipeline (setup_pipeline): it writes a GitHub Actions workflow into your repository and asks you for the secret to set. From then on, every deploy runs the tests and the result shows on the pull request. EveryRole never touches your repository itself.
As the app grows, new pages are explored after deploys and suggested; your agent (or you) adds the right ones. Before it says a change is done, it runs the tests on the change’s preview deploy.
Alerts
EveryRole tells you when something changes: a test that worked stopped working (“stopped working since run #12”), or a broken test works again. A run where everything stays the same sends nothing. “Couldn’t check” on its own never sends anything either: it isn’t your problem.
You always see alerts under Notifications. Email and Slack alerts are set in Settings → Alerts and in each app’s Settings tab. The first full run of an app is its starting point, so it only shows up in the app.
The check-up email is the calm one: proof your apps are watched, even on a day when all is well. When each app was checked (date and time), how each kind of user did, what stopped working or works again, and what waits for you; never your credits. Choose every day, every week, only when something changes, or never, the hour and time zone, and which kinds of user, in Settings → Alerts.
Credits
One test run as one kind of user uses 1 credit, one page opened for the access grid 0.25, and one exploration 1 credit per page it explores, as each kind of user. Sign-ins are free, and so is everything that couldn’t be checked because of a problem on our side.
When a job starts, the most it can use is set aside; when it ends, you pay only for what was really checked. Your plan’s credits start fresh every month and don’t carry over; bought credits never expire. See pricing.
Security
- Passwords are encrypted at rest. Test-account emails and passwords are stored encrypted (Fernet: AES with an HMAC) with a key only the server holds. They are decrypted only to hand them to the test run that signs in, as environment variables of that one process, and are never written to disk in clear.
- AI models never see them. Passwords are typed by EveryRole itself; models only ever see the name of a value (“password”), and every request to a model passes a filter that replaces any secret with a placeholder.
- Exploring can’t change your app. While exploring, every request that would save (POST, PUT, PATCH, DELETE, form submissions, beacons) is blocked in the browser and recorded.
- Tests only do what you approved. Tests that change data run only in the environments you chose, and never-click words are enforced before every click.
- Your account. Passwords are hashed with Argon2id, sessions use HttpOnly cookies, every form is protected against cross-site requests, and each workspace can only see its own apps.
- Known limits. The save-blocker works inside the page: it doesn’t see service-worker traffic, websockets, or apps that change data on a plain page load. Keep risky words in the never-click list and run data-changing tests on preview deploys.