Skip to content

Excellent verifies agent work
before it ships

It runs beside the AI agents you already use, checks each task against independent evidence, and keeps a receipt of what passed.

Works with the agent you already use

  • OpenClaw
  • Claude Code
  • Codex
  • Cursor
  • Hermes
  • Cline
  • Grok
  • Perplexity
  • OpenCode
  • Antigravity
  • Gemini
  • GitHub Copilot
  • Pi
  • Muse
  • Instinct

Your agent says
the task is done

Command:

Checked in 8.4s · 6 checks

Returned — 1 check failed

Excellent says
not yet

receipt · csv-export

A receipt, and so much more

Gates that run before anything is trusted

Run the focused proof for each change: tests, lint, build, browser, API, CI, and deployment. Name the checks a task must pass before its result counts.

Tell your agent

  • “Require a browser check on this task”
  • “Do not accept this until the build passes”
  • “Re-run the failing check and show me the diff”
  • “What did you skip, and why?”

Learn more about checks

Keep the record on a server you run

Run a shared primary your team controls, at an address like proof.your-company.dev. It is a small box on your own infrastructure — Docker, a $4–12/month VPS billed by your own provider — and nothing passes through us. Team sync is beta and hands-on: a teammate joins from the command line, not the in-app switch, and you paste the device line they send you into the server's allowlist file by hand. There is no roster API. Deleting that line revokes the device, with no restart. Verification evidence does not sync yet, so a teammate can see the work, not the proof behind it.

  • invoices-csv-export
    MKClaude Code
  • auth-rate-limit
    JTCodex
  • checkout-webhook
    SRGrok
  • search-pagination
    ALHermes
  • schema-migration
    DPOpenClaw
  • upload-validation
    EWPi
  • session-expiry
    NBClaude Code
  • invoice-pdf-export
    RGCodex
  • rate-limit-headers
    TCGrok

Tell your agent

  • “Set up our team server”
  • “Add this device to the allowlist”
  • “Revoke the laptop I just lost”

How team sync works

Example receipts

We build Excellent with Excellent

We have not published a verified improvement yet — the current build status is in the changelog. The receipts below are examples of what the system produces, not results from your repository.

  • ReturnedExample
    The export ignored the active invoice filter. Expected 12 rows, found 47.
    Claude Codeinvoices-csv-export1 of 6 checks failed
  • VerifiedExample
    Second attempt passed all six checks, including the browser check on the filtered view.
    Claude Codeinvoices-csv-export6 checks
  • ReturnedExample
    Tests passed, but the build failed on a type error in the new route.
    Codexsearch-paginationbuild
  • ReturnedExample
    The rate limit header was set but never read. The browser check caught the retry loop.
    Cursorauth-rate-limitbrowser
  • VerifiedExample
    Migration applied cleanly and the rollback path was exercised before accept.
    Claude Codeschema-migration8 checks
  • UncheckedExample
    No deploy target is configured, so the deployment check could not run.
    Codexcheckout-webhook1 skipped
  • ReturnedExample
    Upload accepted a zero-byte file. A validation error was expected.
    Clineupload-validation2 failed
  • VerifiedExample
    Session expiry now clears the refresh token. Checked against the live API.
    Geminisession-expiry5 checks
  • ReturnedExample
    The agent called it done with the dev server still holding the port.
    OpenCodeinvoice-pdf-export1 failed
  • VerifiedExample
    Lint and tests passed, and the screenshot diff showed no unintended visual change.
    GitHub Copilotrate-limit-headers7 checks
  • VerifiedExample
    It worked, but only on the third attempt. The earlier failures are kept on the receipt.
    Claude Codewebhook-retry3 attempts
  • UncheckedExample
    Staging database was unreachable, so the integration check is unproven, not passing.
    Codexorder-sync2 skipped
  • RecommendedExample
    Candidate prompt beat the current setup on held-out work, with no hard regressions.
    Claude Codecandidate-promptheld-out
  • RecommendedExample
    A smaller model for repo reading held quality at a lower cost per verified result.
    Codexmodel-splitcost
  • Not promotedExample
    Candidate looked better on average but regressed two hard cases. Kept the current setup.
    Claude Codecandidate-prompt-bregression
  • RecommendedExample
    Browser checks after each UI change caught failures the end-of-run check missed.
    Cursorcheck-policyquality
  • RecommendedExample
    A new retry rule needed less human help per task on held-out work.
    Claude Coderetry-rulehelp
  • Not promotedExample
    A larger context window cost more for the same verified rate. Rejected.
    Geminicontext-sizecost
  • Not promotedExample
    The held-out set surfaced a regression the summary score had averaged away.
    OpenCodecandidate-prompt-cregression
  • VerifiedExample
    The agent setup approved at the time of the change is recorded on the receipt.
    Claude Coderelease-receiptreceipt
  • MixedExample
    The same task ran on three agents. Two verified, one returned.
    Cursorcross-agent-check3 runs
  • VerifiedExample
    Evidence bundle exported, then checked offline a week later. Still matched.
    GitHub Copilotaudit-exportoffline
  • ReturnedExample
    A reviewer note on the returned attempt records the risk that was knowingly accepted.
    Clinelegacy-importaccepted risk
  • VerifiedExample
    The git tree matched the receipt exactly. No untracked changes at accept time.
    Claude Codebilling-proration6 checks

FAQ

A verification system for agentic tasks. It runs beside the agent you already use, checks the work with independent evidence, and leaves a receipt you can inspect later.

No. Your agent does the work. Excellent starts the attempt, records what happened, checks the result against your repo and tools, and compares different ways of running it.

The one you already use — Claude Code, Codex, Cursor, Copilot, Cline, Gemini, OpenCode, or any other agent that does work you can check. Excellent observes the run instead of replacing it.

Whatever can prove the task: the git tree, tests, lint, the build, browser behavior, API responses, CI, and the deployment. Git and test evidence are the most settled; browser, CI, and deployment observers are newer. When a check cannot run, Excellent says so instead of implying a pass.

The attempt is returned with the exact failure — expected behavior and observed behavior — and sent back to the agent. The original context and the retry history are kept on the receipt.

None of it reaches us. Excellent keeps your repository, attempts, evidence, and receipts in a local database on your machine, and sends none of it to Excellent. Model calls are the exception worth naming: they go from your machine to the provider you chose, on your own account — the same traffic your agent already makes. The security page has both paths in full.

Yes. You install it and run it on your own machine, alone or as a team. There are no seats to buy and nothing metered for local verification.

You pay your model provider directly, never us. Excellent uses the API key or the agent login you already have — read from your environment or sealed with AES-256-GCM in your OS keychain, never from a plaintext config file — and the calls bill to your account. You can also run through your existing agent CLI and give Excellent no key at all.

Yes, in beta, and it is off until you turn it on. Teammates sync to a small server you provision and operate yourself, so nothing passes through us. Two caveats we would rather say first. Joining is hands-on, and it happens at the command line: the in-app switch takes a relay address but no workspace id, so it cannot join an existing workspace. A teammate runs sync setup --team-bundle with the line you send them, then sync pair, and sends you back the device line it prints — which you paste into the server's device allowlist by hand, because there is no roster API. Deleting that line revokes the device, with no restart. And verification evidence and receipts do not sync yet, so a teammate cannot open the proof behind your result. The security page says exactly what does and does not travel.

We are building Excellent with Excellent, and we have not published a verified improvement yet. The results page says so plainly and will carry the numbers when one exists. Whatever we publish will be our results on our repository, not a promise about yours.

Try it on a real repo