Excellent verifies agent work
before it ships
It runs beside the AI agents you already use, checks each task against independent evidence, and keeps a receipt of what passed.
Works with the agent you already use
Your agent says
the task is done
Excellent says
not yet
A receipt, and so much more
Gates that run before anything is trusted
Run the focused proof for each change: tests, lint, build, browser, API, CI, and deployment. Name the checks a task must pass before its result counts.
Tell your agent
Bring proof from your stack
Connect the systems that can prove real engineering work: git, CI, browsers, deployments, and your issue tracker.
Tell your agent
Verified, returned, or unchecked
Every attempt lands as a result you can act on. Failures keep their detail, and work that could not be checked is never counted as a pass.
Tell your agent
Proof attached to the work
Commands, exit codes, logs, screenshots, diffs, and CI runs are captured against the exact attempt that produced them.
Tell your agent
Every agent run captured
What the agent was asked to do, what it claimed, what it changed, and how many passes it took. Retries keep the original context.
Tell your agent
The task before the attempt
Define what the agent is about to attempt: scope, acceptance criteria, repository context, and the checks that must prove it.
Tell your agent
Keep the record on a server you run
Run a shared primary your team controls, at an address like proof.your-company.dev. It is a small box on your own infrastructure — Docker, a $4–12/month VPS billed by your own provider — and nothing passes through us. Team sync is beta and hands-on: a teammate joins from the command line, not the in-app switch, and you paste the device line they send you into the server's allowlist file by hand. There is no roster API. Deleting that line revokes the device, with no restart. Verification evidence does not sync yet, so a teammate can see the work, not the proof behind it.
- invoices-csv-exportMK
- auth-rate-limitJT
- checkout-webhookSR
- search-paginationAL
- schema-migrationDP
- upload-validationEW
- session-expiryNB
- invoice-pdf-exportRG
- rate-limit-headersTC
Example receipts
We build Excellent with Excellent
We have not published a verified improvement yet — the current build status is in the changelog. The receipts below are examples of what the system produces, not results from your repository.
- ReturnedExample
The export ignored the active invoice filter. Expected 12 rows, found 47.
invoices-csv-export1 of 6 checks failed
- VerifiedExample
Second attempt passed all six checks, including the browser check on the filtered view.
invoices-csv-export6 checks
- ReturnedExample
Tests passed, but the build failed on a type error in the new route.
search-paginationbuild
- ReturnedExample
The rate limit header was set but never read. The browser check caught the retry loop.
auth-rate-limitbrowser
- VerifiedExample
Migration applied cleanly and the rollback path was exercised before accept.
schema-migration8 checks
- UncheckedExample
No deploy target is configured, so the deployment check could not run.
checkout-webhook1 skipped
- ReturnedExample
Upload accepted a zero-byte file. A validation error was expected.
upload-validation2 failed
- VerifiedExample
Session expiry now clears the refresh token. Checked against the live API.
session-expiry5 checks
- ReturnedExample
The agent called it done with the dev server still holding the port.
invoice-pdf-export1 failed
- VerifiedExample
Lint and tests passed, and the screenshot diff showed no unintended visual change.
rate-limit-headers7 checks
- VerifiedExample
It worked, but only on the third attempt. The earlier failures are kept on the receipt.
webhook-retry3 attempts
- UncheckedExample
Staging database was unreachable, so the integration check is unproven, not passing.
order-sync2 skipped
- RecommendedExample
Candidate prompt beat the current setup on held-out work, with no hard regressions.
candidate-promptheld-out
- RecommendedExample
A smaller model for repo reading held quality at a lower cost per verified result.
model-splitcost
- Not promotedExample
Candidate looked better on average but regressed two hard cases. Kept the current setup.
candidate-prompt-bregression
- RecommendedExample
Browser checks after each UI change caught failures the end-of-run check missed.
check-policyquality
- RecommendedExample
A new retry rule needed less human help per task on held-out work.
retry-rulehelp
- Not promotedExample
A larger context window cost more for the same verified rate. Rejected.
context-sizecost
- Not promotedExample
The held-out set surfaced a regression the summary score had averaged away.
candidate-prompt-cregression
- VerifiedExample
The agent setup approved at the time of the change is recorded on the receipt.
release-receiptreceipt
- MixedExample
The same task ran on three agents. Two verified, one returned.
cross-agent-check3 runs
- VerifiedExample
Evidence bundle exported, then checked offline a week later. Still matched.
audit-exportoffline
- ReturnedExample
A reviewer note on the returned attempt records the risk that was knowingly accepted.
legacy-importaccepted risk
- VerifiedExample
The git tree matched the receipt exactly. No untracked changes at accept time.
billing-proration6 checks
FAQ
A verification system for agentic tasks. It runs beside the agent you already use, checks the work with independent evidence, and leaves a receipt you can inspect later.
No. Your agent does the work. Excellent starts the attempt, records what happened, checks the result against your repo and tools, and compares different ways of running it.
The one you already use — Claude Code, Codex, Cursor, Copilot, Cline, Gemini, OpenCode, or any other agent that does work you can check. Excellent observes the run instead of replacing it.
Whatever can prove the task: the git tree, tests, lint, the build, browser behavior, API responses, CI, and the deployment. Git and test evidence are the most settled; browser, CI, and deployment observers are newer. When a check cannot run, Excellent says so instead of implying a pass.
The attempt is returned with the exact failure — expected behavior and observed behavior — and sent back to the agent. The original context and the retry history are kept on the receipt.
None of it reaches us. Excellent keeps your repository, attempts, evidence, and receipts in a local database on your machine, and sends none of it to Excellent. Model calls are the exception worth naming: they go from your machine to the provider you chose, on your own account — the same traffic your agent already makes. The security page has both paths in full.
Yes. You install it and run it on your own machine, alone or as a team. There are no seats to buy and nothing metered for local verification.
You pay your model provider directly, never us. Excellent uses the API key or the agent login you already have — read from your environment or sealed with AES-256-GCM in your OS keychain, never from a plaintext config file — and the calls bill to your account. You can also run through your existing agent CLI and give Excellent no key at all.
Yes, in beta, and it is off until you turn it on. Teammates sync to a small server you provision and operate yourself, so nothing passes through us. Two caveats we would rather say first. Joining is hands-on, and it happens at the command line: the in-app switch takes a relay address but no workspace id, so it cannot join an existing workspace. A teammate runs sync setup --team-bundle with the line you send them, then sync pair, and sends you back the device line it prints — which you paste into the server's device allowlist by hand, because there is no roster API. Deleting that line revokes the device, with no restart. And verification evidence and receipts do not sync yet, so a teammate cannot open the proof behind your result. The security page says exactly what does and does not travel.
We are building Excellent with Excellent, and we have not published a verified improvement yet. The results page says so plainly and will carry the numbers when one exists. Whatever we publish will be our results on our repository, not a promise about yours.