How to Trust AI-Written Tests
A green test written by the same agent that wrote the code only shows the two agree. To trust AI-written tests, run QA as a closed feedback loop: write tests against current, independently stated requirements, check what is actually covered, and have a named person verify each test. TestChimp does this with Verified Tests and coverage signals.
Last updated: 2026-10-01
Why is a green test not evidence?
Because a test only tells you that the code and the test agree with each other. When one AI agent writes both, a misread requirement gets baked into the code and into the test in the same way. The suite stays green while the product does the wrong thing.
Before agents, "who wrote this test?" was a workable stand-in for "who checked it?". When agents write tests at speed, that stand-in stops working. Green is an execution fact. A link from a test to a requirement is a mapping. Neither one says that anybody checked whether the test verifies the behavior.
What are the common failure modes of AI-written tests?
| Failure mode | What it looks like |
|---|---|
| Mirror assertions | The test asserts whatever the code currently returns, so it can never disagree with the code |
| Boundaries defined in code, not in the ticket | Limits, rounding rules and error cases come from the implementation, not from what the team decided |
| Assertions weakened to make a fix pass | A failing test is "repaired" by loosening the check instead of fixing the product |
| Coverage numbers that mislead | A high percentage counts lines or tests executed, while important behaviors have no meaningful assertion |
| Linked but never proven | A test is tagged to a scenario but never asserts the expected outcome |
| Unowned claims | A dashboard says "covered" and nobody can say who looked at the script |
These are patterns practitioners report, not a TestChimp-specific list. The last two are the ones TestChimp's Verified Tests feature was built around.
Does mutation testing solve it?
Mutation testing is a real and useful answer, and it is worth using. It changes your code in small ways and checks whether your tests fail. If a test still passes when the code is broken, that test is weak. It is a good check on whether a test is sensitive to change.
What it cannot tell you:
- Whether the test checks the right behavior. A test can fail on every mutation and still assert something the team never asked for.
- Whether the requirement itself is current and correct.
- Who looked at the test and decided it covers the scenario.
So mutation testing and human verification answer different questions. A sensible stack uses both: mutation testing for "would this test notice a break?", and a recorded human check for "is this the test we meant to have?".
How do the options compare?
| Approach | Independent of the code? | Human sign-off recorded? | Traceable to a requirement? | State visible later? |
|---|---|---|---|---|
| Self-graded AI tests (green run only) | No | No | Only if someone links them | Pass or fail only |
| Mutation testing | Partly (checks sensitivity, not intent) | No | No | A score per run |
| Human review in pull requests | Yes, if the reviewer knows the requirement | In the PR approval, not tied to coverage | Only by convention | Lives in PR history |
| Verified Tests in TestChimp | Tests link to scenarios kept as Markdown in your repo | Yes, a named person and time | Yes, through the scenario annotation | Badge on the test and on coverage views |
None of these is a proof of correctness. They are layers.
What can you check today, without any tool?
Use this as a review checklist for any AI-written test:
- Find the requirement. Can you point to a written story or scenario the test is meant to cover, written before or independently of the code?
- Break the code. Change the behavior on purpose. Does the test fail? (Mutation testing automates this.)
- Read the assertions. Does the test assert the expected outcome from the requirement, or only that "something rendered"?
- Check the boundaries. Do the limits and error cases come from the requirement or from the implementation?
- Look at the diff of any "fix". If a test was changed to go green, was the assertion weakened?
- Look for neighbours. Is this a near-duplicate of another test, or a green run for a nearby flow counted as coverage of a harder one?
- Name an owner. Can you say who looked at this test and when?
How does TestChimp run the QA loop so tests can be trusted?
TestChimp runs the QA process as a closed feedback loop. It surfaces product risks and coverage gaps, agents work through them, and the results flow back in. Humans stay in control of intent (what the product should do), and agents do the execution and upkeep against that intent.
In practice:
- Plans live in your repo. Stories and scenarios are Markdown files in a mapped
plans/folder. Plans and tests sync both ways between TestChimp and your git repo, and changes arrive as pull requests or merge requests you review. - Per pull request, run QA. After a change, a developer runs the
/testchimp run QAworkflow, which chains planning, environment, test and smoke steps under one scope. - In CI, results are reported back. The
@testchimp/playwrightplugin reports runs and traces to TestChimp. - On a cadence, upkeep.
/testchimp upkeepcloses coverage gaps found from requirements and real user behavior.
How do requirements stay current? (Meet-Bots and AgentWatch)
Tests can only be aimed at the right target if the requirements do not rot. Meet-Bots and AgentWatch keep the QA surface current from the places where decisions are made, so tests are written against up-to-date requirements:
- Meet-Bots join Google Meet, Zoom, Microsoft Teams and Webex calls, transcribe the meeting and keep summaries and notes. Decisions can flow into stories and scenarios.
- AgentWatch (in TestChimp Studio) watches local coding-agent sessions such as Cursor and Claude Code, spots when business rules were decided in chat, and drafts updates to your stories and scenarios.
Both propose and people approve, so every requirement change is a reviewable diff in git. See Meet-Bots and AgentWatch.
How do Verified Tests work?
A verified test in TestChimp is a test a project member has inspected and confirmed actually covers a linked scenario, not merely that it ran green or is annotated. The reviewer opens a test execution, looks at the screen captures and steps from the run, reads the test code, and then uses the check badge next to the test name.
Badge states on a test:
| State | Meaning |
|---|---|
| Hollow | None of the linked scenarios are verified |
| Grey | Some, but not all, linked scenarios are verified |
| Blue | Every linked scenario for that test is verified |
| Blue with a red "!" (stale) | Verified earlier, but an agent later patched the test. By default, a reported fix makes the badge stale until someone re-verifies it. This can be changed in project settings |
On coverage views, each scenario shows a filled badge when a covering test has been verified and a hollow badge when tests exist but nobody has inspected them. Hovering a verified badge shows who verified it. Un-verifying is recorded, so the audit trail stays.
Verified Tests status is stored in TestChimp, not in the test source. It is the only thing in TestChimp that does not also live in your git repo. Details: Verified Tests.
Which governance signals show what is covered at release time?
Verified Tests answers "did a person check this test?". Three more signals answer "what is covered at all?":
- Semantic coverage. Stories, scenarios, tests, issues and TrueCoverage events are placed in one embedding space (the Semantic Canvas), so you can see what is close in meaning, whether it has a linked test, and which tests are near-duplicates. See Semantic Canvas.
- API schema coverage (API contract coverage). Connect your OpenAPI specs from git, run tests with API capture enabled, and see per-operation coverage plus request and response field coverage. Traffic that hits paths or fields the spec never declared is marked as undocumented. See API contract governance.
- TrueCoverage. User events from production and from test runs are overlaid, so you can see which real user pathways have no test coverage. See TrueCoverage.
Before a release, a short review maps to what the platform shows: uncovered scenarios, failing or unexecuted scenarios, hollow badges, stale badges, API and user-behavior gaps, and pending requirement updates from Meet-Bots or AgentWatch. See Release management.
On plans: according to the pricing page, TrueCoverage and API contract coverage are part of the Growth plan, and Team does not list them. Meet-Bots usage is bundled (20 hours on Team, 80 hours on Growth) and AgentWatch is included in both.
What does this look like in a real team?
A financial risk analytics company with 20 engineers and 1 QA engineer uses TestChimp this way. Developers run the /testchimp run QA workflow after each pull request. The QA engineer writes the stories and scenarios. The team reports 98% coverage across more than 1,000 scenarios in about 10 weeks.
This is customer-reported, and "coverage" here is the customer's own measure. It shows that a small QA function can set the intent while the loop does the execution, not that any coverage number replaces verification.
What does this not solve?
- Verification is a human sanity check, not a proof of correctness. A reviewer can still approve a weak test.
- The loop assumes the scenarios are right. A wrong requirement stays wrong until someone notices.
- Capture depends on the team using Meet-Bots and AgentWatch. Decisions made elsewhere are not seen.
- Some signals depend on setup: API capture needs to be turned on, and TrueCoverage needs the SDK instrumented in your app.
Frequently asked questions
Can AI write its own tests?
Yes, and it often writes good ones. The risk is that the same agent grades its own work. Keep intent with people, link each test to a written scenario, and have a person verify the tests that matter.
What is a Verified Test?
A test that a project member has inspected, confirming it actually covers a linked scenario. The status is recorded with who verified it and when, and it goes stale if an agent later patches the test.
Does mutation testing replace review?
No. Mutation testing shows whether a test would notice a break. It does not show whether the test checks the behavior the team wanted, or who approved it. Use both.
Who should sign off on AI-written tests?
A project member who can judge the requirement, such as a QA lead or the engineer who owns the feature.
What happens when the code changes after a test is verified?
By default, when an agent reports a fix to an existing test, the badge turns stale until someone re-verifies it. Teams can change this behavior in project settings.
Do I need TestChimp to follow this advice?
No. The checklist above works with any stack. TestChimp adds the loop that keeps requirements current, the coverage signals, and a recorded verification status.