Who Checks the Tests Your AI Agent Wrote? The QA Loop and Verified Tests
Last updated: 2026-09-30
TL;DR: When an AI agent writes the code and the tests, green proves little, because the agent is checking its own work. TestChimp runs the QA process as a feedback loop: it surfaces product risks and coverage gaps, agents work through them, and results flow back in. Governance signals (Verified Tests, TrueCoverage, semantic coverage and API contract coverage) show what is actually covered and who verified it. Meet-Bots and AgentWatch keep requirements current, so tests are written against up-to-date intent.
Why do agent-written green tests prove nothing?
Because the agent grades its own work. If the model misreads a ticket, its tests misread it the same way, and the suite stays green while the product does the wrong thing. Practitioners describe this as the agent "marking its own homework".
Three failure modes show up once agents write both sides (they are described in the Verified Tests announcement):
| Failure mode | What goes wrong |
|---|---|
| Linked is not proven | The test is tagged to a scenario but never asserts the expected outcome |
| Happy-path impersonation | A green run for a neighbouring flow is counted as coverage of the hard one |
| Unowned claims | Coverage says "covered", and nobody can say who looked at the script |
Green is an execution fact. A scenario link is a mapping. Neither tells you that anyone checked whether the test verifies the behaviour. Before agents, "who wrote this test?" was a workable stand-in for "who checked it?". With agents writing tests at speed, that stand-in no longer holds.
The answer has three parts: a QA loop that keeps agents working on the real risks, governance signals that show what is covered and who verified it, and requirements that stay current so tests are aimed at the right target.
How does TestChimp run the QA loop?
TestChimp is a QA workflow layer for agents. Per the docs, it orchestrates a tight feedback loop so coding agents such as Claude continuously understand product risk, coverage gaps and where users actually spend time, and then act on them. Humans stay in control of intent (what the product should do), and agents do the execution and upkeep against that intent.
In practice the loop looks like this:
- Plans live in your repo. Stories and scenarios are markdown files in a mapped
plans/folder. Plans and tests sync both ways between TestChimp and your repo, and changes arrive as pull requests or merge requests you review. - Per PR, run QA. The
/testchimp run QAworkflow chains plan, environment, SmartTest and smoke steps, with optional exploratory and TrueCoverage steps, under one scope. - In CI, results are reported back. The
@testchimp/playwrightplugin reports runs and traces to TestChimp and tags user events with test identity for TrueCoverage. - On a cadence, upkeep.
/testchimp upkeepcloses coverage gaps found from requirements and TrueCoverage, explores high-signal paths and cleans up.
The web app (or TestChimp Studio on the desktop) is the control plane where the risk list is visible: uncovered scenarios, coverage gaps, bugs found by exploration and workflow executions. Details: What is TestChimp? and Upkeep.
Green run alone vs green run with governance signals
| Green run on its own | With governance signals | |
|---|---|---|
| What it tells you | The tests that exist passed | Which scenarios, API operations and user events are covered, and by which tests |
| What a coverage report counts | Tests | Scenarios with linked tests, API operations and schema fields exercised, production events seen in test runs |
| Who looked at the test | Unknown | A Verified Tests badge shows who inspected it |
| Near-duplicate tests | Look like extra coverage | Surfaced by the semantic graph |
| Gaps | Found by accident | Listed and actionable: uncovered scenarios, uncovered API operations, production events with no test coverage |
The next sections cover each signal.
What is a Verified Test?
A verified test in TestChimp is a SmartTest a project member has inspected and confirmed actually covers a linked scenario, not merely that it ran green or is annotated.
(A SmartTest is a standard Playwright test in your repo with scenario links added.) To verify one, a reviewer opens a test execution, looks at the screen captures and steps from the run, reads the test code with View Test, and then uses the check badge next to the test name. Status is stored in TestChimp, not in the Playwright source, so agents keep authoring tests while humans record that they inspected them.
The badge has three states on a test:
- Hollow: none of the linked scenarios are verified.
- Grey: some, but not all, linked scenarios are verified.
- Blue: every linked scenario for that test is verified.
On coverage views, each scenario shows a filled badge when a covering test has been verified, a hollow badge when tests exist but nobody has inspected them, and no badge when there are no linked tests. Hovering a verified badge shows who verified it. Un-verifying is recorded as "manually unverified", so the audit trail stays.
Two honest details matter for governance:
- Stale markers. A blue badge with a red "!" means the test was verified, but an agent later patched it. By default, when an agent reports a fix to an existing test, the badge goes stale until someone re-verifies it. This can be changed in project settings.
- Inspection is deliberate. Clicking the coverage-side badge does not verify anything. It points you to the execution viewer, so verification stays an inspection action and not a bulk paint on a dashboard.
Details: Verified Tests and Link tests to scenarios.
What other signals show what is actually covered?
Verified Tests answer "did a person check this test?". Three more signals answer "what is covered at all?":
- API contract coverage (API schema coverage). Connect your OpenAPI specs from git, run SmartTests with API capture enabled, and TestChimp shows per-operation coverage plus request, query and response field and response-code coverage. Traffic that hits paths or fields the spec never declared is marked as undocumented. Each gap can become an issue, a create-tests prompt for your agent, or a deliberate ignore. See API contract governance.
- TrueCoverage. The same user-event taxonomy is overlaid from production (via RUM SDKs) and from test runs, so you see which real user pathways have no test coverage, with metrics such as demand, duration, drop-off and depth. See TrueCoverage.
- Semantic coverage. Stories, scenarios, SmartTests, issues and TrueCoverage events are placed in one embedding space in the Semantic Canvas, so you can see what is close in meaning, whether it has a linked test, and which tests are near-duplicates. See Semantic Canvas.
Together with requirement coverage (next sections), these show where coverage is real, where it is thin, and where it is duplicated.
How do requirements stay current? (Meet-Bots, AgentWatch)
Tests can only be aimed at the right target if the requirements do not rot. Much of what a team knows is decided out loud or in chat, then never written down. Meet-Bots and AgentWatch keep the QA surface current from the two places those decisions now happen, so the tests agents write are written against up-to-date requirements:
- Meet-Bots join Google Meet, Zoom, Microsoft Teams and Webex calls, transcribe once, and keep the meeting notes as tribal knowledge. After the call, the transcript can be handed to a workflow that drafts user stories and scenarios from what the team agreed. Team members can also @mention the bot for answers grounded in the stories, scenarios and test results. See Meet-Bots and Meetings, PII, and follow-ups.
- AgentWatch (in TestChimp Studio) watches local coding-agent sessions such as Cursor and Claude Code, spots when business rules or UX expectations were decided in chat, and proposes story and scenario updates. You review and approve the proposed plan, the
plans/folder updates, and you raise a PR. See AgentWatch and Enable and configure AgentWatch.
Both mechanisms propose; people approve. Stories and scenarios are markdown files with frontmatter in your Git repo (see Requirement planning), and plans sync both ways between TestChimp and the repo, so every requirement change is a reviewable diff. The tests are then authored and maintained against those scenarios, for example by the create-tests workflow, which writes tests for requirements or API coverage gaps. See Author plans and Create tests.
How does traceability connect tests to scenarios?
Traceability is code-native. A test declares which scenario it covers with a Playwright annotation on the test:
test('successful login', {
annotation: [{ type: 'scenario', description: '#TS-101' }],
}, async ({ page }) => {
// ...
});
TestChimp links the test to the scenario by its #TS-<n> ordinal and tracks execution results from CI runs (via the @testchimp/playwright reporter). One test can list several scenarios. In Test Planning, the Insights tab shows requirement coverage for any folder, filterable by environment, release and time range, including scenarios with no linked tests, scenarios with no recent executions, and failing scenarios. Renaming a scenario does not reset verification as long as its ordinal stays the same; linking a different scenario is a new claim and starts unverified.
Put together, the chain is: scenario (kept current) -> linked test -> execution -> human verification. Each link is visible, and the last one has a named person and a time.
Details: Requirement traceability and Traditional traceability vs TestChimp.
What should a team review before release?
A short checklist that maps to what the platform shows:
- Uncovered scenarios. Any high-priority scenario with no linked test (Insights, Requirement Coverage, filtered to the release scope).
- Failing or unexecuted scenarios. Scenarios that have tests but no recent execution, or whose latest run is red.
- Hollow badges. Scenarios with tests that nobody has inspected. Decide which ones need a human look before this release.
- Stale badges. Tests verified earlier that an agent has since patched.
- API and user-behaviour gaps. Uncovered API operations or fields, and production events with no test coverage in TrueCoverage.
- Requirement drift. Pending AgentWatch plans or meeting follow-ups that have not been reviewed yet, so the scenarios may not reflect what was decided.
- Release evidence. The release view brings together test runs, requirement coverage, ExploreChimp findings and release checks; a gate can also be queried programmatically with an API key.
See Release management, Release intelligence and Programmatic release gating.
What does this not solve?
- Verification is a human sanity check, not a proof of correctness. A reviewer can still approve a weak test.
- The loop assumes the scenarios are right. TestChimp helps keep them current and can score their quality (see Requirement quality governance), but a wrong requirement stays wrong until someone notices.
- Capture depends on the team using the bots and enabling AgentWatch. Decisions made outside meetings and agent sessions are not seen.
- Coverage signals depend on setup: API capture is off by default, and TrueCoverage needs the RUM SDK instrumented in your app.
Frequently asked questions
How do I trust tests written by an AI agent?
Do not rely on a green run alone. Look at what is actually covered (linked scenarios, API contract coverage, TrueCoverage, semantic coverage), and have a person inspect the run and the code and record that the test verifies the scenario.
Who should review AI-generated tests?
A project member who can judge the requirement, such as a QA lead or the engineer who owns the feature. In TestChimp they open the test execution, inspect the screen captures, steps and code, and mark the test as verified.
Is a passing test the same as a verified test?
No. A passing test is an execution fact, and a scenario link is a coverage claim, while a verified test has been inspected by a person who confirmed it covers the linked scenario.
What happens to the Verified badge when an agent edits the test later?
By default the badge shows a stale marker when an agent reports a fix to the test, so someone can re-verify it. Teams can change this in project settings.
What does TestChimp do beyond verifying tests?
It runs the QA loop: it orchestrates a feedback loop so agents keep seeing product risk and coverage gaps and act on them, while people own intent. Meet-Bots and AgentWatch keep requirements current, so tests are written against up-to-date requirements.
What coverage signals does TestChimp show for AI-written tests?
Requirement coverage by folder, environment, release and time range; API contract coverage for OpenAPI operations and fields; TrueCoverage for production user events versus test runs; and semantic coverage through the Semantic Canvas. Verified Tests adds who inspected each test.
Related docs
- What is TestChimp?
- Verified Tests (product guide)
- Verified Tests: When Agents Write the Coverage, Trust but Verify (announcement)
- Link tests to scenarios
- Requirement traceability
- API contract governance
- TrueCoverage
- Semantic Canvas
- Requirement planning as code
- Meet-Bots
- AgentWatch
- Release management
- Changelog
- Next article: Exit-safe by design: your tests and plans live in your git repo
- Plans and pricing: testchimp.io/pricing
