Pro

AgentVera Verify: not “code was written” but “the work is done”

An agent saying “done” doesn’t mean the work is done. Verify pulls the task and its acceptance criteria from the message you send the agent; when the agent’s turn ends, a separate verifier looks at the repository, runs the tests and marks every criterion with evidence. There is no form to fill in.

This feature is in the Pro and Team plans. On the free plan it shows as locked; sign in with your AgentVera account and upgrade to unlock it. Compare plans

Criteria come from your message

A message you send an agent from the app is read by Claude Haiku. Each task gets 1 to 7 concrete, checkable outcomes: behaviour, passing tests, UI states, effects on data. Implementation steps don’t count as criteria and nothing beyond your message is added.

  • Independent jobs in one message become separate tasks; questions and chat produce no task.
  • Follow-ups like “go on” or “also fix this” are added to the current task, and failed criteria go back to pending.
  • Short replies like “ok”, “yes” or “continue” and messages under 12 characters are skipped.
  • You can edit the criteria, add your own or remove the check.

An independent check at turn end

When the agent ends its turn and actually worked since your last message, a verifier running on Claude Sonnet gets to work in the agent’s folder or worktree. Its instructions are strict: a criterion passes because it is really met, not because code was written; no evidence means it fails. The verifier can’t edit files.

  • Its tools: Read, Glob, Grep and allowlisted commands (git status, diff, log, show; npm, pnpm and yarn test; vitest, jest, pytest, go test, cargo test, php artisan test, phpunit).
  • Each criterion gets a pass or fail and an evidence note.
  • Failed criteria are checked again when the agent’s next turn ends; “Verify now” runs the check right away.

In the pane and in analytics

A small badge on the agent pane shows the state: extracting criteria, verifying, or a count like 3/5. Click it to see every criterion as passed, failed or pending, with its evidence note. With notifications on, the result also arrives as a desktop notification. The AI ROI tab works out success rate and first-check passes from Verify results.

How it works

  1. 1

    Tell the agent what to do and what counts as done; there is no separate form.

  2. 2

    The Verify badge on the pane shows the extracted criteria; edit them if you like.

  3. 3

    When the agent ends its turn the verifier runs and the badge shows how many passed.

  4. 4

    If some failed, have the agent fix them; they are checked again at the end of the next turn.

Questions

Are failed criteria sent to the agent automatically?

No. The result shows on the badge and in a notification, and you decide what to tell the agent. When it works again and ends its turn, the criteria are checked again.

Which agents does it work with?

Claude Code, Codex and the other CLI agents you message from the app. It doesn’t run on plain terminal panes.

Can the verifier change my code?

No. It can only use read tools and the allowlisted git and test commands.

Which models does it use?

Claude Haiku to extract criteria and Claude Sonnet to verify them, both through the Claude Code login on your computer.

Which plan includes it?

Verify is tied to the same Pro feature as Research, plan, build and the Task line, so it is on in the Pro and Team plans.

Related features

Related terms:Agentic codingVibe coding

All featuresSee the 18 supported coding agents

Bring your agents to one desk.

Start on the free plan. Download it, add a project folder and open your first agent.