You can build a SaaS with AI today: a coding agent like Claude Code or Codex can write sign-up, subscriptions, a dashboard and an API in days. But a working MVP needs more than “build me a SaaS”: a scope cut down to one core flow, a written spec the agent reads every session, and tests on the paths where money and data isolation live. This post walks through the process I follow, from scope to deploy.
The goal is a guide that someone who codes little or not at all can follow. Getting a prototype out with vibe coding is easy; keeping a product alive once someone types in a credit card is the hard part. So I’ll mark where you can move fast and where you need to slow down.
1. Cut the scope to one core flow
MVPs rarely die because of code. They die because of scope. The agent will write anything you ask for; the problem is that every feature you add has to be tested, maintained and debugged.
Before you start, write one sentence: “The user does ___ and gets ___ in return.” For example: “An agency uploads a client’s social posts, and the client approves them from a single link.” Everything outside that sentence goes to version two.
A first version usually needs only:
- Sign-up, login, password reset
- An organization (tenant) with invited members
- One core flow (the sentence above)
- One paid plan, monthly subscription
- A simple settings page and a link to the billing portal
Notification center, a role matrix, multiple languages, a mobile app, an admin panel: later. If customers pay for the core flow, they’ll tell you what else they need.
2. Write a spec the agent follows
Agents don’t carry memory between conversations. Claude Code reads CLAUDE.md at the project root every session; Codex and many other agents read AGENTS.md. Write your decisions there and every new session starts from the same rules. If you use both agents, keep the content in one file and point the other to it.
A useful spec is short and concrete:
# Project: ApproveLink (MVP)
## Product
Agencies upload posts; the client approves or comments from one link.
## Stack
- Next.js (App Router, TypeScript), Tailwind
- Supabase: Postgres, Auth, Storage
- Stripe: one plan, monthly subscription, Checkout + Customer Portal
- Tests: Vitest (unit), Playwright (end to end)
## Rules
- Every table has organization_id; RLS is on for every table.
- Subscription status changes only from the Stripe webhook.
- Database changes only through migration files.
- Secrets live in .env.local; never in code or logs.
- Ask before adding a new package.
- After every task: pnpm lint && pnpm test
## Out of scope (for now)
Roles, i18n, mobile, admin panel.
The “out of scope” section matters. Agents try to be helpful and add things you didn’t ask for; a written boundary cuts that down noticeably.
3. Pick the stack: choose boring
When you work with AI, the best stack is the one the agent has seen the most examples of. For this kind of MVP I’d go with:
- Next.js + TypeScript: Frontend and API in one project. TypeScript catches part of the agent’s mistakes at build time.
- PostgreSQL (Supabase or another managed Postgres): Relational data, row-level security (RLS) and migrations. Supabase also gives you auth and file storage in one place.
- Ready-made auth: Supabase Auth, Auth.js or Clerk. Don’t have the agent write password storage and session handling from scratch.
- Stripe subscriptions: Checkout for the payment page, Customer Portal for plan changes and cancellation. You never build a card form.
- Vercel or a similar platform: Deploys on every push and gives each branch a preview URL.
What to install: Node.js (LTS), pnpm, Git, a GitHub account, the Supabase CLI, the Stripe CLI and a coding agent (claude or codex).
pnpm create next-app@latest approvelink --typescript --tailwind --app
cd approvelink
git init && git add -A && git commit -m "chore: empty project"
Put the spec at the root of this folder and commit it too.
4. The first prompt: plan first, code later
My first message doesn’t ask for code; it asks for a plan. Reading and fixing a plan is much cheaper than reverting 40 wrongly written files.
Read CLAUDE.md. Don't write code yet.
Produce an implementation plan for this MVP:
1. Database tables (fields, relations, RLS policies)
2. Pages and API routes
3. The Stripe flow (which webhook events, where subscription status lives)
4. 6–8 small steps in order, each with what gets tested at the end
List anything unclear as questions at the end.
Read the plan, answer the questions, update the spec if needed. Then have it do the steps one at a time:
Implement step 1 of the plan: write the migration for the organizations,
memberships and projects tables and add the RLS policies. When done,
apply the migration to the local database, write a test proving two
different users can't see each other's data, and run it. Then stop.
“Then stop” is deliberate. After each step I look at the diff and commit, so when the agent drifts there’s a clean point to go back to.
5. Multi-tenant data isolation: the most dangerous part
The worst SaaS bug is “Customer A saw Customer B’s data.” The agent’s code usually looks fine because you test it with one user while building. The problem shows up when the second customer arrives.
My rules:
organization_idon every table. No exceptions, child tables included.- Let the database enforce the filter. On Supabase/Postgres, turn RLS on for every table and write the policy around the organizations the user belongs to. A
where organization_id = ...in app code isn’t enough on its own; the agent can forget it in the next query it writes. - The service role key stays on the server. It bypasses RLS. It must never reach code that ships to the browser; in Next.js, never put it in a
NEXT_PUBLIC_variable. - Test with two users. Create two organizations and, signed in as one, try to fetch the other’s record by its ID. This test should be automated and run on every change.
Ask the agent directly: “List every place in this project that bypasses RLS.” The answer should be server code using the service role and the webhooks; if there’s more, ask why.
6. Migrations: never change the database by hand
When the agent wants to add a column, it should do it with a migration file, not by clicking in a dashboard. With the Supabase CLI:
supabase start # local Postgres (needs Docker)
supabase migration new add_projects # empty migration file
supabase db reset # apply all migrations from scratch
supabase db push # apply to the remote database
Run db reset regularly: if migrations don’t apply cleanly from zero, production will have problems too. Once you’re live, read every migration the agent writes. If you see DROP COLUMN, DROP TABLE, an UPDATE without WHERE, or a locking index on a large table, stop and ask why. A migration that deletes data can’t be undone.
7. Test the money paths
You don’t need tests for everything; MVP time is short. But these paths shouldn’t reach production untested:
- Sign-up → create organization → core flow. End to end with Playwright.
- Stripe webhooks. Does subscription status update correctly on
checkout.session.completed,customer.subscription.updated,customer.subscription.deletedandinvoice.payment_failed? - Access without payment. When a subscription is canceled or a payment fails, is the paid feature actually closed?
- Data isolation. The two-user test from above.
Try Stripe webhooks locally like this:
stripe login
stripe listen --forward-to localhost:3000/api/webhooks/stripe
stripe trigger checkout.session.completed
A prompt for the agent:
Review the Stripe webhook route. Is the signature verified? What happens
if the same event arrives twice (idempotency)? Which table holds
subscription status? Write tests for these four events:
checkout.session.completed, customer.subscription.updated,
customer.subscription.deleted, invoice.payment_failed.
Run them and show me the result.
A common mistake: the agent activates the subscription on the “success” page Stripe redirects to. Anyone can open that URL by hand. Status must change only from a signature-verified webhook.
8. Deploy and monitoring
On a platform like Vercel the flow is simple: push the repo to GitHub, connect the project, enter environment variables in the dashboard. Every pull request gets its own preview URL; try changes there before production.
My checklist before going live:
- Environment variables set in production; Stripe live and test keys not mixed up
- A live webhook endpoint configured in Stripe, its signing secret in an environment variable
- Migrations applied to the production database
- Supabase Auth redirect URLs set for the production domain
- Database backups on
For monitoring, add an error tracking service such as Sentry so you see errors before a user writes “it’s broken.” Telling the agent this is enough: “Add Sentry to the Next.js project, capture server and client errors, and don’t attach user emails or secret values to events.” Add a simple uptime check as well.
What the AI gets wrong
- Inventing APIs that don’t exist. Especially with new library versions. When the build or tests break, hand the error to the agent as is and ask it to check the official docs.
- Changing the test to make it pass. If a test file shows up in the diff, ask why.
- Adding packages you don’t need. The “ask before adding a package” rule in the spec pays off here.
- Swallowing errors. Empty
catch {}blocks make payment failures invisible. - Authorization in the client. Hiding a button isn’t security; the check belongs on the server and in the database.
- Expanding scope. Revert features that arrive with “I also added…” or move them to a separate branch.
- Losing context in long sessions. As a conversation grows, the agent forgets early decisions. Start a fresh session per step; lasting decisions already live in the spec. Closing a session before the context window fills protects both quality and token cost.
Security: short but serious
- Secrets go in
.env.local. Check that it’s in.gitignore. If a key was ever committed, revoke it and issue a new one; deleting it from history isn’t enough. - Never paste API keys into the chat. The agent doesn’t need to see the key, only the variable name.
- Read the commands the agent runs. Especially
rm,git push --force, database commands andcurl … | sh. Auto-approve only safe actions like reading files and running tests. - Don’t point the agent at the production database. Let it work on a local or test database.
- Validate input on the server. Have every API input checked with a schema library like Zod.
- Audit dependencies. Hand
pnpm auditoutput to the agent now and then.
When to bring in a human developer
If you’ve hit any of these, a few hours of review by an experienced developer is worth it:
- Before taking the first real payment (webhooks, subscription status, refunds)
- When customer data includes personal data (GDPR and similar obligations)
- When an enterprise customer sends a security questionnaire
- When you and the agent have failed to fix the same bug three sessions in a row
- When performance problems start (slow queries, missing indexes)
Keep the spec, the migrations folder and test output ready to make the review fast. A project written with an agent but committed in small steps and covered by tests is one a developer can understand quickly.
How I do this with AgentVera
All of the above works in a single terminal. I use AgentVera to manage several agents at once; these are the parts that help:
- Frontend and backend in separate worktrees. In the workspace I open a Claude Code agent for the API and webhooks and a Codex agent for the UI, each in its own git worktree. They don’t touch the same files, so they don’t break each other. The method is covered in git worktrees for parallel agents.
- The diff before merging. The Git panel shows each worktree’s changes, and code review checks every finished turn. What to look for is in AI code review.
- Seeing the database. With the database manager I open the tables in local Postgres and check what the agent actually wrote; DB Guard flags risky migrations at the end of a turn.
- Previews. Preview environments give each worktree’s dev server its own address, so I can try two branches side by side.
- Watching tokens. Each pane shows live context size; when it balloons, I close the session and start fresh.
These speed things up but don’t replace the rules above: the spec, the tests and your own review still matter.
Wrapping up
Building a SaaS with AI is less about generating code and more about writing decisions down and protecting the critical paths. Cut scope to one flow, put decisions in CLAUDE.md or AGENTS.md, pick a boring, mainstream stack, enforce tenant isolation in the database, test the money paths, and set up monitoring before you deploy. Have a human look before real money and real data go in. If you want to run agents in parallel, download AgentVera for free and open your first two worktrees.