Working with AI agents
How ShipKit is set up for Claude Code, Codex and Cursor — the rule files, the skills for recurring jobs, and the tests that catch an agent's mistakes.
ShipKit is written to be changed by a coding agent as much as by hand. The conventions are written down where an agent reads them, the recurring setup jobs are step-by-step procedures an agent can follow, and the rules that are easy to break are checked by tests rather than left to memory. None of this is required: it is all plain Markdown and TypeScript, and a human can read and follow it the same way.
The rule files
| File | Read by | What it holds |
|---|---|---|
CLAUDE.md | Claude Code, automatically | Commands, a map of the codebase, and the hard rules |
src/CLAUDE.md | Claude Code, when working under src/ | The design language: colour, radius, buttons, borders |
AGENTS.md | Codex, Cursor, GitHub Copilot, Windsurf, opencode, Zed, Amp | A pointer to the two files above, plus the skill table |
DESIGN.md | Anyone who asks why | The reasoning behind the architecture |
There is one set of rules, kept in CLAUDE.md. AGENTS.md does not repeat
them; it tells other agents to read CLAUDE.md first, so the two never drift
apart.
The hard rules are the ones that cost the most when broken: authorization is a permission check, never a role check; all client data goes through one channel (route loader, query, server function, Drizzle); D1 has no interactive transactions; every feature is a deletable vertical slice; payments arrive as events; server functions never localize; money is integer cents. Each has its own page in these docs — start with Architecture and Permissions.
When you change a convention in your own project, change it in CLAUDE.md
too. An agent follows what the file says, not what you meant.
Skills
The recurring jobs that involve more than code — accounts, dashboards,
secrets — are written as skills under .claude/skills/<name>/SKILL.md. In
Claude Code each one is a slash command. Any other agent can open the file
and follow it; AGENTS.md tells them to.
| Skill | What it does |
|---|---|
/setup | Installs dependencies, creates .dev.vars with a generated secret, migrates the local database, starts the dev server and makes your first admin |
/google-oauth | Walks you through the Google Console and wires the credentials in (required: env validation fails without them) |
/stripe | Builds the Stripe catalog, sets keys and the webhook secret, fills in src/config/plans.ts, and runs a local purchase through the Stripe CLI |
/creem | The same for Creem, the alternate provider |
/deploy | Provisions D1 and R2, applies remote migrations, pushes secrets, deploys, and finishes the post-deploy wiring |
/add-feature | Scaffolds a new vertical slice with every registry line and marker in place |
/delete-feature | Removes an optional module with its machine-verified recipe |
/seo-audit | Runs Lighthouse and checks the SEO artefacts of a production build |
/admin-api | Queries the read-only admin data API and turns the JSON into analysis |
The skills split the work honestly: the agent runs every command and edits every file, and you do what only a person can — clicking through the Google Console or the Stripe dashboard, pasting values back, confirming anything that costs money.
/admin-api is portable on purpose. Copy .claude/skills/admin-api/ into any
other agent's skills directory, or point the agent at
/api/v1/admin/skill.md, which your deployed app serves verbatim. See
API keys.
Guard rails
An agent that misreads a rule usually produces code that looks right. These checks turn the common misreadings into a failing command:
| Check | What it catches |
|---|---|
src/core/authz/registry.test.ts | An admin route without requirePermission, an adminFn export without assertPermission, a feature's permissions.ts missing from the registry, a database helper added to the client-bundled core/server/fn.ts |
src/core/settings/handlers.test.ts | A feature's settings.ts that imports server-only code (it would drag cloudflare:workers into the browser) or is not registered |
src/core/settings/i18n.test.ts | An operator setting shipped without translated labels |
src/config/private-paths.test.ts | A new route outside the marketing pages that is missing from the private-path list, which feeds robots.txt |
src/config/plans.test.ts | Plan and credit-pack prices that break the pricing rules stated at the top of plans.ts |
src/core/payment/money-path.test.ts | Payment handlers that misbehave against a real in-memory D1 with every migration applied |
bun run verify:deletion | A feature whose wiring is no longer cleanly removable |
| CI "Schema and migrations agree" | A schema change committed without its migration |
Some rules are enforced by the build instead of a test. Importing
cloudflare:workers anywhere except src/core/server/cf.ts, for example,
breaks the client bundle, and a missing m.* message key fails typecheck.
A workflow that works
-
Describe the outcome, name the skill. "Add a changelog feature with
/add-feature" gives the agent the checklist; "add a changelog" leaves it to guess which registries to touch. -
Let the skill drive setup work. For Google, Stripe, Creem and deploys, run the skill instead of asking for the steps. It verifies each stage with a command rather than trusting a log line.
-
Make it prove the change. Before you accept anything, have the agent run:
bun run typecheck bun run check bun run testAfter touching routes, navigation, auth guards or the app shell, add
bun run e2e. After adding or wiring a feature, addbun run verify:deletion <feature>. -
Read the diff yourself. Look in particular for secrets in committed files, a new
role !== 'admin', adb.transaction()call, and copy hard-coded in a component instead ofmessages/en.jsonandmessages/zh.json. -
Migrations are yours to run. An agent can generate a migration with
bun run db:generate, but applying it to production (bun run db:migrate:remote) is a step to take deliberately. See Database.
Limits
The tests check structure, not intent. They will tell you an admin function
asks for a permission; they cannot tell you it asks for the right one. The
design rules in src/CLAUDE.md are not tested at all, so review visual
changes in the browser. And an agent's confidence is not evidence: "done"
means the commands above passed, not that the agent said so.