Working with AI agents

How ShipKit is set up for Claude Code, Codex and Cursor — the rule files, the skills for recurring jobs, and the tests that catch an agent's mistakes.

ShipKit is written to be changed by a coding agent as much as by hand. The conventions are written down where an agent reads them, the recurring setup jobs are step-by-step procedures an agent can follow, and the rules that are easy to break are checked by tests rather than left to memory. None of this is required: it is all plain Markdown and TypeScript, and a human can read and follow it the same way.

The rule files

FileRead byWhat it holds
CLAUDE.mdClaude Code, automaticallyCommands, a map of the codebase, and the hard rules
src/CLAUDE.mdClaude Code, when working under src/The design language: colour, radius, buttons, borders
AGENTS.mdCodex, Cursor, GitHub Copilot, Windsurf, opencode, Zed, AmpA pointer to the two files above, plus the skill table
DESIGN.mdAnyone who asks whyThe reasoning behind the architecture

There is one set of rules, kept in CLAUDE.md. AGENTS.md does not repeat them; it tells other agents to read CLAUDE.md first, so the two never drift apart.

The hard rules are the ones that cost the most when broken: authorization is a permission check, never a role check; all client data goes through one channel (route loader, query, server function, Drizzle); D1 has no interactive transactions; every feature is a deletable vertical slice; payments arrive as events; server functions never localize; money is integer cents. Each has its own page in these docs — start with Architecture and Permissions.

When you change a convention in your own project, change it in CLAUDE.md too. An agent follows what the file says, not what you meant.

Skills

The recurring jobs that involve more than code — accounts, dashboards, secrets — are written as skills under .claude/skills/<name>/SKILL.md. In Claude Code each one is a slash command. Any other agent can open the file and follow it; AGENTS.md tells them to.

SkillWhat it does
/setupInstalls dependencies, creates .dev.vars with a generated secret, migrates the local database, starts the dev server and makes your first admin
/google-oauthWalks you through the Google Console and wires the credentials in (required: env validation fails without them)
/stripeBuilds the Stripe catalog, sets keys and the webhook secret, fills in src/config/plans.ts, and runs a local purchase through the Stripe CLI
/creemThe same for Creem, the alternate provider
/deployProvisions D1 and R2, applies remote migrations, pushes secrets, deploys, and finishes the post-deploy wiring
/add-featureScaffolds a new vertical slice with every registry line and marker in place
/delete-featureRemoves an optional module with its machine-verified recipe
/seo-auditRuns Lighthouse and checks the SEO artefacts of a production build
/admin-apiQueries the read-only admin data API and turns the JSON into analysis

The skills split the work honestly: the agent runs every command and edits every file, and you do what only a person can — clicking through the Google Console or the Stripe dashboard, pasting values back, confirming anything that costs money.

/admin-api is portable on purpose. Copy .claude/skills/admin-api/ into any other agent's skills directory, or point the agent at /api/v1/admin/skill.md, which your deployed app serves verbatim. See API keys.

Guard rails

An agent that misreads a rule usually produces code that looks right. These checks turn the common misreadings into a failing command:

CheckWhat it catches
src/core/authz/registry.test.tsAn admin route without requirePermission, an adminFn export without assertPermission, a feature's permissions.ts missing from the registry, a database helper added to the client-bundled core/server/fn.ts
src/core/settings/handlers.test.tsA feature's settings.ts that imports server-only code (it would drag cloudflare:workers into the browser) or is not registered
src/core/settings/i18n.test.tsAn operator setting shipped without translated labels
src/config/private-paths.test.tsA new route outside the marketing pages that is missing from the private-path list, which feeds robots.txt
src/config/plans.test.tsPlan and credit-pack prices that break the pricing rules stated at the top of plans.ts
src/core/payment/money-path.test.tsPayment handlers that misbehave against a real in-memory D1 with every migration applied
bun run verify:deletionA feature whose wiring is no longer cleanly removable
CI "Schema and migrations agree"A schema change committed without its migration

Some rules are enforced by the build instead of a test. Importing cloudflare:workers anywhere except src/core/server/cf.ts, for example, breaks the client bundle, and a missing m.* message key fails typecheck.

A workflow that works

  1. Describe the outcome, name the skill. "Add a changelog feature with /add-feature" gives the agent the checklist; "add a changelog" leaves it to guess which registries to touch.

  2. Let the skill drive setup work. For Google, Stripe, Creem and deploys, run the skill instead of asking for the steps. It verifies each stage with a command rather than trusting a log line.

  3. Make it prove the change. Before you accept anything, have the agent run:

    bun run typecheck
    bun run check
    bun run test
    

    After touching routes, navigation, auth guards or the app shell, add bun run e2e. After adding or wiring a feature, add bun run verify:deletion <feature>.

  4. Read the diff yourself. Look in particular for secrets in committed files, a new role !== 'admin', a db.transaction() call, and copy hard-coded in a component instead of messages/en.json and messages/zh.json.

  5. Migrations are yours to run. An agent can generate a migration with bun run db:generate, but applying it to production (bun run db:migrate:remote) is a step to take deliberately. See Database.

Limits

The tests check structure, not intent. They will tell you an admin function asks for a permission; they cannot tell you it asks for the right one. The design rules in src/CLAUDE.md are not tested at all, so review visual changes in the browser. And an agent's confidence is not evidence: "done" means the commands above passed, not that the agent said so.