Operations

What runs in production besides requests — the nightly cron jobs, logs and error pages, rate limits, security headers, CI and backups.

This page covers what keeps a deployed ShipKit healthy once it is live: scheduled maintenance, where errors show up, what protects the public endpoints, and what CI checks before it ships.

Cron jobs

One cron trigger in wrangler.jsonc fires every day at 00:00 UTC ("crons": ["0 0 * * *"]). The Worker's scheduled handler in src/server.ts calls runJobs(), which runs every job registered through registerJob in src/core/jobs/events.ts. Each feature registers its jobs with one import line in src/core/jobs/handlers.ts, so deleting a feature removes its jobs with it.

Jobs run one after another. A job that throws is logged and the rest still run. Every job must be idempotent, because a cron can fire twice and you can fire it by hand.

JobWhat it does
payment.prune-webhooksDelete stored webhook payloads older than the retention setting (90 days by default)
auth.purge-expiredDelete expired sessions, verification rows and expired API keys, which better-auth never removes on its own
auth.remind-deleted-accountsEmail a closed account 3 days before it is erased, once
auth.purge-deleted-accountsErase closed accounts whose retention window (30 days by default) has passed, 25 per night
billing.expire-stale-subscriptionsExpire subscriptions whose paid period ended more than the grace period ago (3 days by default) with no webhook — a safety net for missed events
affiliate.mature-commissionsMark pending commissions as payable once their refund holdback has elapsed
credits.monthly-grantGrant each user's monthly allowance, once per calendar month
audit.pruneDelete audit log rows older than the retention setting (180 days by default)
email_log.pruneDelete email log rows older than the retention setting (90 days by default)
demo.purge-visitorsErase day-old public-demo accounts. Matches nothing unless demo mode was used

The retention and grace periods are operator settings, editable at Admin → Settings without a deploy (see Configuration).

The time of day matters for one job: the previous month's credit allowance expires at exactly 00:00 UTC, and credits.monthly-grant runs at that instant so balances do not read zero for hours. If you move the cron, keep that in mind.

To run the jobs locally while bun run dev is up:

curl "localhost:3000/cdn-cgi/handler/scheduled"

To add a job, see Adding features: register it in your feature's server/jobs.ts and add the import line to src/core/jobs/handlers.ts.

Logs and errors

wrangler.jsonc enables Workers Logs ("observability": { "enabled": true }) at a 100% sample rate. Request logs, exceptions and anything written with console.* appear in the Cloudflare dashboard under your Worker, with no code. On very high traffic, lower head_sampling_rate. To follow logs live from a terminal:

bunx wrangler tail

Useful lines to search for:

  • [jobs] <name>: … — each job's result, such as pruned=12, or its failure
  • stripe webhook <type> <id> failed / creem webhook … failed — a payment handler threw; the event was released so the provider's retry gets a second attempt
  • [email] RESEND_API_KEY not set — an email was skipped

Users never see a stack trace. src/components/error-pages.tsx holds the 404 and 500 pages, wired as the router's default not-found and error components; production shows generic copy and the detail goes to the log.

The admin console also records what happened at the business level: the audit log, the email log with every send attempt and its status, and stored webhook payloads. See Admin console.

Rate limits

Two Cloudflare rate-limiting bindings are declared in wrangler.jsonc, both per client IP unless a key is given:

BindingLimitUsed for
PUBLIC_RATE_LIMIT5 per minuteSign-in code requests (per IP and per recipient address), avatar uploads, file uploads (per user), feedback screenshots, the demo sign-in
ANALYTICS_RATE_LIMIT200 per minuteThe first-party PostHog proxy at /api/ph/*, which carries several calls per page view

Call isRateLimited(scope, options) from src/core/server/rate-limit.ts and return tooManyRequests() (a 429 with retry-after: 60) on any new public endpoint that costs real resources. Pass key to count against something other than the IP, as the sign-in code limit does for the recipient's address. If the binding is missing the check fails open, so a config gap never takes an endpoint down. Separately, API keys carry better-auth's own limit of 300 requests per minute.

Security headers

Every response carries X-Content-Type-Options, X-Frame-Options: DENY, Referrer-Policy, Permissions-Policy, Cross-Origin-Opener-Policy and, over HTTPS, Strict-Transport-Security. They come from two places, and the two lists must stay in sync:

  • src/core/server/security-headers.ts for everything the Worker answers, applied in src/server.ts
  • public/_headers for prerendered pages and static assets, which Cloudflare serves without running the Worker

e2e/security-headers.spec.ts checks pages, redirects and API responses.

There is no Content-Security-Policy by default. TanStack Start hydrates through inline scripts, so a useful CSP needs per-request nonces; add one when you have that in place. The opener policy is same-origin-allow-popups rather than same-origin because Google One Tap needs its popup to keep window.opener.

Server functions have one more guard: src/server.ts refuses a call to /_serverFn/… that a browser marks as cross-site, with a 403, before it is decoded. API routes are not covered, because webhooks and API-key clients come from elsewhere by design.

CI

.github/workflows/ci.yml runs on every push to main, on pull requests, and by hand (workflow_dispatch):

JobWhenWhat
Typecheck · lint · test · buildAlwaysbuild, typecheck, check, test, then db:generate must write nothing — a schema change without its migration fails here
End-to-endAlwaysThe Playwright suite against a dev server
Verify deletionPull requestsOne parallel job per recipe from scripts/verify-deletion.ts
DeployPush to main or manual run, after the first two passbun run deploy

CI builds with placeholder secrets written from .dev.vars.example; real keys never go into CI. The Worker reads its real secrets at runtime from wrangler secret put.

The deploy job needs CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID in the repository's production environment. Without them it skips and says so. It never applies migrations. It checks that the production D1 has no pending migrations and refuses to deploy if it has, so the order is always: run bun run db:migrate:remote yourself, then push (or re-run the workflow). See Deploy.

Backups

ShipKit ships no backup job. Your data lives in two Cloudflare services:

  • D1 keeps a point-in-time history of the database (Time Travel), managed with wrangler d1 time-travel; how far back it reaches depends on your Cloudflare plan. For a copy you hold yourself, wrangler d1 export DB --remote --output backup.sql writes the database as SQL.
  • R2 holds uploaded files and profile photos. There is no automatic copy; if you need one, set it up on the Cloudflare side.

Before a risky migration, take an export first. Migrations are applied by hand for exactly this reason: a schema change should be a step you watch.