Deploying Kartotek
For operators. This doc says what a host has to provide to run Kartotek, independent of any hosting provider: how the process is shaped, the port and health checks, the one directory that must survive restarts, how upgrades and passive instances behave, and which configuration is secret. A ready-made container image definition and an example Docker Compose file ship with the source, described below. For what to back up and how to restore, see Backup and disaster recovery; for first steps once the site is up, see Administrator quick start.
The examples use kartotek.info as the site's hostname; substitute your own.
Process model
One long-running Node.js process serves both the HTTP API and the built frontend. There is no separate frontend host, no separate API host, and no function-per-request model. The process is expected to stay up between deploys rather than spin down when idle: due-Post publishing, daily backup export, off-site backup sync, web-archival jobs, link-health checks and the resumable corpus-wide Backfills all run on the process's own timers, checked once at start and then on an interval. A platform that suspends or scales an idle process to zero would stop that work.
Run exactly one instance of this process per site. Two instances on one database would each publish due Posts and submit archival jobs.
Start the compiled entry point directly, node server/dist/index.js, rather than through npm start. The npm wrapper leaves two extra npm processes alive for the life of the instance, which costs memory on a small host and does nothing else. The command may be run from the repository root or from server/.
One deliberate exception: the visual worker. Visual classification (scene and face classification with ONNX models) runs in a second, separate long-running process, node server/dist/visualWorkerServer.js, which should be reachable only from the main process, never from the public internet. It exists so that an inference-driven memory spike can take down only that process, not the site. The main process finds it through VISUAL_MODEL_SERVICE_URL, and both processes must share the same VISUAL_MODEL_SERVICE_TOKEN. A host that does not run the worker still runs the site normally; visual classification then reads as unavailable, exactly as it does when the model files are missing.
Port and health checks
The process listens on one HTTP port, read from PORT (it falls back to 3000 if unset, which is convenient locally; set it explicitly in a real deployment). The API, under /_api, and the built frontend are served from that single port. Put a TLS-terminating reverse proxy in front of it and route your hostname, for example kartotek.info, to it.
Two unauthenticated endpoints suit a health check:
GET /_api/healthanswers{ok, build, passive}once the process is listening, which is after the database has opened and migrated. It reads only the environment, so it is the right container or load-balancer probe.buildis the value of theRENDER_GIT_COMMITenvironment variable (the name is historical; set it at runtime to the commit your image was built from, orbuildreadsnull); the Admin and public client compare their own build against it.passivesays whether the instance is passive (below).GET /_api/admin/statusanswers a small JSON body with a deploy timestamp (deployedAt), the most recent durable database write (lastContentUpdateAt) and whetherffprobeis available (ffprobeAvailable). It is deliberately unauthenticated, because the Admin top bar reads it before a session exists, and it is cheap: no database writes and a cached probe. A monitor that inspects the body, not just the status code, can use it to distinguish "process up" from "process up and database current".
The visual worker answers GET /health on its own port (10000 by default).
A host that attaches the data directory as a disk that can be mounted by only one instance at a time must stop the old instance before starting the new one, so a deploy there is not zero-downtime. The health check still earns its place: a deploy whose process never answers is marked failed instead of live.
The persistent data directory
Kartotek needs exactly one writable directory that survives restarts and redeploys. It is controlled by one environment variable, DB_PATH, which names the SQLite database file; every other persistent path is derived from that file's parent directory at start. Mount durable storage at that one directory and nothing else needs provisioning. In the container image it is /data, with DB_PATH=/data/app.db.
Under that directory Kartotek creates and manages, without further configuration:
- the SQLite database itself;
- uploaded Media and their derivatives (thumbnails, transcodes);
- a cache for fetched and embedded external media;
- the static export directories for generated Post PDFs and JSON exports;
- the archive of captured web pages (Link PDFs and WARCs, and Post WARCs);
- the Media Archive's thumbnails;
- the daily database exports and the
kartotek-export/folder described in Backup and disaster recovery.
None of these is mounted, sized or backed up separately: one durable directory covers them all. A host that gives the process an ephemeral filesystem will silently lose the database and every uploaded file on the next restart. That is a hard requirement, not an optimization.
PDF and WARC capture drive a headless Chromium. Outside the container image, install a Chromium and point PUPPETEER_EXECUTABLE_PATH at it, or let Puppeteer download its own into a cache directory given as an absolute path in PUPPETEER_CACHE_DIR, the same location at build time and at run time. If neither works the site still boots and serves normally; only static PDF and WARC capture degrades.
Upgrades and migrations
There is no migrate command. Schema migrations are idempotent and run automatically during startup, before the process accepts requests, so an upgrade is "stop the old version, start the new one". A migration is a one-way door: there is no rollback tooling, so rolling back across a schema change means restoring a saved copy of the data directory (see Backup and disaster recovery). Take or confirm a recent backup before upgrading.
Passive instances
A second instance running on a restored copy of the database (a rehearsal before moving hosts, or a standby) inherits the live site's settings and credentials along with its content. Started as-is it would act as the site: sync into the live backup folder, mail login links that point at the live site, publish due Posts and submit Wayback captures.
Set KARTOTEK_PASSIVE=true in such an instance's environment. It is read from the environment only, so a restored database can neither turn it off nor needs editing for it to hold. While it is set:
- Dropbox backup sync never runs and a cleanup delete is refused (reading Dropbox for a restore still works);
- no mail is sent (each suppressed message is logged by subject);
- scheduled publishing, the archival sweep and the quarterly Wayback re-submission do not start, and Wayback archival reads as disabled;
- the public site URL comes from
PUBLIC_ORIGINonly, never from the database, so an instance withoutPUBLIC_ORIGINmints no login link at all.
Everything local keeps running: the daily export, Admin-started Backfills, link health and sessions. GET /_api/health reports passive: true, and the process logs one line saying so at start.
To promote an instance to the live site, remove the variable (or set it to anything but true) and restart, at the same moment the previous live instance stops. Two non-passive instances on copies of one database is the failure this switch exists to prevent.
Backups
A complete backup is two things: the data directory above, and your own record of the environment variables below. Kartotek keeps a daily-rotated, validated SQLite snapshot and a set of small JSON extracts inside that same directory, and can mirror them off-site to Dropbox, but that is a convenience, not a substitute: a volume that is never copied elsewhere takes the daily snapshot down with it. Copy the data directory (or at least database-exports/ and kartotek-export/) somewhere outside the host's primary storage on whatever cadence matches the data loss you can accept. Restoring is replacing the directory's contents with a saved copy and restarting. See Backup and disaster recovery for the details.
Configuration and secrets
Two mechanisms configure Kartotek, and a host only provisions the first:
- Environment variables: a small set needed before the process can do anything else: the database path, the port, the single admin email address, the public site URL, and every external-service credential. Supply them through your host's environment or secret mechanism; Kartotek never reads a checked-in file for them. Treat all of them as sensitive, including the admin email, which is real personal information and the sole check on who may ever log in.
- Admin settings: the much larger set of ordinary tunables (timeouts, batch sizes, feature toggles), edited in the running Admin interface and stored in the same database. A host does not provision these.
The variables most deployments set:
| Variable | Purpose |
|---|---|
DB_PATH |
The SQLite file; its directory is the persistent data directory. |
PORT |
The HTTP port. |
ADMIN_EMAIL |
The one address allowed to log in to Admin. Without it, no login link can be sent. |
PUBLIC_ORIGIN |
The site's public origin, for example https://kartotek.info: absolute https, no path, query or fragment. |
RESEND_API_KEY |
Sends the emailed login link. |
VISUAL_MODEL_SERVICE_URL, VISUAL_MODEL_SERVICE_TOKEN |
Where the visual worker is, and the secret both processes share. |
KARTOTEK_PASSIVE |
true marks a passive instance (above). |
RENDER_GIT_COMMIT |
The build identifier reported by GET /_api/health. |
The public site URL is required once Admin login is configured. The emailed login link carries a live token, so it is built from PUBLIC_ORIGIN (or the equivalent Admin setting) and never from the request Host header, which a forged request could point at an attacker. If ADMIN_EMAIL is set but no valid origin is, the server refuses to start. An install with no admin address has no login flow to protect and starts normally, as does a local development server.
Optional integrations (DeepL translation, Internet Archive credentials, the Dropbox backup connection and its BACKUP_CREDENTIAL_ENCRYPTION_KEY, visit analytics) are each off until configured and fail with a clear "not configured" message on first use; nothing at boot requires them.
The container image
The Dockerfile at the repository root is a multi-stage build with two runtime targets:
runtimeis the web application: the API, the built frontend, the static exports, and PDF and WARC capture. It is based onnode:26.10.0-bookworm-slimand has Chromium and the fonts it needs baked in (PUPPETEER_EXECUTABLE_PATH=/usr/bin/chromium, so nothing is downloaded at run time). It runs as the non-root userapp(uid and gid 1001), listens onPORT=3000, expects the data volume at/data(owned byapp), and startsnode server/dist/index.js. ItsHEALTHCHECKcallsGET /_api/health.worker-runtimeis the visual worker, with the ONNX model files baked into the image. It runs as the same non-root user, listens onPORT=10000, keeps no state, and startsnode server/dist/visualWorkerServer.js. ItsHEALTHCHECKcallsGET /health.
Build each target from the repository root:
docker build --target runtime -t kartotek:1.0.0 .
docker build --target worker-runtime -t kartotek-visual-worker:1.0.0 .
The built image carries the documentation and the stylesheet catalogue as part of the application, because the public /_docs/ site and the Admin CSS catalogue read them from the deployed files at run time. Do not strip docs/ or the client styles from the image.
A named volume created by Docker is owned by root; before the first start, make the volume's contents owned by uid 1001 so the app user can write to /data.
Example: Docker Compose
deploy/compose.example.yml in the source is a generic Compose file for any Docker host behind a reverse proxy. It defines:
- a
webservice built from theruntimetarget, with the health check above, an.env.productionfile for secrets, and an external named volumekartotek-datamounted at/data; - a
visual-workerservice built from theworker-runtimetarget, on a private network only, never reachable from the proxy; restart: unless-stopped, boundedjson-filelogging, and commented-out memory and CPU limits to adjust to your host;KARTOTEK_PASSIVEdefaulting totrue, so a forgotten setting gives a passive instance, never an accidentally live one. Set it tofalsein the environment for the live site.
Create the volume once (docker volume create kartotek-data), put your secrets in .env.production, point your reverse proxy at web:3000 on the proxy network (for example for kartotek.info), and run docker compose -f compose.example.yml up -d. Keep the web service at one replica.