Production health and operations
This runbook applies to a live foundry serve process (including standalone,
service, and Docker modes). Static foundry build output has no running
health or metrics endpoints: monitor the CDN or web server that hosts it separately. Run
commands from the project directory with the same environment and config overlays used
by the server. Start with foundry doctor before deployment.
Health checks
GET /__health returns JSON with status, version,
commit, managed, instance_id when present,
admin_ready, generated_at, and checks.
HEAD /__health performs the same checks and returns no body. Other methods
return 405 with Allow: GET, HEAD. Responses have
Cache-Control: no-store.
- 200 means all checks pass (
status: healthy). -
503 means at least one check failed (
status: degraded). Inspectchecks. - Connection refusal or timeout means the probe cannot reach the process.
The checks cover admin configuration readiness and read/write access to
content, data, public, themes, and
plugins. Storage checks create, write, and remove a temporary file in each
directory. They expose logical names such as storage.data and generic
errors rather than filesystem paths. admin_ready means admin is enabled and
its path is configured; it does not verify a login, external dependencies, plugin
behavior, or rendering. Disabling admin currently makes health degraded. This is a
readiness probe, not a separate process-only liveness endpoint. Avoid aggressive restart
loops on storage faults; alert and investigate mounts, permissions, capacity, and
configuration first.
curl --fail --max-time 5 http://127.0.0.1:8080/__health
curl --head --max-time 5 http://127.0.0.1:8080/__health
Probe every 30 seconds with a 5-second timeout and allow startup time. The production Compose file uses this endpoint. Test a representative public page and authenticated admin login separately. The endpoint is unauthenticated and includes version and instance metadata; restrict it at your reverse proxy if that metadata should remain private. Health traffic follows the same server routing as other requests.
Metrics export
Set a strong random FOUNDRY_METRICS_TOKEN in the
server process environment and restart to enable
GET /__metrics. An unset or whitespace-only token keeps the reserved
endpoint at 404 and disables request instrumentation. Supply
Authorization: Bearer <token>; missing or incorrect credentials
return 401. Authenticated HEAD returns headers only; other methods return
405. Metrics responses use Prometheus text exposition format 0.0.4 and
Cache-Control: no-store. No plugin or extra Go dependency is required.
# Inject the token through your service's protected environment file or secrets system.
curl --fail --max-time 5 \
-H "Authorization: Bearer $FOUNDRY_METRICS_TOKEN" \
http://127.0.0.1:8080/__metrics
| Metric | Type | Meaning |
|---|---|---|
foundry_http_requests_total{status_class="2xx"} |
Counter |
Completed requests, with fixed 1xx–5xx status classes and other for
nonstandard codes
|
foundry_http_request_duration_seconds_total |
Counter | Sum of elapsed request time in seconds |
foundry_http_requests_in_flight |
Gauge | Requests currently being handled |
foundry_go_goroutines |
Gauge | Current Go goroutines |
foundry_go_heap_alloc_bytes |
Gauge | Allocated Go heap bytes |
foundry_go_gc_cycles_total |
Counter | Completed Go garbage collections |
HTTP metrics include public pages, assets, health, and mounted admin/plugin routes. Metrics scrapes (including failed authentication) are excluded. Counters reset on process restart. Duration includes streaming connections and is recorded when a request finishes; it is an aggregate, not a latency histogram or percentile. Metrics contain no URL, query, username, token, or filesystem labels. Runtime memory collection happens only on an authenticated GET scrape.
Example Prometheus scrape configuration:
scrape_configs:
- job_name: foundry
metrics_path: /__metrics
scrape_interval: 30s
scrape_timeout: 5s
authorization:
type: Bearer
credentials_file: /etc/prometheus/secrets/foundry-token
static_configs:
- targets: ['127.0.0.1:8080']
Use the address reachable from Prometheus (a container's loopback is its own). Keep the
token file readable only by the scraper account. Use TLS when scraping across hosts and
restrict access at the proxy/firewall. Rotate the token and restart Foundry, then update
the scraper's secret. For Compose, set the token in the deployment environment;
docker-compose.prod.yml forwards it to Foundry.
Alert on sustained scrape failures (up{job="foundry"} == 0),
readiness failures, and elevated
rate(foundry_http_requests_total{status_class="5xx"}[5m]).
Average completed-request duration can be calculated as the rate of the duration counter
divided by sum(rate(foundry_http_requests_total[5m])); guard against a zero
denominator. Also monitor disk space, inode availability, backup age, and host/container
resources through your infrastructure monitoring.
Structured logging
FOUNDRY_LOG accepts debug, info (default),
warn/warning, and error.
FOUNDRY_LOG_FORMAT=json selects newline-delimited JSON for Foundry's
slog logs on stderr. The default and unrecognized formats use readable
text. JSON entries include time, level, msg, and
event-specific typed attributes. Configure these variables before starting the process;
changes require restart. CLI progress output and third-party output can still be plain
text.
FOUNDRY_LOG=info FOUNDRY_LOG_FORMAT=json foundry serve
The production Compose file defaults to JSON and info:
docker compose -f docker-compose.prod.yml logs --follow --tail 100 foundry
Send stderr to your collector or service log file. Set retention and rotation at the
service/container logging layer and alert on repeated errors. Foundry's installed
services write .foundry/run/service.log; rotate it with a strategy
compatible with the manager's open file handles. foundry logs -f is the
standalone log viewer. JSON formatting does not redact event attributes. Avoid debug
mode in normal production: verbose request diagnostics can include queries, remote
addresses, user agents, and filesystem context. Restrict access to logs and apply
collector redaction where needed.
Backup and restore drills
Built-in ZIP and Git snapshots cover the content tree, including its
config and media. They do not provide a complete machine recovery image. Independently
back up data/ (admin accounts, sessions, audit/runtime state),
themes/, plugins/, external secret material, service
definitions, and the executable version. public/ can be rebuilt; include it
if your recovery-time target requires it. Treat backups containing configuration and
accounts as sensitive.
foundry backup create
foundry backup list
foundry backup git-snapshot "before deployment"
foundry backup git-log 10
Managed ZIPs default to .foundry/backups; the production Docker overlay
uses data/backups on the persistent data volume.
backup.on_change is debounced and is not a scheduled complete-system
backup. Retention can prune local ZIPs. Copy verified archives off the host and keep an
independent retention policy. Never rely on a backup stored only on the disk being
protected. Quiesce writers for consistent content snapshots and stop the service during
full filesystem snapshots, especially when copying account/session data.
Perform this drill on an isolated recovery project, never the live content tree:
-
Record the application revision, config overlays, backup timestamp, archive checksum
(
shasum -a 256 snapshot.zip), and expected recovery point/time. Copy the archive off-host and verify its checksum after retrieval. - Prepare a disposable project with the same executable, themes, plugins, and directory names. Restore independently saved data and secrets where needed; disable managed callbacks, remote Git pushes, and other production integrations. Use a distinct listener port and local-only admin access.
-
Place the ZIP in that project's configured
backup.dir. Check archive integrity withunzip -t snapshot.zip. Start the recovery server and sign in to its admin UI. Use the Platform → Operations backup controls to restore the selected ZIP. There is currently nofoundry backup restoreCLI command. - The restore snapshots the existing recovery content first, extracts the archive, and replaces the content tree. It requires backup space as well as temporary extraction space. It restores content only, and can replace config; inspect the restored configuration and reapply isolation settings before restarting the recovery server.
-
Run
foundry doctorandfoundry buildin the recovery project with the intended overlays. Restart it, inspect/__health, verify representative pages/media, and test admin login and a disposable content edit. Compare document/media counts and selected checksums with the expected recovery point. - Record elapsed recovery time, recovered data age, missing files, and corrective actions. Retain drill evidence and repeat after storage/auth/plugin changes and on a regular schedule suited to your recovery targets.
For actual recovery, stop traffic and writers, retain the damaged state for rollback/analysis, restore the verified complete backup set, rebuild, and run these checks before reopening traffic. A green health response alone does not prove content or account recovery.
Service-manager examples
Run the built-in commands from the project root as the deployment user:
foundry service install
foundry service status
foundry service restart
foundry service stop
foundry service start
Linux installation creates a user systemd unit in ~/.config/systemd/user/;
macOS creates a LaunchAgent in ~/Library/LaunchAgents/. Names and paths are
printed by installation/status and stored in .foundry/run/service.json.
Installation enables and starts the service. Linux users may need
loginctl enable-linger "$USER" for logout/reboot persistence. A
macOS LaunchAgent runs in the user's login session; use Docker or a separately managed
system service for deployments that must operate without that session.
The generated definitions run serve from the project root and do not
capture shell environment variables or command-line overlays. Configure a manager
provided environment explicitly. For Linux, add an override using the actual unit name
from service status:
systemctl --user edit <installed-unit-name>.service
[Service]
Environment=FOUNDRY_LOG=info
Environment=FOUNDRY_LOG_FORMAT=json
EnvironmentFile=/absolute/path/to/protected/foundry.env
# Optional explicit production overlay; use the executable printed by service status.
ExecStart=
ExecStart=/absolute/path/to/foundry --config-overlay content/config/site.production.yaml serve
Store FOUNDRY_METRICS_TOKEN=... and required admin secrets in that
protected file, with permissions limited to the service account. Then run:
systemctl --user daemon-reload
systemctl --user restart <installed-unit-name>.service
systemctl --user status <installed-unit-name>.service
For macOS, add an EnvironmentVariables dictionary to the installed plist,
containing FOUNDRY_LOG, FOUNDRY_LOG_FORMAT, and required
secrets. Restrict the plist's permissions, and unload/reload the agent using its
installed path:
launchctl bootout "gui/$(id -u)" /absolute/path/to/installed.plist
launchctl bootstrap "gui/$(id -u)" /absolute/path/to/installed.plist
Use the plist's ProgramArguments array to add any required config overlay
before serve. Reinstalling can regenerate the plist; preserve your
settings. Check service status, health, metrics authentication, and log collection after
every restart and host reboot. Test these manager-specific steps on your host before
relying on unattended operation.
Build identity and update eligibility
foundry version --json and Admin → Operations distinguish the running build
from the latest release. Tagged releases, snapshots, modified builds, and unknown
provenance have different update eligibility. Containers update through deployment
tooling; mounted site Git history does not establish binary provenance. See
update guidance and
managed runtime boundaries.
