Files
fluxer/fluxer_docs/src/content/docs/operator/upgrading.mdx
T

397 lines
25 KiB
Plaintext

---
# SPDX-License-Identifier: AGPL-3.0-or-later
title: Upgrading
description: Upgrading and rolling back an instance with the installer, what the backup covers, and how to restore it.
---
:::tip[Plutonium and donations fund Fluxer]
We are grateful to everyone supporting the project through [Fluxer Plutonium](https://fluxer.app/plutonium) or [donations](https://fluxer.app/donate). All of our code is free and open source on [GitHub](https://github.com/fluxerapp/fluxer). The Operator Pass is coming, and adds a direct line to the team for help and feedback.
:::
An upgrade moves an instance to a newer release. `install.sh --update` does all of it: back up, download the new stack files, pull the new images, recreate the containers that changed. The stack files are the compose files, the `Caddyfile` and `.env.example`. Allow about twenty minutes, most of it the backup and the pull.
| Requirement | Value |
| --- | --- |
| Working directory | The one holding `.env` and the stack files |
| Compose commands | Behind your own reverse proxy, set `COMPOSE_FILE` in `.env` once and every command picks up the overlay |
| Free disk | Room for one database dump, one copy of the uploads, and the new images beside the old ones |
| Downtime | A minute or two for the recreate, plus however long the uploads copy takes with the stack stopped |
Read the release notes between the version you run and the one you are moving to, at [Releases](https://github.com/fluxerapp/fluxer/releases).
## Match the images to the stack files
The images come from `FLUXER_IMAGE_TAG` in `.env`. The stack files come from a git ref, which is a branch or a tag in the Fluxer repository. The installer derives that ref from the image tag: `main` for `v1` or `latest`, and the tag string itself for anything else. `--ref` overrides the derivation. `docker compose pull` never updates a stack file, and refreshing a stack file never moves an image.
The repository holds no ref by the name of a pinned image tag. A release tags each image on its own, as `[email protected]`, so pass `--ref` with that tag or with the commit it points at. Without that override the download fails with exit 4.
The stack files on `main` can run ahead of the `v1` images. The current `docker-compose.yml` passes an optional setting with no Compose default through empty when `.env` leaves it out, and only images built from the same change or later read an empty value as unset. Older images reject some of those empty values and `api`, `worker` and `media-proxy` fail to start. Refresh the stack files only once the images the tag names are at least as new, or pin both with `--ref`.
## What the script does
`--update` runs these steps in this order, and stops on the first one that fails:
| Step | What it does |
| --- | --- |
| Mint | Writes any key the refreshed stack requires that `.env` does not hold, listed under [Run the upgrade](#run-the-upgrade) |
| Record | Writes the image references, the image ID each container has, and the tag `.env` names, into a new record directory |
| Save | Copies `.env` and the stack files into that record |
| Dump | Dumps the database with the stack still serving, and checks the custom-format header on the result |
| Copy | Stops the stack, copies the uploads volume, and starts it again on the images it was already running |
| Fetch | Downloads the stack files at the ref into a staging directory inside the working directory |
| Guard | Refuses when the refreshed `docker-compose.yml` moves Postgres to a new major version |
| Place | Moves every downloaded file into the working directory at once |
| Pull | Pulls the images while the old containers still serve |
| Recreate | Runs `docker compose up -d --remove-orphans`, then restarts `edge` when the `Caddyfile` changed |
| Verify | Polls Compose state until every service is ready, then probes `/_health` on the hostname |
A failed run leaves the instance running on the images it already had.
`docker compose up -d` does not notice a changed `Caddyfile`. The script restarts `edge` by name to pick it up.
An instance set up before the edge served LiveKit at `/livekit` needs one edit that the upgrade does not make. Until you make it, voice does not connect. [Voice signalling moved to /livekit](#voice-signalling-moved-to-livekit) has the edit.
The refreshed `docker-compose.yml` renames the `caddy` service to `edge`, and the two publish the same host ports. `--remove-orphans` takes the old container down in the same call, so the new one can bind them. Without it the step stops with `Bind for 0.0.0.0:443 failed`. It also removes any container in the project whose service the loaded Compose files no longer define, so keep `COMPOSE_FILE` the same across an upgrade. [Keep a local compose change](#keep-a-local-compose-change) has the file layout that survives one. [When the compose file list names a missing file](#when-the-compose-file-list-names-a-missing-file) has the error Compose prints when `COMPOSE_FILE` names a file that does not exist.
The rename also moves the `caddy-data` and `caddy-config` volumes to `edge-data` and `edge-config`, so the old two are left unused and the edge requests its certificate again on the first start.
Upgrades change `FLUXER_IMAGE_TAG` and add missing required secrets. Existing values are preserved, except for the unusable upload-relay placeholder described below.
The refreshed stack requires `FLUXER_ERLANG_COOKIE`. `--update` generates a missing value as 64 hex characters. Without it, Compose commands fail with `set FLUXER_ERLANG_COOKIE in .env`.
To write it by hand, run the generator in a shell:
```bash
printf '\nFLUXER_ERLANG_COOKIE=%s\n' "$(openssl rand -hex 32)" >> .env
```
Do not paste `FLUXER_ERLANG_COOKIE=$(openssl rand -hex 32)` into `.env`. Compose does not execute shell commands in that file.
`api` and `media-proxy` require `FLUXER_MEDIA_PROXY_UPLOAD_RELAY_SECRET_BASE64` to decode to at least 32 bytes. `--update` replaces a missing or `CHANGE_ME` value with a generated secret. Both services must use the same value.
To write it by hand, run the generator in a shell the same way:
```bash
printf '\nFLUXER_MEDIA_PROXY_UPLOAD_RELAY_SECRET_BASE64=%s\n' "$(openssl rand -base64 32)" >> .env
```
## The script is the reference
Read the [Linux and macOS installer](https://fluxer.dev/install.sh) or [Windows installer](https://fluxer.dev/install.ps1) before running it.
## Keep a local compose change
An upgrade replaces the stack files, including direct edits to `docker-compose.yml`, `docker-compose.proxy.yml`, `tunnel.compose.yml`, `external-object-store.compose.yml` or the `Caddyfile`. Use `--dry-run` to see which files will change.
The supported way to hold a local choice is a separate file, listed in `COMPOSE_FILE` in `.env`:
```ini
COMPOSE_FILE=docker-compose.yml:docker-compose.proxy.yml:local.compose.yml
```
Compose merges the files left to right, so `local.compose.yml` wins over the ones before it. An upgrade refreshes the stack files and nothing else, so a file outside that set survives every upgrade untouched. A record copies `.env` and those files, and no other file from the working directory. An override file is never backed up with them, so keep it wherever the rest of the configuration lives.
Keep environment settings in `.env`. Upgrades preserve them apart from the tag and required-secret changes described above.
[Enable the overlay](/operator/reverse-proxy/#enable-the-overlay) has the overlay this is most often used for. [Docker labels](/operator/reverse-proxy/#docker-labels) has a worked third file.
## When the compose file list names a missing file
`COMPOSE_FILE` names files relative to the working directory, and Compose refuses to load the set when one of them is absent:
```
stat /srv/fluxer/docker-compose.proxy.yml: no such file or directory
```
Every `docker compose` command stops there. This can affect older installations whose `.env` names an overlay that has not been downloaded. The installer reports the missing file and a command to restore it.
Download the missing file at the ref the upgrade moves to, then run the upgrade:
```bash
curl -fsSL --proto '=https' --tlsv1.2 -o docker-compose.proxy.yml \
https://raw.githubusercontent.com/fluxerapp/fluxer/main/deploy/self-hosting/docker-compose.proxy.yml
```
`main` is the ref for `FLUXER_IMAGE_TAG=v1` and for `latest`. [Match the images to the stack files](#match-the-images-to-the-stack-files) has the ref a pinned tag needs.
Removing the line from `.env` also lets the upgrade run, and it is the wrong fix. Without the overlay, the edge publishes host ports `80` and `443`, which the reverse proxy already on the host uses.
## Run the upgrade
Keep the installer beside the stack files. [Get started](/operator/get-started/#step-4-bring-up-the-instance) has the download and the checksum check for a fresh copy.
See the plan first:
```bash
cd ~/fluxer
sh install.sh --update --dry-run
```
It prints:
- The images the stack runs now
- The backup it intends to take
- The stack files the ref changes
- The services those changes make it restart
- Any refusal that would stop the run
It writes nothing outside a temporary directory it removes on exit.
Then run it:
```bash
sh install.sh --update
```
On Windows:
```powershell
.\install.ps1 -Update
```
The run ends by naming the record it wrote and the command that rolls back to it. It checks its own work with the readiness poll and one probe of `/_health`. Run the public probes in [Check that it works](/operator/get-started/#step-5-check-that-it-works) afterwards.
| Flag | Windows | Meaning |
| --- | --- | --- |
| `--update` | `-Update` | Record, back up, refresh, pull, recreate, verify |
| `--rollback` | `-Rollback` | Put back the images and stack files of the newest record |
| `--dry-run` | `-DryRun` | Print the plan. Change nothing |
| `--backup-dir <path>` | `-BackupDir` | Where records go. Default `<dir>/backups` |
| `--no-volume-backup` | `-NoVolumeBackup` | Take the database dump and skip the uploads copy |
| `--no-volume-compression` | `-NoVolumeCompression` | Copy the uploads as a plain `.tar`. Less downtime, more disk |
| `--skip-backup-accept-data-loss` | `-SkipBackupAcceptDataLoss` | Upgrade with no backup at all |
`--dir` and `--ref` mean what they mean on an install, and the installer derives an omitted `--ref` from the `FLUXER_IMAGE_TAG` line in `.env`. Every failure prints a sentence on standard error and exits non-zero, and the script header lists what each code means. A Postgres major version change is exit 3.
## Voice signalling moved to /livekit
The edge serves LiveKit at `/livekit/*`. An instance set up before that routing served it under `/gateway/livekit`, and that path is gone. Voice stops connecting after the upgrade and the run still reports success, because the readiness poll and the `/_health` probe never reach voice.
Before:
```
wss://chat.example.com/gateway/livekit
```
After:
```
wss://chat.example.com/livekit
```
Both `https` and `wss` work. Update `FLUXER_LIVEKIT_URL` in `.env` to the new path, or remove it to use the derived address, then run `docker compose up -d api`.
When `api` starts, it writes the new address to the existing `default-server-1` server in the `default` region. For any other server, change **Endpoint** under **Voice Servers** in the admin dashboard. Correct `.env` first, or a later restart can restore the old address on the default server.
Voice media is unaffected. It never went through the edge, and `7881/tcp` and `7882/udp` still reach the host directly.
## Short-lived data expires on Postgres
On the Postgres backend, data the schema keeps for a limited time expires the way it does on Cassandra, counted from its last write. After the upgrade, these are the differences most people notice.
- The mentions inbox lists the last 7 days of mentions.
- A device that has not opened Fluxer in 90 days gets no push notifications until Fluxer is opened on it again.
- An account's ended sessions in the admin dashboard go back 30 days.
Within minutes of the upgrade, `worker` starts deleting existing data that is past that lifetime. A rollback does not bring it back. [Restore a backup](#restore-a-backup) does. `worker` checks again once a day, so data written while a rollback was in place gets the same treatment after the next upgrade.
Instances on Cassandra already behave this way.
## Check the svc request ceiling
`docker-compose.yml` used to pass `FLUXER_SVC_MAX_CONCURRENT_REQUESTS` only to the `users` and `messages` routers and shards. It now reaches every svc container, so a value set in `.env` also caps `snowflakes`, `gifs` and `unfurl`. Unset, `snowflakes` allows `320` requests in flight, `messages` `192` and the rest `64`. An instance that sets the name either raises it to at least `320` or removes it from `.env`, then runs `docker compose up -d`.
## Add the passkey columns on Cassandra
An instance on the Postgres backend needs nothing here. An instance on Cassandra or Scylla adds two columns to `webauthn_credentials` before it starts the new `api` and `worker` images. Replace `fluxer` with the value of `FLUXER_CASSANDRA_KEYSPACE`.
```sql
ALTER TABLE fluxer.webauthn_credentials ADD rp_id text;
ALTER TABLE fluxer.webauthn_credentials ADD superseded_by text;
```
Both columns start empty on every row, and existing passkeys keep working as before. Wait until every node reports the same schema version, as `nodetool describecluster` shows, then upgrade.
:::danger[Upgrading without the columns breaks every connection]
The new release reads both columns whenever it loads passkeys, and every Gateway connection loads them. Without the columns no client finishes connecting, and passkey sign-in fails.
:::
An earlier release ignores the columns, so a rollback leaves them in place.
## What the backup covers
Each upgrade writes one record directory, named `record-` and a UTC stamp, under `backups` or wherever `--backup-dir` points. It holds:
| Artefact | What it covers |
| --- | --- |
| `fluxer.dump` | Every account, message, guild and configuration row, in the Postgres custom format |
| `seaweedfs-data.tgz` | Every upload, avatar, report and harvest, or `seaweedfs-data.tar` under `--no-volume-compression` |
| `.env` | Every secret the instance was built with |
The stack files sit in the record beside them, so a rollback puts back the exact files the instance was running. The record directory is created `0700` and the `.env` copy inside it is `0600`. Keep records wherever you already keep secrets.
Nothing else is copied. The dump covers `postgres-data`. The other volumes either rebuild themselves or hold queued work, and [Volumes and buckets](/operator/configuration/#volumes-and-buckets) lists them all with what each one holds.
Put a dump on a timer as well. On Linux and macOS, this crontab line writes one a night and keeps two weeks of them:
```cron
15 3 * * * cd /srv/fluxer && docker compose exec -T postgres sh -c 'pg_dump -U $POSTGRES_USER -d $POSTGRES_DB --format=custom' > "backups/fluxer-$(date -u +\%Y\%m\%dT\%H\%M\%SZ).dump" && find backups -name 'fluxer-*.dump' -mtime +14 -delete
```
`/srv/fluxer` stands for the directory holding `.env`. Write it as an absolute path, because cron does not expand `~`. An unescaped `%` in a crontab is a newline, which is why every one above has a backslash.
## Run a data store outside the stack
The stack ships its own Postgres and its own object store and points at both by service name. `.env` moves either one somewhere else, and the bundled service is used when the line is absent:
```ini
FLUXER_POSTGRES_HOST=db.example.com
FLUXER_POSTGRES_PORT=5432
FLUXER_POSTGRES_SSL=true
FLUXER_S3_ENDPOINT=https://s3.eu-central-1.amazonaws.com
FLUXER_S3_REGION=eu-central-1
FLUXER_S3_FORCE_PATH_STYLE=false
```
`FLUXER_S3_PUBLIC_ENDPOINT` follows `FLUXER_S3_ENDPOINT` when it is not set on its own, and the credentials stay `FLUXER_S3_ACCESS_KEY` and `FLUXER_S3_SECRET_KEY`. `FLUXER_S3_BUCKET_CDN`, `FLUXER_S3_BUCKET_UPLOADS`, `FLUXER_S3_BUCKET_REPORTS` and `FLUXER_S3_BUCKET_HARVESTS` name the buckets. The bundled object store creates whichever names they hold, and a store outside the stack needs those buckets to exist already.
`FLUXER_KV_URL`, `FLUXER_NATS_URL`, `FLUXER_SEARCH_URL` and `FLUXER_LIVEKIT_INTERNAL_URL` move the other bundled services the same way, and `FLUXER_NATS_JETSTREAM_URL` and `FLUXER_SVC_NATS_URL` follow `FLUXER_NATS_URL` when they are not set on their own.
Pointing the stack elsewhere leaves the bundled service defined and running with nothing reading it. The object store has a shipped overlay for that. `external-object-store.compose.yml` keeps `seaweedfs` and `seaweedfs-init` from starting, and `api`, `worker` and `media-proxy` stop waiting for `seaweedfs-init`. Add it to `COMPOSE_FILE` in `.env`, after any other overlay:
```ini
COMPOSE_FILE=docker-compose.yml:external-object-store.compose.yml
```
Behind a reverse proxy the line is `docker-compose.yml:docker-compose.proxy.yml:external-object-store.compose.yml`. The overlay needs Compose 2.24.4 or newer. An upgrade refreshes it with the other stack files, skips the uploads copy and does not wait for `seaweedfs-init`.
Compose leaves containers that already exist in place, so remove the two old ones once after the switch:
```bash
docker compose rm -sf seaweedfs seaweedfs-init
```
The `seaweedfs-data` volume stays until you delete it. Copy its objects to the new store before you delete it.
The other bundled services have no shipped overlay. Take one out with an override file listed in `COMPOSE_FILE`, which [Keep a local compose change](#keep-a-local-compose-change) describes, rather than by editing `docker-compose.yml`, which the Place step replaces on every upgrade.
The installer backs up only the bundled database and object store. It skips the database dump when the stack defines no `postgres` service, and when `FLUXER_POSTGRES_HOST` or the host in `FLUXER_POSTGRES_URL` names anything other than `postgres`. The bundled service is idle in that shape, and its data directory still holds the role and database it was first started with, so the by-hand dump and restore commands on this page do not work against it either. Arrange separate backups for external stores before upgrading, and keep them with the release's backup record.
## Roll back
A rollback puts the previous release back:
```bash
cd ~/fluxer
sh install.sh --rollback
```
On Windows it is `.\install.ps1 -Rollback`. It takes the newest record, puts back the images and the stack files it holds, recreates, restarts `edge`, and verifies. `--dry-run` prints that plan too.
A rollback never pulls, so it needs the old images still on the host. Run [Reclaim disk](#reclaim-disk) only once the upgrade is known good.
A rollback does not undo database migrations or restore data. If the previous release requires the old schema, restore the database backup separately.
## Restore a backup
Restore both artefacts with the stack stopped, so nothing writes while they are replaced. Name your own record directory instead of the one below.
The database restores from the custom-format dump:
```bash
docker compose stop
docker compose up -d --wait postgres
docker compose exec -T postgres sh -c 'pg_restore -U $POSTGRES_USER -d $POSTGRES_DB --clean --if-exists' \
< backups/record-20260831T120000Z/fluxer.dump
docker compose up -d
```
`--clean` prints notices about objects that do not exist yet, which is expected against a fresh directory. The role and database come from the `postgres` container's environment, which follows `FLUXER_POSTGRES_USERNAME` and `FLUXER_POSTGRES_DATABASE`, the same way the installer's dump reads them. These commands are for the bundled database. A database outside the stack restores with its own tools.
PowerShell has no `<` redirection. On Windows, copy the dump into the container and name it as a file:
```powershell
docker compose stop
docker compose up -d --wait postgres
docker compose cp backups\record-20260831T120000Z\fluxer.dump postgres:/tmp/fluxer.dump
docker compose exec -T postgres sh -c 'pg_restore -U $POSTGRES_USER -d $POSTGRES_DB --clean --if-exists /tmp/fluxer.dump'
docker compose exec -T postgres rm /tmp/fluxer.dump
docker compose up -d
```
The uploads restore through the same helper container the backup used, in reverse:
```bash
docker compose stop
docker run --rm -v fluxer_seaweedfs-data:/data \
-v "$PWD/backups/record-20260831T120000Z:/backup" alpine:3.22 \
sh -c 'find /data -mindepth 1 -delete && tar xzf /backup/seaweedfs-data.tgz -C /data'
docker compose up -d
```
For a `seaweedfs-data.tar`, write `tar xf` in place of `tar xzf`. `tar` extracts over whatever is already on the volume, so the `find` empties it first. In PowerShell write `${PWD}` instead of `$PWD`. Success is an existing attachment URL answering 200 again. The `fluxer_` prefix is the Compose project name, which `docker-compose.yml` sets to `fluxer`.
## Move to a new Postgres major version
`docker-compose.yml` runs `postgres:16-alpine` unless `FLUXER_POSTGRES_IMAGE` in `.env` names another image. A newer major does not read the data directory an older major wrote, so the data moves across through a dump. The upgrade refuses when the refreshed compose file would run a different major from the current one, because the move destroys the volume holding the database. It reads the current major from the image the old file names, and takes `FLUXER_POSTGRES_IMAGE` into account only for a file that reads it.
Take the dump while the old major is still serving, then stop everything and drop the volume:
```bash
docker compose exec -T postgres sh -c 'pg_dump -U $POSTGRES_USER -d $POSTGRES_DB --format=custom' \
> backups/pre-major.dump
docker compose down
docker volume rm fluxer_postgres-data
```
On Windows, run `pg_dump` inside the container with `-f /tmp/pre-major.dump` and copy the file out with `docker compose cp`. Restore it with `docker compose cp` and `pg_restore` as [Restore a backup](#restore-a-backup) shows.
Once that volume is removed, the dump is the only copy of the database.
`FLUXER_POSTGRES_IMAGE` takes effect only once `docker-compose.yml` reads it. An instance on older stack files runs `sh install.sh --update` first, or edits the tag in its `docker-compose.yml` as before.
Set `FLUXER_POSTGRES_IMAGE` in `.env` to the new tag, such as `postgres:17-alpine`, start the database on its own, and restore into the empty directory:
```bash
docker compose up -d --wait postgres
docker compose exec -T postgres sh -c 'pg_restore -U $POSTGRES_USER -d $POSTGRES_DB --clean --if-exists' \
< backups/pre-major.dump
docker compose up -d
```
Run `sh install.sh --update` afterwards to finish the move on the refreshed files. Remove the `FLUXER_POSTGRES_IMAGE` line once the refreshed `docker-compose.yml` runs the same major by default.
## Pin a release
`FLUXER_IMAGE_TAG=v1` tracks the latest compatible release, so every upgrade can move the instance. Set it to a release tag to hold one version and upgrade on your own schedule.
A pinned instance also pins its stack files. [Match the images to the stack files](#match-the-images-to-the-stack-files) has the `--ref` a pinned tag needs.
The tag applies only to the Fluxer images. `caddy`, `postgres`, `valkey`, `nats`, `meilisearch`, `seaweedfs` and `livekit` take their tags from `docker-compose.yml`, and move when an upgrade refreshes that file. An image name in `.env`, such as `FLUXER_POSTGRES_IMAGE`, holds one of them in place instead. [Images](/operator/configuration/#images) lists the names.
## Change the domain or the passkey relying party
`FLUXER_DOMAIN` supplies the default `FLUXER_PASSKEY_RP_ID`, which is the domain a passkey is tied to. `FLUXER_PUBLIC_ORIGIN` does not move it. A passkey works only under the identifier it was created with.
:::danger[Changing either value invalidates every passkey]
Passkeys registered under the old `FLUXER_PASSKEY_RP_ID` stop working and cannot be recovered. Every user has to register a new one, so keep another sign-in method working before you change it.
:::
Changing `FLUXER_DOMAIN` also changes the certificate the edge requests, the CORS origins the API accepts, and the admin OAuth redirect URI.
## Reclaim disk
Each upgrade leaves the previous images behind, which adds up to a few gigabytes over time:
```bash
docker image prune -f
```
On the default moving tag, the pull moves `v1` onto the new images and leaves the ones it replaced untagged, so this removes them. A rollback on a moving tag needs those images. Prune only once the upgrade is known good. A pinned release keeps its own tag and survives the prune. Remove one of those by naming its tag in `docker image rm`.
`docker system prune -a` reclaims more. It also deletes every image that no container references, including ones unrelated to Fluxer.
:::danger[Never add -a to a volume prune]
`docker volume prune -a` removes named volumes no container uses. After a `docker compose down` that is every volume the instance has, which is the whole instance.
:::