Files
fluxer/fluxer_docs/src/content/docs/operator/upgrading.mdx
T

351 lines
22 KiB
Plaintext

---
# SPDX-License-Identifier: AGPL-3.0-or-later
title: Upgrading
description: Upgrading and rolling back an instance with the installer, what the backup covers, and how to restore it.
---
:::tip[Plutonium and donations fund Fluxer]
We are grateful to everyone supporting the project through [Fluxer Plutonium](https://fluxer.app/plutonium) or [donations](https://fluxer.app/donate). All of our code is free and open source on [GitHub](https://github.com/fluxerapp/fluxer). The Operator Pass is coming, and adds a direct line to the team for help and feedback.
:::
An upgrade moves an instance to a newer release. `install.sh --update` does all of it: back up, download the new stack files, pull the new images, recreate the containers that changed. The stack files are the three compose files, the `Caddyfile` and `.env.example`. Allow about twenty minutes, most of it the backup and the pull.
| Requirement | Value |
| --- | --- |
| Working directory | The one holding `.env`, `docker-compose.yml`, `docker-compose.proxy.yml`, `tunnel.compose.yml` and `Caddyfile` |
| Compose commands | Behind your own reverse proxy, set `COMPOSE_FILE` in `.env` once and every command picks up the overlay |
| Free disk | Room for one database dump, one copy of the uploads, and the new images beside the old ones |
| Downtime | A minute or two for the recreate, plus however long the uploads copy takes with the stack stopped |
Read the release notes between the version you run and the one you are moving to, at [Releases](https://github.com/fluxerapp/fluxer/releases).
## Match the images to the stack files
The images come from `FLUXER_IMAGE_TAG` in `.env`. The stack files come from a git ref, which is a branch or a tag in the Fluxer repository. The installer derives that ref from the image tag: `main` when the tag is `v1` or `latest`, and the tag string itself for anything else. `--ref` overrides the derivation. `docker compose pull` never updates a stack file, and refreshing a stack file never moves an image.
The repository holds no ref by the name of a pinned image tag. A release tags each image on its own, as `[email protected]`, so pass `--ref` with that tag or with the commit it points at. Without that override the download fails with exit 4.
## What the script does
`--update` runs these steps in this order, and stops on the first one that fails:
| Step | What it does |
| --- | --- |
| Mint | Writes any key the refreshed stack requires that `.env` does not hold, listed under [Run the upgrade](#run-the-upgrade) |
| Record | Writes the image references, the image ID each container has, and the tag `.env` names, into a new record directory |
| Save | Copies `.env` and the five stack files into that record |
| Dump | Dumps the database with the stack still serving, and checks the custom-format header on the result |
| Copy | Stops the stack, copies the uploads volume, and starts it again on the images it was already running |
| Fetch | Downloads the stack files at the ref into a staging directory inside the working directory |
| Guard | Refuses when the refreshed `docker-compose.yml` moves Postgres to a new major version |
| Place | Moves every downloaded file into the working directory at once |
| Pull | Pulls the images while the old containers still serve |
| Recreate | Runs `docker compose up -d --remove-orphans`, then restarts `edge` when the `Caddyfile` changed |
| Verify | Polls Compose state until every service is ready, then probes `/_health` on the hostname |
A failed run leaves the instance running on the images it already had.
`docker compose up -d` does not notice a changed `Caddyfile`. The script restarts `edge` by name to pick it up.
An instance older than the `/livekit` routing needs one edit no upgrade can make for it. [Voice signalling moved to /livekit](#voice-signalling-moved-to-livekit) has it, and voice stays silent until it is made.
The refreshed `docker-compose.yml` renames the `caddy` service to `edge`, and the two publish the same host ports. `--remove-orphans` takes the old container down in the same call, so the new one can bind them. Without it the step stops with `Bind for 0.0.0.0:443 failed`. It also removes any container in the project whose service the loaded Compose files no longer define, so keep `COMPOSE_FILE` the same across an upgrade. [Keep a local compose change](#keep-a-local-compose-change) has the file layout that survives one. [When the compose file list names a missing file](#when-the-compose-file-list-names-a-missing-file) has the failure a line naming an absent file produces.
The rename also moves the `caddy-data` and `caddy-config` volumes to `edge-data` and `edge-config`, so the old two are left unused and the edge requests its certificate again on the first start.
`FLUXER_IMAGE_TAG` is the only line in `.env` that either upgrade mode rewrites once the instance holds every key the stack requires. An upgrade also writes a key the refreshed stack requires that `.env` does not hold at all, listed below, and it leaves every value already set as it is.
The refreshed `docker-compose.yml` requires `FLUXER_ERLANG_COOKIE`. An `.env` written by an earlier installer holds no such line, and `--update` writes one into `.env` before it reads anything, so an upgrade needs no edit for it. The value is 64 hex characters, which is what a fresh install writes. Until `.env` sets it, every Compose command against the stack stops with `set FLUXER_ERLANG_COOKIE in .env`.
To write it by hand, run the generator in a shell:
```bash
printf '\nFLUXER_ERLANG_COOKIE=%s\n' "$(openssl rand -hex 32)" >> .env
```
Run that as a command. Pasting `FLUXER_ERLANG_COOKIE=$(openssl rand -hex 32)` into `.env` as text stores those characters as the value, because Compose reads a line literally and runs nothing in it. The leading newline keeps the key on its own line. An editor can leave the last line of `.env` without a newline, and the two would otherwise join.
`api` and `media-proxy` now refuse to start unless `FLUXER_MEDIA_PROXY_UPLOAD_RELAY_SECRET_BASE64` decodes to at least 32 bytes. The line has been in `.env.example` for a while as `CHANGE_ME`, which decodes to 6 bytes, and nothing read it before, so an instance can be running today with the placeholder. `--update` replaces an absent or `CHANGE_ME` value with a generated one before it reads anything. Both services read the same value from one `.env` line. A `CHANGE_ME` value stops `api` at boot with `FLUXER_MEDIA_PROXY_UPLOAD_RELAY_SECRET_BASE64 must decode to at least 32 bytes`, and an empty one with `FLUXER_MEDIA_PROXY_UPLOAD_RELAY_SECRET_BASE64 is required for the API`.
To write it by hand, run the generator in a shell the same way:
```bash
printf '\nFLUXER_MEDIA_PROXY_UPLOAD_RELAY_SECRET_BASE64=%s\n' "$(openssl rand -base64 32)" >> .env
```
## The script is the reference
Every step above sits in the script beside a comment holding the command that does that step alone and the reason the step exists. Read [https://fluxer.dev/install.sh](https://fluxer.dev/install.sh), or [https://fluxer.dev/install.ps1](https://fluxer.dev/install.ps1) for Windows.
## Keep a local compose change
The Place step replaces all five stack files with the copies at the ref. An edit made directly in `docker-compose.yml`, `docker-compose.proxy.yml`, `tunnel.compose.yml` or the `Caddyfile` is gone once that step runs. The run prints one line for the whole step and never names the files it replaced. `--dry-run` names every one of them, as `changes`, `unchanged` or `is new`.
The supported way to hold a local choice is a separate file, listed in `COMPOSE_FILE` in `.env`:
```ini
COMPOSE_FILE=docker-compose.yml:docker-compose.proxy.yml:local.compose.yml
```
Compose merges the files left to right, so `local.compose.yml` wins over the ones before it. An upgrade refreshes the five stack files and nothing else, so a file outside that set survives every upgrade untouched. A record copies `.env` and those five files, and no other file from the working directory. An override file is never backed up with them, and it belongs wherever the rest of the configuration lives.
`.env` needs no such file. The only value either upgrade mode replaces there is `FLUXER_IMAGE_TAG`, and the Mint step adds a required key the file does not hold at all.
[Enable the overlay](/operator/reverse-proxy/#enable-the-overlay) has the overlay this is most often used for. [Docker labels](/operator/reverse-proxy/#docker-labels) has a worked third file.
## When the compose file list names a missing file
`COMPOSE_FILE` names files relative to the working directory, and Compose refuses to load the set when one of them is absent:
```
stat /srv/fluxer/docker-compose.proxy.yml: no such file or directory
```
Every `docker compose` command stops there. An upgrade checks the list before it runs any of them, so it stops with a sentence naming the file and the `curl` that puts it back, and writes nothing.
This happens on an instance older than the installer. `docker-compose.proxy.yml` and `tunnel.compose.yml` are files the upgrade downloads, so a `COMPOSE_FILE` line naming one of them in a directory that never held it fails long before the Fetch step that would have supplied it.
Download the missing file at the ref the upgrade moves to, then run the upgrade:
```bash
curl -fsSL --proto '=https' --tlsv1.2 -o docker-compose.proxy.yml \
https://raw.githubusercontent.com/fluxerapp/fluxer/main/deploy/self-hosting/docker-compose.proxy.yml
```
`main` is the ref for `FLUXER_IMAGE_TAG=v1` and for `latest`. [Match the images to the stack files](#match-the-images-to-the-stack-files) has the ref a pinned tag needs.
Removing the line from `.env` also gets the run moving and is the wrong fix. Without the overlay the edge publishes `80` and `443` and takes them from the reverse proxy already on the host.
## Run the upgrade
Keep the installer beside the stack files. [Get started](/operator/get-started/#step-4-bring-up-the-instance) has the download and the checksum check for a fresh copy.
See the plan first:
```bash
cd ~/fluxer
sh install.sh --update --dry-run
```
It prints:
- The images the stack runs now
- The backup it intends to take
- The stack files the ref changes
- The services those changes make it restart
- Any refusal that would stop the run
It writes nothing outside a temporary directory it removes on exit.
Then run it:
```bash
sh install.sh --update
```
On Windows:
```powershell
.\install.ps1 -Update
```
The run ends by naming the record it wrote and the command that rolls back to it. It checks its own work with the readiness poll and one probe of `/_health`. Run the public probes in [Check that it works](/operator/get-started/#step-5-check-that-it-works) afterwards.
| Flag | Windows | Meaning |
| --- | --- | --- |
| `--update` | `-Update` | Record, back up, refresh, pull, recreate, verify |
| `--rollback` | `-Rollback` | Put back the images and stack files of the newest record |
| `--dry-run` | `-DryRun` | Print the plan. Change nothing |
| `--backup-dir <path>` | `-BackupDir` | Where records go. Default `<dir>/backups` |
| `--no-volume-backup` | `-NoVolumeBackup` | Take the database dump and skip the uploads copy |
| `--skip-backup-accept-data-loss` | `-SkipBackupAcceptDataLoss` | Upgrade with no backup at all |
`--dir` and `--ref` mean what they mean on an install, and the installer derives an omitted `--ref` from the `FLUXER_IMAGE_TAG` line in `.env`. Every failure prints a sentence on standard error and exits non-zero, and the script header lists what each code means. A Postgres major version change is exit 3.
## Voice signalling moved to /livekit
The edge serves LiveKit at `/livekit/*`. An instance set up before that routing served it under `/gateway/livekit`, and that path is gone. Voice stops connecting after the upgrade and the run still reports success, because the readiness poll and the `/_health` probe never reach voice.
Before:
```
wss://chat.example.com/gateway/livekit
```
After:
```
wss://chat.example.com/livekit
```
Only the path is wrong. `docker-compose.yml` derives `https://chat.example.com/livekit` when `.env` sets no `FLUXER_LIVEKIT_URL`, and the client rewrites a leading `http` to `ws` itself, so `https` and `wss` both work.
The address browsers dial is a column on a voice server row in the database. No upgrade migrates that row, and it is corrected in two places.
First `.env`. When it sets `FLUXER_LIVEKIT_URL`, put the new path in that line and run `docker compose up -d api`. `api` writes that value onto the default voice server row at every start, so a row corrected in the dashboard while `.env` still names the old path goes back to the old path on the next restart. Removing the line is also correct, and Compose then derives the URL from `FLUXER_PUBLIC_ORIGIN`, or from the scheme and the domain.
Then the dashboard, at `https://chat.example.com/admin` under **Voice Servers**. The **Endpoint** field on that page holds the address. `api` writes one row only, the one in region `default` with server id `default-server-1`, and leaves every other row as it is. That row corrects itself on the first `api` start after `.env` is right, so the dashboard is for every other row. When that region or that server is absent, its log says `skipping config sync` and it writes nothing. A region added by hand, or one left from an older layout under another id, keeps the endpoint it already holds until that field is edited.
Voice media is unaffected. It never went through the edge, and `7881/tcp` and `7882/udp` still reach the host directly.
## What the backup covers
Each upgrade writes one record directory, named `record-` and a UTC stamp, under `backups` or wherever `--backup-dir` points. It holds:
| Artifact | What it covers |
| --- | --- |
| `fluxer.dump` | Every account, message, guild and configuration row, in the Postgres custom format |
| `seaweedfs-data.tgz` | Every upload, avatar, report and harvest |
| `.env` | Every secret the instance was built with |
The five stack files sit in the record beside them, so a rollback puts back the exact files the instance was running. The record directory is created `0700` and the `.env` copy inside it is `0600`. Keep records wherever you already keep secrets.
The dump costs no downtime. The uploads copy stops the stack, and the script starts the stack again on the old images before it goes any further.
Nothing else is copied. The dump covers `postgres-data`. The other five volumes either rebuild themselves or hold queued work, and [Volumes and buckets](/operator/configuration/#volumes-and-buckets) lists all seven with what each one holds.
Put a dump on a timer as well. On Linux and macOS, this crontab line writes one a night and keeps two weeks of them:
```cron
15 3 * * * cd /srv/fluxer && docker compose exec -T postgres pg_dump -U fluxer -d fluxer --format=custom > "backups/fluxer-$(date -u +\%Y\%m\%dT\%H\%M\%SZ).dump" && find backups -name 'fluxer-*.dump' -mtime +14 -delete
```
`/srv/fluxer` stands for the directory holding `.env`. Write it as an absolute path, because cron does not expand `~`. An unescaped `%` in a crontab is a newline, which is why every one of them has a backslash.
## Run a data store outside the stack
The stack ships its own Postgres and its own object store and points at both by service name. `.env` moves either one somewhere else, and the bundled service is used when the line is absent:
```ini
FLUXER_POSTGRES_HOST=db.example.com
FLUXER_POSTGRES_PORT=5432
FLUXER_POSTGRES_SSL=true
FLUXER_S3_ENDPOINT=https://s3.eu-central-1.amazonaws.com
FLUXER_S3_REGION=eu-central-1
FLUXER_S3_FORCE_PATH_STYLE=false
```
`FLUXER_S3_PUBLIC_ENDPOINT` follows `FLUXER_S3_ENDPOINT` when it is not set on its own, and the credentials stay `FLUXER_S3_ACCESS_KEY` and `FLUXER_S3_SECRET_KEY`. `FLUXER_S3_BUCKET_CDN`, `FLUXER_S3_BUCKET_UPLOADS`, `FLUXER_S3_BUCKET_DOWNLOADS`, `FLUXER_S3_BUCKET_REPORTS` and `FLUXER_S3_BUCKET_HARVESTS` name the five buckets. The bundled object store creates whichever names they hold, and a store outside the stack needs those buckets to exist already.
`FLUXER_KV_URL`, `FLUXER_NATS_URL`, `FLUXER_SEARCH_URL` and `FLUXER_LIVEKIT_INTERNAL_URL` move the other four bundled services the same way, and `FLUXER_NATS_JETSTREAM_URL` and `FLUXER_SVC_NATS_URL` follow `FLUXER_NATS_URL` when they are not set on their own.
Pointing the stack elsewhere leaves the bundled service defined and running with nothing reading it. Take it out with an override file listed in `COMPOSE_FILE`, which [Keep a local compose change](#keep-a-local-compose-change) describes, rather than by editing `docker-compose.yml`, which the Place step replaces on every upgrade.
An upgrade backs up what the stack holds. The Dump step reaches into the `postgres` service and the Copy step reaches into the `seaweedfs` volume, so a stack that defines neither has neither step to run. The run says which one it skipped and goes on. Backing up a store outside the stack belongs to whoever runs it, and the record the upgrade writes holds no dump of it, so take one before upgrading if a rollback would need it.
## Roll back
A rollback puts the previous release back:
```bash
cd ~/fluxer
sh install.sh --rollback
```
On Windows it is `.\install.ps1 -Rollback`. It takes the newest record, puts back the images and the stack files it holds, recreates, restarts `edge`, and verifies. `--dry-run` prints that plan too.
A rollback never pulls, so it needs the old images still on the host. Run [Reclaim disk](#reclaim-disk) only after an upgrade satisfies you.
The database does not move. `api`, `worker`, `users-shard` and `messages-shard` apply schema work in place while they start, and an older image does not undo it. Across a release that changed the schema, the dump in the record is the only way back, and putting it back is a decision you make.
## Restore a backup
Both artifacts go back with the stack stopped, so nothing writes while they are replaced. Name your own record directory in place of the one below.
The database restores from the custom-format dump:
```bash
docker compose stop
docker compose up -d --wait postgres
docker compose exec -T postgres pg_restore -U fluxer -d fluxer --clean --if-exists \
< backups/record-20260831T120000Z/fluxer.dump
docker compose up -d
```
`--clean` prints notices about objects that do not exist yet, which is expected against a fresh directory.
PowerShell has no `<` redirection. On Windows, copy the dump into the container and name it as a file:
```powershell
docker compose stop
docker compose up -d --wait postgres
docker compose cp backups\record-20260831T120000Z\fluxer.dump postgres:/tmp/fluxer.dump
docker compose exec -T postgres pg_restore -U fluxer -d fluxer --clean --if-exists /tmp/fluxer.dump
docker compose exec -T postgres rm /tmp/fluxer.dump
docker compose up -d
```
The uploads restore through the same helper container the backup used, in reverse:
```bash
docker compose stop
docker run --rm -v fluxer_seaweedfs-data:/data \
-v "$PWD/backups/record-20260831T120000Z:/backup" alpine:3.22 \
sh -c 'find /data -mindepth 1 -delete && tar xzf /backup/seaweedfs-data.tgz -C /data'
docker compose up -d
```
`tar` extracts over whatever is already on the volume, so the `find` empties it first. In PowerShell write `${PWD}` in place of `$PWD`. Success is an existing attachment URL answering 200 again. The `fluxer_` prefix is the Compose project name, which `docker-compose.yml` sets to `fluxer`.
## Move to a new Postgres major version
`docker-compose.yml` pins `postgres:16-alpine`. A newer major does not read the data directory an older major wrote, so the data moves across through a dump. The upgrade refuses when the refreshed compose file changes that pin, because the move destroys the volume holding the database.
Take the dump while the old major is still serving, then stop everything and drop the volume:
```bash
docker compose exec -T postgres pg_dump -U fluxer -d fluxer --format=custom \
> backups/pre-major.dump
docker compose down
docker volume rm fluxer_postgres-data
```
On Windows, take that dump the way [Restore a backup](#restore-a-backup) moves one, with `-f /tmp/pre-major.dump` and `docker compose cp`, and restore it the same way.
Once that volume is removed, the dump is the only copy of the database.
Put the new `postgres` tag in `docker-compose.yml`, start the database on its own, and restore into the empty directory:
```bash
docker compose up -d --wait postgres
docker compose exec -T postgres pg_restore -U fluxer -d fluxer --clean --if-exists \
< backups/pre-major.dump
docker compose up -d
```
Run `sh install.sh --update` afterwards to finish the move on the refreshed files.
## Pin a release
`FLUXER_IMAGE_TAG=v1` tracks the latest compatible release, so every upgrade can move the instance. Set it to a release tag to hold one version and upgrade on your own schedule.
A pinned instance also pins its stack files. [Match the images to the stack files](#match-the-images-to-the-stack-files) has the `--ref` a pinned tag needs.
The tag applies only to the eleven Fluxer images. `caddy`, `postgres`, `valkey`, `nats`, `meilisearch`, `seaweedfs`, and `livekit` are pinned inside `docker-compose.yml` and move only when an upgrade refreshes that file.
## Change the domain or the passkey relying party
`FLUXER_DOMAIN` supplies the default `FLUXER_PASSKEY_RP_ID`, which is the domain a passkey is tied to. A passkey works only under the identifier it was created with.
:::danger[Changing either value invalidates every passkey]
Passkeys registered under the old `FLUXER_PASSKEY_RP_ID` stop working and cannot be recovered. Every user has to register a new one, so keep another sign-in method working before you change it.
:::
Changing `FLUXER_DOMAIN` also changes the certificate the edge requests, the CORS origins the API accepts, and the admin OAuth redirect URI.
## Reclaim disk
Each upgrade leaves the previous images behind, which adds up to a few gigabytes over time:
```bash
docker image prune -f
```
On the default moving tag, the pull moves `v1` onto the new images and leaves the ones it replaced untagged, so this removes them. A rollback on a moving tag needs those images. Prune only once the upgrade satisfies you. A pinned release keeps its own tag and survives the prune. Remove one of those by naming its tag in `docker image rm`.
`docker system prune -a` reclaims more. It also deletes every image that no container references, including ones unrelated to Fluxer.
:::danger[Never add -a to a volume prune]
`docker volume prune -a` removes named volumes no container uses. After a `docker compose down` that is every volume the instance has, which is the whole instance.
:::