Restoring platform databases

If you need to restore your database (for instance, if it becomes corrupted), you can do so using the backups within App Platform. Before proceeding, familiarize yourself with the CNPG documentation on PostgreSQL database recovery. The steps here are specific to App Platform installed on LKE.

Recovery options

Regular recovery

Use this procedure when a database is in an unhealthy state, for example due to volume filesystem corruption. To revert a database when there have been undesired changes, use point-in-time recovery instead.

  1. Note the name of the Backup resource you intend to restore from, if available.
  2. Prepare access to object storage for recovery.
  3. Prepare direct access to ArgoCD in case single sign-on becomes unavailable during recovery.
  4. Pause automated sync operations for apl-operator and ArgoCD.
  5. Shut down any service that accesses the database.
  6. Update the backup configuration and edit the Application manifest of the database accordingly.
  7. Delete the database Cluster resource and reactivate ArgoCD sync.
  8. Restart applications using the database and apl-operator.

List backup resources

List the backups that are available for application's database (Gitea, Harbor, or Keycloak). Only use completed backups for recovery. Backup name timestamps are in UTC.

kubectl get backup -n gitea

Prepare access to object storage

The database recovery requires an ObjectStore resource pointing to the backup location. When backups are enabled, this resource already exists in the cluster and can be copied and modified for recovery.

  1. Dump the existing resource to a file:

    kubectl get objectstores.barmancloud.cnpg.io -n gitea gitea-db -oyaml > recovery.yaml
  2. Edit the file and make the following changes:

    • Remove the status section.
    • In metadata, keep only namespace and change name, for example to keycloak-db-recovery.
    • Leave the spec section unchanged.

    It should look like the following (here provided for Keycloak). If the resource is not available, you can also use this as a template:

    apiVersion: barmancloud.cnpg.io/v1
    kind: ObjectStore
    metadata:
      # Replace name. Note this needs to correspond with the "source" value later.
      name: keycloak-db-recovery
      # Replace namespace with "gitea" or "harbor" as applicable.
      namespace: keycloak
    spec:
      configuration:
        data:
          compression: gzip
        # Replace the bucket name, "keycloak" with the application name
        # (pathSuffix in previous backup configuration)
        destinationPath: s3://bucket-name/keycloak
        # Replace "nl-ams-1" with your storage region
        endpointURL: https://nl-ams-1.linodeobjects.com
        s3Credentials:
          accessKeyId:
            key: S3_STORAGE_ACCOUNT
            name: linode-creds
          secretAccessKey:
            key: S3_STORAGE_KEY
            name: linode-creds
        wal:
          compression: gzip
      instanceSidecarConfiguration:
        env:
        - name: AWS_REQUEST_CHECKSUM_CALCULATION
          value: WHEN_REQUIRED
        - name: AWS_RESPONSE_CHECKSUM_VALIDATION
          value: WHEN_REQUIRED
        logLevel: info
        resources:
          limits:
            cpu: 300m
            memory: 256Mi
          requests:
            cpu: 50m
            memory: 48Mi
        retentionPolicyIntervalSeconds: 1800
      retentionPolicy: 7d
  3. Store this resource in the cluster by either running kubectl apply -f recovery.yaml (replace file name as needed) or committing it to the values repository at env/manifests/<namespace>/<resource-name>.yaml. In the latter case, ArgoCD will synchronize this ObjectStore resource to the cluster.

Access to ArgoCD

During recovery of platform-critical services, such as Keycloak, single sign-on becomes unavailable. Retaining direct access to ArgoCD lets you synchronize changes and monitor status throughout the process. This step is not required for Harbor.

  1. Retrieve the ArgoCD admin password:

    kubectl get secrets -n argocd argocd-initial-admin-secret -ojson | jq -r '.data.password' | base64 -d
  2. Configure port forwarding to access ArgoCD at http://localhost:8080 (change the port if there is a conflict):

    kubectl port-forward -n argocd svc/argocd-server 8080:80

Deactivate sync and apl-operator

While making temporary manual changes, disable automated reconciliation to prevent conflicting modifications.

  1. First, scale down apl-operator and remove its ArgoCD auto-sync policy:

    kubectl patch applications.argoproj.io -n argocd apl-operator-apl-operator --patch '[{"op": "remove", "path": "/spec/syncPolicy/automated"}]' --type=json
    kubectl scale --replicas=0 -n apl-operator deployment apl-operator
  2. Then, disable auto-sync for the database ArgoCD application.

    kubectl patch applications.argoproj.io -n argocd gitea-gitea-otomi-db --patch '[{"op": "remove", "path": "/spec/syncPolicy/automated"}]' --type=json

Shut down services

Shut down the application before backing up or restoring the database to prevent write operations that could cause inconsistencies.

Commands for temporarily disabling Gitea:

## Disable ArgoCD auto-sync during the changes
kubectl patch applications.argoproj.io -n argocd gitea-gitea --patch '[{"op": "remove", "path": "/spec/syncPolicy/automated"}]' --type=json
## Scale Gitea deployment to zero
kubectl patch deploy -n gitea gitea --patch '[{"op": "replace", "path": "/spec/replicas", "value": 0}]' --type=json
## Verify that pods are shut down
kubectl get deploy -n gitea gitea  # Should show READY 0/0

Adjustments to the backup configuration

After recovery, CNPG creates new backups. To prevent accidental mixing or overwriting of backups, CNPG does not allow the new backup destination and the recovery source to share the same location. The path suffix (the directory inside the object storage bucket) must be changed.

  1. In env/settings.yaml, set pathSuffix under platformBackups.database.<app> (where <app> is gitea, harbor, or keycloak) to a new value that does not already exist in object storage, for example <app>-1:

    # ...
    platformBackups:
        database:
            gitea:
                # ...
                pathSuffix: gitea-1
    # ...
  2. Because ArgoCD sync is currently disabled, apply this change directly to the application manifest as well. The most convenient way to do this is through the ArgoCD UI. Note that helm.values is a multiline string.

    In the helm.values structure, update backup.linode.destinationPath to match the new path suffix, and set clusterSpec.bootstrap.recovery and clusterSpec.externalClusters as shown below.

    🚧

    Important: Save these changes but do not synchronize yet.

    Example for Gitea (application gitea-gitea-otomi-db):

    # ...
    source:
      # ...
      helm:
        values: >
          backup:
            enabled: true
            linode:
              # Modify suffix according to new pathSuffix in values
              destinationPath: s3://bucket-name/gitea-1
          # ...
          clusterSpec:
              bootstrap:
                  recovery:
                      backup:
                          # Name of Backup resource from step 1
                          # If not set, will restore to latest backup available in bucket
                          name: ...
                      database: gitea
                      owner: gitea
                      # To correspond with name in externalClusters
                      source: recovery-obj
              externalClusters:
                - name: recovery-obj
                  plugin:
                      enabled: true
                      isWALArchiver: false
                      name: barman-cloud.cloudnative-pg.io
                      parameters:
                          # Name of ObjectStore resource from step 2
                          barmanObjectName: gitea-db-recovery
                          serverName: gitea-db
      # ...

Restore the database

After shutting down applications and preparing the configuration changes, delete the database Cluster. ArgoCD then re-creates the cluster and restores it from the backup.

The following kubectl commands are equivalent to the same operations in the ArgoCD UI.

Commands for Gitea:

## Delete the cluster
kubectl delete cluster -n gitea gitea-db
## Re-enable ArgoCD auto-sync
kubectl patch application -n argocd gitea-gitea-otomi-db --patch '[{"op": "add", "path": "/spec/syncPolicy/automated", "value": {"prune": true, "allowEmpty": true}}]' --type=json

Restart services

Once maintenance is complete, restart apl-operator:

kubectl scale deploy -n apl-operator apl-operator --replicas=1

apl-operator eventually reverts the manual edits to Application manifests without disrupting the database. It also re-enables automated synchronization and restarts services. To bring services back online faster, use the following commands.

Commands for Gitea:

## Re-enable ArgoCD auto-sync, which should also change the Gitea statefulset to scale up
kubectl patch applications.argoproj.io -n argocd gitea-gitea --patch '[{"op": "add", "path": "/spec/syncPolicy/automated", "value": {"prune": true, "allowEmpty": true}}]' --type=json
## Optional: scale up, for not having to wait for re-sync of ArgoCD
kubectl patch statefulset -n gitea gitea --patch '[{"op": "replace", "path": "/spec/replicas", "value": 1}]' --type=json

Point-in-time recovery

To restore a database to a specific point in time, edit the application manifest. The most convenient way to do this is through the ArgoCD UI. Note that helm.values is a multiline string.

Add a recoveryTarget to the recovery section following the CloudnativePG docs. The example below restores Gitea to the state just before a change made after 2023-03-06 08:00:39 CET.

# ...
source:
  # ...
  helm:
    values: >
      # ...
      clusterSpec:
          bootstrap:
              recovery:
                  database: gitea
                  owner: gitea
                  source: recovery-obj  # To correspond with name in externalClusters
                  recoveryTarget:
                      # Time base target for the recovery
                      targetTime: "2023-03-06 08:00:39+01"
# ...

You can also specify a named backup resource in the backup.name field (step 1). The backup must be from before the chosen recovery target timestamp, noting that backup names use UTC timestamps.

The targetTime format is not standard ISO 8601. Instead, the date and time are separated by a space and the timezone is written as an explicit offset, such as +00 (UTC) or +01 (CET without DST). For all valid formats, see the PostgreSQL documentation.

Emergency backup and restore

Use pg_dump and pg_restore when the CNPG operator is unavailable. These tools can also serve as an additional safety measure before any of the procedures above. Backups are stored locally on the machine where the commands run and require a stable connection to the database pods throughout the operation.

  1. Scale the application using the database cluster to zero. See the Shut down services section above.
  2. Run the backup or restore commands below.
  3. Restart application processes (See the Restart services section above).
🚧

Important: The pg_restore commands below include the --clean flag, which drops existing tables before importing. This differs from the CNPG documentation and is necessary because the database is typically not empty after application startup. Use this flag with care.

In each command, replace the -n suffix of the pod name (for example gitea-db-n) with the actual primary pod instance number (for example gitea-db-1).

  1. Determine the primary instance:

    kubectl get cluster -n gitea gitea-db
  2. Backup:

    kubectl exec -n gitea gitea-db-n postgres \
      -- pg_dump -Fc -d gitea > gitea.dump
  3. Restore:

    kubectl exec -i -n gitea gitea-db-n postgres \
      -- pg_restore --no-owner --role=gitea -d gitea --verbose --clean < gitea.dump

Did this page help you?