Restoring platform databases
If you need to restore your database (for instance, if it becomes corrupted), you can do so using the backups within App Platform. Before proceeding, familiarize yourself with the CNPG documentation on PostgreSQL database recovery. The steps here are specific to App Platform installed on LKE.
Recovery options
- Regular recovery: Restore a database after corruption or when it's in an unhealthy state.
- Point-in-time recovery: Restore a database from a specific point in time before undesired changes.
- Emergency backup and restore: Use
pg_dumpandpg_restorewhen the CNPG operator is unavailable or before running the procedures above.
Regular recovery
Use this procedure when a database is in an unhealthy state, for example due to volume filesystem corruption. To revert a database when there have been undesired changes, use point-in-time recovery instead.
- Note the name of the
Backupresource you intend to restore from, if available. - Prepare access to object storage for recovery.
- Prepare direct access to ArgoCD in case single sign-on becomes unavailable during recovery.
- Pause automated sync operations for
apl-operatorand ArgoCD. - Shut down any service that accesses the database.
- Update the backup configuration and edit the
Applicationmanifest of the database accordingly. - Delete the database
Clusterresource and reactivate ArgoCD sync. - Restart applications using the database and
apl-operator.
List backup resources
List the backups that are available for application's database (Gitea, Harbor, or Keycloak). Only use completed backups for recovery. Backup name timestamps are in UTC.
kubectl get backup -n giteaPrepare access to object storage
The database recovery requires an ObjectStore resource pointing to the backup location. When backups are enabled, this resource already exists in the cluster and can be copied and modified for recovery.
-
Dump the existing resource to a file:
kubectl get objectstores.barmancloud.cnpg.io -n gitea gitea-db -oyaml > recovery.yaml -
Edit the file and make the following changes:
- Remove the
statussection. - In
metadata, keep onlynamespaceand changename, for example tokeycloak-db-recovery. - Leave the
specsection unchanged.
It should look like the following (here provided for Keycloak). If the resource is not available, you can also use this as a template:
apiVersion: barmancloud.cnpg.io/v1 kind: ObjectStore metadata: # Replace name. Note this needs to correspond with the "source" value later. name: keycloak-db-recovery # Replace namespace with "gitea" or "harbor" as applicable. namespace: keycloak spec: configuration: data: compression: gzip # Replace the bucket name, "keycloak" with the application name # (pathSuffix in previous backup configuration) destinationPath: s3://bucket-name/keycloak # Replace "nl-ams-1" with your storage region endpointURL: https://nl-ams-1.linodeobjects.com s3Credentials: accessKeyId: key: S3_STORAGE_ACCOUNT name: linode-creds secretAccessKey: key: S3_STORAGE_KEY name: linode-creds wal: compression: gzip instanceSidecarConfiguration: env: - name: AWS_REQUEST_CHECKSUM_CALCULATION value: WHEN_REQUIRED - name: AWS_RESPONSE_CHECKSUM_VALIDATION value: WHEN_REQUIRED logLevel: info resources: limits: cpu: 300m memory: 256Mi requests: cpu: 50m memory: 48Mi retentionPolicyIntervalSeconds: 1800 retentionPolicy: 7d - Remove the
-
Store this resource in the cluster by either running
kubectl apply -f recovery.yaml(replace file name as needed) or committing it to the values repository atenv/manifests/<namespace>/<resource-name>.yaml. In the latter case, ArgoCD will synchronize thisObjectStoreresource to the cluster.
Access to ArgoCD
During recovery of platform-critical services, such as Keycloak, single sign-on becomes unavailable. Retaining direct access to ArgoCD lets you synchronize changes and monitor status throughout the process. This step is not required for Harbor.
-
Retrieve the ArgoCD admin password:
kubectl get secrets -n argocd argocd-initial-admin-secret -ojson | jq -r '.data.password' | base64 -d -
Configure port forwarding to access ArgoCD at
http://localhost:8080(change the port if there is a conflict):kubectl port-forward -n argocd svc/argocd-server 8080:80
Deactivate sync and apl-operator
While making temporary manual changes, disable automated reconciliation to prevent conflicting modifications.
-
First, scale down
apl-operatorand remove its ArgoCD auto-sync policy:kubectl patch applications.argoproj.io -n argocd apl-operator-apl-operator --patch '[{"op": "remove", "path": "/spec/syncPolicy/automated"}]' --type=json kubectl scale --replicas=0 -n apl-operator deployment apl-operator -
Then, disable auto-sync for the database ArgoCD application.
kubectl patch applications.argoproj.io -n argocd gitea-gitea-otomi-db --patch '[{"op": "remove", "path": "/spec/syncPolicy/automated"}]' --type=json
Shut down services
Shut down the application before backing up or restoring the database to prevent write operations that could cause inconsistencies.
Commands for temporarily disabling Gitea:
## Disable ArgoCD auto-sync during the changes
kubectl patch applications.argoproj.io -n argocd gitea-gitea --patch '[{"op": "remove", "path": "/spec/syncPolicy/automated"}]' --type=json
## Scale Gitea deployment to zero
kubectl patch deploy -n gitea gitea --patch '[{"op": "replace", "path": "/spec/replicas", "value": 0}]' --type=json
## Verify that pods are shut down
kubectl get deploy -n gitea gitea # Should show READY 0/0Adjustments to the backup configuration
After recovery, CNPG creates new backups. To prevent accidental mixing or overwriting of backups, CNPG does not allow the new backup destination and the recovery source to share the same location. The path suffix (the directory inside the object storage bucket) must be changed.
-
In
env/settings.yaml, setpathSuffixunderplatformBackups.database.<app>(where<app>isgitea,harbor, orkeycloak) to a new value that does not already exist in object storage, for example<app>-1:# ... platformBackups: database: gitea: # ... pathSuffix: gitea-1 # ... -
Because ArgoCD sync is currently disabled, apply this change directly to the application manifest as well. The most convenient way to do this is through the ArgoCD UI. Note that
helm.valuesis a multiline string.In the
helm.valuesstructure, updatebackup.linode.destinationPathto match the new path suffix, and setclusterSpec.bootstrap.recoveryandclusterSpec.externalClustersas shown below.Important: Save these changes but do not synchronize yet.
Example for Gitea (application
gitea-gitea-otomi-db):# ... source: # ... helm: values: > backup: enabled: true linode: # Modify suffix according to new pathSuffix in values destinationPath: s3://bucket-name/gitea-1 # ... clusterSpec: bootstrap: recovery: backup: # Name of Backup resource from step 1 # If not set, will restore to latest backup available in bucket name: ... database: gitea owner: gitea # To correspond with name in externalClusters source: recovery-obj externalClusters: - name: recovery-obj plugin: enabled: true isWALArchiver: false name: barman-cloud.cloudnative-pg.io parameters: # Name of ObjectStore resource from step 2 barmanObjectName: gitea-db-recovery serverName: gitea-db # ...
Restore the database
After shutting down applications and preparing the configuration changes, delete the database Cluster. ArgoCD then re-creates the cluster and restores it from the backup.
The following kubectl commands are equivalent to the same operations in the ArgoCD UI.
Commands for Gitea:
## Delete the cluster
kubectl delete cluster -n gitea gitea-db
## Re-enable ArgoCD auto-sync
kubectl patch application -n argocd gitea-gitea-otomi-db --patch '[{"op": "add", "path": "/spec/syncPolicy/automated", "value": {"prune": true, "allowEmpty": true}}]' --type=jsonRestart services
Once maintenance is complete, restart apl-operator:
kubectl scale deploy -n apl-operator apl-operator --replicas=1apl-operator eventually reverts the manual edits to Application manifests without disrupting the database. It also re-enables automated synchronization and restarts services. To bring services back online faster, use the following commands.
Commands for Gitea:
## Re-enable ArgoCD auto-sync, which should also change the Gitea statefulset to scale up
kubectl patch applications.argoproj.io -n argocd gitea-gitea --patch '[{"op": "add", "path": "/spec/syncPolicy/automated", "value": {"prune": true, "allowEmpty": true}}]' --type=json
## Optional: scale up, for not having to wait for re-sync of ArgoCD
kubectl patch statefulset -n gitea gitea --patch '[{"op": "replace", "path": "/spec/replicas", "value": 1}]' --type=jsonPoint-in-time recovery
To restore a database to a specific point in time, edit the application manifest. The most convenient way to do this is through the ArgoCD UI. Note that helm.values is a multiline string.
Add a recoveryTarget to the recovery section following the CloudnativePG docs. The example below restores Gitea to the state just before a change made after 2023-03-06 08:00:39 CET.
# ...
source:
# ...
helm:
values: >
# ...
clusterSpec:
bootstrap:
recovery:
database: gitea
owner: gitea
source: recovery-obj # To correspond with name in externalClusters
recoveryTarget:
# Time base target for the recovery
targetTime: "2023-03-06 08:00:39+01"
# ...You can also specify a named backup resource in the backup.name field (step 1). The backup must be from before the chosen recovery target timestamp, noting that backup names use UTC timestamps.
The targetTime format is not standard ISO 8601. Instead, the date and time are separated by a space and the timezone is written as an explicit offset, such as +00 (UTC) or +01 (CET without DST). For all valid formats, see the PostgreSQL documentation.
Emergency backup and restore
Use pg_dump and pg_restore when the CNPG operator is unavailable. These tools can also serve as an additional safety measure before any of the procedures above. Backups are stored locally on the machine where the commands run and require a stable connection to the database pods throughout the operation.
- Scale the application using the database cluster to zero. See the Shut down services section above.
- Run the backup or restore commands below.
- Restart application processes (See the Restart services section above).
Important: The
pg_restorecommands below include the--cleanflag, which drops existing tables before importing. This differs from the CNPG documentation and is necessary because the database is typically not empty after application startup. Use this flag with care.
In each command, replace the -n suffix of the pod name (for example gitea-db-n) with the actual primary pod instance number (for example gitea-db-1).
-
Determine the primary instance:
kubectl get cluster -n gitea gitea-db -
Backup:
kubectl exec -n gitea gitea-db-n postgres \ -- pg_dump -Fc -d gitea > gitea.dump -
Restore:
kubectl exec -i -n gitea gitea-db-n postgres \ -- pg_restore --no-owner --role=gitea -d gitea --verbose --clean < gitea.dump
Updated 14 days ago
