Viya 4 를 LTS 사이로 업그레이드하면서 내장 PostgreSQL 의 메이저 버전이 바뀌는 릴리스가 있다. 이때 Crunchy PostgreSQL Operator(PGO) 가 PGUpgrade 커스텀 리소스를 만들어 pg_upgrade 를 돌린다. 업그레이드 Job 이 한 번 실패하면 파드가 올라오지 못하고 계속 같은 자리에서 멈춘다.
kubectl -n <namespace> get pod -l postgres-operator.crunchydata.com/cluster
kubectl -n <namespace> describe pod -l postgres-operator.crunchydata.com/cluster \
| egrep -i 'Reason|Error|Warning|Mount|Backoff'
kubectl -n <namespace> get postgresclusters.postgres-operator.crunchydata.com -o wide
kubectl -n <namespace> logs deploy/<crunchy-postgres-operator> -c operator --tail=200
실패 원인은 대개 이 중 하나다.
| 원인 | 신호 |
|---|---|
| 이미지 pull 실패 | ImagePullBackOff. 서비스 어카운트에 imagePullSecrets 가 안 붙었거나 미러에 태그가 없다 |
| PVC 용량 부족 | 새 메이저 버전 데이터 디렉터리를 새로 만들어야 하므로 기존 사용량의 두 배 가까이 필요하다 |
| 업그레이드 잔재 | 실패한 이전 시도가 남긴 디렉터리 때문에 다음 시도가 곧바로 실패한다 |
| CRD · 오퍼레이터 버전 불일치 | 업그레이드 자산의 CRD 가 적용되지 않았다 |
SAS 가 제공하는 업그레이드 번들에 postgres-upgrade.sh 와 update-target-manifests.yaml 이 들어 있다. 후자는 직접 작성하는 파일이 아니라 번들에서 그대로 가져와야 한다. 스크립트가 라벨로 골라 적용하기 때문에 라벨을 임의로 바꾸면 사전 점검에서 멈춘다.
find . -iname 'update-target-manifests.yaml' -o -ipath '*crunchy*pgupgrade*'
파일 안에는 PGO CRD, 오퍼레이터 구성(ServiceAccount · Role · RoleBinding · Deployment), 이미지 pull 시크릿, PGUpgrade CR, 그리고 새 버전으로 갱신된 PostgresCluster CR 이 들어 있다.
yq '. | select(.kind=="PGUpgrade") | {name:.metadata.name, spec:.spec}' update-target-manifests.yaml
yq '. | select(.kind=="PostgresCluster") | {ver:.spec.postgresVersion, image:.spec.image}' update-target-manifests.yaml
업그레이드가 한 번 실패하면 데이터 볼륨에 새 버전용 디렉터리가 남는다. PGO 의 업그레이드 Job 은 대상 디렉터리가 없다는 전제로 시작하므로, 잔재가 있으면 1 단계에서 즉시 실패한다.
볼륨 내용은 대개 이런 모양이다.
pg12 pg12_wal pg16_stale_<epoch> pg16_wal pgbackrest
| 항목 | 처리 |
|---|---|
pg12, pg12_wal, pgbackrest |
지우지 않는다. 원본 데이터와 백업이다 |
pg16_stale_* |
실패 흔적. 공간이 모자랄 때만 삭제 |
pg16_wal |
비우거나 이름을 바꾼다. 남아 있으면 init 컨테이너가 WAL 심볼릭 링크를 만들지 못하고 죽는다 |
PVC 를 직접 만지기보다 헬퍼 파드로 처리하는 편이 권한 문제를 피한다.
apiVersion: v1
kind: Pod
metadata:
name: pgdata-fix
spec:
restartPolicy: Never
containers:
- name: sh
image: busybox:1.36
command: ["sh","-c"]
args:
- |
ls -la /pgdata
[ -d /pgdata/pg16 ] && mv /pgdata/pg16 /pgdata/pg16_stale_$(date +%s)
if [ -d /pgdata/pg16_wal ]; then
if [ -z "$(ls -A /pgdata/pg16_wal)" ]; then
rmdir /pgdata/pg16_wal
else
mv /pgdata/pg16_wal /pgdata/pg16_wal_stale_$(date +%s)
fi
fi
ls -la /pgdata
volumeMounts:
- name: pgdata
mountPath: /pgdata
volumes:
- name: pgdata
persistentVolumeClaim:
claimName: <postgres-pgdata-pvc>
정리하는 동안 Job 이 다시 만들어지지 않게 업그레이드 허용 어노테이션을 먼저 떼고 기존 Job 을 지운다.
kubectl -n <namespace> annotate postgrescluster <cluster> \
postgres-operator.crunchydata.com/allow-upgrade-
kubectl -n <namespace> delete job -l postgres-operator.crunchydata.com/pgupgrade=<upgrade-name>
kubectl -n <namespace> annotate postgrescluster <cluster> \
postgres-operator.crunchydata.com/allow-upgrade=<upgrade-name>
kubectl -n <namespace> get jobs,pods | grep -E 'pgupgrade|upgrade-pgdata'
kubectl -n <namespace> logs -f <upgrade-pod> -c pgupgrade
정상이면 Performing Consistency Checks 로 시작해 완료 메시지로 끝난다.
repo-host 파드가 Ready 인지. pgBackRest 저장소가 준비되지 않으면 업그레이드가 진행되지 않는다.kubectl -n <namespace> get pod -l postgres-operator.crunchydata.com/data=pgbackrest -o wide
kubectl -n <namespace> describe pvc <postgres-pgdata-pvc>