CDP Private Cloud Data Services(ECS) 1.5.5 설치 중 istiod helm 릴리스가 failed 로 남는다. 설치 스크립트는 helm install istiod ... --wait --timeout 300s 로 5분을 기다리다 타임아웃하므로, helm 실패는 결과이고 원인은 파드 쪽에 있다.
HELM=/opt/cloudera/parcels/ECS-<ver>/installer/install/bin/linux/helm
export KUBECONFIG=/etc/rancher/rke2/rke2.yaml
KUBECTL=/var/lib/rancher/rke2/bin/kubectl
$HELM list -A
$HELM status istiod -n istio-system
$HELM history istiod -n istio-system
$KUBECTL -n istio-system get pods -o wide
$KUBECTL -n istio-system describe pod -l app=istiod | grep -A 25 Events
$KUBECTL -n istio-system get events --sort-by=.lastTimestamp | tail -30
$KUBECTL get crd | grep istio.io
리소스가 부족하면 Status: Pending 과 FailedScheduling ... Insufficient cpu/memory 가 나오거나, 떴다 죽으면서 Reason: OOMKilled, Exit Code: 137 이 찍힌다. Successfully assigned istio-system/istiod-... to <node> 로 스케줄링이 성공했는데 이미지 풀에서 막혔다면 리소스 문제가 아니다.
$KUBECTL top nodes
$KUBECTL describe nodes | grep -A 5 "Allocated resources"
$KUBECTL describe nodes | grep -E "MemoryPressure|DiskPressure|PIDPressure"
이 사례에서는 조치할 때마다 오류가 바뀌었고, 메시지마다 원인이 달랐다.
| 오류 | 원인 |
|---|---|
unable to read CA cert "/etc/docker/certs.d/<registry-host>:5000/ca.crt": no such file or directory |
그 노드에 레지스트리 CA 가 배포되지 않았다 |
x509: certificate signed by unknown authority |
레지스트리를 재설치해 CA 가 바뀌었는데 노드에 옛 CA 가 남아 있다 |
x509: cannot validate certificate for 10.14.51.12 because it doesn't contain any IP SANs |
인증서에 DNS 이름만 있는데 endpoint 를 IP 로 지정했다 |
code = NotFound ... not found |
TLS 는 성공. 레지스트리에 그 이미지가 없다 |
401 UNAUTHORIZED |
인증 정보 불일치 |
embedded registry 가 다른 노드에 있으면 각 노드에 같은 경로로 CA 를 배포해야 한다. 멀티노드 설치에서 이 단계가 누락되는 일이 잦다.
mkdir -p "/etc/docker/certs.d/<registry-host>:5000"
scp root@<registry-host>:"/etc/docker/certs.d/<registry-host>:5000/ca.crt" \
"/etc/docker/certs.d/<registry-host>:5000/ca.crt"
chmod 640 "/etc/docker/certs.d/<registry-host>:5000/ca.crt"
crictl pull registry.ecs.internal/cloudera_thirdparty/hardened/gloo-mesh/istio-<hash>/pilot:<tag>
echo "exit: $?"
$KUBECTL -n istio-system delete pod -l app=istiod
RKE2 는 /var/lib/rancher/rke2/agent/etc/containerd/certs.d/<registry-host>/hosts.toml 경로도 참조하므로 함께 확인한다.
echo | openssl s_client -connect <registry-host>:<port> 2>/dev/null \
| openssl x509 -noout -text | grep -A 2 "Subject Alternative Name"
echo | openssl s_client -connect <registry-host>:<port> 2>/dev/null \
| openssl x509 -noout -subject -issuer -dates
cat /etc/rancher/rke2/registries.yaml
SAN 에 DNS: 만 있으면 IP 로 접속할 수 없다. registries.yaml 의 endpoint 를 인증서에 있는 FQDN 으로 바꾸고 /etc/hosts 에 매핑을 추가한 뒤 RKE2 를 재시작한다.
mirrors:
registry.ecs.internal:
endpoint:
- "https://<cert-fqdn>:<port>"
configs:
"<cert-fqdn>:<port>":
tls:
ca_file: /etc/docker/certs.d/<cert-fqdn>:<port>/ca.crt
systemctl restart rke2-server # 서버 노드
systemctl restart rke2-agent # 에이전트 노드
ECS 의 helm 차트는 이미지를 항상 registry.ecs.internal/... 로 참조한다. 이 이름이 실제 어느 레지스트리로 가는지는 RKE2/containerd 설정이 정한다. Embedded 모드는 로컬 docker registry 로, External 모드는 지정한 외부 레지스트리로 연결된다. 따라서 Harbor 에 이미지를 다 넣어 두어도 파드가 여전히 embedded 를 보고 있으면 NotFound 가 난다.
$KUBECTL -n istio-system get pod -l app=istiod -o jsonpath='{.items[*].spec.containers[*].image}'; echo
getent hosts registry.ecs.internal
grep registry.ecs.internal /etc/hosts
cat /etc/rancher/rke2/registries.yaml
find /var/lib/rancher/rke2/agent/etc/containerd/certs.d -name hosts.toml -exec echo "=== {} ===" \; -exec cat {} \;
레지스트리에 이미지가 실제로 있는지는 카탈로그나 스토리지 경로로 확인한다.
curl -sk "https://<registry>/v2/_catalog?n=5000" | tr ',' '\n' | grep -iE "istio|pilot|gloo"
curl -sk "https://<registry>/v2/cloudera_thirdparty/hardened/gloo-mesh/istio-<hash>/pilot/tags/list"
# embedded registry 스토리지 직접 확인
ls /data/docker/docker-registry/docker/registry/v2/repositories/
find /data/docker/docker-registry/docker/registry/v2/repositories/ -maxdepth 6 -type d \
\( -iname "*istio*" -o -iname "*pilot*" -o -iname "*gloo*" \)
이 사례에서는 스토리지에 cloudera 만 있고 cloudera_thirdparty 네임스페이스가 통째로 없었다.
TLS 가 활성화된 커스텀 레지스트리만 지원된다. 순서가 중요하다. CM 마법사의 Configure Docker Repository 단계에서 레지스트리 주소와 인증서를 넣고 copy-docker 스크립트를 생성해 먼저 끝까지 실행해 Harbor 를 채운 뒤에야 설치를 이어간다(이미지 복사는 4~5시간 걸린다). copy-docker 는 pull 이 실패해도 멈추지 않고 넘어가므로 일부만 복사된 채 설치가 진행되는 사고가 난다.
docker login container.repository.cloudera.com -u <cloudera-user> # 소스(인터넷 설치 시)
docker login <harbor>:58443 -u <push-user> # 타깃
bash -x copy-docker.txt 2>&1 | tee /tmp/copy-docker.log
grep -iE "error|denied|unauthorized|not found|fail|x509|manifest unknown" /tmp/copy-docker.log
curl -sk -u <user>:${PASSWORD} "https://<harbor>:58443/v2/_catalog?n=5000" | tr ',' '\n' | grep -iE "istio|pilot|gloo"
Cloudera Manager → ECS 서비스 → Configuration 에서 exter 로 검색해 다음을 채운다.
| 항목 | 값 |
|---|---|
| Enable External Container Registry | 체크 |
| External Container Registry User / Password | Harbor push 권한 계정 |
| External Container Registry | <harbor-host>:58443 — 호스트:포트만, https:// 와 슬래시 없이 |
| Base path | copy-docker 로 이미지를 넣은 Harbor 프로젝트명. 비워 두면 NotFound 가 난다 |
| External Docker Registry Certificate (PEM) | Harbor 인증서 PEM 전체. 비워 두면 TLS 실패 |
Base path 가 비면 ECS 는 <registry>/cloudera_thirdparty/... 를 찾는데 Harbor 는 반드시 프로젝트 아래에 저장하므로 실제 위치인 <registry>/<project>/cloudera_thirdparty/... 와 어긋난다. 이것이 이 사례에서 NotFound 가 계속 난 원인이었다.
curl -sk -u admin:${PASSWORD} "https://<harbor>:58443/api/v2.0/projects" | tr ',' '\n' | grep -i '"name"'
echo | openssl s_client -connect <harbor>:58443 -showcerts 2>/dev/null | openssl x509 -outform PEM
Harbor CA 는 모든 ECS 노드의 /etc/docker/certs.d/<harbor-host>:<port>/ca.crt 와 RKE2 containerd certs.d 양쪽에 배포한다. Harbor 를 지웠다 다시 설치하면 CA 가 바뀌므로 노드의 옛 CA 를 지우고 다시 배포해야 한다.
openssl x509 -in "/etc/docker/certs.d/<host>:<port>/ca.crt" -noout -fingerprint -dates
echo | openssl s_client -connect <host>:<port> 2>/dev/null | openssl x509 -noout -fingerprint -dates
# 두 fingerprint 가 다르면 재배포
ImagePullBackOff 는 파드가 생긴 뒤 이미지를 못 받는 것이고, 파드가 아예 안 생기면 Deployment → ReplicaSet → Pod 체인의 위쪽 문제다.
$KUBECTL -n istio-system get all
$KUBECTL -n istio-system describe rs | tail -20
$KUBECTL get validatingwebhookconfigurations | grep -i istio
$KUBECTL get mutatingwebhookconfigurations | grep -i istio
재설치를 반복하면 죽은 istiod 를 가리키는 웹훅이 남아 파드 생성을 막는 경우가 있다(failed calling webhook). 잔재 웹훅을 지우고 재설치한다.
Canal CNI 는 노드 간 오버레이 통신에 VXLAN UDP 를 쓴다. 이 포트가 막히면 파드 네트워크가 노드를 넘지 못해 istiod readiness 나 webhook 통신이 실패한다.
| 포트 | 프로토콜 | 용도 |
|---|---|---|
| 4789 | UDP | Canal VXLAN |
| 8472 | UDP | Flannel VXLAN(구버전) |
| 51820 | UDP | WireGuard(암호화 사용 시) |
| 6443 / 9345 / 10250 / 2379-2380 | TCP | API server / RKE2 노드 등록 / kubelet / etcd |
for port in 4789 8472; do
iptables -I INPUT -p udp --dport $port -j ACCEPT
iptables -I FORWARD -p udp --dport $port -j ACCEPT
done
iptables-save > /etc/sysconfig/iptables
# firewalld 라면
firewall-cmd --permanent --add-port=4789/udp && firewall-cmd --reload
ip -d link show type vxlan
$KUBECTL get pods -n kube-system | grep canal
Check Prerequisites 단계에서 자주 걸리는 두 가지는 이름 해석과 스토리지다. CM URL 에 쓴 호스트명, 관리 콘솔 주소, localhost 가 모두 풀려야 하고 /etc/nsswitch.conf 의 hosts: 에 files 가 있어야 한다. /var/lib 와 /data/docker 는 각각 300 GiB 이상이어야 하며 /var/lib/longhorn 은 심볼릭 링크면 안 된다.
hostname -f
getent hosts localhost
grep "^hosts" /etc/nsswitch.conf
df -h /var/lib /data/docker
ls -ld /var/lib/longhorn
istiod 가 실패 목록에서 빠진 것은 확인했으나 Running 상태까지 명시적으로 확인하지 못했고, 같은 시점에 GPU 가 없는 노드에서 nvidia-device-plugin 이 CrashLoopBackOff 인 문제는 별도 이슈로 남았다.