0/12 nodes are available:
1 Insufficient nvidia.com/gpu
3 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: true}
8 node(s) didn't match Pod's node affinity/selector
노드 12 대 전부가 탈락한 이유가 항목별로 나뉘어 있다. 숫자가 가장 큰 항목이 핵심이다. 위 사례에서는 8 대가 affinity 불일치이므로, GPU 워커 노드에 Pod 이 요구하는 레이블이 붙어 있지 않을 가능성이 높다.
| 항목 | 의미 |
|---|---|
Insufficient nvidia.com/gpu |
GPU 를 요청했으나 남은 GPU 가 없음 |
| untolerated taint | control-plane 등 taint 가 걸린 노드. toleration 없으면 배치 불가 |
| didn't match node affinity/selector | nodeSelector · nodeAffinity 조건에 맞는 레이블이 없음 |
# GPU 할당 현황
kubectl describe nodes | grep -A 5 "nvidia.com/gpu"
kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\\.com/gpu
# Pod 이 요구하는 조건
kubectl get pod <pod> -n <ns> -o yaml | grep -A 20 -E "affinity|nodeSelector|tolerations"
# 노드 레이블 · taint
kubectl get nodes --show-labels
kubectl describe node <gpu-node> | grep -A 5 Taints
kubectl describe node <gpu-node> | grep -A 10 "Allocated resources"
kubectl label node <node> node-type=gpuresources.limits."nvidia.com/gpu" 를 줄인다. GPU 는 공유되지 않으므로 요청 수만큼 물리 GPU 가 필요하다.kubectl get pods -n gpu-operator 에서 device plugin 이 Running 이어야 nvidia.com/gpu 가 allocatable 에 잡힌다.