CAI 세션과 CDE 작업이 Pending 또는 ApplicationRejected에 머무는 증상이 나타난다.
YuniKorn 내부 상태가 실제 클러스터와 어긋나 앱을 수락한 뒤에도 allocation이 진행되지 않았다. OutOfcpu·OutOfmemory가 함께 나타날 수 있으며, 이 기록에서 제시한 조치는 우회 방법이다.
yunikorn-scheduler Pod를 재시작해 내부 캐시와 클러스터 상태를 다시 동기화한다.
명령 예시는 문서 작성 시 보강한 것이며 대상 시스템에서 실행 검증하지 않았다.
kubectl get deployment -A
kubectl -n '<namespace>' get deployment 'yunikorn-scheduler' -o yaml > resource-before.yaml
kubectl -n '<namespace>' rollout restart deployment/'yunikorn-scheduler'
kubectl -n '<namespace>' rollout status deployment/'yunikorn-scheduler' --timeout=300s
kubectl -n '<namespace>' get pods -o wide
<namespace>는 첫 명령에서 확인한 값으로 바꾼다. 생성된 이름에 접두사·접미사가 있으면 yunikorn-scheduler도 조회 결과로 바꾼다.실제 리소스가 StatefulSet 등 다른 형태이면 조회 결과의 controller를 사용한다. 이미 실행 중인 작업 영향과 scheduler 로그를 함께 확인한다.
새 세션/작업이 Pending에서 Running으로 전환되는지 확인한다.