| 구성 요소 | 역할 |
|---|---|
| Driver | main() 을 실행하고 SparkSession 을 만든다. DAG 를 만들고 태스크를 나눠 보내며 결과를 모은다 |
| Executor | 태스크를 실제로 실행하고 데이터를 캐시한다. 애플리케이션이 끝날 때까지 살아 있다 |
| Cluster Manager | 자원을 배분한다. YARN · Kubernetes · Standalone · Mesos |
드라이버가 죽으면 애플리케이션 전체가 끝난다. 익스큐터가 죽으면 그 태스크만 다른 익스큐터에서 다시 실행된다. 이 차이가 배포 모드 선택의 근거가 된다.
| 항목 | client | cluster |
|---|---|---|
| 드라이버 위치 | spark-submit 을 실행한 장비 |
클러스터 안(YARN 이면 AM 컨테이너) |
| 제출 단말 종료 시 | 애플리케이션도 끝난다 | 계속 실행된다 |
| 로그 확인 | 콘솔에 바로 나온다 | yarn logs 로 받는다 |
| 적합한 용도 | 대화형 셸, 개발, Thrift Server | 배치 잡, 운영 스케줄러 |
spark-submit --master yarn --deploy-mode cluster \
--class com.example.MyJob \
--driver-memory 4g --executor-memory 8g --num-executors 10 \
/path/my-job.jar arg1 arg2
spark-submit --master yarn --deploy-mode client \
--class com.example.MyJob /path/my-job.jar
spark-shell 과 pyspark 는 항상 client 모드다. 드라이버가 로컬에 있어야 대화형 입력을 받을 수 있기 때문이다.
익스큐터 하나에 코어를 너무 많이 주면 GC 가 길어지고, 너무 적게 주면 컨테이너 수가 늘어 오버헤드가 커진다. 실무에서는 익스큐터당 코어 4~5, 메모리 8~32GB 범위에서 출발한다. YARN 이 컨테이너에 붙이는 오버헤드(spark.executor.memoryOverhead)까지 더한 값이 yarn.scheduler.maximum-allocation-mb 를 넘지 않아야 한다.
spark-submit \
--conf spark.executor.cores=4 \
--conf spark.executor.memory=16g \
--conf spark.executor.memoryOverhead=2g \
--conf spark.sql.shuffle.partitions=400 \
...
동적 할당을 쓰면 부하에 따라 익스큐터 수가 조정된다. 셔플 서비스 설정이 함께 필요하다.
--conf spark.dynamicAllocation.enabled=true \
--conf spark.shuffle.service.enabled=true \
--conf spark.dynamicAllocation.minExecutors=2 \
--conf spark.dynamicAllocation.maxExecutors=50
collect() 나 toPandas() 는 모든 결과를 드라이버 메모리로 가져온다. 큰 결과에서는 드라이버가 OOM 으로 죽는다. 한도를 넘으면 미리 실패하도록 spark.driver.maxResultSize 가 걸려 있다. 이 값을 올리기 전에 정말 전부 모아야 하는지 다시 본다. 대부분은 파일로 쓰거나 집계 후 가져오면 된다.
yarn application -list -appStates RUNNING
yarn logs -applicationId application_1234567890123_0001 | head -200
실행 중에는 Spark UI(드라이버의 4040 포트, YARN 에서는 ApplicationMaster 링크)에서 스테이지별 시간과 셔플량을 본다. 끝난 애플리케이션은 History Server 에서 같은 화면을 볼 수 있다.