kuberenets의 POD 안에서 spark-submit을 실행해서 minio 안의 데이터를 출력
참고: https://www.youtube.com/watch?v=ZzFdYm_DqEM&t=307s
spark operator가 결과를 확인하기 가장 좋은 방법이라고 말함
기본 이미지로 실행하면 hadoop2.7과 dup나서 에러가 발생한다고 함
- 구글 오픈소스로, 쿠버네티스에 스파크를 실행하는데 쉬움
- https://github.com/GoogleCloudPlatform/spark-on-k8s-operator/blob/master/examples/spark-pi.yaml
$ cd spark_home
$ helm install sparkoperator spark-operator/spark-operator
NAME: sparkoperator
LAST DEPLOYED: Tue Sep 14 15:54:30 2021
NAMESPACE: default
STATUS: deployed
REVISION: 1
TEST SUITE: None
$ wget https://mirror.navercorp.com/apache/spark/spark-3.1.2/spark-3.1.2-bin-hadoop3.2.tgz
./bin/docker-image-tool.sh -r <repo> -t my-tag -p ./kubernetes/dockerfiles/spark/bindings/python/Dockerfile build
./bin/docker-image-tool.sh -r spark-base -t 1.0.0 -p ./kubernetes/dockerfiles/spark/bindings/python/Dockerfile build
-r registry 이름으로 이미지가 만들어짐 ex) [spark-base]/spark-py:1.0.0
EC2에서 빌드해서 테스트하니 됨
[ec2-user@ip-172-31-28-78 spark-3.1.2-bin-hadoop3.2]$ ./bin/docker-image-tool.sh -r encore_spark_base -t 1.0.0 -p ./kubernetes/dockerfilngs/python/Dockerfile build
Sending build context to Docker daemon 255.5MB
Step 1/18 : ARG java_image_tag=11-jre-slim
Step 2/18 : FROM openjdk:${java_image_tag}
---> e4beed9b17a3
Step 3/18 : ARG spark_uid=185
---> Using cache
---> 2bfe6aacdcd3
Step 4/18 : RUN set -ex && sed -i 's/http:\/\/deb.\(.*\)/https:\/\/deb.\1/g' /etc/apt/sources.list && apt-get update && ln -&& apt install -y bash tini libc6 libpam-modules krb5-user libnss3 procps && mkdir -p /opt/spark && mkdir -p /opt/spark/examdir -p /opt/spark/work-dir && touch /opt/spark/RELEASE && rm /bin/sh && ln -sv /bin/bash /bin/sh && echo "auth required se_uid" >> /etc/pam.d/su && chgrp root /etc/passwd && chmod ug+rw /etc/passwd && rm -rf /var/cache/apt/*
---> Using cache
---> e856ef214626
Step 5/18 : COPY jars /opt/spark/jars
---> 54d8127376d1
Step 6/18 : COPY bin /opt/spark/bin
---> a958a432a5bb
Step 7/18 : COPY sbin /opt/spark/sbin
$ docker ps
deet1107/spark-py 1.0.0 86aed3057138 4 hours ago 880MB
$ docker push deet1107/sparkpy:1.0.0
- m1에서 빌드하면 arm64여서 오류가 남
- 내부망에서는 네트워크 관련 에러가 나올 수도 있음
$ ./bin/docker-image-tool.sh -r spark-base -t 1.0.0 -p ./kubernetes/dockerfiles/spark/bindings/python/Dockerfile build
Sending build context to Docker daemon 356.8MB
Step 1/18 : ARG java_image_tag=11-jre-slim
Step 2/18 : FROM openjdk:${java_image_tag}
---> e4beed9b17a3
Step 3/18 : ARG spark_uid=185
---> Using cache
---> b098f4c33b7f
Step 4/18 : RUN set -ex && sed -i 's/http:\/\/deb.\(.*\)/https:\/\/deb.\1/g' /etc/apt/sources.list && apt-get update && ln -s /lib /lib64 && apt install -y bash tini libc6 libpam-modules krb5-user libnss3 procps && mkdir -p /opt/spark && mkdir -p /opt/spark/examples && mkdir -p /opt/spark/work-dir && to uch /opt/spark/RELEASE && rm /bin/sh && ln -sv /bin/bash /bin/sh && echo "auth required pam_wheel.so use_uid" >> /etc/pam.d/su && chgrp root /etc/passwd && chmod ug+rw /etc/passwd && rm -rf /var/cache/apt/*
---> Running in b7962b7dc9af
+ sed -i s/http:\/\/deb.\(.*\)/https:\/\/deb.\1/g /etc/apt/sources.list
+ apt-get update
Err:1 https://deb.debian.org/debian bullseye InRelease
Could not handshake: Error in the pull function. [IP: 146.75.50.132 443]
Err:2 http://security.debian.org/debian-security bullseye-security InRelease
Connection failed [IP: 151.101.2.132 80]
Err:3 https://deb.debian.org/debian bullseye-updates InRelease
Could not handshake: Error in the pull function. [IP: 146.75.50.132 443]
Reading package lists...
W: Failed to fetch https://deb.debian.org/debian/dists/bullseye/InRelease Could not handshake: Error in the pull function. [IP: 146.75.50.132 443]
W: Failed to fetch http://security.debian.org/debian-security/dists/bullseye-security/InRelease Connection failed [IP: 151.101.2.132 80]
W: Failed to fetch https://deb.debian.org/debian/dists/bullseye-updates/InRelease Could not handshake: Error in the pull function. [IP: 146.75.50.132 443]
W: Some index files failed to download. They have been ignored, or old ones used instead.
+ ln -s /lib /lib64
+ apt install -y bash tini libc6 libpam-modules krb5-user libnss3 procps
WARNING: apt does not have a stable CLI interface. Use with caution in scripts.
$ docker login -u deet1107
$ docker pull deet1107/spark-py:1.0.0
$ docker logout
$ docker login vharbor.encore.encore -u tedkim
$ docker tag deet1107/spark-py:1.0.0 vharbor.encore.encore/library/spark-py:1.0.0
$ Dockerfile
FROM encore.encore/project-private/spark-py-base:1.0.0
USER root
WORKDIR /app
COPY main.py .
$ docker build -t vharbor.encore.encore/test-kube/test-k8s:v1.0.0 .
$ docker push vharbor.encore.encore/test-kube/test-k8s:v1.0.0
$ vi k8s_ext.yml
kind: Service
apiVersion: v1
metadata:
name: vminio01
spec:
type: ExternalName
externalName: vminio01.encore.encore
$ vi k8s.yml
#Copyright 2018 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#
# Support for Python is experimental, and requires building SNAPSHOT image of Apache Spark,
# with `imagePullPolicy` set to Always
apiVersion: "sparkoperator.k8s.io/v1beta2"
kind: SparkApplication
metadata:
name: sparkjobextbminio01 #pyspark-pi
namespace: default
spec:
type: Python
pythonVersion: "3"
mode: cluster
imagePullSecrets:
- regcred
#image: "vharbor.encore.encore/test-kube/test-spark-k8s-vminio01:v1.2.0" #spark-py:v3.1.1
image: "encore.encore/project-private/spark-py-hello:1.0.0" #spark-py:v3.1.1
imagePullPolicy: Always
mainApplicationFile: local:///app/main.py #local:///opt/spark/examples/src/main/python/pi.py
sparkVersion: "3.1.2" #"3.1.1"
restartPolicy:
type: Never
driver:
cores: 1
coreLimit: "1000m"
memory: "512m"
labels:
version: 3.1.2
#serviceAccount: sparkoperator
#serviceAccount: sparkoperator-spark
serviceAccount: sparkoperator-spark-operator
executor:
cores: 1
instances: 1
memory: "512m"
labels:
version: 3.1.2
deps:
jars:
# - https://repo1.maven.org/maven2/org/apache/hadoop/hadoop-aws/3.2.0/hadoop-aws-3.2.0.jar
# - https://repo1.maven.org/maven2/com/amazonaws/aws-java-sdk-bundle/1.11.375/aws-java-sdk-bundle-1.11.375.jar
- http://10.240.29.37/files/sparkjob/hadoop-aws-3.2.0.jar
- http://10.240.29.37/files/sparkjob/aws-java-sdk-bundle-1.11.375.jar
$ kubectl apply -f k8s.yml
# spark-job 확인
$ kubectl get pods
# 10초뒤 spark-job의 로그 확인
$ kubectl logs spark-job
(결과)
+-----------+
|avg(amount)|
+-----------+
| 2800.0|
+-----------+