CM 7.13.1 / Runtime 7.3.1 기준으로 HDFS Balancer(노드 간), HDFS Disk Balancer(노드 내 디스크 간), Kudu Rebalancer 를 실행하는 방법과 권장 파라미터다. Kudu 리밸런서는 디스크 사용률이 아니라 태블릿 서버 간 replica 개수를 맞추며, HDFS 와 Kudu 가 같은 디스크를 쓰면 Kudu 를 먼저, HDFS 를 나중에 돌린다.
hdfs dfsadmin -report | grep -E "Name|DFS Used%" # 노드별 편차 10% 이상이면 Balancer
df -h /data*/dfs/dn # 노드 내 마운트별 편차면 Disk Balancer
Balancer 역할은 Clusters → HDFS → Instances 에 있어야 Actions → Rebalance 가 보인다. 상시 프로세스가 아니라 health 가 None 인 것이 정상이다. 임계값은 HDFS → Configuration → Scope: Balancer → Rebalancing Threshold(기본 10%) 이다.
dfs.datanode.balance.max.concurrent.moves 는 DataNode 와 Balancer 양쪽 Safety Valve(hdfs-site.xml)에 넣어야 하며 CDP 7.x 기본값은 50 이고 DataNode 쪽은 재시작이 필요하다.
<property>
<name>dfs.datanode.balance.max.concurrent.moves</name>
<value>50</value>
</property>
| 속성 | 기본값 | Background | Fast |
|---|---|---|---|
| dfs.datanode.balance.max.concurrent.moves (DataNode) | 50 | 디스크 수 × 4 | 디스크 수 × 4 |
| dfs.datanode.balance.bandwidthPerSec | 10 MB | 유지 | 10 GB |
| dfs.datanode.balance.max.concurrent.moves (Balancer) | 50 | 디스크 수 | 디스크 수 × 4 |
| dfs.balancer.moverThreads | 1000 | 유지 | 20000 |
| dfs.balancer.max-size-to-move | 10 GB | 1 GB | 100 GB |
| dfs.balancer.getBlocks.min-block-size | 10 MB | 유지 | 100 MB |
sudo -u hdfs hdfs dfsadmin -setBalancerBandwidth 104857600 # 재시작 없이 즉시 반영
sudo -u hdfs hdfs balancer -threshold 10 # Kerberos 면 hdfs 키탭으로 kinit 후
진행 상황은 Running Commands 또는 Instances → Balancer → Log Files → Stderr Log 에서 본다. 언제든 중단해도 안전하며 No block has been moved for 5 iterations 로 끝난다. CM 에는 정기 스케줄 기능이 없으므로 cron 에서 CM API(POST /api/v54/clusters/<cluster>/services/hdfs/commands/hdfsRebalance)를 호출하거나 CLI 를 건다. 블록이 수백만 개 이상이면 block report 가 128MB 를 넘어 잘리므로 ipc.maximum.data.length 를 256M 이상으로 올린다.
UI 전용 버튼이 없다. HDFS Service Advanced Configuration Snippet (Safety Valve) for hdfs-site.xml 에 dfs.disk.balancer.enabled=true 를 넣고 DataNode 를 재시작한 뒤 CLI 로 실행한다.
hdfs diskbalancer -plan <datanode-fqdn>
hdfs diskbalancer -execute /system/diskbalancer/<date>/<datanode>.plan.json
hdfs diskbalancer -query <datanode-fqdn>
관련 값은 dfs.disk.balancer.max.disk.throughputInMBperSec(10), dfs.disk.balancer.max.disk.errors(5), dfs.disk.balancer.plan.threshold.percent(10) 이다.
Clusters → Kudu → Actions → Run Kudu Rebalancer Tool 은 기본 플래그로만 돈다. 플래그를 주려면 CLI 를 쓴다.
kinit -kt /var/run/cloudera-scm-agent/process/$(ls -rt /var/run/cloudera-scm-agent/process | grep KUDU | tail -n1)/kudu.keytab kudu/$(hostname -f)
kudu cluster rebalance master-01:7051,master-02:7051,master-03:7051 --report_only --output_replica_distribution_details
kudu cluster rebalance master-01:7051,master-02:7051,master-03:7051 --max_moves_per_server=10
kudu cluster ksck master-01:7051,master-02:7051,master-03:7051
--tables, --max_run_time_sec, --ignored_tservers, --move_replicas_from_ignored_tservers(decommission), --enable_range_rebalancing 이 자주 쓰인다. 모든 tserver 가 살아 있어야 하고 하나라도 내려가면 도구가 종료되며, 재실행하면 이어서 진행한다. 신규 tserver 는 새 태블릿 생성 때만 쓰이므로 증설 직후 반드시 돌린다. Kudu 에는 노드 내 디스크 밸런서가 없고 데이터 디렉터리 제거도 불가하므로 디스크 치우침은 서비스 중지 후 디렉터리를 통째로 옮기는 방식뿐이다.
배치 시간대를 피해 야간·주말에 돌린다. 블록 이동 후 데이터 로컬리티가 일시 저하된다. 랙별 노드 수가 크게 불균형하면 랙 인식 정책 때문에 임계값에 도달하지 못할 수 있다. HDFS 와 Kudu 리밸런서를 동시에 돌리지 않는다.