노드 사이에 데이터가 치우쳤을 때 고르게 옮긴다. Kerberos 가 켜져 있으면 티켓부터 발급한다.
# kinit -kt hdfs.keytab hdfs@KRB.HAEDONGG.NET
kinit hdfs@KRB.HAEDONGG.NET
# Balancer 가 쓸 대역폭을 올린다. 예시는 10GiB/s
hdfs dfsadmin -setBalancerBandwidth 10737418240
# 노드 간 차이가 5% 이내가 될 때까지
hdfs balancer -policy datanode -threshold 5
# 1% 이내까지 빠르게 밀어붙일 때
hdfs balancer -Ddfs.balancer.movedWinWidth=5400000 \
-Ddfs.balancer.moverThreads=1000 \
-Ddfs.balancer.dispatcherThreads=200 \
-Ddfs.datanode.balance.max.concurrent.moves=100 \
-Ddfs.datanode.balance.bandwidthPerSec=10737418240 \
-Ddfs.balancer.max-size-to-move=10737418240 \
-threshold 1
# 끝나면 다른 작업에 영향이 없도록 대역폭을 낮춘다. 예시는 100MiB/s
hdfs dfsadmin -setBalancerBandwidth 104857600
| 키 | 기본값 | 뜻 |
|---|---|---|
dfs.balancer.movedWinWidth |
5400000 | ms. 최근에 옮긴 블록을 기억하는 창. 90분 |
dfs.balancer.moverThreads |
1000 | 블록을 옮기는 스레드 수 |
dfs.balancer.dispatcherThreads |
200 | 작업을 분배하는 스레드 수 |
dfs.datanode.balance.max.concurrent.moves |
100 | 한 데이터노드가 동시에 처리하는 이동 수 |
dfs.datanode.balance.bandwidthPerSec |
104857600 | byte/s. Balancer 가 쓰는 대역폭 |
dfs.balancer.max-size-to-move |
10737418240 | byte. 반복 한 번에 옮길 최대 크기 |
-threshold |
10 | %. 평균 사용률과의 차이가 이 값 이내가 될 때까지 수행. 설정 키가 아니라 CLI 옵션이다 |
키 이름을 틀리면 조용히 무시되고 기본값으로 돈다. moveWinWidth 가 아니라 movedWinWidth 이고, dfs.balance.bandwidthPerSec 가 아니라 dfs.datanode.balance.bandwidthPerSec 다[1]. 앞의 두 표기는 옛 문서에 흔히 남아 있는 잘못된 이름이다.
# 사용예
hdfs balancer -policy datanode -threshold 5
# 데이터 노드 간 밸런싱, 노드간 데이터 차이가 5% 이하가 될 때 까지 밸런싱
sudo -u hdfs hdfs balancer
[-policy (policy)] [-threshold (threshold)] [-blockpools (comma-separated list of blockpool ids)]
[-include [-f (hosts-file) | (comma-separated list of hosts)]]
[-exclude [-f (hosts-file) | (comma-separated list of hosts)]]
[-idleiterations (idleiterations)] [-runDuringUpgrade]
| -policy <policy> | blockpool/datanode 중 하나의 정책으로 hdfs balance를 수행한다. 'datanode'는 각 노드들의 사용량을 balancing하는 것 이라면, 'blockpool' 은 각 node의 pool까지 balancing 하는 것이다. 기본 값은 'datanode' 이다. |
| -threshold <threshold> | 1.0 ~100.0 사이의 수를 입력하여 어느 정도까지 node 간 balancing을 수행 할 것인지 설정한다. 기본 값은 10.0 으로 각 노드들을 10% 미만으로 차이가 날 때까지 balancing 수행한다. |
| -blockpools <comma-separated list of blockpool ids> | hdfs balancer가 명시된 리스트의 blockpool만 balancing 수행한다. 만약 list가 비워져 있다면 모든 block pool을 balancing 한다. 기본 값은 공백이다. |
| -include [-f <hosts-file> - <comma-separated list of hosts>] | balancing을 수행할 host를 명시한다. -f 옵션으로 host들의 리스트를 가지고 있는 파일을 지정하던가, 콤마(,)로 구분된 여러 host들을 지정한다. 만약 공백이라면 모든 datanode에 대해서 진행하며, 기본 값은 공백이다. |
| -exclude [-f <hosts-file> - <comma-separated list of hosts>] | -include 옵션과 반대로 balacing에서 제외할 host들만 입력한다. 공백값은 어떠한 node들도 제외되지 않는 것이고 기본 값은 공백이다. |
| -idleiterations <idleiterations> | hdfs balacer가 더 이상 balancing할 block이 없을 때까지 반복적으로 balancer를 수행한다. 기본 값은 5인데, 이동할 block이 없더라도 5번의 검사를 진행한다. |
| -runDuringUpgrade | 만약 이 옵션이 추가 되어 있다면 HDFS 업그레이드 수행 중에 balancer를 수행한다. (하지만 일반적으로 hdfs 업그레이드 중 balancer를 수행하는 것은 권고하지 않는다. 지속적으로 삭제되는 hdfs block들이 hdfs 내부 trash 공간을 빠르게 채울 것 이기 때문이다. ) |
표 안의 /는 | 이다.
| dfs.disk.balancer.enabled | This parameter controls if diskbalancer is enabled for a cluster. if this is not enabled, any execute command will be rejected by the datanode.The default value is false. |
| dfs.disk.balancer.max.disk.throughputInMBperSec | This controls the maximum disk bandwidth consumed by diskbalancer while copying data. If a value like 10MB is specified then diskbalancer on the average will only copy 10MB/S. The default value is 10MB/S. |
| dfs.disk.balancer.max.disk.errors | sets the value of maximum number of errors we can ignore for a specific move between two disks before it is abandoned. For example, f a plan has 3 pair of disks to copy between , and the first disk set encounters more than 5 errors, then we abandon the first copy and start the second copy in the plan. The default value of max errors is set to 5. |
| dfs.disk.balancer.block.tolerance.percent | The tolerance percent specifies when we have reached a good enough value for any copy step. For example, f you specify 10% then getting close to 10% of the target value is good enough. |
| dfs.disk.balancer.plan.threshold.percent | The percentage threshold value for volume Data Density in a plan. If the absolute value of volume Data Density which is out of threshold value in a node, it means that the volumes corresponding to the disks should do the balancing in the plan. The default value is 10. |
| Property | Default | Background Mode | Fast Mode |
| dfs.datanode.balance.max.concurrent.moves | 5 | # of disks | 4 x (# of disks) |
| dfs.balancer.moverThreads | 1000 | use default | 20,000 |
| dfs.balancer.max-size-to-move | 10737418240(10 GB) | 1073741824 (1GB) | 107374182400 (100 GB) |
| dfs.balancer.getBlocks.min-block-size | 10485760 (10 MB) | use default | 104857600 (100 MB) |
hdfs balancer -Ddfs.balancer.max-size-to-move=107374182400 -idleiterations -1 -policy datanode -threshold 1
옵션 이름은 -idleiteration 이 아니라 -idleiterations 이며, -1 은 옮길 블록이 없어질 때까지 무제한 반복한다는 뜻이다. 기본값은 5 다[2].
키 이름과 기본값은 Hadoop DFSConfigKeys · hdfs-default.xml 기준 — 2026-09-20 확인. https://hadoop.apache.org/docs/stable/hadoop-project-dist/hadoop-hdfs/hdfs-default.xml ↩︎
hdfs balancer 명령 옵션 — 2026-09-20 확인. https://hadoop.apache.org/docs/stable/hadoop-project-dist/hadoop-hdfs/HDFSCommands.html ↩︎