hbase org.apache.hadoop.hbase.mapreduce.Export 가 만든 파일은 Hadoop SequenceFile 이다. 파일 앞부분을 보면 매직 문자열과 키·값 클래스 이름이 보인다.
SEQ...org.apache.hadoop.hbase.io.ImmutableBytesWritable...org.apache.hadoop.hbase.client.Result
키는 행 키를 감싼 ImmutableBytesWritable, 값은 한 행의 모든 셀을 담은 Result 다. 텍스트 파일이 아니므로 cat 이나 head 로 읽으면 알아볼 수 없다.
hdfs dfs -ls /backup/mytable
hdfs dfs -cat /backup/mytable/part-m-00000 | head -c 200 | xxd | head
가장 간단한 방법은 Hadoop 의 -text 다. SequenceFile 을 인식해 사람이 읽을 수 있는 형태로 바꿔 준다.
hdfs dfs -text /backup/mytable/part-m-00000 | head -20
Result 의 toString() 결과가 찍히므로 셀이 많으면 잘린다. 값 전체를 정확히 봐야 하면 아래의 프로그램 방식을 쓴다.
원래 목적이 복원이면 Import 를 쓴다. 대상 테이블은 미리 만들어 두어야 한다.
hbase org.apache.hadoop.hbase.mapreduce.Import mytable /backup/mytable
다른 이름의 테이블로 넣거나 컬럼 패밀리 이름을 바꾸려면 매핑 옵션을 준다.
hbase -Dhbase.import.version=2 \
org.apache.hadoop.hbase.mapreduce.Import \
-Dimport.filter.class=... newtable /backup/mytable
대량이면 Import 로 HFile 을 만든 뒤 벌크 로드하는 편이 빠르다.
hbase org.apache.hadoop.hbase.mapreduce.Import \
-Dimport.bulk.output=/tmp/hfiles newtable /backup/mytable
hbase org.apache.hadoop.hbase.tool.LoadIncrementalHFiles /tmp/hfiles newtable
내용을 가공해야 하면 SequenceFile 리더로 직접 연다.
Configuration conf = HBaseConfiguration.create();
Path path = new Path("hdfs:///backup/mytable/part-m-00000");
try (SequenceFile.Reader reader =
new SequenceFile.Reader(conf, SequenceFile.Reader.file(path))) {
ImmutableBytesWritable key = new ImmutableBytesWritable();
Result value = new Result();
while (reader.next(key, value)) {
System.out.println(Bytes.toString(key.get()));
for (Cell cell : value.rawCells()) {
System.out.printf(" %s:%s = %s%n",
Bytes.toString(CellUtil.cloneFamily(cell)),
Bytes.toString(CellUtil.cloneQualifier(cell)),
Bytes.toString(CellUtil.cloneValue(cell)));
}
}
}
hbase classpath 로 클래스패스를 얻어 실행한다.
javac -cp "$(hbase classpath)" ReadExport.java
java -cp "$(hbase classpath):." ReadExport
Spark 로 읽는 방법도 있다.
val rdd = sc.sequenceFile[ImmutableBytesWritable, Result]("hdfs:///backup/mytable")
| 형식 | 만든 도구 | 읽는 법 |
|---|---|---|
| SequenceFile | Export |
hdfs dfs -text, Import |
| HFile | 리전 데이터, 벌크 로드 산출물 | hbase hfile -p -f <path> |
| 스냅샷 | snapshot · ExportSnapshot |
hbase snapshot info, restore_snapshot |
hbase hfile -p -f /hbase/data/default/mytable/<region>/<cf>/<file> | head
Export 산출물은 셀 값과 타임스탬프를 그대로 담는다. Import 로 복원하면 원본 타임스탬프가 유지되므로, TTL 이 걸린 테이블에서는 넣자마자 만료될 수 있다.Export 는 MapReduce 작업이라 스캔 부하가 원본 테이블에 그대로 걸린다. 큰 테이블은 스냅샷 기반 내보내기를 검토한다.part-m-* 여러 개로 나뉜다. 개별 파일만 보고 전체라고 판단하지 않는다.