Azure Data Lake Storage Gen2 는 계층 네임스페이스(Hierarchical Namespace)를 켠 Blob Storage 다. 같은 데이터를 두 가지 API 로 다룰 수 있고, 목적에 따라 고르면 된다.
| 패키지 | 성격 |
|---|---|
azure-storage-file-datalake |
디렉터리·파일 개념을 그대로 쓴다. 디렉터리 이름 변경, ACL 설정처럼 Gen2 전용 기능이 있다 |
azure-storage-blob |
컨테이너와 blob 으로 다룬다. 단순히 파일 하나를 올리고 내리는 정도면 충분하다 |
azure-identity |
인증 자격 증명을 만드는 패키지. 위 둘과 함께 쓴다 |
python3 -m venv ~/venvs/azure && source ~/venvs/azure/bin/activate
pip install azure-storage-file-datalake azure-identity
ModuleNotFoundError: No module named 'azure.identity' 는 패키지가 없거나, 설치한 인터프리터와 실행하는 인터프리터가 다른 것이다. python3 -m pip show azure-identity 로 실제 설치 경로를 확인한다.
DefaultAzureCredential 은 환경 변수 → 관리 ID → Azure CLI 로그인 순으로 자격 증명을 찾는다. 키를 코드나 설정 파일에 두지 않아도 되므로 이쪽을 먼저 고려한다.
from azure.identity import DefaultAzureCredential
from azure.storage.filedatalake import DataLakeServiceClient
account = "mystorageaccount"
service = DataLakeServiceClient(
account_url=f"https://{account}.dfs.core.windows.net",
credential=${MASKED}),
)
이 방식으로 쓰려면 해당 주체에 Storage Blob Data Reader 또는 Storage Blob Data Contributor 역할이 있어야 한다. 구독 수준의 Owner 나 Contributor 만으로는 데이터 평면 접근이 되지 않는다. 이 점을 놓쳐 AuthorizationPermissionMismatch 로 막히는 일이 잦다.
from azure.storage.filedatalake import DataLakeServiceClient
service = DataLakeServiceClient(
account_url="https://mystorageaccount.dfs.core.windows.net",
credential="${AZURE_STORAGE_ACCOUNT_KEY}",
)
계정 키는 포털의 저장소 계정 → 보안 + 네트워킹 → 액세스 키에 있다. 키 값은 코드나 저장소에 넣지 않고 환경 변수나 Key Vault 에서 읽는다.
Connection string missing required connection details 는 연결 문자열의 형식이 맞지 않을 때 난다. 온전한 형태는 다음과 같다.
DefaultEndpointsProtocol=https;AccountName=<계정>;AccountKey=<키>;EndpointSuffix=core.windows.net
Gen2 엔드포인트가 blob.core.windows.net 이 아니라 dfs.core.windows.net 이라는 점도 확인한다. DataLakeServiceClient 에는 dfs 를 준다.
import os
from azure.identity import DefaultAzureCredential
from azure.storage.filedatalake import DataLakeServiceClient
ACCOUNT = "mystorageaccount"
FILESYSTEM = "mycontainer"
REMOTE = "raw/2026/report.csv"
LOCAL = "/data/download/report.csv"
service = DataLakeServiceClient(
account_url=f"https://{ACCOUNT}.dfs.core.windows.net",
credential=${MASKED}),
)
fs = service.get_file_system_client(FILESYSTEM)
file_client = fs.get_file_client(REMOTE)
os.makedirs(os.path.dirname(LOCAL), exist_ok=True)
with open(LOCAL, "wb") as f:
file_client.download_file().readinto(f)
readall() 은 파일 전체를 메모리에 올린다. 큰 파일은 위처럼 readinto() 로 스트림에 바로 쓴다.
with open(LOCAL, "rb") as f:
fs.get_file_client("processed/report.csv").upload_data(f, overwrite=True)
overwrite=True 를 빠뜨리면 같은 경로가 이미 있을 때 실패한다.
for path in fs.get_paths(path="raw/2026", recursive=True):
print(path.name, path.is_directory, path.content_length)