Atlas 는 메타데이터를 JanusGraph 에 저장하고, 검색 인덱스를 Solr 에 둔다. 필요한 컬렉션은 세 개다.
| 컬렉션 | 용도 |
|---|---|
vertex_index |
엔티티(그래프 정점) 속성 인덱스 |
edge_index |
관계(그래프 간선) 인덱스 |
fulltext_index |
전문 검색 인덱스 |
셋 중 하나라도 없으면 Atlas 가 기동 중 실패하거나 검색이 동작하지 않는다.
curl -s "http://<solr-host>:8983/solr/admin/collections?action=LIST&wt=json"
curl -s "http://<solr-host>:8983/solr/admin/collections?action=CLUSTERSTATUS&wt=json" | head -50
SolrCloud 의 설정 세트는 ZooKeeper 에 있다.
solrctl instancedir --list
/opt/cloudera/parcels/CDH/lib/solr/bin/solr zk ls /configs -z <zk-host>:2181/solr
Atlas 배포본에 설정 세트 템플릿이 들어 있다. 이것을 ZooKeeper 에 올린 뒤 컬렉션을 만든다.
ATLAS_HOME=/opt/cloudera/parcels/CDH/lib/atlas
ZK=zk1:2181,zk2:2181,zk3:2181/solr
solrctl --zk "$ZK" instancedir --create atlas_configs "${ATLAS_HOME}/conf/solr"
solrctl --zk "$ZK" collection --create vertex_index -s 2 -r 2 -c atlas_configs
solrctl --zk "$ZK" collection --create edge_index -s 2 -r 2 -c atlas_configs
solrctl --zk "$ZK" collection --create fulltext_index -s 2 -r 2 -c atlas_configs
-s 는 샤드 수, -r 은 복제 계수다. 복제 계수는 Solr 서버 대수를 넘길 수 없다. Cloudera 클러스터가 아니면 solr 스크립트를 직접 쓴다.
solr create_collection -c vertex_index -d atlas_configs -shards 2 -replicationFactor 2
컬렉션을 만들 때 쓰는 설정 세트는 Atlas 가 제공하는 것이어야 한다. 기본 설정 세트(_default)로 만들면 스키마가 맞지 않아 색인이 실패한다.
atlas-application.properties 에서 검색 백엔드를 Solr 로 지정한다.
atlas.graph.index.search.backend=solr
atlas.graph.index.search.solr.mode=cloud
atlas.graph.index.search.solr.zookeeper-url=zk1:2181,zk2:2181,zk3:2181/solr
atlas.graph.index.search.solr.zookeeper-connect-timeout=60000
atlas.graph.index.search.solr.zookeeper-session-timeout=60000
ZooKeeper 주소에 chroot 경로(/solr)까지 포함해야 한다. 이것을 빠뜨려 컬렉션을 찾지 못하는 경우가 가장 흔하다.
설정을 바꾼 뒤 Atlas 를 재시작한다.
systemctl restart atlas
tail -f /var/log/atlas/application.log
curl -s "http://<solr-host>:8983/solr/vertex_index/select?q=*:*&rows=0&wt=json"
numFound 가 0보다 크면 색인이 쌓이고 있는 것이다. Atlas UI 의 검색이 결과를 내는지도 함께 본다.
색인이 비어 있는데 엔티티는 존재한다면, 인덱스를 다시 만들어야 한다. Atlas 는 그래프 저장소가 정본이므로 인덱스는 재생성이 가능하다. 재색인 절차는 버전마다 다르므로 해당 버전 문서를 확인한다. (확인 필요)
ranger_audits 컬렉션과 같은 Solr 인스턴스를 쓰면 자원 경합이 생긴다. 규모가 커지면 분리한다.curl 에 --negotiate -u : 를 붙여야 하고, 그 전에 kinit 이 필요하다.