Product: Semarchy Data Platform Self-Hosted
Version: 1.3.0 or later (object names checked against 1.4.1)
Author: Hélène Zosym



Need


A first installation of SDP Self-Hosted can stop part-way, for example because a Secret was wrong, a database or search engine was unreachable, or Helm timed out. By then the setup Jobs have already created objects in Kubernetes, PostgreSQL, Keycloak and the search engine. Their Terraform state is stored in Kubernetes Secrets in the installation namespace.


If you delete the namespace without cleaning the external services, the next attempt finds objects that its new Terraform state does not know about, and fails again. This article describes how to return the environment to the state it was in just after the preparation steps, so that the next installation starts clean.


Use this procedure only for a failed first installation, or for a non-production environment you want to rebuild. It deletes the Keycloak database, and may delete the Data Management database. On an environment that holds business data, take a backup first and contact Semarchy Support.


Summarized Solution

  1. Check whether you need a cleanup at all. If you are past the first job - you do, at least on the indexes.
  2. Back up the Secrets you created, then delete the installation namespace. This stops every SDP pod and removes the Terraform state.
  3. Drop and recreate the keycloak database, if you have keycloak-0 pod up and running.
  4. If the previous attempt reached the dm-setup-main-tf-apply Job, drop and recreate the selfhosted-dm database as well.
  5. Delete the search engine indexes, index template and lifecycle policy created for this release. Do not delete all indexes.
  6. Recreate the namespace and the Secrets as described in the documentation, run the prechecks, and install again without --wait or --atomic.


Detailed Solution


1. Decide whether a cleanup is needed

The setup Jobs are Terraform applies that keep their state in the namespace. However, some jobs will create stateful elements in the keycloak database and in the Search Engine. 

  1. If the job log-explorer-service-setup-tf-apply completed successfully - you need to clean up indexes in Search Engine. 
  2. If the keycloak-0 pod has been started - you need to clean up keycloak database
  3. In all cases it is best to clean up the namespace.


2. Set the environment variables

The procedure works on every supported platform: Amazon Web Services, Microsoft Azure, Rancher RKE2 and Red Hat OpenShift. Set these variables once, in the shell you use for the whole procedure. On OpenShift, you can use oc wherever this article uses kubectl.

export NAMESPACE=<SDP_NAMESPACE>
export RELEASE=<HELM_RELEASE_NAME>          # for example: sdp
export CLIENT_NS=<NAMESPACE_FOR_CLIENT_PODS> # for example: default

# PostgreSQL administrator (the account you used to prepare the databases)
export POSTGRES_HOST=<POSTGRES_HOST>
export POSTGRES_PORT=5432
export POSTGRES_ADMIN_USER=<POSTGRES_ADMIN_USER>
export POSTGRES_ADMIN_PASSWORD=<POSTGRES_ADMIN_PASSWORD>

# Search engine administrator
export SEARCH_ENGINE_TYPE=<opensearch|elasticsearch>   # same value as global.searchEngine.type
export SEARCH_ENGINE_URL=https://<SEARCH_ENGINE_HOST>:<PORT>
export SEARCH_ADMIN_USER=<SEARCH_ENGINE_ADMIN_USER>
export SEARCH_ADMIN_PASSWORD=<SEARCH_ENGINE_ADMIN_PASSWORD>

The client pods used in steps 4 and 6 only need network access to PostgreSQL and the search engine. On OpenShift, the postgres:16 image may not start under the restricted security context constraint: run the psql and curl commands from a project where it is allowed, or from any workstation or bastion host that can reach these services.

3. Back up your Secrets and delete the namespace

Delete the namespace before touching the databases. Otherwise Keycloak and Data Management keep running and may recreate tables in the databases you have just emptied.

Optionally, save the Secrets you created so that you can restore them in step 8 without typing them again (requires jq):

for s in semarchy-harbor keycloak-postgres dm-postgres dm-postgres-datasource-1 dm-postgres-datasource-2 \
         kafka-keycloak dm-kafka "${SEARCH_ENGINE_TYPE}-provider" mail-secret; do
  kubectl get secret "$s" -n "$NAMESPACE" -o json \
    | jq 'del(.metadata.uid, .metadata.resourceVersion, .metadata.creationTimestamp, .metadata.managedFields, .metadata.annotations, .metadata.labels)'
done > sdp-input-secrets.json

Add your TLS Secrets and, for a private CA, ca-http to the list if you created them manually. If you use another image pull secret name for an air-gapped registry, use that name instead of semarchy-harbor. The file contains credentials: store it securely and delete it once the installation succeeds.

Delete the namespace and wait until it is gone:

kubectl delete namespace "$NAMESPACE"
kubectl wait --for=delete namespace/"$NAMESPACE" --timeout=10m

This removes every SDP workload, the Helm release record and the tfstate-* Secrets that hold the Terraform state. Deleting the namespace is preferable to helm uninstall here: after a failed installation, the uninstall hooks may themselves fail or wait for a component that never started.

4. Recreate the Keycloak database

Start a temporary PostgreSQL client in the cluster:

kubectl run psql-client --rm -it --image=postgres:16 -n "$CLIENT_NS" \
  --env="PGHOST=$POSTGRES_HOST" \
  --env="PGPORT=$POSTGRES_PORT" \
  --env="PGUSER=$POSTGRES_ADMIN_USER" \
  --env="PGPASSWORD=$POSTGRES_ADMIN_PASSWORD" \
  -- bash

Inside the pod, drop and recreate the database with the same settings as in the preparation steps:

psql postgres

DROP DATABASE "keycloak" WITH (FORCE);
CREATE DATABASE "keycloak" ENCODING = 'UTF8' LC_COLLATE = 'C' LC_CTYPE = 'C' TEMPLATE template0 CONNECTION LIMIT = -1;
ALTER DATABASE "keycloak" OWNER TO "keycloak";
GRANT ALL PRIVILEGES ON DATABASE "keycloak" TO "keycloak";
\c keycloak
GRANT ALL ON SCHEMA public TO keycloak;
\c postgres

Keep the pod open if you also need step 5. The keycloak role and its password are kept, so the keycloak-postgres Secret does not change.


If DROP DATABASE fails with must be owner of database, the admin user is not a member of the owner role. This is common on managed PostgreSQL services such as Amazon RDS and Azure Database for PostgreSQL. Grant the membership, then drop again: GRANT "keycloak" TO CURRENT_USER;


5. Recreate the Data Management database, if the previous attempt reached dm-setup-main-tf-apply

Check how far the previous attempt went. If you still have its output or a diagnostic bundle, look for the Job <release>-dm-setup-main-tf-apply. It is a post-install hook with weight 6 (see the SDP Helm Chart description article for the full sequence).

The previous attempt…What is left in selfhosted-dmAction
stopped before dm-setup-main-tf-apply ranOnly the extensions schema with uuid-ossp and fuzzystrmatch, created by the pre-install Job dm-setup-extensions-tf-apply. The next attempt reuses it as is.Nothing to do.
ran dm-setup-main-tf-apply or laterGrants on the repository schema, the repository tables created by the Data Management pods when they started. These refer to the Keycloak realm and clients you deleted in step 4.Drop and recreate repository schema as described below

In the same psql session, connected to selfhosted-dm:

DROP SCHEMA IF EXISTS  "selfhosted-dm-repo-user" CASCADE;

-- Repository schema and read-only access
CREATE SCHEMA "selfhosted-dm-repo-user";
ALTER SCHEMA "selfhosted-dm-repo-user" OWNER TO "selfhosted-dm-repo-user";
GRANT CONNECT ON DATABASE "selfhosted-dm" TO "selfhosted-dm-repo-user-ro";
GRANT USAGE ON SCHEMA "selfhosted-dm-repo-user" TO "selfhosted-dm-repo-user-ro";

\q
exit

6. Clean the search engine

The setup Jobs create the following objects, all prefixed with the Helm release name:

ObjectName (release sdp)
Log indexes and aliasessdp-log-explorer-index-000001, … with aliases sdp-log-explorer-index-write and sdp-log-explorer-index-read
Index templatesdp-log-explorer-index-template
Lifecycle policysdp-log-explorer-index-policy: an ISM policy on OpenSearch, an ILM policy on Elasticsearch
Billing indexsdp-billing-metrics

Start a temporary curl client in the cluster:

kubectl run curl-client --rm -it --image=curlimages/curl -n "$CLIENT_NS" \
  --env="SEARCH_ENGINE_URL=${SEARCH_ENGINE_URL}" \
  --env="SEARCH_ADMIN_USER=${SEARCH_ADMIN_USER}" \
  --env="SEARCH_ADMIN_PASSWORD=${SEARCH_ADMIN_PASSWORD}" \
  --env="RELEASE=${RELEASE}" \
  -- sh

Inside the pod, list the indexes of this release, then delete them by name. Their aliases are deleted with them. These commands are the same for OpenSearch and Elasticsearch:

AUTH="${SEARCH_ADMIN_USER}:${SEARCH_ADMIN_PASSWORD}"
INDEXES="${RELEASE}-log-explorer-index-*,${RELEASE}-billing-metrics"

# List
curl -s -u "$AUTH" "${SEARCH_ENGINE_URL}/_cat/indices/${INDEXES}?h=index"

# Delete, one index at a time
for i in $(curl -s -u "$AUTH" "${SEARCH_ENGINE_URL}/_cat/indices/${INDEXES}?h=index"); do
  curl -s -u "$AUTH" -X DELETE "${SEARCH_ENGINE_URL}/${i}"; echo
done


Then delete the index template and the lifecycle policy, using the commands for your search engine.


OpenSearch (Amazon OpenSearch Service, or OpenSearch on RKE2 and OpenShift):

# Index template: the API depends on the OpenSearch version, so try both
curl -s -u "$AUTH" -X DELETE "${SEARCH_ENGINE_URL}/_index_template/${RELEASE}-log-explorer-index-template"; echo
curl -s -u "$AUTH" -X DELETE "${SEARCH_ENGINE_URL}/_template/${RELEASE}-log-explorer-index-template"; echo

# ISM policy
curl -s -u "$AUTH" -X DELETE "${SEARCH_ENGINE_URL}/_plugins/_ism/policies/${RELEASE}-log-explorer-index-policy"; echo

# Check that nothing is left
curl -s -u "$AUTH" "${SEARCH_ENGINE_URL}/_cat/indices/${RELEASE}-*?v"
curl -s -u "$AUTH" "${SEARCH_ENGINE_URL}/_plugins/_ism/policies/${RELEASE}-log-explorer-index-policy"; echo

exit


Elasticsearch (Elastic Cloud on Azure):

# Index template
curl -s -u "$AUTH" -X DELETE "${SEARCH_ENGINE_URL}/_index_template/${RELEASE}-log-explorer-index-template"; echo

# ILM policy
curl -s -u "$AUTH" -X DELETE "${SEARCH_ENGINE_URL}/_ilm/policy/${RELEASE}-log-explorer-index-policy"; echo

# Check that nothing is left
curl -s -u "$AUTH" "${SEARCH_ENGINE_URL}/_cat/indices/${RELEASE}-*?v"
curl -s -u "$AUTH" "${SEARCH_ENGINE_URL}/_ilm/policy/${RELEASE}-log-explorer-index-policy"; echo

exit

A 404 response on a delete means the object did not exist, which is fine. If your search engine uses a certificate from a private CA, add --cacert with your CA file, or -k for a test environment only.

Keep the search engine role and user you created during preparation: the opensearch-provider or elasticsearch-provider Secret still uses them.

8. Recreate the namespace and the Secrets

kubectl create namespace "$NAMESPACE"

If you saved the Secrets in step 3, restore them:

jq -s '{apiVersion: "v1", kind: "List", items: .}' sdp-input-secrets.json | kubectl apply -n "$NAMESPACE" -f -

Otherwise, create them again exactly as described in the Configure Kubernetes page for your platform:

SecretPurposeChanges after this cleanup?
semarchy-harbor (or your internal registry Secret)Image pull credentialsNo
keycloak-postgresKeycloak databaseNo: role and password are kept
dm-postgresData Management repositoryNo
dm-postgres-datasource-1, dm-postgres-datasource-2Initial datasourcesNo
kafka-keycloak, dm-kafkaKafka, Amazon MSK or Event HubsOnly if you recreated the event hub in step 7
opensearch-provider or elasticsearch-providerSearch engineNo
mail-secretSMTP serverNo
TLS Secrets, ca-httpIngress certificates, private CANo, unless cert-manager issues them

While recreating them, check the values that most often cause a first installation to fail:

  • the SASL mechanism and login module match your messaging service: PLAIN for Azure Event Hubs, SCRAM for Kafka and Amazon MSK;
  • the PostgreSQL TLS settings (sslmode in the JDBC URLs and in values.yaml) match what your database server requires. Managed services such as Amazon RDS and Azure Database for PostgreSQL usually require TLS;
  • every key listed in the documentation is present: the prechecks in step 9 report missing keys.

9. Install again, following the best practices

Run the prechecks

The prechecks validate the Secrets and their keys, connectivity to PostgreSQL, Kafka, the search engine and SMTP, DNS, and certificates, before anything is installed. Run them after creating the Secrets:

helm template "$RELEASE" -n "$NAMESPACE" \
  oci://registry.na.semarchy.net/semarchy-release/semarchy-data-platform --version <CHART_VERSION> \
  --set global.prechecks.enabledPreflights=true \
  -f values.yaml | kubectl preflight -

Fix every Fail before installing. The preflight plugin is installed with Krew; see Run deployment prechecks.

Install without --wait or --atomic, with a longer timeout

helm upgrade --install "$RELEASE" \
  -n "$NAMESPACE" \
  oci://registry.na.semarchy.net/semarchy-release/semarchy-data-platform \
  --version <CHART_VERSION> \
  --timeout 20m \
  -f values.yaml
  • Do not use --wait. Keycloak, billing and Data Management mount Secrets and ConfigMaps that post-install Jobs create. Helm with --wait waits for these pods before it runs the post-install Jobs, so it can only time out. The chart checks readiness itself with its rollout-check Job.
  • Do not use --atomic. It rolls back an installation that was progressing normally, and leaves behind the external objects the Jobs had already created: exactly the situation this article cleans up.
  • Use --timeout 20m, or more on slow networks. Helm's default of 5 minutes is often too short for the first image pulls.
  • Keep the same release name, namespace, global.site_id and global.site_name as in your preparation. They are part of resource and state names.


While the installation runs, pods in CreateContainerConfigError or ContainerCreating are expected. They wait for a Secret or ConfigMap that a later Job creates. See the SDP Helm Chart description article to know which Job each pod waits for.


10. Troubleshooting tools

If the installation fails again, find the root cause before cleaning up again.

  1. Find the first failing Job or pod. Later failures are usually consequences of an earlier one.
kubectl get jobs -n "$NAMESPACE"
kubectl get pods -n "$NAMESPACE"
kubectl logs job/<JOB_NAME> -n "$NAMESPACE"
kubectl describe pod <POD_NAME> -n "$NAMESPACE"
  1. Map it to its stage with the SDP Helm Chart description article: each Job and pod is listed with what it needs and its frequent issues.
  2. Run the built-in diagnostic tool. It collects logs, Job statuses, selected Secrets (values redacted), Keycloak realms and clients, and a PostgreSQL connection test, and analyses them:
export CJ=$(kubectl get cronjob -n "$NAMESPACE" -l k8s.semarchy.net/diagnostic-install-job=true -o jsonpath='{.items[0].metadata.name}')
kubectl create job -n "$NAMESPACE" --from=cronjob/$CJ semarchy-diagnostic-install-job
export JOB=$(kubectl get job -n "$NAMESPACE" -l k8s.semarchy.net/diagnostic-install-job=true -o jsonpath='{.items[0].metadata.name}')
kubectl wait --for=condition=complete job/$JOB -n "$NAMESPACE" --timeout 10m
export POD=$(kubectl get pod -n "$NAMESPACE" -l k8s.semarchy.net/diagnostic-sidecar=true -o jsonpath='{.items[0].metadata.name}')
export LATEST=$(kubectl exec -n "$NAMESPACE" "$POD" -- sh -c 'ls -1t /diagnostic-volume/support-bundle*.tar.gz 2>/dev/null | head -n1')
kubectl cp -n "$NAMESPACE" "$POD:$LATEST" ./support-bundle.tar.gz

Start with analysis.json in the bundle. Run the tool before deleting the namespace, and attach the bundle to your Support ticket.


Checklist

Step
☐Root cause of the failure identified (first failing Job or pod, diagnostic bundle saved)
☐Retry without cleanup attempted, or cleanup confirmed as necessary
☐Input Secrets backed up, namespace deleted and fully removed
☐keycloak database dropped and recreated
☐selfhosted-dm repository schema dropped and recreated, if dm-setup-main-tf-apply had run
☐Search engine indexes, template and lifecycle policy of this release deleted, by name
☐Namespace and Secrets recreated as in the documentation for your platform, TLS and CA Secrets included
☐Prechecks pass
☐Installed with --timeout 20m, without --wait or --atomic

Related articles