Product: Semarchy Data Platform Self-Hosted
Version: 1.3.0 or later (checked against 1.4.1)
Platforms: Amazon Web Services, Microsoft Azure, Rancher RKE2, Red Hat OpenShift
Author: Hélène Zosym
Need
When helm upgrade --install returns, you need to confirm that the platform is usable, not only that Helm finished. This article lists the checks to run after an installation or an upgrade, in order, from the Kubernetes cluster to the first administrator using Data Management. For each check it gives the expected result and where to look when it fails.
Summarized Solution
- The Helm release status is
deployed. - Every pod is
Runningand ready. - Every Job left in the namespace is
Complete. - The Welcome page opens at
https://<site_name>.<domain>and redirects to the login page. - The first administrator receives the activation email.
- The administrator completes the activation and can log in.
- The administrator can open Platform Administration. A 404 error points to the ingress path in
values.yaml. - The administrator grants themselves the Data Management roles, logs out and back in, and sees the Data Products section and the Data Management tools.
Detailed Solution
Set these variables first. <site_name> is the value of global.site_name (default selfhosted) and <domain> the value of global.domain. On OpenShift, you can use oc wherever this article uses kubectl.
export NAMESPACE=<SDP_NAMESPACE> export RELEASE=<HELM_RELEASE_NAME> # for example: sdp
1. The deployment is successful
helm list -n "$NAMESPACE" helm status "$RELEASE" -n "$NAMESPACE"
Expected: STATUS: deployed with the chart version you installed.
Helm reports deployed only when every hook Job has succeeded, including the rollout-check Job that waits for all Deployments to be ready.
| Status | What to do |
|---|---|
failed | A hook Job failed or Helm timed out. Go to step 3 to find the failing Job. |
pending-install or pending-upgrade | Helm is still running, or was interrupted. Wait for the command to end. If it was interrupted, run the same helm upgrade --install command again. |
2. All pods are Running
kubectl get pods -n "$NAMESPACE" # Only the pods that need attention kubectl get pods -n "$NAMESPACE" --field-selector=status.phase!=Running,status.phase!=Succeeded
Expected: every pod is Running, with all its containers ready (for example 2/2), and the restart count stays stable. The release normally runs:
| Pod | Replicas |
|---|---|
<release>-semarchy-iam-keycloak-0 | 1 |
<release>-dm-core-active-* | 1 |
<release>-dm-core-passive-* | 2 by default (dm.core.passive.replicaCount) |
<release>-billing-service-*, <release>-log-explorer-service-* | 1 each |
<release>-welcome-*, <release>-site-admin-*, <release>-user-profile-*, <release>-log-explorer-* | 1 each |
<release>-reloader-*, <release>-semarchy-data-platform-diag-sidecar-* | 1 each |
For a pod that is not ready:
kubectl describe pod <POD_NAME> -n "$NAMESPACE" # Events: image pull, missing Secret, probe failures kubectl logs <POD_NAME> -n "$NAMESPACE" --all-containers
| Status | Usual cause |
|---|---|
ImagePullBackOff | Registry credentials (semarchy-harbor), registry not reachable, or an image that does not follow your internal registry in an air-gapped installation. |
CreateContainerConfigError, ContainerCreating | A Secret or ConfigMap the pod mounts does not exist. After a successful installation this should not happen: check that the Job that creates it completed (see the SDP Helm Chart description article). |
CrashLoopBackOff, not ready | The application cannot reach a dependency: PostgreSQL, Kafka, the search engine, or Keycloak through its external URL. Read the pod logs. |
Pending | Not enough CPU or memory in the cluster, or no node matches the scheduling constraints. |
3. All Jobs are Completed
kubectl get jobs -n "$NAMESPACE"
Expected: every Job listed shows 1/1 completions and the status Complete.
Most setup Jobs are Helm hooks that are deleted once they succeed, so only a few remain in the namespace. A Job that is no longer listed after a deployed release has succeeded. You typically still see <release>-semarchy-iam-create-perm-admin-job-1, <release>-semarchy-iam-apply-config-1 and <release>-semarchy-iam-finalize-tf-apply, plus any diagnostic Job you created.For a Job that is not complete, read its logs:
kubectl describe job <JOB_NAME> -n "$NAMESPACE" kubectl logs job/<JOB_NAME> -n "$NAMESPACE"
Start with the first Job that failed: later failures are usually its consequences. The SDP Helm Chart description article lists every Job with what it needs and its frequent issues.
Optional: run the diagnostic tool
The built-in diagnostic tool runs these checks for you (Kubernetes version, Secrets, Job completions, pod health, a PostgreSQL connection test, Keycloak realms and clients) and saves the result in a bundle. Run it after every installation and keep the bundle; attach it to any Support ticket.
export CJ=$(kubectl get cronjob -n "$NAMESPACE" -l k8s.semarchy.net/diagnostic-install-job=true -o jsonpath='{.items[0].metadata.name}')
kubectl create job -n "$NAMESPACE" --from=cronjob/$CJ semarchy-diagnostic-install-job
export JOB=$(kubectl get job -n "$NAMESPACE" -l k8s.semarchy.net/diagnostic-install-job=true -o jsonpath='{.items[0].metadata.name}')
kubectl wait --for=condition=complete job/$JOB -n "$NAMESPACE" --timeout 10m
export POD=$(kubectl get pod -n "$NAMESPACE" -l k8s.semarchy.net/diagnostic-sidecar=true -o jsonpath='{.items[0].metadata.name}')
export LATEST=$(kubectl exec -n "$NAMESPACE" "$POD" -- sh -c 'ls -1t /diagnostic-volume/support-bundle*.tar.gz 2>/dev/null | head -n1')
kubectl cp -n "$NAMESPACE" "$POD:$LATEST" ./support-bundle.tar.gzReview analysis.json in the bundle first.
4. The Welcome page is accessible
Open https://<site_name>.<domain> in a browser, for example https://selfhosted.platform.example.com.
Expected: the browser redirects to the login page served by Keycloak at https://<domain>/auth/…, with a valid certificate.
| Symptom | Check |
|---|---|
| Name not resolved | DNS records for <domain> and *.<domain> (or <site_name>.<domain>) pointing to the ingress controller: nslookup <site_name>.<domain> |
| Certificate warning | The TLS Secret named in the ingress covers both hostnames; with cert-manager, check that the Certificate is Ready: kubectl get certificate -n "$NAMESPACE" |
| 404 from the ingress controller, or the wrong application | Ingress class and hosts: kubectl get ingress -n "$NAMESPACE". Every host must use <site_name>.<domain>, or <domain> for Keycloak, and every ingress must have the class of your controller. |
| Redirect loop or error after the redirect | global.domain and global.externalProtocol in values.yaml match the URL you use, and the Keycloak pod is ready. |
5. The activation email is received
At the end of the installation, the invite-site-admin Job asks Keycloak to send an activation email to the first administrator: the address set in tenant-settings.user_creation.email.
Expected: the email arrives within a few minutes. Its link is valid for 72 hours.
If it does not arrive:
- Check the spam or quarantine folder, and the logs of your SMTP relay.
- Check the
mail-secretvalues: host, port, user, password, sender address, and thesmtpSslandsmtpStartTlsflags. Some services only accept a verified sender address, for example Amazon SES. - Check that the address in
tenant-settings.user_creation.emailis correct. - To send the email again, fix the cause and run the same
helm upgrade --installcommand: the invitation Job runs on every upgrade and sends the email again as long as the administrator has not verified their address. Follow it withkubectl logs -f job/<release>-semarchy-data-platform-invite-site-admin -n "$NAMESPACE"while it runs.
6. The administrator can log in
Open the link in the email and complete the requested actions:
- verify the email address;
- set a password;
- configure multi-factor authentication with an authenticator app (Authy, Google Authenticator, Microsoft Authenticator…) when the realm requires it;
- accept the terms and conditions, if prompted.
Expected: you are logged in and redirected to Platform Administration.
If the link has expired, send the email again as described in step 5. If login fails after the actions, check the Keycloak pod logs.
7. The administrator can access Platform Administration
Open https://<site_name>.<domain>/admin, or the Platform Administration link from the Welcome page.
Expected: the Platform Administration page renders, with its Users section.
A 404 error on Platform Administration comes from the ingress path configured in values.yaml. Compare the site-admin ingress with the starter values file for your platform: the host must be <site_name>.<domain>, with the path /admin and pathType: Prefix.
site-admin:
ingress:
className: "<YOUR_INGRESS_CLASS_NAME>"
hosts:
- host: "<site_name>.<YOUR_GLOBAL_DOMAIN>"
paths:
- path: /admin
pathType: Prefix
With another path type, such as ImplementationSpecific, or with no paths entry, the ingress controller does not route the /admin/… pages to the application. Check what was deployed with kubectl get ingress <release>-site-admin -n "$NAMESPACE" -o yaml. The same check applies to / (welcome), /user-profile and /log-explorer. Fix values.yaml and run helm upgrade --install again.
8. The administrator can use Data Management
The first administrator manages the platform, but does not have Data Management permissions yet. Grant them, then check that Data Management is available:
- In Platform Administration, open Users and select the administrator user.
- Grant the DM User and DM Admin permissions, and save.
- Log out, then log in again so the new permissions are applied.
- Open the Welcome page.
Expected:
- the main page shows the Data Products section;
- the Tools section shows DM Administration, Dashboard Builder and DM REST API.
If these are missing after logging in again, Data Management is not running correctly. The most likely cause is the installation of the Data Management repository:
- Read the logs of the Data Management pods, starting with the active one:
kubectl logs deploy/<release>-dm-core-active -n "$NAMESPACE" --all-containers kubectl logs deploy/<release>-dm-core-passive -n "$NAMESPACE" --all-containers
- Look for PostgreSQL errors: connection, TLS, permission denied, or missing schema or extension.
- Read the logs of the
<release>-dm-setup-main-tf-applyJob, which creates the repository schema, its grants and the Data Management clients in Keycloak. The Job is deleted once it succeeds; to see its logs, runhelm upgrade --installagain and follow it while it runs:
kubectl logs -f job/<release>-dm-setup-main-tf-apply -n "$NAMESPACE"
- Check that the
dm-postgresSecret and the database preparation follow the documentation: the repository role owns theselfhosted-dmdatabase and its schema, and theuuid-osspandfuzzystrmatchextensions are allowed (on Azure, throughazure.extensions).
Checklist
| Check | Expected | |
|---|---|---|
| ☐ | helm status | deployed |
| ☐ | kubectl get pods | All Running and ready, stable restart count |
| ☐ | kubectl get jobs | Every remaining Job 1/1, Complete |
| ☐ | Diagnostic bundle (optional) | No failed check in analysis.json |
| ☐ | https://<site_name>.<domain> | Redirects to the login page, valid certificate |
| ☐ | Activation email | Received by the first administrator |
| ☐ | Activation and login | Password and MFA set, administrator logged in |
| ☐ | https://<site_name>.<domain>/admin | Platform Administration renders, no 404 |
| ☐ | DM User and DM Admin granted, logged out and in | Data Products section; DM Administration, Dashboard Builder and DM REST API in Tools |
If you want your team to follow the documentation and have a check list of what is already done and what's left, you can use the attached checklist as guidance.
Related articles
- SDP Helm Chart description
- How to clean up the environment after a failed SDP installation
- Create the admin account (documentation)
- Troubleshoot a deployment (documentation)