Open Ask AI (⌘/Ctrl+I)

DX Cloud Alerts

DX Cloud alerts give a closer look at the status of your cluster(s). If there is an active alert, you see a notification at the top of the Cockpit.

Alerts notification in the Cockpit top bar

There are three categories of alerts:

For runbooks on specific alert names, see Troubleshooting alerts in DX Cloud Operations.

Select desired cluster

Select your desired cluster from the dropdown menu at the top of the Cockpit.

Select desired cluster in the Cockpit

Active alerts

Active alerts are currently firing for your cluster and should be investigated.

Active alerts in the Cockpit

Pending alerts

Pending alerts mean that at least one time series returned by the evaluation engine is Pending.

Pending alerts in the Cockpit

Inactive alerts

Inactive alerts have already occurred. Use them for information and history.

Inactive alerts in the Cockpit

Alert types

The following are the current alerts for DX Cloud. Linked names open the Operations troubleshooting page when one exists.

AlertDescription
CustomerMagnoliaContainerOomKilledMagnolia container has been OOMKilled at least once in the last 10 minutes.
CustomerMagnoliaDownMagnolia instance has not been running for at least the last 30 minutes.
CustomerMagnoliaServiceErrorsMagnolia service has application errors (503) for the last 5 minutes.
CustomerMagnoliaSlowResponseMagnolia author average response time for 90% of requests is more than 2 seconds for the last 15 minutes. Also covers CustomerMagnoliaPublicSlowResponse (public: more than 800 milliseconds for 90% of requests over 15 minutes).
CustomerClusterHighMemoryPressureA node in a customer cluster is under heavy memory pressure. Available memory is under 5% and there is a high rate of major page faults.
CustomerClusterFileSystemAlmostFullThe filesystem on a customer cluster has less than 1% space available.
CustomerClusterFileSystemFillingUpThe filesystem on a customer cluster has less than 30% space available and is projected to be full within 8 hours.
CustomerDatabasePersistentVolumeFillingUpA database persistent volume used by Magnolia has less than 1% space available. You may not be able to successfully add or publish content.
CustomerDatabasePersistentVolumeAlmostFullA database persistent volume on a customer cluster has less than 10% space available.
CustomerMagnoliaHomePersistentVolumeAlmostFullA Magnolia home persistent volume has less than 10% space available. Try removing search indexes on the affected volume to free space.
CustomerCertificateExpiringThe certificate for a host/ingress on a customer cluster is expiring in 14 days.
CustomerTomcatHighLoadTomcat is using more than 20% of its thread pool for the last 10 minutes.
CustomerMagnoliaCrashLoopingMagnolia instance is being restarted frequently by Kubernetes.
CustomerBackupCentralDownThe central backup server in your cluster is not running or not responding for at least 1 hour. Backups may not be saved.
CustomerElevatedServiceErrorsA service on your production cluster is returning errors (5xx) for more than 10% of requests for the last 10 minutes.