Job Execution Failure

Incident Report for Coalesce

Resolved

Between approximately 10:00 and 12:00 UTC today, some customers experienced delayed or stuck job runs for our Transform customers in North America.

During a routine, provider-managed Kubernetes upgrade in one of our US regions, our cloud provider was unable to provision replacement compute capacity. We experienced a similar issue on August 17th as described here:

https://status.coalesce.io/incidents/43zcl3trbjxx

We took steps after that prior incident to prevent provider upgrades from effecting our customers but those changes were not sufficient. We are currently making the following changes:

- Expanded Capacity Buffers: We are adding supplementary node pools (using C3/C3D machine types) alongside our primary pools to act as an immediate capacity buffer if our primary hardware type is unavailable.

- Automated Job Cleanup: We deployed an improvement to our job lifecycle management where delayed or stuck jobs now cleanly self-terminate, instantly freeing up resources for healthy jobs.

- Resilient Compute Architecture: We are actively implementing GKE Compute Classes, which will allow our infrastructure to automatically draw from up to six different machine types, eliminating our reliance on a single hardware type during cloud provider shortages.
Posted Sep 01, 2026 - 18:28 UTC
This incident affected: Transform - GCP (us-central1).