We have resolved the outage. A terraform bug in the rotation logic for one of our client certificates caused the cluster to enter an unhealthy state. This was resolved through an internal fix to our terraform configs by setting up a new resource. We are monitoring the terraform bug and have reported upstream. We apologise for the inconvenience and do please reach out to our team on slack or support@speakeasy.com with any questions. We are here to assist.
We will be adding additional validation in our infrastructure tests and monitors to catch similar bugs in cert rotation in the future.
Resolved
We have resolved the outage. A terraform bug in the rotation logic for one of our client certificates caused the cluster to enter an unhealthy state. This was resolved through an internal fix to our terraform configs by setting up a new resource. We are monitoring the terraform bug and have reported upstream. We apologise for the inconvenience and do please reach out to our team on slack or support@speakeasy.com with any questions. We are here to assist.
We will be adding additional validation in our infrastructure tests and monitors to catch similar bugs in cert rotation in the future.
Monitoring
All systems are operational and the downtime is resolve. The Team is investigating the root cause and will post an update soon.
Investigating
We are experiencing downtime on all services. Dashboard is currently inaccessible. The team is working on a resolution and we'll have an update soon.