When the cloud API blocks operations

A Hetzner Cloud API and Console incident on 6 October prevented some resource actions for 103 minutes. The event did not make every hosted application unavailable. It did expose a more important question: who can distinguish provider trouble from application failure and keep operating safely while the control plane is impaired?

Author
Intercube
Published
Reading time
8 min read

01

What happened on 6 October

Hetzner reported a fault affecting its Cloud API and Console from 08:20 to 10:03 UTC on 6 October 2026. Incorrect limit calculations prevented some actions, including creating resources and moving a resource. The provider listed load balancers, networks, cloud servers, object storage and cloud volumes among the affected systems, then confirmed that the problem was fixed.

The notice describes failed control-plane actions. It does not say that every running server or website went offline. That distinction matters. An application can continue serving traffic while operators temporarily lose the ability to create capacity, change networking or execute a recovery step through the provider API.

The practical impact therefore depends on what the application needed during those 103 minutes. A stable workload may have required no change. A deployment, scale event, disaster-recovery action or infrastructure migration could have encountered a blocked operation at exactly the wrong moment.

02

Control-plane failure changes safe operations

When a provider control plane is impaired, repeated automation can make diagnosis harder. Failed create or move operations may be cleanly rejected, delayed or partially represented in local state. Operators need to pause non-essential changes, inspect the provider status, reconcile actual resources with infrastructure state and retry only when the outcome of the first request is known.

Recovery plans must also account for the availability of the tools used to execute them. A snapshot or backup is useful only when the team can locate it, understand what data it contains and restore it through an available path. Hetzner documents that server backups and snapshots copy a server disk but do not include attached volumes, which makes application-aware recovery planning necessary.

This is why provider status, application health and deployment state belong in one incident view. Without that context, a team can spend the first part of an incident debugging its own code, or attempt infrastructure changes that cannot complete while the provider API is unavailable.

  • Separate application availability from provider control-plane availability.
  • Pause non-essential automation until failed operations have a known outcome.
  • Reconcile infrastructure state before retrying create, move or network changes.
  • Test recovery with every data store and attached volume in scope.

03

A cloud account does not provide an operator

Cloud providers operate physical infrastructure and their own service control planes. The customer still owns application architecture, deployment behaviour, provider-specific automation, monitoring, backups and the decision process during an incident. A low-cost cloud server does not include someone who understands what a failed provider action means for an order flow, queue or database.

That gap often remains hidden while everything is healthy. It appears when a release coincides with provider trouble, capacity has to change quickly or recovery depends on a resource action that is temporarily unavailable. Developers then become the infrastructure response team, even if that work was never part of the product plan.

A managed hosting relationship makes the boundary explicit. The application team owns the application and releases. The operator owns the deeper infrastructure lifecycle, follows provider incidents, protects routine changes and coordinates recovery with knowledge of the actual workload.

04

Move the operating burden, not only the server

Changing infrastructure providers does not remove this responsibility. Every provider has maintenance windows, quotas, service incidents and failure domains. The durable improvement is to design around the application, automate repeatable operations and assign an operator who understands both the provider and the workload.

Intercube provides that managed layer. We provision and maintain the infrastructure, connect deployments, monitor the platform and handle provider-level operations while the customer team retains control of its code and releases. When an underlying service has an incident, the customer has one technical party responsible for translating it into application impact and action.

If a team is currently responsible for a raw cloud account, migration to Intercube managed hosting is included at no additional cost within the agreed onboarding scope. The move covers the target environment, application and data transfer, deployment preparation, production cutover and fallback planning where the workload requires it.

What this incident should prompt

  • Map which deployment and recovery actions depend on the provider control plane.
  • Confirm that backups cover attached data and can be restored through a tested path.
  • Define how automation pauses and reconciles after uncertain provider operations.
  • Give one operator responsibility for infrastructure, monitoring and recovery.
  • Move to managed hosting when cloud operations keep falling back to developers.

Primary sources

Take cloud operations off the product team.

We can assess the current environment, design the managed target and move the application without an additional migration fee.

Plan a managed migration