Skip to content
9 Oct  2026

 

Most organisations would never design a critical system around a single point of failure.

Yet many still run their AWS and cloud operations that way.

One or two people understand the environment. They know why certain decisions were made, where the risks are, which systems are fragile, what changed last week, and what to do when something goes wrong.

While those people are available, the arrangement can appear to work.

The real test comes when they are not.

The holiday test

A useful way to assess the resilience of your cloud operating model is to ask a simple question:

What happens when your most knowledgeable cloud person takes a week off?

Do incidents take longer to resolve?

Do changes get delayed because nobody else feels confident approving them?

Do colleagues contact that person anyway, even though they are on leave?

Does the business start hoping that nothing serious happens until they return?

If the answer to any of those questions is yes, the issue is not simply staffing capacity. It is key-person risk.

The business has become dependent on knowledge, judgement and operational context that exists primarily in one person’s head.

The hidden impact on your people

Key-person dependency creates an obvious operational risk, but it also creates a significant people problem.

When a cloud specialist is unavailable, the pressure does not disappear. It moves to the rest of the team.

Colleagues who may not fully understand the environment are left trying to diagnose issues, make decisions and manage risk with incomplete context. This increases stress and can lead to slower resolution, unnecessary escalation or overly cautious decision-making.

At the same time, the unavailable employee often remains psychologically on call.

They may be contacted while on holiday, during evenings, while caring for family, or at times when they are supposed to be recovering from work. Even when they are not contacted, they may feel unable to disconnect because they know nobody else has the same level of knowledge.

Over time, this creates strain on both professional and personal relationships.

Annual leave stops being genuine recovery time. The individual feels responsible for keeping everything running, while the wider team becomes increasingly reliant on them.

What can look like dedication is often an unhealthy operating model.

When inconvenience becomes business impact

In many cases, key-person dependency leads to frustration, delay and additional pressure.

In a serious incident, however, it can lead to genuine business impact.

If a critical service fails and the only person who understands the relevant configuration cannot be contacted, the organisation may face:

  • Longer outages

  • Slower incident response

  • Increased security exposure

  • Delayed customer communication

  • Higher recovery costs

  • Greater risk of making an incorrect change under pressure

  • The most concerning part is that the problem may remain hidden until the organisation experiences a major incident, resignation, extended sickness absence or unexpected departure.

By that point, it is too late to start transferring knowledge.

This is not resilient governance

Cloud governance is often discussed in terms of policies, controls, cost management and security standards.

Those things matter, but governance must also be operationally resilient.

If your cloud operations fall apart when one person takes leave, you do not have effective governance. You have a single point of failure.

Resilient governance means the organisation can continue to operate when individuals are unavailable. Responsibilities are clear, decisions are documented, processes are repeatable, and knowledge is accessible to more than one person.

It also means that incidents can be managed without relying on personal memory or informal relationships.

A policy document alone does not create resilience. The operating model behind it does.

Moving beyond the individual hero

Many organisations rely on highly capable individuals who repeatedly step in, solve difficult problems and keep services running.

These people are often described as indispensable.

That may feel like a compliment, but it should also be treated as a warning.

A resilient cloud function should not depend on individual heroics. It should be designed around shared ownership, documented processes and repeatable operational practices.

That includes:

  • Clear runbooks for common incidents and operational tasks

  • Accurate documentation of architecture, dependencies and known risks

  • Defined escalation paths

  • Shared access to monitoring, alerts and operational history

  • Consistent change and incident management

  • Regular knowledge transfer

  • Coverage outside normal working hours

  • A clear understanding of who owns each service and decision

The objective is not to make skilled people less important.

It is to stop the organisation from becoming dangerously dependent on them.

How Managed Services reduces key-person risk

A well-designed Managed Services model provides more than additional technical capacity.

It creates a shared operational capability around the customer environment.

Instead of relying on one person’s memory, knowledge is captured through documentation, service onboarding, incident records, change history, monitoring data and ongoing operational engagement.

Responsibility is distributed across a wider team with defined processes and escalation routes.

This creates continuity during holidays, sickness, periods of high workload and staff turnover.

Managed Services can provide:

Shared operational ownership

Cloud operations become the responsibility of a service team rather than a single individual. There is a broader pool of engineers who can respond, investigate and escalate issues.

Documented knowledge

Environment-specific information is recorded and maintained rather than retained informally. This helps ensure that operational knowledge remains available even when people change roles or leave the organisation.

24/7 coverage

Incidents do not always occur during office hours. Continuous coverage reduces the expectation that internal staff must remain informally available during evenings, weekends or holidays.

Consistent incident and change processes

Defined processes improve response quality and reduce the likelihood of rushed or undocumented decisions during high-pressure situations.

Access to broader experience

An internal cloud engineer may only see the challenges within one organisation. A Managed Services team works across multiple environments and can bring wider operational experience to incidents, architecture decisions and recurring problems.

AI-assisted operational context

AI Ops can strengthen continuity by analysing monitoring data, incidents, alerts and operational patterns within a specific environment.

Over time, this can help identify recurring behaviour, surface relevant historical context and support faster investigation.

AI does not replace engineering judgement. Its value is in helping to capture and retrieve context that might otherwise depend on one person remembering what happened six months ago.

Used effectively, it becomes another layer of organisational memory.

Managed Services should complement your people

Adopting Managed Services does not mean replacing the internal cloud team.

In many organisations, the strongest model is a partnership.

Internal teams retain strategic ownership, business context and control over priorities. The Managed Services provider contributes operational coverage, specialist expertise, governance, tooling and continuity.

This allows internal cloud specialists to spend less time acting as the permanent emergency contact and more time on architecture, optimisation, innovation and business improvement.

It also gives them permission to take leave without worrying that the environment will become unmanageable in their absence.

That is not just better for the individual. It is better for retention, morale and long-term organisational resilience.

A question every leadership team should ask

The real question is not whether your cloud team is capable.

It is whether your cloud operating model can function without them.

Could the business manage a serious incident tomorrow if its most experienced cloud engineer was on a flight, unwell or no longer employed by the organisation?

Would the wider team know where to look, who to contact and what actions to take?

Would your processes support them, or would they immediately reach for the absent person’s phone number?

If business continuity depends on one person answering that call, the organisation is exposed.

Build resilience before it is tested

Key-person risk rarely feels urgent while the key person is available.

That is why it is so often ignored.

But holidays, sickness and resignations are not unusual events. They are normal parts of running a business.

A resilient cloud operating model should be designed with that reality in mind.

Managed Services provides a way to distribute knowledge, strengthen governance, improve coverage and reduce the pressure placed on individual employees.

Because the measure of a mature cloud operation is not how well it performs when everyone is available.

It is how well it continues when they are not.

 

Is Your AWS Environment Dependent On One Person?

Most organisations don't realise how much operational risk is concentrated in a handful of individuals until they're unavailable.

A Cloud Governance Snapshot helps you identify:

✓ Key-person dependencies

✓ Governance gaps

✓ Operational risks

✓ Areas where knowledge and ownership need strengthening

✓ Opportunities to improve resilience

Get Your Governance Snapshot Within 48 Hours

 

No cost. No obligation. No downside.

 

 

 

Share: