Skip to content
Managed cloud operations on Datadog

One engineer covers 21% of the year.
Production runs the other 79%.

A year is 8,760 hours. One full time engineer covers about 1,856 of them. That arithmetic is why continuous cover starts at five people, and why the dedicated hire never quite gets justified.

Critical Support is the alternative. We operate your AWS and Azure environment, own incidents 24x7 with a contractual 15 minute response for SEV-1 and SEV-2, and we do it inside the Datadog you already pay for. No migration onto monitoring of ours. It stays your environment, with your access.

Critical Support, running on the Datadog you already own

Most cloud MSPs bring their own monitoring and expect you to move onto it. The telemetry then belongs to the provider, the view is whatever they choose to expose, and the Datadog work you have already paid for is written off. We took the other route and built the managed service on Datadog itself.

15 min
SEV-1 and SEV-2 response, contractual
24x7x365
Incident ownership, not office hours
40 to 60%
Lower MTTR in production
AWS and Azure
The whole environment, not just the tooling
01

Two problems arrive together

The first is the rota. Somebody is holding the pager tonight, and it is usually a developer who was never hired to carry one. The team is too small to staff a proper rotation, so cover becomes goodwill, and goodwill has a half life.

A second hire does not fix it. Two people cannot cover a year of nights, weekends, leave and notice periods between them. The arithmetic runs out long before the enthusiasm does.

The second problem shows up the moment an MSP is brought in to solve the first. Most of them arrive with their own monitoring stack and a migration plan attached. The Datadog instrumentation, the dashboards, the alert tuning, the tribal knowledge of which signal actually means something: all of it gets rebuilt on somebody else's platform, and the view of production becomes whatever the provider decides to show.

So the choice looks like sacrificing the observability investment in order to get the cover. It is a false choice, and it exists because most managed services were never built on a platform the customer could own.

02

What continuous cover costs to build

The headcount is not a matter of opinion. A year contains 8,760 hours. One full time engineer, after holiday and public holidays, covers roughly 1,856 of them. Everything else follows from those two numbers.

Rota calculator

Set the cover you need and the salary you would pay. The result is what the rota costs to run in house.

£80,000
£45k £130k
Engineers needed
5
To cover the hours, before leave
Annual cost
£520,000
Salary plus 30% employment on costs
One hire covers
21%
Of the hours required

An estimate, and a generous one. It assumes 1,856 productive hours per engineer per year, a flat 30% for employer national insurance, pension and benefits, and nothing at all for recruitment, tooling, management time, training, or the cost of the rota being one resignation away from collapse. It excludes Datadog licensing, which is billed by Datadog on usage. Salary is whatever you set it to, not a figure we are asserting.

Critical Support delivers that cover at around 75%* less than building it, at the standard of an in house 24/7 team. A scoping call gets to the figure for a real environment.

03

What we actually run

Critical Support operates the cloud environment itself, not just the monitoring on top of it. Infrastructure provisioning through Terraform and landing zone standards, access control and security posture, observability, cost, and the incident cover across all of it. Your team approves changes and keeps IAM and admin control. We implement, operate and improve.

Incidents, end to end

Every incident runs the same five stages, whatever time it fires.

  • Monitoring. Datadog telemetry, alerting and noise reduction, with synthetic monitors and anomaly detection to catch problems before a customer reports them.
  • Triage. Severity classification from SEV-1 to SEV-4, blast radius assessment and ownership assignment.
  • Response. Runbooks, safe workarounds and rollback procedures run by on-call engineers, with notification inside the contracted window.
  • Escalation. On-call routing, cloud provider escalation, vendor coordination and customer communication, tracked in Datadog Incident Management.
  • Recovery and review. Fix or rollback with validation, a blameless postmortem on every SEV-1, and the actions fed back into the engineering backlog.

Response time is a firm contractual commitment: 15 minutes to first engineer contact for SEV-1 and SEV-2, 24x7x365. Recovery time is a target, because recovery depends on the nature of the incident, not just our speed. Covered incidents include service outages, performance degradation, operational triage and containment of security alerts, integration and API failures, and cloud provider incidents affecting the environment.

The work between incidents

Incident cover on its own leaves the estate exactly as it was. Improvement engineering runs every month across six pillars: reliability and resilience, security and compliance, cost and FinOps, performance and scalability, automation, and governance and observability. This is the work that never has a deadline attached and therefore never gets done: the monitors nobody trusts, the runbooks that were never written, the permissions nobody has reviewed, the costs nobody owns.

Runbooks, standards and knowledge stay in your environment. A single resignation should not remove the ability to run production, which is the failure mode a one person rota is built on.

What procurement will ask for

Security review is usually where a new supplier stalls, so the paperwork exists before anybody asks. The data processing agreement is embedded in the standard MSA rather than available on request. Data residency is UK, with log retention aligned to your policy. We hold ISO 27001 and Cyber Essentials Plus and are registered on the NHS DSPT, and the security questionnaire responses and evidence packs are already written.

Your Datadog, operated

Every Critical Support customer keeps direct access to their own Datadog environment: infrastructure, APM, logs, traces, security signals, cloud cost and LLM monitoring. The environment our engineers work from is the environment you see. There is no proprietary layer in between and nothing is abstracted away from you.

The boundary

We operate the stack. You own the product. We never touch your application code, your model or your business logic. Your team keeps IAM and admin control and signs off material changes. Agents own the analysis. Humans own the outcome.

04

Why this claim is checkable

Critical Cloud reduces mean time to resolve incidents by 40 to 60 per cent. The reason the managed service can be built on Datadog rather than bolted beside it is that Datadog audited it and said so: Critical Cloud is the world's first Powered by Datadog accredited managed service provider, and a Datadog Advanced Partner. Datadog certified one managed provider globally. We are it.

The accreditation is not a logo for a reseller agreement. It requires Advanced status plus a technical review by Datadog engineering covering architecture, onboarding, governance maturity and live customer implementations.

Powered by Datadog

The world's first accredited MSP. No other provider in EMEA holds it.

ISO 27001 and Cyber Essentials Plus

Independently audited, not self assessed. Evidence packs available for procurement.

UK and Ireland engineering

Cardiff, London and Dublin. AWS Partner and Microsoft Partner.

05

Three ways to buy it

The full service is not right for everybody, and saying so is cheaper for both of us than discovering it in month three.

Critical Support The full managed cloud service. 24x7 incident ownership across AWS and Azure with a contractual 15 minute SEV-1 and SEV-2 response, plus monthly improvement engineering across all six pillars. Three plans: Core, Standard and Advanced, scoped to the environment.
Critical Support Lite Right sized for earlier stage or single cloud products where daytime or out of hours cover is sufficient and enterprise SLA commitments are limited. Steps up to the full service as the platform grows.
Critical Response An incident response retainer. Rapid detection, clear escalation and fast recovery, without the monthly improvement engineering. For teams who have the platform capability and need the cover.
06

How this differs from Managed Datadog

Critical Support is the cloud managed service, and it is the service this page is about. It runs the cloud environment and uses Datadog to do it. Managed Datadog is a separate service that manages the Datadog platform itself. The two get confused often enough to be worth separating plainly.

Critical Support The service on this page. Runs the cloud environment on AWS and Azure and owns incidents 24x7, using Datadog as the observability platform to do it. The right service when nobody is holding the pager. It is the service the Powered by Datadog accreditation recognises.
Managed Datadog A different service. Manages the Datadog platform itself, every month: monitors, dashboards, tagging standards, signal quality, SLOs and cost. The customer carries on running their own cloud. The right service when the cloud operations are already covered and it is Datadog that has drifted.
Both together Common enough. The cloud is operated under Critical Support and the Datadog platform is kept clean under Managed Datadog. They are scoped and bought separately.

Who this is not for

Teams with a real platform function already staffed for out of hours do not need us to hold the pager, and will get more from Catalyst or a HealthScan than from a managed service. Teams looking for a provider to take the product roadmap as well as the operations will find the boundary frustrating: we operate the stack, and the product stays yours.

Teams who took this route

OPX

A lean DevOps team migrating from Heroku to Azure, with Datadog first observability and proactive monitoring that reduced noise.

Read the OPX case study

CETA

A Datadog trial turned into a production rollout through FETCH and then HyperCare, with a small technology team throughout.

Read the CETA case study

Zero

Strengthening Azure for fintech scale, where the regulatory floor and the engineering floor are the same floor.

Read the Zero case study

Delivered by Critical Cloud, a Powered by Datadog accredited managed provider and Datadog Advanced Partner. Datadog certified one managed provider globally. We are it. ISO 27001 and Cyber Essentials Plus certified.

Common questions

Is this the same as Managed Datadog?

No. Managed Datadog is the recurring management of the Datadog platform itself: monitors, dashboards, tagging standards, signal quality and cost. Critical Support is the managed cloud service. It operates the AWS and Azure environment and owns incidents 24x7, and it uses Datadog as the observability platform to do that. They are separate services with separate scopes. Some customers buy one, some buy both.

Do we have to move off our own Datadog account?

No. It stays your account, your data and your access. We operate inside the environment you already own rather than migrating you onto proprietary monitoring of ours. If the commercial relationship ends, the environment and its history remain yours.

How many engineers does continuous cover actually take?

A year is 8,760 hours. One full time engineer covers about 1,856 of them, which is around 21 per cent of the year. Continuous cover therefore starts at five people, before leave, training, notice periods and attrition are counted.

We are too small for a full managed service. What then?

Critical Support Lite is the right sized version for earlier stage or single cloud products where daytime or out of hours cover is sufficient. Critical Response is an incident response retainer without the monthly improvement engineering. Both step up to full Critical Support as the platform and the commercial requirements grow.

Does an AI agent run our incidents?

No. Agents own the analysis. Humans own the outcome. Diagnostic tooling assists triage and speeds up investigation, and a Critical Cloud engineer approves every production, security and cost change.

Talk to us

Hear it from your monitoring, not your customers.

A scoping call covers the environment, the cover that exists today, and what the current arrangement is costing. It takes half an hour.

  • We come back within one working day.
  • An engineer is on the first call, not just an account manager.
  • Taking over from an existing provider is a normal starting point, including the disorderly kind.

Prefer not to use a form? Email hello@criticalcloud.ai or call +44 (0)204 538 1116.

* The comparison is against a four person in house team providing continuous cover, costed at a median UK engineer salary plus employer national insurance, pension and benefits, and set against the annual cost of Critical Support at the tier that delivers a comparable standard of cover on a committed term. Four is the conservative end of the comparison: the arithmetic above shows that covering 8,760 hours at 1,856 hours per engineer needs five. Tier and commitment term both move the figure, and a pay as you go term moves it furthest, which is why we scope it rather than publish it.