One engineer covers 21% of the year.
Production runs the other 79%.
A year is 8,760 hours. One full time engineer covers about 1,856 of them. That arithmetic is why continuous cover starts at five people, and why the dedicated hire never quite gets justified.
Critical Support is the alternative. We operate your AWS and Azure environment, own incidents 24x7 with a contractual 15 minute response for SEV-1 and SEV-2, and we do it inside the Datadog you already pay for. No migration onto monitoring of ours. It stays your environment, with your access.
Critical Support, running on the Datadog you already own
Most cloud MSPs bring their own monitoring and expect you to move onto it. The telemetry then belongs to the provider, the view is whatever they choose to expose, and the Datadog work you have already paid for is written off. We took the other route and built the managed service on Datadog itself.
Two problems arrive together
The first is the rota. Somebody is holding the pager tonight, and it is usually a developer who was never hired to carry one. The team is too small to staff a proper rotation, so cover becomes goodwill, and goodwill has a half life.
A second hire does not fix it. Two people cannot cover a year of nights, weekends, leave and notice periods between them. The arithmetic runs out long before the enthusiasm does.
The second problem shows up the moment an MSP is brought in to solve the first. Most of them arrive with their own monitoring stack and a migration plan attached. The Datadog instrumentation, the dashboards, the alert tuning, the tribal knowledge of which signal actually means something: all of it gets rebuilt on somebody else's platform, and the view of production becomes whatever the provider decides to show.
So the choice looks like sacrificing the observability investment in order to get the cover. It is a false choice, and it exists because most managed services were never built on a platform the customer could own.
What continuous cover costs to build
The headcount is not a matter of opinion. A year contains 8,760 hours. One full time engineer, after holiday and public holidays, covers roughly 1,856 of them. Everything else follows from those two numbers.
Set the cover you need and the salary you would pay. The result is what the rota costs to run in house.
An estimate, and a generous one. It assumes 1,856 productive hours per engineer per year, a flat 30% for employer national insurance, pension and benefits, and nothing at all for recruitment, tooling, management time, training, or the cost of the rota being one resignation away from collapse. It excludes Datadog licensing, which is billed by Datadog on usage. Salary is whatever you set it to, not a figure we are asserting.
Critical Support delivers that cover at around 75%* less than building it, at the standard of an in house 24/7 team. A scoping call gets to the figure for a real environment.
What we actually run
Critical Support operates the cloud environment itself, not just the monitoring on top of it. Infrastructure provisioning through Terraform and landing zone standards, access control and security posture, observability, cost, and the incident cover across all of it. Your team approves changes and keeps IAM and admin control. We implement, operate and improve.
Incidents, end to end
Every incident runs the same five stages, whatever time it fires.
- Monitoring. Datadog telemetry, alerting and noise reduction, with synthetic monitors and anomaly detection to catch problems before a customer reports them.
- Triage. Severity classification from SEV-1 to SEV-4, blast radius assessment and ownership assignment.
- Response. Runbooks, safe workarounds and rollback procedures run by on-call engineers, with notification inside the contracted window.
- Escalation. On-call routing, cloud provider escalation, vendor coordination and customer communication, tracked in Datadog Incident Management.
- Recovery and review. Fix or rollback with validation, a blameless postmortem on every SEV-1, and the actions fed back into the engineering backlog.
Response time is a firm contractual commitment: 15 minutes to first engineer contact for SEV-1 and SEV-2, 24x7x365. Recovery time is a target, because recovery depends on the nature of the incident, not just our speed. Covered incidents include service outages, performance degradation, operational triage and containment of security alerts, integration and API failures, and cloud provider incidents affecting the environment.
The work between incidents
Incident cover on its own leaves the estate exactly as it was. Improvement engineering runs every month across six pillars: reliability and resilience, security and compliance, cost and FinOps, performance and scalability, automation, and governance and observability. This is the work that never has a deadline attached and therefore never gets done: the monitors nobody trusts, the runbooks that were never written, the permissions nobody has reviewed, the costs nobody owns.
Runbooks, standards and knowledge stay in your environment. A single resignation should not remove the ability to run production, which is the failure mode a one person rota is built on.
What procurement will ask for
Security review is usually where a new supplier stalls, so the paperwork exists before anybody asks. The data processing agreement is embedded in the standard MSA rather than available on request. Data residency is UK, with log retention aligned to your policy. We hold ISO 27001 and Cyber Essentials Plus and are registered on the NHS DSPT, and the security questionnaire responses and evidence packs are already written.
Your Datadog, operated
Every Critical Support customer keeps direct access to their own Datadog environment: infrastructure, APM, logs, traces, security signals, cloud cost and LLM monitoring. The environment our engineers work from is the environment you see. There is no proprietary layer in between and nothing is abstracted away from you.
The boundary
We operate the stack. You own the product. We never touch your application code, your model or your business logic. Your team keeps IAM and admin control and signs off material changes. Agents own the analysis. Humans own the outcome.
Why this claim is checkable
Critical Cloud reduces mean time to resolve incidents by 40 to 60 per cent. The reason the managed service can be built on Datadog rather than bolted beside it is that Datadog audited it and said so: Critical Cloud is the world's first Powered by Datadog accredited managed service provider, and a Datadog Advanced Partner. Datadog certified one managed provider globally. We are it.
The accreditation is not a logo for a reseller agreement. It requires Advanced status plus a technical review by Datadog engineering covering architecture, onboarding, governance maturity and live customer implementations.
Powered by Datadog
The world's first accredited MSP. No other provider in EMEA holds it.
ISO 27001 and Cyber Essentials Plus
Independently audited, not self assessed. Evidence packs available for procurement.
UK and Ireland engineering
Cardiff, London and Dublin. AWS Partner and Microsoft Partner.
Three ways to buy it
The full service is not right for everybody, and saying so is cheaper for both of us than discovering it in month three.
How this differs from Managed Datadog
Critical Support is the cloud managed service, and it is the service this page is about. It runs the cloud environment and uses Datadog to do it. Managed Datadog is a separate service that manages the Datadog platform itself. The two get confused often enough to be worth separating plainly.
Who this is not for
Teams with a real platform function already staffed for out of hours do not need us to hold the pager, and will get more from Catalyst or a HealthScan than from a managed service. Teams looking for a provider to take the product roadmap as well as the operations will find the boundary frustrating: we operate the stack, and the product stays yours.
Teams who took this route
OPX
A lean DevOps team migrating from Heroku to Azure, with Datadog first observability and proactive monitoring that reduced noise.
CETA
A Datadog trial turned into a production rollout through FETCH and then HyperCare, with a small technology team throughout.
Zero
Strengthening Azure for fintech scale, where the regulatory floor and the engineering floor are the same floor.
Common questions
Is this the same as Managed Datadog?
No. Managed Datadog is the recurring management of the Datadog platform itself: monitors, dashboards, tagging standards, signal quality and cost. Critical Support is the managed cloud service. It operates the AWS and Azure environment and owns incidents 24x7, and it uses Datadog as the observability platform to do that. They are separate services with separate scopes. Some customers buy one, some buy both.
Do we have to move off our own Datadog account?
No. It stays your account, your data and your access. We operate inside the environment you already own rather than migrating you onto proprietary monitoring of ours. If the commercial relationship ends, the environment and its history remain yours.
How many engineers does continuous cover actually take?
A year is 8,760 hours. One full time engineer covers about 1,856 of them, which is around 21 per cent of the year. Continuous cover therefore starts at five people, before leave, training, notice periods and attrition are counted.
We are too small for a full managed service. What then?
Critical Support Lite is the right sized version for earlier stage or single cloud products where daytime or out of hours cover is sufficient. Critical Response is an incident response retainer without the monthly improvement engineering. Both step up to full Critical Support as the platform and the commercial requirements grow.
Does an AI agent run our incidents?
No. Agents own the analysis. Humans own the outcome. Diagnostic tooling assists triage and speeds up investigation, and a Critical Cloud engineer approves every production, security and cost change.
Hear it from your monitoring, not your customers.
A scoping call covers the environment, the cover that exists today, and what the current arrangement is costing. It takes half an hour.
- We come back within one working day.
- An engineer is on the first call, not just an account manager.
- Taking over from an existing provider is a normal starting point, including the disorderly kind.
Prefer not to use a form? Email hello@criticalcloud.ai or call +44 (0)204 538 1116.
* The comparison is against a four person in house team providing continuous cover, costed at a median UK engineer salary plus employer national insurance, pension and benefits, and set against the annual cost of Critical Support at the tier that delivers a comparable standard of cover on a committed term. Four is the conservative end of the comparison: the arithmetic above shows that covering 8,760 hours at 1,856 hours per engineer needs five. Tier and commitment term both move the figure, and a pay as you go term moves it furthest, which is why we scope it rather than publish it.