Skip to content
Assess, Cloud HealthScan

Independent AWS and Azure health check
that tells you what to fix first, with evidence.

Nobody has reviewed the account properly since it was built. Permissions accumulated, backups were configured once and never tested, and the bill goes up every quarter without anyone able to say which service is driving it. Everyone involved suspects the platform has drifted. What is missing is the evidence to say where it drifted, what the risk actually is, and which fixes are worth doing first.

Read the full context

Cloud HealthScan is an automated, read-only assessment of your AWS or Azure environment, measured against the AWS and Azure Well-Architected Frameworks and Critical Cloud operational best practice. It diagnoses, reports and recommends remediations. Nothing is changed in your environment, and you do not need Datadog or any other observability platform to have one.

The outcome

Know what it takes to run your cloud properly, before you commit budget.

A health scorecard by domain, a prioritised remediation plan with impact and effort ratings, and findings you can put in front of stakeholders, all evidenced from your own AWS or Azure account.

The boundary

Read-only. We diagnose, we do not change anything.

Cloud HealthScan reports and recommends. Doing the remediation work, and operating the platform afterwards, is Critical Support. Datadog is neither required nor assumed.

Read-only
No changes to your environment
Automated
Diagnose, report, recommend
AWS + Azure
Well-Architected Frameworks
Independent
Evidence, not sales collateral
Quick facts
ScopeAWS, Azure or both
AccessRead-only access to the cloud account
Measured againstAWS and Azure Well-Architected Frameworks, plus Critical Cloud operational best practice
DatadogNot required
Best whenYou need evidence of operational risk before committing budget or choosing a provider
Natural path forward
  • Critical Support: to deliver the remediations and operate the platform afterwards
  • Datadog HealthScan: if what you want assessed is an existing Datadog environment
01
The problem

You suspect the platform has drifted. You cannot prove it or prioritise it.

Cloud environments do not fail all at once. They accumulate. A permission granted for one migration and never revoked. A backup policy that was correct for the architecture two rewrites ago. An alert that fires every night and gets muted rather than fixed. Each one is defensible on its own, and together they are the reason an ordinary Tuesday becomes an incident.

The obstacle is rarely willingness. It is that thirty known problems all look about the same size written down, and nothing on the list says which one takes the platform down first. So the list stays a list.

A Cloud HealthScan is how you get the evidence. It assesses what you actually run against published vendor standards and against how a 24×7 operations team expects a platform to behave, then ranks what it finds by impact and effort. The output is a decision you can defend, not a list of opinions.

02
What we assess

Vendor best practice, plus the operational standards the frameworks leave to you

The AWS and Azure Well-Architected Frameworks are the industry reference for how a cloud platform should be built. They are deliberately silent on how it should be run day to day. We assess against both.

Well-Architected pillars
Operational excellence
Change process, automation coverage, infrastructure as code, drift between what is declared and what is deployed.
Security
Identity and permission sprawl, least privilege, network exposure, secrets handling, logging and audit coverage.
Reliability
Failure domains, redundancy, backup configuration and whether recovery has ever actually been tested.
Performance efficiency
Resource sizing against real usage, scaling behaviour, and the bottlenecks that only appear under load.
Cost optimisation
Unused and oversized resources, commitment coverage, and which workloads are driving the trend in the bill.
Sustainability
Workload efficiency and the resource footprint of how the platform is currently configured.
Critical Cloud operational best practice
Incident readiness
Whether a SEV-1 at 3am has an owner, a path and a runbook, or whether it has a group chat.
Signal quality
Whether monitoring would tell you first, and whether the alerts that fire are ones anyone still reads.
Access governance
Who can reach production, how that is granted and reviewed, and what is left behind when someone leaves.
Recovery testing
The difference between a backup that exists and a restore that has been proven to work.

We operate AWS and Azure platforms 24×7 for customers in regulated sectors, and we hold ISO 27001 and Cyber Essentials Plus. The operational half of this assessment is drawn from that, not from a checklist.

03
Deliverables

Findings you can act on and present

Everything below is written to survive being forwarded. The engineer who has to do the work and the person who has to approve the budget can both read the same document.

Health scorecard
Green, amber or red for every domain assessed, so the shape of the problem is visible in one page.
Executive summary
What the material risks are and what they mean commercially, in language that does not require an engineer to translate it.
Environment snapshot
What is actually running, where, and how it is configured. Frequently the first accurate inventory a team has had.
Prioritised remediation plan
Every finding rated by impact and effort, with the evidence behind it, ordered so the first item is the one worth doing first.
Findings session
A walkthrough with the engineers who ran the assessment, where you can push back on any conclusion in it.

The remediation plan is yours. Act on it in-house, hand it to your current provider, or ask us to deliver it. A Cloud HealthScan that ends with you fixing four things yourself is a Cloud HealthScan that worked.

04
How it runs

Diagnose, report, recommend

The collection and analysis are automated, which is what keeps the assessment consistent and quick to start. The judgement about what matters is not.

01

Diagnose

You grant read-only access. Our tooling collects configuration, telemetry and usage across the accounts in scope and assesses them against every domain above.

02

Report

Findings are scored and evidenced. Our engineers review the output, discard what is noise in your context, and write the summary.

03

Recommend

You get the remediation plan, ranked by impact and effort, and a session to work through it. What happens next is your call.

Automation applies to the collection and the analysis. The order you act in is set by engineers, and so is the remediation work.

05
Evidence

What this looks like on a real platform

Zero

A Microsoft Azure Well-Architected Review for a fintech, run through Azure Lighthouse. We assessed the environment, evidenced what we found, and delivered a roadmap ranked by impact and effort. The company decided what to do with it, where the regulatory floor and the engineering floor are the same floor.

Read the Zero case study →

FAQ

Do we need Datadog to have a Cloud HealthScan?

No. Cloud HealthScan assesses the AWS or Azure platform itself, not a Datadog environment. It is designed for teams who have not chosen an observability platform yet. If you are already running Datadog and want that environment reviewed instead, the Datadog HealthScan is the right service.

Does Critical Cloud change anything during a Cloud HealthScan?

No. The assessment is read-only. We collect configuration and telemetry, assess it, and deliver findings. Nothing is configured, modified or deployed, and your environment is unchanged at the end of it.

What is the assessment measured against?

Two things. First, vendor best practice: the AWS Well-Architected Framework and the Azure Well-Architected Framework. Second, Critical Cloud operational best practice, which covers the SRE and run-the-platform standards the frameworks leave to you, such as incident readiness, alert quality, on-call structure and recovery testing.

What do we get at the end?

A health scorecard by domain, an executive summary, an environment snapshot and a prioritised remediation plan with impact and effort ratings, plus a findings session with our engineers. The remediation plan is written so you can act on it yourself, hand it to your current provider, or ask us to deliver it.

Does a Cloud HealthScan commit us to anything?

No. It is an assessment, not an onboarding step. Some findings show the platform is in better shape than expected. Where it does surface material risk, you decide whether to fix it yourself or ask us to operate the platform through Critical Support.

What comes after a Cloud HealthScan

Critical Support

Hand the platform to an accountable team: 24×7 incident cover with a contractual 15-minute SEV-1 response, and improvement engineering every month that works through findings like these.

Critical Support →

Critical Response

Your engineers keep running the platform day to day. We own incident response in the coverage window you choose.

Critical Response →

Observability, powered by Datadog

Where the assessment shows you cannot see enough to operate safely, Datadog is usually the answer. We are the world's first Powered by Datadog accredited MSP.

Datadog services →

Delivered by Critical Cloud, a Powered by Datadog accredited managed provider and Datadog Advanced Partner. Datadog certified one managed provider globally. We are it. ISO 27001 and Cyber Essentials Plus certified.

Book a Cloud HealthScan

Share a few details about your environment and we will come back with scheduling and scoping. Read-only access, no disruption to your team, no Datadog required.

  • Independent. Findings and evidence, not a sales document.
  • Safe. Read-only access throughout, and nothing changes in your environment.
  • Yours. The remediation plan is yours to act on, whoever ends up doing the work.

Need an independent view of the cloud you already run?

Start with the evidence. Decide what to do with it afterwards.