Independent AWS and Azure health check
that tells you what to fix first, with evidence.
Nobody has reviewed the account properly since it was built. Permissions accumulated, backups were configured once and never tested, and the bill goes up every quarter without anyone able to say which service is driving it. Everyone involved suspects the platform has drifted. What is missing is the evidence to say where it drifted, what the risk actually is, and which fixes are worth doing first.
Read the full context
Cloud HealthScan is an automated, read-only assessment of your AWS or Azure environment, measured against the AWS and Azure Well-Architected Frameworks and Critical Cloud operational best practice. It diagnoses, reports and recommends remediations. Nothing is changed in your environment, and you do not need Datadog or any other observability platform to have one.
Know what it takes to run your cloud properly, before you commit budget.
A health scorecard by domain, a prioritised remediation plan with impact and effort ratings, and findings you can put in front of stakeholders, all evidenced from your own AWS or Azure account.
Read-only. We diagnose, we do not change anything.
Cloud HealthScan reports and recommends. Doing the remediation work, and operating the platform afterwards, is Critical Support. Datadog is neither required nor assumed.
- Critical Support: to deliver the remediations and operate the platform afterwards
- Datadog HealthScan: if what you want assessed is an existing Datadog environment
You suspect the platform has drifted. You cannot prove it or prioritise it.
Cloud environments do not fail all at once. They accumulate. A permission granted for one migration and never revoked. A backup policy that was correct for the architecture two rewrites ago. An alert that fires every night and gets muted rather than fixed. Each one is defensible on its own, and together they are the reason an ordinary Tuesday becomes an incident.
The obstacle is rarely willingness. It is that thirty known problems all look about the same size written down, and nothing on the list says which one takes the platform down first. So the list stays a list.
A Cloud HealthScan is how you get the evidence. It assesses what you actually run against published vendor standards and against how a 24×7 operations team expects a platform to behave, then ranks what it finds by impact and effort. The output is a decision you can defend, not a list of opinions.
Vendor best practice, plus the operational standards the frameworks leave to you
The AWS and Azure Well-Architected Frameworks are the industry reference for how a cloud platform should be built. They are deliberately silent on how it should be run day to day. We assess against both.
We operate AWS and Azure platforms 24×7 for customers in regulated sectors, and we hold ISO 27001 and Cyber Essentials Plus. The operational half of this assessment is drawn from that, not from a checklist.
Findings you can act on and present
Everything below is written to survive being forwarded. The engineer who has to do the work and the person who has to approve the budget can both read the same document.
The remediation plan is yours. Act on it in-house, hand it to your current provider, or ask us to deliver it. A Cloud HealthScan that ends with you fixing four things yourself is a Cloud HealthScan that worked.
Diagnose, report, recommend
The collection and analysis are automated, which is what keeps the assessment consistent and quick to start. The judgement about what matters is not.
Diagnose
You grant read-only access. Our tooling collects configuration, telemetry and usage across the accounts in scope and assesses them against every domain above.
Report
Findings are scored and evidenced. Our engineers review the output, discard what is noise in your context, and write the summary.
Recommend
You get the remediation plan, ranked by impact and effort, and a session to work through it. What happens next is your call.
Automation applies to the collection and the analysis. The order you act in is set by engineers, and so is the remediation work.
What this looks like on a real platform
Zero
A Microsoft Azure Well-Architected Review for a fintech, run through Azure Lighthouse. We assessed the environment, evidenced what we found, and delivered a roadmap ranked by impact and effort. The company decided what to do with it, where the regulatory floor and the engineering floor are the same floor.
FAQ
Do we need Datadog to have a Cloud HealthScan?
No. Cloud HealthScan assesses the AWS or Azure platform itself, not a Datadog environment. It is designed for teams who have not chosen an observability platform yet. If you are already running Datadog and want that environment reviewed instead, the Datadog HealthScan is the right service.
Does Critical Cloud change anything during a Cloud HealthScan?
No. The assessment is read-only. We collect configuration and telemetry, assess it, and deliver findings. Nothing is configured, modified or deployed, and your environment is unchanged at the end of it.
What is the assessment measured against?
Two things. First, vendor best practice: the AWS Well-Architected Framework and the Azure Well-Architected Framework. Second, Critical Cloud operational best practice, which covers the SRE and run-the-platform standards the frameworks leave to you, such as incident readiness, alert quality, on-call structure and recovery testing.
What do we get at the end?
A health scorecard by domain, an executive summary, an environment snapshot and a prioritised remediation plan with impact and effort ratings, plus a findings session with our engineers. The remediation plan is written so you can act on it yourself, hand it to your current provider, or ask us to deliver it.
Does a Cloud HealthScan commit us to anything?
No. It is an assessment, not an onboarding step. Some findings show the platform is in better shape than expected. Where it does surface material risk, you decide whether to fix it yourself or ask us to operate the platform through Critical Support.
What comes after a Cloud HealthScan
Critical Support
Hand the platform to an accountable team: 24×7 incident cover with a contractual 15-minute SEV-1 response, and improvement engineering every month that works through findings like these.
Critical Response
Your engineers keep running the platform day to day. We own incident response in the coverage window you choose.
Observability, powered by Datadog
Where the assessment shows you cannot see enough to operate safely, Datadog is usually the answer. We are the world's first Powered by Datadog accredited MSP.
Book a Cloud HealthScan
Share a few details about your environment and we will come back with scheduling and scoping. Read-only access, no disruption to your team, no Datadog required.
- Independent. Findings and evidence, not a sales document.
- Safe. Read-only access throughout, and nothing changes in your environment.
- Yours. The remediation plan is yours to act on, whoever ends up doing the work.
Need an independent view of the cloud you already run?
Start with the evidence. Decide what to do with it afterwards.