Independent Datadog audit and health check
that tells you what's actually wrong, with evidence.
The alerts stopped meaning anything months ago. Nobody opens the dashboards. The bill goes up every quarter and no single person can say which product is driving it. Everyone involved knows the environment has drifted. What is missing is the evidence to say where it drifted, how much it is costing, and which fixes are worth doing first.
Read the full context
HealthScan is an independent, read-only Datadog audit of your existing environment: a health scorecard, a prioritised improvement backlog, and stakeholder-ready findings, delivered in 1-2 weeks. Nothing is changed in your environment.
Know where your observability is failing you, and what to fix first.
A health scorecard, a prioritised improvement backlog and stakeholder-ready findings, all evidenced from your own environment.
Read-only. We diagnose, we do not change anything.
Implementation is Catalyst or an Accelerator. To assess the AWS and Azure platform rather than the Datadog environment, see the Cloud HealthScan.
- Catalyst: to deliver the improvements HealthScan identifies
- Managed Datadog: for ongoing platform management
You know something's wrong but you can't prove it or prioritise it
Engineering teams accumulate a general sense that their Datadog environment has problems, but general sense doesn't move budget. The specific symptoms are observable but unquantified. The fixes are obviously needed but unordered. The case for investment hasn't been made.
- Alert noise is high enough that engineers have started ignoring pages, but nobody has audited the monitors or measured the false-positive rate
- Dashboard usage is low, but nobody has mapped which views are actually used versus which were created and never touched again
- Datadog costs have grown faster than coverage, but the cost profile hasn't been reviewed against what's actually generating signal
- Ownership mapping has drifted as the team changed, services route alerts to the wrong people or nowhere at all
- An improvement initiative has been proposed, but leadership needs evidence before approving the budget
- A new CTO or platform lead wants an independent baseline before committing to a direction
Alert noise is usually the first symptom
Monitors get added after incidents and rarely get removed. Thresholds set for last year's traffic keep firing at this year's. Each one was reasonable on the day it was created. Together they produce a pager that goes off often enough for on-call engineers to stop reading it, and an alerts channel everyone has muted.
A HealthScan audits the monitors behind that noise: the ones that fire and are never acted on, the ones that have never fired at all, the ones that route to nobody, and the important services that would not page anyone if they degraded. Each finding goes into the backlog with a recommended action and an impact and effort rating, so the noise gets fixed in the order that matters.
Every domain of the Datadog deployment
HealthScan covers eight domains of a Datadog environment. Each is assessed against what good looks like and scored green, amber, or red.
Findings you can act on and present
HealthScan closes with a findings readout and a complete deliverable package, structured for both the engineering team and a leadership audience.
What’s not included
HealthScan is an assessment and roadmap. Implementation changes can be delivered separately via HyperCare or Managed Datadog.
Sample report format
What the deliverable looks like. The report opens with an executive summary and scorecard, then works through findings and a prioritised roadmap.
Evidence, not opinion
Your platform is stable. Somebody is making it look that way. A HealthScan shows you what your environment is quietly relying on: the monitors that would page nobody, the runbooks that were never written, the costs nobody owns. The gap we find most often is the document that matters most: what to do when the serious alert fires.
Every finding is evidenced from your own telemetry, prioritised, and mapped to an action your engineers can pick up. If the environment is in good shape, the report says that too.
How to assess your Datadog environment yourself
You do not need us to start. If you want a view of where your Datadog deployment stands before talking to anyone, these are the questions we work through, in the order we work through them. They are also a reasonable way to decide whether an independent assessment is worth your time.
Coverage
Pick your three most important services. For each, can you name the monitor that would fire first if it degraded? If not, coverage is the gap, not tooling.
Tag consistency
Can you group cost, latency and error rate by the same service name? Where tags disagree across infrastructure, APM and logs, every cross-cutting question becomes manual work.
Alert quality
What proportion of pages in the last month resulted in an action? Two more measurable versions: how many of your monitors have never fired, and what share of your alerts comes from your top ten monitors. A team that has learned to ignore alerts will also ignore the one that mattered.
Runbooks
Open your five most important monitors: is a runbook linked from any of them? A monitor without one hands the responder a puzzle at 3am instead of a procedure.
Cost drivers
Can you name your top three cost drivers by product line? If not, start with log indexing, custom metric cardinality and APM host count. The most common indexing strategy we find is no strategy at all: everything ingested gets indexed. On cardinality, Datadog’s own guidance on custom metrics governance notes that customers using Metrics without Limits on unqueried metrics often see up to a 70% reduction in custom metrics usage without losing critical visibility.
Commitment
When did anyone last put your usage page next to your contract? Usage and commitment drift apart quietly, and renewal is the most expensive place to find out.
Ownership
Who owns a monitor when it fires at 3am, and who owns the dashboard nobody has opened in six months? Unowned configuration is where drift accumulates.
Retention and access
Does retention match your policy rather than the default, and is access federated through your identity provider? Both tend to be set once and never revisited.
Working through these usually surfaces plenty on its own. Where a HealthScan adds value is the weighted scorecard, the comparison against how other environments are configured, and a prioritised backlog you can put in front of a budget holder. See also Datadog pricing and cost optimisation if cost is the immediate concern.
FAQ
Questions about HealthScan.
Does Critical Cloud make any changes during a HealthScan?
No. HealthScan is entirely read-only. We review your environment and deliver findings, nothing is configured, modified, or deployed. Your environment is unchanged at the end of the engagement.
What does the HealthScan deliverable look like?
A health scorecard (green/amber/red by domain), an executive summary, an environment snapshot, and a prioritised improvement backlog with impact/effort ratings. Two included sessions: kickoff and findings readout with Q&A.
How new does our Datadog environment need to be for HealthScan to be useful?
HealthScan is most valuable for environments that have been running for at least six months, enough time for patterns of debt to emerge. For very new environments, HyperCare or LaunchPad are typically the better starting point.
Does HealthScan always lead to Catalyst?
Not necessarily. Sometimes findings show the environment is in better shape than expected. Sometimes they surface the need for Managed Datadog rather than a one-time fix. HealthScan gives you an honest picture, what happens next depends on what it finds.
What comes after HealthScan
The natural paths from a completed assessment.
Catalyst
The natural follow-on to HealthScan. Catalyst delivers the improvements the assessment identifies, practitioners who complete the approved backlog, not consultants who hand it back.
Managed Datadog
If HealthScan reveals the environment needs ongoing management rather than a one-time fix, Managed Datadog is the recurring engagement that keeps it clean as the platform grows.
UK Datadog partner
HealthScan is delivered by the world's first Powered by Datadog accredited MSP. The assessors are the same engineers who run Datadog in production every day.
Book a HealthScan
Share a few details about your environment and we will come back with scheduling and scoping. Read-only access, 1-2 weeks, no disruption to your team.
- Independent. Findings and evidence, not a sales document.
- Fast. A clear picture of what's working and what isn't, in 1-2 weeks.
- Safe. Read-only access throughout.
Need an independent view of your Datadog environment?
HealthScan delivers a clear picture of what's working and what isn't, in 1-2 weeks, read-only, with no disruption to your team or environment.