Datadog for UK financial services
observability that produces evidence.
Regulated firms need observability that satisfies an auditor as well as an on-call engineer. This is how we configure and operate Datadog to do both: coverage mapped to important business services, retention matched to policy, and an incident record you can defend.
Regulated firms need observability to produce evidence, not just alerts
In most businesses observability exists to keep the service up. In a regulated firm it also has to demonstrate that you were in control.
An unregulated engineering team can run a monitoring stack that only its own engineers understand, because the only audience is its own engineers. A regulated firm has three audiences: the on-call engineer, the internal risk function, and eventually a supervisor or auditor asking what happened and when you knew.
That third audience changes the design. Monitor coverage has to be traceable to the services that matter to customers, not just the ones that page most often. Retention has to satisfy a policy rather than a preference. Access has to be attributable. Configuration change has to leave a trail. And the answer to what was the impact needs to be reconstructable months later, not just visible on a live dashboard.
This page is about how we configure and operate Datadog to meet that standard. If you are looking for the wider service and the commercial picture, our financial services and fintech cloud operations page covers it.
We operate and evidence the runtime controls that support your compliance obligations. We are not a regulatory advisory firm and we do not interpret regulation on your behalf.
The Datadog configuration decisions that matter in a regulated estate
Five areas where a regulated deployment diverges from a standard one.
Coverage Monitors mapped to important business services Coverage traceable to the services that matter to customers, not the ones that happen to be noisy.
Operational resilience frameworks turn on the idea of important business services and the tolerance for disruption to each. Observability that cannot be mapped back to those services cannot evidence anything about them.
We build the monitor set from that mapping outwards: for each important business service, the components it depends on, the signals that indicate degradation, and the alert that fires before the tolerance is breached rather than after. The output is a coverage matrix you can hand to a risk function, not just a folder of monitors.
Evidence Retention and evidence export Retention configured to your policy, with the ability to reconstruct an incident months later.
Datadog retention defaults are set for convenience, not for regulated retention policy. We configure indexes, retention tiers and archive destinations to match what your policy actually requires, which is often longer for some data classes and considerably shorter for others.
Log archives to your own cloud storage give you a durable, cheap record independent of the indexed tier, which is usually the right answer for data you must retain but rarely query.
Access Access control and attribution SSO, role mapping and an audit trail that shows who changed what.
Access is federated through your identity provider rather than local Datadog accounts, with role mapping that reflects the separation of duties your risk function expects. Read-only, editor and admin boundaries are set deliberately.
Datadog audit trail records configuration change, so a monitor that was silenced or a retention setting that was reduced is attributable. In a regulated context this matters as much as the telemetry itself.
Incident Incident timeline and post-incident record A defensible record of detection, escalation and resolution.
The question after a material incident is rarely only what broke. It is when you detected it, who was told, what you did and how long each step took. We configure incident management so that record is produced as a by-product of responding, rather than reconstructed afterwards from chat logs.
Our 24x7 incident response service, Critical Support, is where that record is generated in practice, with a named team and defined escalation.
Residency Site region and data residency A decision made before ingestion, not after.
Datadog has historically operated a US site and an EU1 site, with EU1 storing data in Frankfurt, and has been expanding its data residency options. For firms with an explicit UK-only mandate this needs settling before any data is ingested, because changing site region later means re-instrumenting and losing history.
Our guide to Datadog UK data residency, DORA and NIS2 covers the current position in more detail.
How we start with a regulated firm
Usually with an assessment, because the estate almost always predates the requirement.
Most regulated firms we work with already have Datadog, or already have something. The problem is rarely a blank page. It is a deployment that grew service by service, has coverage nobody has mapped, retention nobody has checked against policy, and a monitor set that pages too often to be trusted.
HealthScan is the usual first step: an independent, read-only assessment producing a health scorecard and a prioritised improvement backlog. Read-only matters here, because it means the assessment can happen without a change request.
From there the work is either a project, through LaunchPad or Catalyst, or ongoing operation through Managed Observability and Critical Support. What we do not do is recommend replacing a platform because it is misconfigured. That is an expensive way to solve a configuration problem.
Financial services and fintech cloud operations
The full picture: what we operate for regulated financial platforms, and the commercial and delivery model behind it.
See the full service →Frequently asked questions
Direct answers to the questions we are asked most often.
Q Does Critical Cloud make our firm compliant?
No, and any provider claiming otherwise should be treated with caution. Compliance is your obligation and depends on far more than your observability stack. What we do is operate and evidence the runtime controls that support your compliance obligations: monitor coverage mapped to important business services, retention configured to your policy, attributable access, and a defensible incident record. We are not a regulatory advisory firm and we do not interpret regulation on your behalf.
Q How does observability relate to operational resilience requirements?
Operational resilience frameworks are built around important business services and the tolerance for disruption to each. Demonstrating you can stay within a tolerance requires knowing when a service is degrading, which is an observability question. In practice that means monitor coverage traceable to each important business service, alerting that fires before a tolerance is breached rather than after, and a record that lets you reconstruct what happened. Configuring Datadog to produce that is the substance of this work.
Q Can Datadog data be kept in the UK?
Datadog has historically operated a US site and an EU1 site, with EU1 storing data in Frankfurt, and has been expanding its data residency options. For many UK financial services firms EU1 is acceptable. Where there is an explicit UK-only data residency mandate, the site region must be decided before ingestion begins, because moving it later means re-instrumenting and losing historical data. This is one of the first questions we settle.
Q We already have Datadog but it pages constantly. Where do we start?
With an assessment rather than a rebuild. HealthScan is a read-only review of the existing deployment that produces a health scorecard and a prioritised backlog, so you can see what is causing the noise and what to fix first. Read-only means it needs no change to your environment. Alert fatigue in a regulated firm is a control weakness as much as an engineering annoyance, because a team that has learned to ignore alerts will ignore the one that mattered.
Observability your risk function can rely on
If your Datadog deployment cannot currently answer what broke, when you knew and what the impact was, that is a fixable problem. Usually a configuration one.