# Critical Cloud, full content > The prose of Critical Cloud's core service and company pages in one file, > for AI agents and answer engines. Generated at build time from the site itself. > Curated menu version: https://criticalcloud.ai/llms.txt Managed Runtime Assurance for AI-era software. Critical Cloud is the world's first "Powered by Datadog" accredited MSP, a Datadog Advanced Partner (Premier is a higher tier in Datadog's standard partner network), ISO/IEC 27001:2022 certified and Cyber Essentials Plus certified. Cardiff (HQ), London and Dublin. --- ## / Source: https://criticalcloud.ai/ Managed Runtime Assurance for AI-era software # Ship AI fast. Stay in control. Critical Cloud operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software. Your team owns the product. We own the operating layer that keeps it reliable, secure, cost-controlled and evidence-ready, powered by Datadog. Talk to us about runtime assurance -> See Critical Support Scroll 01 The operating model ## Managed Runtime Assurance is the operating model behind serious software. Software is becoming faster to create, but production is not becoming easier to operate safely. Every application, AI feature and agentic workflow creates runtime risk: reliability, security, cost, resilience, evidence and human approval. Critical Cloud brings those responsibilities together as one managed outcome. For AI workloads and agentic workflows, we deliver this as managed AI operations . Production stays healthy, incidents are handled, controls are enforced and evidence is ready. What we own - -> Observability and signal quality - -> Incident response and escalation - -> Cloud runtime operations - -> Security operations and access governance - -> Cost control and optimisation - -> Runtime evidence and assurance reporting - -> Human governance for AI-assisted operations What you own - -> Product idea - -> Application code - -> Model behaviour - -> Business logic - -> Customer experience - -> Product roadmap We operate the stack. You own the product. What is Managed Runtime Assurance? -> 02 The boundary ## We operate the stack. You own the product. We operate, secure, and govern the stack your AI runs on. We never touch your app, your model, or your business logic. That boundary is what makes us a trustworthy, impartial layer: we have no agenda over your product, so we can stand behind whether your operations are sound. 15 min Incident response target 24×7 Always-on coverage 200+ Cloud projects delivered Certified ISO 27001 + Cyber Essentials Plus 03 human-in-loop ## Agents own the analysis. Humans own the outcome. Agents are good at the labour: root-cause analysis across telemetry no human can hold in their head, surfacing correlations, drafting fixes. What does not automate is ownership of the outcome. A human evaluates the plan, weighs the context the agent lacks, and stays accountable for accuracy, trust, and compliance. In control of failures Datadog Bits AI detection and remediation, under our governance In control of the attack surface AI Guard and runtime protection, operated by us In control in production Agent Observability and the Agent Console The foundation The full Datadog platform we are accredited to operate Bits AI SRE · investigation Agents investigate, humans approve. Bits AI SRE assembles hypotheses and remediation; our engineers own the call. 04 Cloud · Datadog · AI ## Cloud, Datadog, AI. We operate cloud platforms using Datadog to deliver unified observability across infrastructure, APM, logs, traces, security signals, cloud cost insight, and LLM monitoring. Our services group simply around how we deliver that outcome. Datadog · Cloudcraft observability map One view of the whole estate. Live agent and signal coverage across AWS and Azure: infrastructure, APM, security, and cost in a single Datadog plane we operate for you. 01 Cloud AWS and Azure platforms designed for modern and AI-driven workloads: infrastructure as code, least-privilege access, deep observability. ### Critical Support Datadog-powered cloud managed services for AWS and Azure, combining 24×7 incident management with improvement engineering. - -> Always-on coverage for your cloud platform, with clear ownership and escalation. - -> Real engineers embedded alongside your team, not ticket-only support. - -> Improvement hours every month so reliability, security, and cost control improve over time. - -> Datadog-first visibility for faster diagnosis and less noise. - -> Transparent operations : you retain access to your operational data while we manage and optimise the observability layer. AWS + Azure CloudOps / SRE Shared responsibility 02 Datadog Implementation, optimisation, and managed Datadog, delivered by engineers who run Datadog as the backbone of our own managed services. ### Datadog expertise Adopt Datadog cleanly, reduce alert noise, and keep your observability estate healthy as you scale. - -> FETCH™ : fast, structured implementation that gets you to meaningful value from your Datadog trial. - -> LaunchPad™ : fully managed, end-to-end Datadog deployment delivered by Critical Cloud. - -> HyperCare™ : stabilisation after go-live, noise reduction, SLOs, and runbooks that match real operations. - -> Managed Datadog : ongoing hygiene, improvements, and platform evolution by engineers who live in Datadog daily. FETCH™ LaunchPad™ HyperCare™ Managed Datadog 03 AI AI infrastructure on AWS and Azure, plus AI Factory deployments, built with human-in-the-loop controls, auditability, and cost guardrails. ### AI, powered by Datadog We use Datadog to give your AI workloads full visibility: LLM observability, agent tracing, GPU monitoring. And we use Datadog’s own AI to run your cloud operations leaner and faster. Datadog for AI LLM and agent observability, quality evaluation, and GPU fleet monitoring. Observable, evaluated, and cost-controlled in production. AI for Datadog Bits AI SRE, Watchdog, and Security Analyst inside our managed service. Engineers arrive at incidents with context, not a blank screen. LLM tracing Agent observability GPU monitoring AI Factory 05 the differentiator ## Powered by Datadog, not locked behind it. Many traditional cloud MSPs rely on locked-down proprietary monitoring that prioritises provider efficiency over customer insight, exposing a one-size-fits-all view and keeping customers dependent. Critical Cloud takes a different approach: bespoke managed services built on Datadog, the industry-leading observability platform. Every Critical Support customer has direct access to their own Datadog environment, with full-fidelity visibility across infrastructure, APM, logs, traces, security signals, cloud cost insight, and LLM monitoring, tailored to their AWS and Azure architecture. Datadog is embedded into our 24×7 operational model, driving real-time alerting, faster diagnosis, and disciplined incident response, so issues are detected early, understood in context, and resolved decisively. Read the full approach -> Datadog · unified asset visibility Your data stays yours. Full-fidelity coverage in a Datadog environment you keep direct access to. We manage and optimise it, you never lose visibility. observe -> respond -> improve ## How we operate 01 ### Instrument the platform Datadog foundations: tagging, dashboards, alert hygiene, SLOs and ownership so the signals are trustworthy. 02 ### Operate 24×7 Incident ownership with clear escalation. Fast diagnosis, controlled remediation, and structured communication. 03 ### Improve continuously Monthly engineering to reduce repeat incidents, strengthen security posture, and control cloud cost. 06 proof ## Case studies View all case studies -> Cloud · Critical Support · Azure ### OPX: Driving observability with Datadog Datadog underpinned Azure migration, delivering visibility, faster response, proactive monitoring. Datadog · FETCH™ · SaaS ### CETA: Getting more from a Datadog trial FETCH delivered rapid Datadog value; HyperCare accelerated adoption with observability. Datadog · HyperCare™ · Azure ### EIP: Rapid, reliable Datadog onboarding HyperCare enabled fast Datadog rollout across Azure supporting 140+ services. ## Experience you can verify Critical Cloud delivers Datadog-powered cloud managed services for AWS and Azure. Our work is led by practitioners with deep production experience in modern cloud operations, incident response, and observability. 60% Lower MTTR in production 75% Lower than an in-house team 200+ Cloud projects delivered 10+ yrs Datadog experience Around 75% lower than building an in-house team. Budget reallocation, not net new spend: recruiting, paying, tooling, and rota-ing an internal team for 24×7 cloud operations, redirected into a managed service that already runs at that standard. Our capabilities 24×7 Incident Management CloudOps / SRE AWS & Azure Platforms Datadog Platform Engineering Security & Compliance Cloud Cost Insight Automation & Runbooks AI Workloads & LLM Monitoring ## Built for regulated industries. Our specialism is sectors where failure has consequences and where compliance obligations mean the operating model, evidence trail, and security posture of the MSP matter as much as uptime. Financial Services & Fintech Healthcare & Healthtech SaaS & Technology Retail & E-commerce All industries -> 07 trust ## Partnerships and compliance Officially accredited. Independently certified. Built for trust. Powered by Datadog and an Advanced Partner in the UK, with AWS and Microsoft partnerships. ISO 27001 and Cyber Essentials Plus underpin secure, auditable delivery. ## FAQ The questions we get most often from tech-led teams considering Critical Support or Datadog services. **What is Critical Cloud’s trust layer for AI operations? +** The accountable layer that lets a company ship autonomous systems fast and stay in control of them in production. We operate, secure, and govern the stack the AI runs on, so the team can stay focused on the product. **What is Critical Support? +** Critical Support is our 24×7 cloud managed service for AWS and Azure. We take incident ownership and deliver improvement engineering every month so reliability, security, and cost control improve over time. **Do we keep access to our Datadog data and dashboards? +** Yes. We aim for transparent operations. You retain access to your operational data and visibility, while we build, manage, and continuously optimise the observability layer and operating practices. **What happens in the first 30 days? +** We onboard access safely, establish operational ownership and escalation, baseline dashboards and alerting, and agree the first improvement plan. The goal is to stabilise quickly and then move into continuous improvement. **How fast do you respond to incidents? +** Our target is a 15-minute incident response, with clear escalation. We’ll confirm the exact targets and communication model during onboarding to match your platform and risk profile. **Can you help with Datadog even if we don’t need a full MSP? +** Yes. We offer implementation and stabilisation packages (e.g., FETCH™ and HyperCare™) as well as ongoing managed Datadog. Ideal if you want Datadog done properly without committing to full 24×7 operations. **Do you support AWS, Azure, or both? +** Both. We specialise in AWS and Azure, and can support single-cloud or multi-cloud depending on how your product and risk profile evolve. ## Talk to us about your runtime layer. Start with the cloud, Datadog, incident or AI runtime problem you have today. Build toward a Managed Runtime Assurance model that supports where your software is going. Book a runtime assurance call -> What is Managed Runtime Assurance? --- ## /managed-runtime-assurance/ Source: https://criticalcloud.ai/managed-runtime-assurance/ Managed Runtime Assurance # Managed Runtime Assurance for AI-era software. Managed Runtime Assurance is the accountable operation of production applications, cloud platforms and AI systems so they stay observable, secure, resilient, cost-controlled and evidence-ready. Critical Cloud delivers that outcome through Critical Support, Managed Observability, Runtime Evidence and AI Runtime Operations, powered by Datadog. Talk to us about runtime assurance -> See Critical Support Scroll 01 01: Why now ## Software is easier to build. Production is harder to operate safely. AI-assisted development, agentic workflows and faster release cycles increase the amount of software reaching production. The bottleneck is no longer only building the thing. It is keeping the runtime observable, secure, resilient, cost-controlled and evidence-ready once customers depend on it. 02 02: The definition ## What Managed Runtime Assurance means Customers are not buying a tool, a ticket queue or a traditional support wrapper. They are buying one managed outcome: production stays healthy, incidents are handled, controls are enforced and evidence is ready when customers, auditors or regulators ask for it. 03 03: The boundary ## We operate the stack. You own the product. We never touch your app, your model or your business logic. That boundary is what makes us trustworthy: we have no agenda over your product, so we can stand behind whether the operating layer is sound. You own - -> Product idea - -> Application code - -> Model behaviour - -> Business logic - -> Customer experience - -> Product roadmap Critical Cloud owns - -> Observability - -> Cloud runtime operations - -> Incident response - -> Security operations - -> Cost control - -> Runtime evidence - -> Runbooks and escalation - -> Human governance for AI-assisted operations 04 04: What is included ## The operating layer behind mission-critical software **01 ### Managed Observability Signal quality, alert hygiene and incident readiness, powered by Datadog and operated by us. +** Signal quality, alert hygiene, dashboard ownership and incident readiness. Powered by Datadog, operated by us. See Managed Datadog -> **02 ### Runtime Operations Cloud platform operations for AWS and Azure, covering cost control, access governance and runbooks. +** Cloud platform operations for AWS and Azure. Cost control, access governance, change management and runbooks. **03 ### Incident Response Owned incident management where engineers respond, investigate, resolve and improve. +** Owned incident management with clear escalation. Engineers respond, investigate, resolve and improve. See Critical Support -> **04 ### Runtime Evidence Incident records, access logs and assurance packs that support customer, auditor and regulatory obligations. +** Incident records, access logs, change context, recovery evidence, SLO reporting and assurance packs. Supports customer, auditor and regulatory obligations. **05 ### AI Runtime Operations Observability, governance and human accountability for AI workloads and agentic workflows in production. +** Observability, governance and human accountability for AI workloads and agentic workflows in production. See AI Runtime Operations -> **06 ### Governed Automation AI-assisted operations under human governance, where agents own the analysis and humans own the outcome. +** AI-assisted operations under human governance. Agents own the analysis. Humans own the outcome. 05 05: The flagship service ## Critical Support is the flagship service. Critical Support delivers Managed Runtime Assurance today for AWS and Azure environments. It combines incident ownership, observability, improvement engineering, security operations, cost control, governance and operational reporting. It is more than a ticket queue. It is the operating model behind your production platform: Datadog-powered observability, incident response, runbooks, escalation, cost control, security operations, evidence and monthly improvement engineering. Across the production environments we operate, Critical Cloud has reduced mean time to resolve incidents by 60%. See Critical Support -> 06 06: The substrate ## Datadog is the operating substrate. Managed Runtime Assurance is the business. Datadog provides telemetry, dashboards, traces, logs, alerts, security signals, incident workflows, cost visibility and AI observability capabilities. Critical Cloud turns those signals into an accountable operating model. Datadog certified one managed provider globally. We are it. Our Datadog credentials -> 07 07: Runtime Evidence ## Runtime Evidence turns operations into assurance. Good operations should produce evidence. Critical Cloud helps maintain incident records, operational reports, access governance, recovery evidence, change context and SLO reporting so customers can show how their runtime is controlled. We operate and evidence the runtime controls that support your compliance obligations. Continuous Runtime Security Validation, with Tarian Labs -> 08 08: AI Runtime Operations ## AI Runtime Operations extends the model. AI workloads create new runtime questions: what happened, what did it cost, which system was involved, what data was touched, what changed and who approved the action. AI Runtime Operations brings observability, governance, cost control and human accountability to those production workflows. Agents own the analysis. Humans own the outcome. See AI Runtime Operations -> 09 09: Who it is for ## Built for tech-led SMBs and regulated scaleups. - -> SaaS and software companies - -> Fintech and financial services platforms - -> Healthtech and healthcare software providers - -> Insurtech and insurance platforms - -> AI-native and AI-adopting teams - -> Regulated or compliance-sensitive B2B software companies See industry pages -> FAQ ## Common questions **What is Managed Runtime Assurance? +** Managed Runtime Assurance is the accountable operation of production applications, cloud platforms and AI systems so they stay observable, secure, resilient, cost-controlled and evidence-ready. **Is Managed Runtime Assurance the same as cloud managed services? +** No. Cloud managed services usually focus on infrastructure and support. Managed Runtime Assurance focuses on the operating layer behind mission-critical software: observability, incidents, controls, evidence, resilience, cost and human governance. **Does Critical Cloud take over our product? +** No. Your team owns the application, model, business logic and product roadmap. Critical Cloud operates the stack around it: cloud, observability, incidents, security operations, evidence and governed automation. **How does Datadog fit? +** Datadog is the operating substrate. It provides the telemetry, alerts, traces, logs, incident workflows, security signals and AI observability capabilities. Critical Cloud turns that platform into a managed operating model. **Can this support compliance obligations? +** Critical Cloud does not make customers compliant. We operate and evidence the runtime controls that support customer, auditor and regulatory obligations. **Do we need to be AI-native to use it? +** No. Many customers start with cloud, observability, incident or evidence problems today, then build toward AI Runtime Operations as their software and risk profile evolves. 10: How to start ## Start with the problem you have today. You do not need to be AI-native to need Managed Runtime Assurance. Start with Datadog adoption, noisy alerts, incident response, cloud operations, runtime evidence or AI workloads moving into production. The destination is the same: an operating model that keeps production trustworthy. Book a runtime assurance call -> See Critical Support Explore Datadog services --- ## /datadog/uk-datadog-partner/ Source: https://criticalcloud.ai/datadog/uk-datadog-partner/ UK Datadog Partner # UK Datadog Partner, with the world's first Powered by Datadog accreditation. The world's first Powered by Datadog accredited MSP. Critical Cloud is also a Datadog Advanced Partner in the UK . Our managed service is built on Datadog, not bolted on top of it. We implement Datadog, stabilise it after go-live, and run it 24×7 for tech-led businesses on AWS and Azure, from offices in Cardiff, London, and Dublin. Talk to a UK Datadog partner Datadog services in detail Scroll ## What is a UK Datadog partner, and what does Critical Cloud do? Critical Cloud is a UK Datadog partner: a Datadog Advanced Partner and the world's first Powered by Datadog accredited MSP. From Cardiff, London and Dublin, we implement, stabilise and run Datadog 24x7 on AWS and Azure for tech-led businesses, including those in regulated sectors. Critical Cloud holds Datadog Advanced Partner status. Premier is a higher tier in Datadog's standard partner network. Our distinguishing credential is the Powered by Datadog accreditation, which Datadog awards only to MSPs whose managed service is operationally built on the platform. World's first Powered by Datadog accredited MSP Advanced Datadog partner tier in the UK ISO 27001 + Cyber Essentials Plus 24×7 UK-based Datadog operations Our Datadog accreditation Two independent signals that set us apart from resellers and project firms. - -> Powered by Datadog , world's first. Datadog's recognition that our managed service is operationally built on the platform, not a monitoring add-on. - -> Datadog Advanced Partner , certified engineers, active delivery, and independently validated customer outcomes. - -> Cardiff · London · Dublin , UK-headquartered, UK contractual base, UK engineering team. - -> ISO 27001 + Cyber Essentials Plus , independently audited. Available on request for procurement. 01 ## What "Powered by Datadog" actually means Most companies calling themselves a Datadog partner are resellers or project firms. Critical Cloud is different: our managed service is built on Datadog as the operational platform. That is what the accreditation recognises and why it matters when you're choosing a UK-based Datadog partner for the long term. ### Datadog reseller vs partner vs MSP Datadog reseller A Datadog reseller licenses the Datadog platform to you under commercial terms and may help with initial setup. After procurement, your team owns configuration and day-to-day operations. Datadog partner A Datadog implementation partner or consultancy delivers Datadog as a time-limited project: integrations, dashboards, alerting, then handover. Ongoing operational accountability ends with the engagement. Datadog MSP A Datadog MSP takes ongoing operational responsibility for running Datadog on your behalf: 24×7 monitoring, incident response, alert tuning and continuous improvement. Powered by Datadog is the accreditation Datadog reserves for MSPs whose managed service is genuinely built on the platform. Critical Cloud is the world's first holder. For a UK business buying a Datadog-powered managed service, this distinction matters. You want a partner who runs Datadog every day, not one who sells it or visits occasionally. Full comparison: Datadog partner vs reseller vs MSP, the difference explained . ### What it means for your team - -> Direct access to your Datadog environment , no black-box monitoring layer. Full visibility into your own data. - -> Advice from operational engineers , not sales engineers or architects who hand off after delivery. Our people handle alerts and incidents daily, and the same discipline underpins how we run AI operations . - -> Platform currency , we stay current with every Datadog release because we have to. You benefit without tracking it yourself. - -> Standards that hold at scale , tagging, ownership, SLOs, and dashboards designed to survive growth, not just the initial rollout. - -> UK accountability , UK-headquartered, UK contracts, UK engineering team. No offshore escalation paths or ambiguous SLAs. ### Datadog and AI operations The same operational model now extends to AI workloads. As the world's first Powered by Datadog accredited MSP, Critical Cloud is a Datadog partner for AI operations: we make AI workloads, agentic workflows and model-serving infrastructure observable, governed and production-safe. Explore AI Runtime Operations . 02 ## What we do as your UK Datadog partner We cover the full Datadog adoption lifecycle, from first trial through to 24×7 managed operations. Every engagement is delivered by a UK-based team of certified Datadog engineers with hands-on production experience. Across the production environments we operate, Critical Cloud has reduced mean time to resolve incidents by 60% and cut log volume by 77%. **01 FETCH™, Trial enablement Get to first meaningful Datadog signals before your trial ends, with integrations, cost guardrails, and an adoption plan. +** Get to first meaningful Datadog signals before your trial ends. We configure integrations, cost guardrails, and early dashboards, and build an adoption plan so your team can evaluate Datadog on real data, not just installed agents. 14-day focused plan Integration setup Cost guardrails End-of-trial review FETCH service detail → **02 LaunchPad™, Managed rollout A fully managed Datadog implementation, standardising tagging, dashboards, alerting, SLOs, and governance so Datadog is production-grade from day one. +** A fully managed Datadog implementation from end to end. We run delivery across assessment, pilot, and scale phases, standardising tagging, dashboards, alerting, SLOs, and governance so Datadog is production-grade from day one. Assess · Pilot · Scale Tagging standards Dashboards + SLOs Governance LaunchPad service detail → **03 HyperCare™, Post go-live stabilisation Critical Cloud engineers embed alongside your team after go-live to reduce alert noise, expand coverage, and build Datadog into daily workflows. +** The first months after go-live determine whether Datadog delivers long-term value or slowly becomes noise. HyperCare embeds Critical Cloud engineers alongside your team to reduce alert noise, expand coverage, fix what was rushed, and build Datadog into daily workflows. Alert tuning APM + log depth Cost hygiene Ownership mapping HyperCare service detail → **04 Managed Datadog Continuous platform management running the backlog of monitor tuning, tagging governance, dashboard evolution, and cost visibility as you scale. +** Continuous Datadog platform management for teams who want their environment to stay clean and useful as they scale. We run the backlog: monitor tuning, tagging governance, dashboard evolution, SLO lifecycle, and cost visibility. Signal quality Platform hygiene Cost visibility Continuous improvement Managed Datadog service detail → **01 HealthScan™, Independent Datadog assessment A read-only assessment delivering a health scorecard, prioritised backlog, and findings readout in 1-2 weeks. +** A read-only assessment of your current Datadog environment. We review signal quality, tagging standards, monitor hygiene, APM coverage, and cost profile and deliver a health scorecard, prioritised backlog, and findings readout in 1-2 weeks. Read-only 1-2 weeks Scorecard + backlog HealthScan service detail → **02 Catalyst, Backlog-led improvement engineering Practitioners complete approved work from an agreed backlog, with no spend above the authorised envelope. +** Practitioners complete approved work, tagging standards, alert tuning, dashboard rationalisation, SLOs, and security signal operations. Backlog agreed before work starts. No spend above the approved envelope without written authorisation. Delivery, not advice Agreed backlog Written authorisation Catalyst service detail → **03 Accelerators, One capability, four weeks Eight fixed-scope, four-week deliveries covering Infrastructure Observability, Code Security, Cloud Security, and more, with working implementation plus handover artefacts. +** Eight fixed-scope, four-week deliveries: Infrastructure Observability, Code Security, Cloud Security, Threat Management, Digital Experience, Software Delivery, Service Management, AI Observability. Working implementation plus handover artefacts. Fixed scope + price 4 weeks 8 to choose from View all accelerators → **04 Managed Datadog, Ongoing platform operations Monthly recurring platform operations across two tiers - see full service detail for tier breakdown, access models, and monthly deliverables. +** Already listed above - see full service detail for the Guided and Proactive tier breakdown, access models, and what Critical Cloud delivers each month. Monthly recurring Two operating tiers Critical Support: 24×7 cloud operations with Datadog at the core Our flagship managed service for tech-led businesses who want Datadog-powered operations running continuously. Critical Support combines 24×7 incident management with monthly improvement engineering- every alert, runbook, and escalation path is built on Datadog. This is the service the Powered by Datadog accreditation recognises. 24×7 incident management Improvement engineering Datadog-native ops AWS + Azure Explore Critical Support → ### Looking for full service detail? The Datadog services page covers every offering in depth- including Splunk → Datadog migration, platform coverage, the full observability maturity model, and the difference between each engagement. All Datadog services → 03 ## UK Datadog case studies Outcomes delivered by Critical Cloud as a Datadog partner in the UK- across implementation, post go-live stabilisation, and managed observability. **01 ### OPX, Datadog-first observability on Azure Critical Cloud made Datadog the operational backbone for OPX's Azure migration, delivering full-stack visibility and monitoring standards that scaled with the platform. +** As OPX migrated to Azure, Critical Cloud made Datadog the operational backbone: full-stack visibility, faster incident response, and monitoring standards that scaled with the platform. Read case study → **02 ### CETA, Getting value from a Datadog trial CETA used FETCH to prove Datadog value fast, then extended with HyperCare for production-ready dashboards, monitors, tagging standards, and APM. +** CETA used FETCH™ to prove Datadog value fast, then extended with HyperCare™ for production-ready dashboards, monitors, tagging standards, and APM. Read case study → **03 ### EIP, Datadog rollout across 140+ services A fast, reliable Datadog rollout across EIP's complex Azure environment, with HyperCare covering automation, enablement, and hands-on support at scale. +** Critical Cloud delivered a fast, reliable Datadog rollout across EIP's complex Azure environment. HyperCare™ covered automation, enablement, and hands-on support at scale. Read case study → View all case studies → 04 ## Trust & credentials Independently certified, UK-headquartered, and operating under an audited information security framework. For a managed Datadog MSP, these aren't optional extras, they're table stakes for regulated industries, enterprise procurement, and any customer taking cloud operations seriously. Datadog accreditation Powered by Datadog, world's first Datadog's recognition that our managed service is operationally built on the platform. Not a reseller tier or a marketing badge, an accreditation earned by demonstrating that Critical Support runs on Datadog as its core operational system. Datadog partner tier Datadog Advanced Partner Requires certified engineers, active customer delivery, and independently validated outcomes- not just contract volumes or logo placement. Information security ISO 27001 certified Independently audited information security management system. Relevant for any UK business or regulated industry where vendor security posture forms part of procurement or due diligence. UK government scheme Cyber Essentials Plus NCSC-backed certification covering technical security controls, independently verified, not self-assessed. Increasingly required for public sector and supply chain engagements. Cloud platforms AWS Partner + Microsoft Partner Active partner status on both major cloud platforms. We implement and operate Datadog on AWS and Azure, including specialist areas such as the AWS European Sovereign Cloud . UK presence Cardiff · London · Dublin UK-headquartered in Cardiff, with offices in London and Dublin. UK contractual base and UK engineering team throughout. We are a UK-based Datadog partner by design, not a global firm with a UK sales office. 05 In brief ## Common questions, answered directly **01 ### What is a UK Datadog partner? A company in the Datadog Partner Network delivering Datadog services in the UK, potentially holding tiered status and the Powered by Datadog accreditation for MSPs. +** A UK Datadog partner is a company accepted into the Datadog Partner Network that delivers Datadog services in the United Kingdom. Partners can hold tiered status based on certified engineers, active delivery, and validated customer outcomes. Some partners also hold the Powered by Datadog accreditation, which Datadog reserves for MSPs whose managed service is operationally built on the platform. Critical Cloud is a Datadog Advanced Partner and the world's first Powered by Datadog accredited MSP. **02 ### What is a Powered by Datadog MSP? A managed service provider accredited by Datadog because its managed service is genuinely built on the platform, distinct from resellers and consulting firms. +** A Powered by Datadog MSP is a managed service provider accredited by Datadog because its managed service is genuinely built on the Datadog platform. This is distinct from resellers, who sell licences, and consulting firms, who deliver time-limited projects. The accreditation means the MSP runs Datadog 24x7 as its operational system, with customers having direct access to their own environment. Critical Cloud is the world's first company to receive this accreditation. **03 ### What does a Datadog implementation partner do? Handles technical deployment of Datadog, from cloud integrations and tagging standards through to monitors, dashboards, SLOs, and ongoing managed operations. +** A Datadog implementation partner handles the technical deployment of Datadog in a customer environment: configuring cloud and infrastructure integrations, establishing tagging standards, building monitors, dashboards and SLOs, and making the platform production-grade. Beyond initial implementation, some partners also offer ongoing managed operations to keep the environment healthy as the business scales. Critical Cloud covers the full lifecycle, from first trial through to 24x7 managed operations. **04 ### Is Critical Cloud the highest Datadog partner tier? No. Advanced is our tier and Premier sits above it. The credential that sets us apart is the Powered by Datadog accreditation, not the partner tier. +** No. Critical Cloud holds Datadog Advanced Partner status. Premier is a higher tier in Datadog's standard partner network. Our distinguishing credential is the Powered by Datadog accreditation, which Datadog awards only to MSPs whose managed service is operationally built on the platform. Critical Cloud is the world's first company to hold it. Partner tier reflects certified engineers and delivery volume. The accreditation reflects how the managed service itself is built, which is the distinction that matters when you are choosing who will run Datadog for you. **05 ### Can you run Datadog across hybrid and multi-cloud environments? Yes. AWS, Azure, on-premises and Kubernetes in one Datadog environment, held together by a single tagging model, ownership map and set of SLOs. +** Yes. Most of the estates we operate are not single-cloud. We implement and run Datadog across AWS and Azure, on-premises and colocated infrastructure, Kubernetes wherever it runs, and the network paths between them. The hard part is not the agents, it is the standards: one tagging model, one ownership map and one set of SLOs, so a service is described the same way whichever platform it sits on. That is what makes a hybrid estate readable in a single Datadog environment instead of three disconnected views, and it is why a migration and a steady-state estate can be measured on the same terms. **06 ### Who can help a UK business set up and manage Datadog? Critical Cloud, from Cardiff, London and Dublin. We implement Datadog, stabilise it after go-live, then run it 24x7 with a UK-based engineering team. +** Critical Cloud does both, from Cardiff, London and Dublin. Setting Datadog up and running it afterwards are different jobs, and most providers only do the first. We implement Datadog end to end: cloud integrations, tagging standards, dashboards, monitors and SLOs. We stabilise it in the months after go-live, when alert noise usually peaks. Then we run it 24x7 as a managed service, with UK-based certified engineers, a UK contractual base and no offshore escalation path. Critical Cloud is a Datadog Advanced Partner and the world's first Powered by Datadog accredited MSP. The fastest way to work out which of those you need is a scoping call . 06 ## FAQ Common questions from teams evaluating a UK-based Datadog partner or a managed Datadog service. **What makes a good UK Datadog partner? +** The most reliable signal is operational depth: does the partner actually run Datadog every day, or do they configure it and hand over? The Powered by Datadog accreditation is Datadog's own mechanism for distinguishing MSPs whose managed service is genuinely built on the platform from resellers and project-delivery firms. Beyond accreditation, look for UK-based certified engineers, an audited security posture (ISO 27001 and Cyber Essentials Plus), and evidence of delivered customer outcomes, not just partner tier logos. For our own part, Critical Cloud has reduced mean time to resolve incidents by 60% and log volume by 77% across the production environments we operate. Ask specifically whether the same team who handles incidents is also the team who implemented Datadog. **What does the "Powered by Datadog" accreditation actually mean? +** Powered by Datadog is an accreditation Datadog awards to managed service providers whose managed service is genuinely built on the Datadog platform, not partners who simply resell licences or deliver stand-alone projects. Critical Cloud is the world's first company to achieve this accreditation. In practice it means our 24×7 managed operations. Critical Support, run on Datadog. The same Datadog environment our engineers operate from is what customers have direct access to. There is no proprietary monitoring layer in between. **Do you offer managed Datadog? +** Yes, two options depending on what you need. Managed Datadog is continuous Datadog platform management: we run the backlog, maintain tagging and dashboard standards, improve signal quality, and evolve SLOs and governance as the platform grows. Your engineers stay focused on product; we keep Datadog clean and useful. Critical Support is our full managed cloud service with 24×7 incident management, where Datadog is embedded into every operational workflow, alerts, escalations, runbooks, and monthly improvement sprints. This is the service the Powered by Datadog accreditation recognises. **Do you work with both AWS and Azure? +** Yes. We are an AWS Partner and a Microsoft Partner, and we implement and operate Datadog on both platforms. Most customers run a mix of AWS and Azure, and we handle Datadog instrumentation, cloud integration, and ongoing observability across the full estate. We also specialise in the AWS European Sovereign Cloud for organisations with data residency or sovereignty requirements. **How quickly can our team get to value with Datadog? +** It depends on where you're starting. During a Datadog trial, FETCH™ gets you to first meaningful signals within days, integrations, cost guardrails, early dashboards, and a plan. The goal is to prove Datadog value before the trial ends, not after. For teams buying Datadog and wanting a production-grade rollout without running it yourselves, LaunchPad™ delivers standardised observability end-to-end in weeks. The fastest path is always a scoping call: we can usually recommend the right engagement within the first conversation. **Are you ISO 27001 and Cyber Essentials Plus certified? +** Yes to both. Critical Cloud holds ISO 27001 certification, an independently audited information security management system and Cyber Essentials Plus, the NCSC-backed UK government scheme with verified technical controls (not self-assessed). Both are maintained and renewable certifications. Certification documentation, security questionnaire responses, and evidence packs are available on request for procurement and due diligence purposes. **Is a Powered by Datadog MSP the same as an AI operations partner? +** A Powered by Datadog MSP runs Datadog as its operational system 24×7. An AI operations partner takes accountable ownership of the stack a company's AI runs on. Critical Cloud is both: the accreditation is the foundation, and AI operations is what it now underpins. **Does Datadog support UK data residency? +** Datadog offers multiple hosting regions, including EU1 in Frankfurt, and has announced plans for a UK data centre presence expected later in 2026. Residency still needs designing: telemetry classification, redaction, retention and access controls all affect where data lands and who can see it. As a UK-based Datadog partner we help teams design deployments around their residency and governance requirements. See our guide to Datadog data residency for the UK public sector . **How much does managed Datadog cost? +** Two costs are involved. Datadog platform licensing is billed by Datadog based on usage: hosts, logs, custom metrics, traces and the products you enable. Our managed service is a monthly recurring subscription scoped to your environment and the operating tier you choose, Guided or Proactive. We do not publish rate cards: the fastest way to an accurate figure is a scoping call , and a Datadog cost review if your current spend is the problem. ## Talk to a UK Datadog partner Whether you're evaluating Datadog, mid-trial, planning a rollout, or looking for a managed Datadog MSP with genuine operational depth, we are the world's first Powered by Datadog accredited partner. Book a call and we'll recommend the simplest next step. Book a call Datadog services --- ## /datadog/ Source: https://criticalcloud.ai/datadog/ Datadog expertise # Adopt Datadog cleanly. Run it as an operating model. Datadog gives you the telemetry. Critical Cloud gives you the operating model. We help teams adopt, tune and run Datadog as part of Managed Runtime Assurance : cleaner signals, better incidents, cost control, governance and continuous improvement. Datadog is the operating substrate. Managed Runtime Assurance is the business. Talk to us View Datadog case studies Scroll By market United Kingdom Ireland EMEA UK Powered by Datadog MSP Advanced Datadog partner tier AWS + Azure cloud specialists 24×7 Datadog at the core of Critical Support Expertise Services Platform coverage Migrations Case studies FAQ Get started Three phases, eight accelerators Adopt → Optimise → Manage covers the full lifecycle. Accelerators deliver one capability fast without a full-platform project. - -> Adopt , FETCH, HyperCare, LaunchPad - -> Optimise , HealthScan, Catalyst - -> Manage , Managed Datadog, Critical Support - -> Accelerators : 8 fixed-scope, four-week deliveries Browse services → Splunk → Datadog 01 Expertise ## Datadog expertise, built from real operations We run Datadog every day inside our managed service (Critical Support). That means our advice is grounded in what works under pressure: incident response, alert tuning, ownership mapping, telemetry cost control, and repeatable standards. - -> Adoption without chaos , consistent rollout patterns across services and environments. - -> Better signals , fewer false positives, clearer ownership, faster triage. - -> Cost clarity , guardrails, scoping, and usage hygiene from the start. - -> Security visibility , turning findings into operational action, not reports. - -> AI readiness , observability for AI workloads, plus LLM monitoring where relevant, delivered as managed AI operations . If you want 24×7 operations with Datadog embedded into the service model, see Critical Support . Signal quality Less noise, clearer ownership Speed Faster diagnosis and triage Standards Naming, tagging, governance Cost insight Visibility that drives action Security signals Operationalising findings AI monitoring Observability for AI workloads 02 Services ## From first trial to always-on managed operations Three phases cover the full Datadog lifecycle. Eight fixed-scope Accelerators add fast capability delivery in one area without a full platform project. Full service detail, including the decision table and individual pages for every service, is at the Datadog services catalogue . Adopt, Get Datadog live properly FETCH (complimentary trial enablement), HyperCare (two-week post-go-live sprint), and LaunchPad (fully managed rollout across Assess → Pilot → Scale). Built right the first time so you don’t rebuild later. FETCH, trial HyperCare, post go-live LaunchPad, full rollout Learn more → Optimise, Make an existing deployment work HealthScan (independent read-only assessment, health scorecard, prioritised backlog in 1-2 weeks) and Catalyst (backlog-led improvement engineering, practitioners deliver, not more reports). HealthScan, assess Catalyst, deliver Learn more → Manage, Ongoing platform operations Managed Datadog (recurring monthly platform ops: backlog execution, governance, signal quality, product expansion) and Critical Support (adjacent 24×7 cloud incident response, Datadog at the core). Managed Datadog, monthly Critical Support: 24×7 Learn more → Accelerators, One capability, four weeks Eight fixed-scope, four-week deliveries across Infrastructure Observability, Code Security, Cloud Security, Threat Management, Digital Experience, Software Delivery, Service Management, and AI Observability. Fixed scope + price 4 weeks 8 to choose from Browse accelerators → Stage 1 Foundations **Detail +** Basic agents and ad-hoc dashboards; minimal tagging and depth. Stage 2 Visibility **Detail +** Complete infra + APM + log coverage, ownership, alerts, dashboards, SLOs. Stage 3 Efficiency **Detail +** Cost optimisation, noise reduction, governance, automation, SLO lifecycle. Stage 4 Performance **Detail +** Predictive insights, synthetic UX, deeper APM optimisation, reliability engineering. 03 Platform coverage ## The integrated platform for monitoring & security We cover the full Datadog platform, with deeper focus on observability , security , digital experience , and AI monitoring . The goal is one consistent, searchable view across infrastructure, applications, logs, security signals, cloud cost, and AI workloads. **01 Observability Infrastructure, containers, logs and traces joined up so engineers can answer “what changed?” fast. +** Infrastructure, containers, logs, traces and service health, joined up so engineers can answer “what changed?” fast. Infra monitoring Kubernetes Logs + traces SLOs **02 Digital experience APM and user experience monitoring that connects customer impact to services and deployments. +** APM and user experience monitoring that connects customer impact to services and deployments. APM RUM Synthetics Error tracking **03 Security Security telemetry that becomes operational: detection, triage, and response workflows aligned to ownership. +** Security telemetry that becomes operational: detection, triage, and response workflows aligned to ownership. Cloud security Threat detection SIEM signals App security **04 AI Observability for AI workloads plus LLM monitoring, so performance and cost stay predictable. +** Observability for AI workloads plus LLM monitoring where relevant, so performance and cost stay predictable. AI workload telemetry LLM monitoring Cost guardrails Runbooks ### Also supported We regularly implement and optimise additional Datadog surfaces-depending on your stack and operating model. Cloud cost monitoring Service catalog Incident workflows Software delivery insights Network + database monitoring Platform governance ## Powered by Datadog, by design Traditional MSPs often rely on proprietary monitoring that limits customer insight. Critical Cloud takes a different approach: bespoke managed services built on Datadog. Customers retain direct access to their Datadog environment, with full-fidelity visibility tailored to your AWS and/or Azure architecture. Datadog partner by market: United Kingdom · Ireland · EMEA overview · All partner credentials Deep dives ## Datadog by platform and sector How a Datadog deployment differs depending on where you run and what you have to evidence. Datadog for AWS in the UK Account integration, ECS, EKS and Lambda coverage, and consolidating CloudWatch without paying for it twice. Datadog for AWS -> Datadog for Azure in the UK Subscription integration, AKS and App Service coverage, and bringing five fragmented native tools into one query surface. Datadog for Azure -> Datadog for UK financial services Coverage mapped to important business services, retention matched to policy, and an incident record you can defend. Regulated observability -> Datadog pricing and cost optimisation 04 Migrations ## Migrations If you’re moving off legacy tooling, we deliver structured migrations with clear scope, low risk, and operational continuity. The most common request: Splunk → Datadog . Splunk → Datadog (logs + dashboards) A predictable migration to Datadog log ingestion: pipeline design, tagging and enrichment, dashboard recreation, side-by-side validation, and a controlled cutover. What we do Discovery & inventory Ingestion design Pipelines + parsing Tagging strategy How we deliver Side-by-side run Cutover plan Runbooks Enablement Result: clean Datadog pipelines, dashboards that work day one, and a cutover without visibility gaps. Discuss a migration → Tool consolidation & observability redesign If you have multiple monitoring tools, duplicated alerting, or inconsistent tagging, we can redesign the observability layer so teams have one clear operational picture. Standards Ownership Dashboards Noise reduction Talk through options → 05 Case studies ## Case studies Examples of Datadog work across implementation, optimisation, and managed outcomes. OPX, Driving observability with Datadog Datadog-first observability patterns to improve visibility and operational confidence. Read case study → CETA, FETCH trial enablement Structured setup and guided adoption to prove value fast during a Datadog trial. Read case study → EIP, Datadog HyperCare Post go-live stabilisation: noise reduction, standards, and stronger operational workflows. Read case study → View all case studies ## FAQ **Do you need access to our infrastructure? +** It depends on the engagement. HyperCare and Managed Datadog can often be delivered with admin access to your Datadog account only. LaunchPad may require infrastructure access to implement agents, integrations, and standard patterns end-to-end. **Can you work with our existing Datadog setup? +** Yes. We can stabilise and improve an existing setup: tagging, dashboards, alert noise reduction, SLOs, cost guardrails, and expanding coverage in controlled waves. **We’re trialling Datadog, what’s the best place to start? +** Start with FETCH if you want structured, low-friction support during the trial. If you already know you want Datadog and you’d rather not run the rollout yourself, LaunchPad is the fastest path to a production-grade setup. **Do you support security and cost monitoring? +** Yes. We commonly help teams operationalise security signals and implement cloud cost visibility and guardrails. The goal is to create actionable workflows (not just more dashboards). **Which clouds do you specialise in? +** AWS and Azure. We implement Datadog in both, and we also operate cloud platforms 24×7 through Critical Support. ## Ready to get Datadog working properly? Tell us what you’re trying to achieve, clean rollout, better signals, security visibility, cost control, or AI observability. We’ll recommend the right service and the simplest next step. Talk to us Browse the service catalogue --- ## /datadog/services/ Source: https://criticalcloud.ai/datadog/services/ Datadog services # Every phase of the Datadog journey, from first trial to always-on operations. Critical Cloud is a UK Datadog partner: a Datadog Advanced Partner and the world's first Powered by Datadog accredited MSP. We deliver this service catalogue from Cardiff, London and Dublin, across AWS and Azure. Datadog creates the most value when it is implemented well, kept clean, and operated continuously. Most teams nail one of those three. Critical Cloud supports all of them, with a defined service for wherever you are in the journey. Talk to us Which service is right for me? Scroll Adopt FETCH · HyperCare · LaunchPad Optimise HealthScan · Catalyst Manage Managed Datadog · Critical Support 8 Accelerators, four-week capability deliveries Adopt Optimise Manage Accelerators Choose a service Full service catalogue Navigate directly to any service. Adopt FETCH™ HyperCare™ LaunchPad™ Optimise HealthScan™ Catalyst Manage Managed Datadog Critical Support Accelerators All 8 Accelerators → 01 Phase 1, Adopt ## Get Datadog live and delivering value Three services for teams at different stages of the Datadog adoption journey- whether you're evaluating, have just gone live, or need a fully managed rollout. FETCH™ Complimentary trial enablement. We work alongside your Datadog account team during the trial- configuring agents, validating telemetry, building early dashboards, and demonstrating value before you commit. The goal is to give you something real to evaluate, not just installed agents. Trial support Complimentary Agent setup Early dashboards Best when: you're evaluating Datadog and want expert hands alongside your account team. FETCH service detail → HyperCare™ Two-week, hands-on post-go-live sprint. After purchase, the first weeks are where early adopters either gain confidence or lose it. HyperCare removes the blockers, misconfigs, noisy alerts, dashboards that don't reflect the actual system, before they cost you momentum. Two-week sprint Post go-live Alert tuning Dashboard build Best when: you've purchased Datadog and need it working properly before the team loses faith in it. HyperCare service detail → LaunchPad™ A fully project-managed rollout to production. Critical Cloud leads delivery end-to-end across three phases, Assess, Pilot, Scale, establishing tagging standards, pipelines, RBAC, dashboards, alerts, and team onboarding. Built right so you don't rebuild later. Assess · Pilot · Scale Tagging standards Dashboards + SLOs Team onboarding Best when: you need a clean, complete Datadog rollout and don't have the internal capacity to run it. LaunchPad service detail → 02 Phase 2, Optimise ## Make an existing Datadog deployment actually work Most teams that have been on Datadog for more than a year have the same problem: they've accumulated technical debt, inconsistent tagging, noisy monitors, dashboards nobody trusts. These two services fix that. HealthScan™ An independent, read-only assessment of your existing Datadog environment. We review agent coverage, metrics and log quality, monitor accuracy, tagging standards, APM configuration, security signals, cost profile, and governance and deliver a health scorecard, prioritised improvement backlog, and a findings readout you can take to stakeholders to justify the next step. Read-only access 1-2 weeks Health scorecard Prioritised backlog Two customer sessions Best when: you need stakeholder-ready evidence of what's broken before committing improvement budget. HealthScan service detail → Catalyst Backlog-led improvement engineering. Once the work is agreed, Critical Cloud practitioners complete it, not more reports, not more recommendations. Catalyst covers tagging and naming standards, monitor redesign and alert tuning, dashboard rationalisation, SLO setup, log and telemetry governance, security signal operations, and incident workflow improvements. Nothing gets added to scope without written authorisation. Agreed backlog Delivery, not advice Tagging + monitors SLOs + governance Best when: you know what's broken and need it delivered, not diagnosed again. Often follows a HealthScan. Catalyst service detail → 03 Phase 3, Manage ## Keep Datadog healthy and your team focused on shipping Two ways to hand ongoing Datadog operations to Critical Cloud, ongoing platform management or full 24×7 cloud incident response with Datadog as the operational foundation. Managed Datadog Recurring monthly Datadog platform management. Critical Cloud runs the backlog- maintaining governance, improving signal quality, expanding product use, and evolving dashboards, SLOs, and alerting as the platform grows, so your team keeps building without the platform falling behind them. Recurring engagement Platform governance Signal quality Continuous improvement Best when: your Datadog environment grows every month and keeping it clean is nobody's primary job. Managed Datadog service detail → Critical Support 24×7 cloud incident management for AWS and Azure, with Datadog as the operational foundation. Critical Support is Critical Cloud's flagship recurring service, engineers on-call around the clock, incidents owned end-to-end, and monthly improvement engineering delivered alongside the operational work. Every alert, runbook, and escalation path is built on Datadog. 24×7 incident management AWS + Azure Datadog-native Improvement engineering Best when: you need continuous operational responsibility, not just tooling management. Critical Support, full service detail → 04 Accelerators ## One Datadog capability. Four weeks. A working implementation. Eight fixed-scope, four-week deliveries, each targeting a specific Datadog capability area and producing a working implementation, not a report. Pick the one that matches your highest-priority gap right now. Infrastructure Observability Core visibility across infra, apps, logs, and data. Operational dashboard and monitor pack delivered. Code Security Security insight inside engineering workflows. Findings baseline, ownership model, remediation backlog. Cloud Security Cloud posture, identity risk, and workload security operational. Prioritised findings and ownership model. Threat Management Detections, investigations, and response workflows made operational. Live detection set on delivery. Digital Experience Frontend and backend customer-impact visibility. Critical journey monitors and DEX dashboards. Software Delivery Build, test, release, and dev-quality signals unified. Delivery dashboards and hotspot identification. Service Management Alerts to coordinated incidents, fast. First live incident workflow and routing model on delivery. AI Observability LLM and AI app performance, cost, quality, and safety visible. AI dashboard pack and first alert pack. Each Accelerator is fixed-scope and delivered in four weeks, working implementation plus handover artefacts plus a next-step recommendation. See all Accelerators in detail → 05 ## Choosing the right service Match your situation to the service designed for it. Most engagements follow a natural path- FETCH or LaunchPad to get live, HealthScan to assess, Catalyst to improve, Managed Datadog or Critical Support to operate. Your situation Start here Evaluating Datadog, want to prove value before committing FETCH™ Just purchased Datadog, early blockers, noisy alerts, dashboards don't work HyperCare™ Need a full rollout but don't have the internal capacity to run it properly LaunchPad™ Already on Datadog, need evidence of what's broken before allocating budget HealthScan™ Know what's wrong with the Datadog setup, need it fixed, not diagnosed again Catalyst Want to operate Datadog month-to-month without building an in-house platform team Managed Datadog Need 24×7 operational coverage, incidents owned, infrastructure managed Critical Support One specific capability area needs delivering in four weeks An Accelerator Not sure? Talk to us , we'll tell you honestly which service fits your situation and what the simplest next step is. 06 In brief ## Common questions, answered directly The questions UK teams ask before they pick a Datadog service. **01 ### Who provides Datadog services in the UK? Critical Cloud, from Cardiff, London and Dublin: a Datadog Advanced Partner and the world's first Powered by Datadog accredited MSP. +** Critical Cloud, from Cardiff, London and Dublin. Critical Cloud is a Datadog Advanced Partner and the world's first Powered by Datadog accredited MSP, the accreditation Datadog awards only to providers whose managed service is operationally built on the platform. Premier is a higher tier than Advanced in Datadog's standard partner network, so it is the accreditation rather than the tier that distinguishes us. Every service in this catalogue is delivered by UK-based certified engineers who run Datadog in production every day, on AWS and Azure, under a UK contractual base. Full detail on the UK Datadog partner page . **02 ### Which Datadog service should we start with? Mid-trial, FETCH. Rolling out, LaunchPad. Newly live, HyperCare. Already running it, HealthScan then Catalyst. Handing it over, Managed Datadog or Critical Support. +** It depends where you are. Mid-trial, start with FETCH. Buying Datadog and planning a rollout, start with LaunchPad. Recently live and drowning in alerts, start with HyperCare. Already running Datadog but unsure what is wrong with it, start with a HealthScan and then Catalyst to deliver the fixes. Wanting someone else to run it from here, start with Managed Datadog or Critical Support. If none of those obviously fits, a short scoping call will settle it faster than a form. **03 ### Can you take over a Datadog deployment another partner implemented? Yes. Usually HealthScan first to assess it independently, Catalyst to work the backlog, then Managed Datadog or Critical Support to operate it. +** Yes, and it is one of the most common ways engagements start. We usually begin with a HealthScan : an independent, read-only assessment of the existing deployment covering coverage gaps, tagging, monitor quality, alert noise and cost drivers, delivered in one to two weeks. Catalyst then works through the resulting backlog, and Managed Datadog or Critical Support takes on ongoing operations. You keep ownership of your own Datadog organisation throughout. We work inside it, not around it. **04 ### What is the difference between Managed Datadog and Critical Support? Managed Datadog runs the platform. Critical Support runs the incidents, 24x7. Many customers take both. +** Managed Datadog is platform operations: we own the Datadog backlog, hold tagging, dashboards, monitors and SLOs to standard, improve signal quality and keep cost drivers visible month to month. Critical Support is cloud incident response: 24×7 cover where our engineers own alerts, triage and escalation across your infrastructure, with Datadog embedded in every workflow. Teams who want the platform kept clean take the first. Teams who want someone accountable at three in the morning take the second. Plenty of customers take both. **05 ### What do Critical Cloud's Datadog services cost? Datadog licensing is billed by Datadog on usage. Our services are either fixed-scope engagements or a monthly subscription. We scope rather than publish rate cards. +** Two costs are involved and they are billed separately. Datadog platform licensing comes from Datadog and is driven by usage: hosts, logs, custom metrics, traces and the products you enable. Critical Cloud's own services are either fixed-scope, fixed-price engagements of a defined duration, such as the four-week Accelerators , or a monthly recurring subscription in the case of Managed Datadog and Critical Support. We do not publish rate cards, because scope varies too much for a published number to mean anything. Describe the estate and we will scope it: talk to us . ### The partner behind the services Critical Cloud is the world's first Powered by Datadog accredited MSP and a Datadog Advanced Partner. Every service in this catalogue is delivered by engineers who run Datadog in production every day, not by consultants who configure it and leave. UK partner credentials EMEA coverage ### Not sure where to start? The most common pattern: a HealthScan first (understand what's actually broken), followed by Catalyst (deliver the fixes), with Managed Datadog or Critical Support for ongoing operations. But every situation is different, talk to us and we'll give you an honest recommendation in the first conversation. Datadog expertise overview ## Ready to talk through which service fits? Tell us where you are in the Datadog journey, we'll recommend the right service and be honest about what the simplest next step actually is. Talk to us Partner page --- ## /cloud/critical-support/ Source: https://criticalcloud.ai/cloud/critical-support/ Managed Runtime Assurance for AWS and Azure # Critical Support: Managed Runtime Assurance for AWS and Azure. Critical Support is our flagship Managed Runtime Assurance service. We own incident response and improvement engineering across reliability, security, cost, performance, automation, governance and observability, so your platform becomes safer to run over time. Talk to us Datadog credentials Scroll Across the production environments we operate, Critical Cloud has reduced mean time to resolve incidents by 60%. 15 min SEV-1 & SEV-2 response time 60% Lower MTTR in production 24×7×365 Always-on coverage Powered by Datadog World's first accredited MSP Improvement pillars Plans Getting started Two services in one Incident Management 24×7 detection → triage → response → escalation → recovery & review. 15-min response time for SEV-1 & SEV-2. Improvement Engineering 16-56 hrs/month across reliability, security, cost, performance, automation & governance. Monthly reporting. Three plans: Core · Standard · Advanced, each priced on scope. Talk to us for pricing. 01 Managed Runtime Assurance ## From cloud support to runtime assurance. Critical Support is more than a ticket queue. It is the operating model behind your production platform: Datadog-powered observability, 24x7 incident response, runbooks, escalation, cost control, security operations, evidence and monthly improvement engineering. What is Managed Runtime Assurance? -> 02 Improvement Engineering ## Six improvement pillars, delivered every month Critical Support isn't just incident cover. Every month our engineers work through an agreed improvement backlog across six pillars, so the platform gets better, not just maintained. Agents own the analysis. Humans own the outcome. **01 Reliability & Resilience Failover design, redundancy, and SLO management to reduce incident frequency and impact. +** Failover design, redundancy improvements, early issue detection, and SLO/SLA management to reduce the frequency and impact of incidents. **02 Security & Compliance Access control reviews, threat detection, and ISO 27001 and Cyber Essentials Plus alignment. +** Access control reviews, vulnerability management, threat detection operationalisation, and alignment to ISO 27001 and Cyber Essentials Plus. **03 Cost Optimisation & FinOps Rightsizing, waste elimination, and cost attribution to give teams financial ownership. +** Rightsizing, waste elimination, reserved instance and savings plan recommendations, and cost attribution to give teams financial ownership. **04 Performance & Scalability Latency diagnosis, scaling improvements, and capacity planning ahead of growth or traffic events. +** Latency diagnosis, scaling policy improvements, database query optimisation, and capacity planning ahead of growth or traffic events. **05 Automation & Efficiency Runbooks as code, IaC improvements, and auto-remediation to reduce the manual operational burden. +** Runbooks as code, IaC improvements, auto-remediation, and reducing the manual operational burden so engineers focus on what matters. **06 Governance & Observability Tagging standards, dashboard quality, alerting hygiene, and governance guardrails that scale with your platform. +** Tagging standards, Datadog dashboard quality, alerting hygiene, reporting cadence, and governance guardrails that scale with your platform. 03 Incident Management ## Five-stage incident lifecycle Every incident follows the same structured process, from first signal to blameless postmortem. Stage 01 Monitoring **Detail +** Datadog telemetry, alerting, and noise reduction keep signal quality high. Synthetic monitors and anomaly detection catch issues before customers report them. Stage 02 Triage **Detail +** Severity classification (SEV-1-4), blast-radius assessment, and ownership assignment. Bits AI SRE assists our engineers, humans confirm before acting. Stage 03 Response **Detail +** Runbooks, safe workarounds, and rollback procedures executed by on-call engineers. Customer notified within the contracted response window. Stage 04 Escalation **Detail +** On-call routing, cloud-provider escalation, vendor coordination, and customer communication, all tracked in Datadog Incident Management. Stage 05 Recovery & Review **Detail +** Fix or rollback with validation, blameless RCA, and improvement actions fed back into the monthly engineering backlog. 60-minute recovery is a target for SEV-1. 04 Plans ## Three plans, Core, Standard, Advanced All plans include 24×7 SEV-1 and SEV-2 incident management with a 15-minute response time and a 60-minute recovery target. Plans differ by platform complexity and monthly improvement engineering hours. Talk to us for pricing. Feature Core Standard Advanced Coverage 24×7 SEV-1 & SEV-2 24×7 SEV-1 & SEV-2 24×7 SEV-1 & SEV-2 Response time 15 min 15 min 15 min Recovery target (SEV-1) 60 min target 60 min target 60 min target Improvement hours/month 16 hrs 32 hrs 56 hrs Cloud scope Single cloud, 1 landing zone (hub + 1-2 spokes) Single cloud, multiple landing zones / accounts AWS and/or Azure, 5+ landing zones / hybrid Improvement pillars covered Reliability, security & cost All six pillars All six pillars, cross-cloud Governance cadence Monthly reporting Fortnightly reporting Weekly review + quarterly strategy Runbooks & RCA Core runbooks, standard reviews Advanced playbooks, full RCA + automation Custom cross-cloud workflows, postmortems Recovery time is a target, not a contractual guarantee. Response time (15 min for SEV-1 & SEV-2) is the contractual commitment. SLAs & response ## What we commit to, precisely Response time is a firm contractual commitment. Recovery time is a target, because recovery depends on the nature of the incident, not just our speed. - -> 15-minute response time for SEV-1 and SEV-2 incidents; this is the contractual commitment. Time to first engineer contact, 24×7×365. - -> 60-minute recovery target for SEV-1; this is a target. We work as fast as technically possible; complex incidents take longer by nature. - -> SEV-1 : complete outage or material risk to business. SEV-2 : significant degradation or partial outage. SEV-3/4 : limited impact, handled in-hours. - -> Blameless postmortem for all SEV-1 incidents. Findings feed the improvement backlog. - -> Incidents covered: service outages, performance degradation, security alerts (operational triage and containment, not SOC/MDR/forensics), integration/API failures, cloud provider incidents affecting your environment. - -> Customer retains: IAM and admin control, access approvals, and all business and release decisions. Material changes need customer sign-off. - -> AI is advisory: Bits AI SRE and Watchdog assist diagnosis, humans approve all production, security, and cost changes. 05 What a good cloud partner looks like ## Five principles we hold ourselves to **01 Transparency You keep direct access to your Datadog environment, your data, and your dashboards at all times. +** You keep access to your Datadog environment, your data, and your dashboards at all times. Nothing is hidden in a proprietary layer. **02 Ownership When an incident fires, we own it through to resolution, not to the first opportunity to hand it back. +** When an incident fires, we own it to resolution, not to the first opportunity to hand it back. Accountability is the baseline. **03 Collaboration Shared backlog, shared visibility. Service reviews are conversations, not status reports. +** Shared backlog, shared visibility. You see what we're working on and why. Service reviews are conversations, not status reports. **04 Integration Improvement work tied to reliability, security, and cost outcomes, not abstract platform activity. +** Improvement work is tied to reliability, security, and cost outcomes, not abstract platform activity. Everything maps to a business metric. **05 Enablement Runbooks and knowledge stay in your environment. You should be less dependent on us over time, not more. +** Runbooks, standards, and knowledge stay in your environment after every engagement. You should be less dependent on us over time, not more. ## We operate the stack. You own the product. We operate, secure, and govern the stack your AI runs on. We never touch your app, your model, or your business logic. That boundary is what makes us a trustworthy, impartial layer: we have no agenda over your product, so we can stand behind whether your operations are sound. 06 Shared responsibility ## You own your apps and decisions. We operate and improve. Hyperscalers provide the platform. Area Customer Critical Cloud Cloud provider Application code & data Owns and controls Supports, does not access data N/A Infrastructure provisioning (Terraform/IaC) Approves changes Implements and improves N/A Monitoring & observability (Datadog) Has full access always Builds, manages, optimises N/A Security, compliance & access control Owns decisions & approvals Operates controls, improves posture Platform primitives Incident management Informed, approves resolution Detects, triages, responds, recovers Provider incident support Global infrastructure & physical security N/A N/A Owns and guarantees 07 Getting started ## Three pathways into Critical Support Whichever path you take, the outcome is the same: 24×7 reliability, observability, and continuous improvement from day one. **Build & Operate ### New platform or product Design and provision with Terraform and best-practice landing zones, then transition directly into 24×7 Critical Support at go-live. +** Design and provision using Terraform and best-practice landing zones, implement Datadog monitoring foundations, then transition directly into 24×7 Critical Support at go-live. **Migrate & Operate ### Move from on-prem or another cloud Execute migration with minimal disruption, then activate incident management and improvement engineering immediately post-migration. +** Plan and execute migration with minimal disruption, align to landing zone standards and Datadog instrumentation, then activate 24×7 incident management and improvement engineering immediately post-migration. **MSP Transfer ### Take over from an existing provider Review configuration, access, and governance, then move onto Critical Support with a service review in week one. +** Review configuration, access, and governance for full transparency. Establish runbooks and Datadog observability baselines. Move onto the Critical Support model with a service review in week one. 08 How we're different ## Critical Support vs hyperscaler support plans AWS Business/Enterprise and Azure Unified/Developer support answer questions. Critical Support owns the environment. Capability Hyperscaler support Critical Support Incident response Advisory guidance; you action We own response and recovery Proactive engineering Not included 16-56 hrs/month across six pillars Observability platform Native CloudWatch / Azure Monitor only Datadog across the full stack (infra, APM, logs, security, cost) Who does the work You, with vendor advice Our SRE team, with your oversight Runbooks & automation You build and maintain We build, own, and improve Blameless postmortems Not standard Included for all SEV-1 incidents Cloud scope Single provider AWS and Azure in one service Powered by Datadog ## Datadog is the operational backbone, not a monitoring add-on Every Critical Support customer has direct access to their own Datadog environment, infrastructure, APM, logs, traces, security signals, cloud cost, and LLM monitoring, all configured to their AWS and/or Azure architecture. You keep full visibility. We operate it. Traditional MSPs rely on proprietary monitoring that limits customer insight. We use Datadog, the same platform our engineers use, in your account, visible to your team at all times. Critical Cloud is the world's first Powered by Datadog accredited MSP → 09 Delivered work ## Case studies ### OPX, Azure + Critical Support Full-stack observability via Datadog across OPX's Azure environment, combined with monthly improvement cycles. Incident noise reduced by more than 60%, with faster root-cause analysis through unified dashboards and alert tuning. Read case study → ### FAW / Hopp Studio, AWS + Critical Support 24×7 incident response plus proactive improvement for coaching systems and public websites. Tighter Datadog monitoring, quicker recovery, and improved resilience during high-traffic events. More case studies → Service family ## Need more flexibility? The full cloud service family. Critical Support is the flagship. If you need lighter cover or incident-response-only, we have options. Incident response only ### Critical Response Detection, response, and recovery, no proactive engineering. Plans: Daytime, Evenings & Weekends, 24×7. For teams that want cover without the full managed service commitment. Critical Response → Start-ups & single-cloud ### Critical Support Lite Right-sized incident cover plus a smaller improvement engineering allocation. Plans: Monitor + Fix, Engineer Assist, Partner Plus. Designed to grow into Critical Support. Critical Support Lite → Platform-specific ### Critical Support by platform The same 24×7 service written through an AWS or Azure lens, with platform-native tooling, architecture patterns, and SEO context for each hyperscaler. AWS → Azure → ## FAQ **What is Critical Support? +** Critical Support is Critical Cloud's flagship managed service: 24×7 incident management combined with monthly improvement engineering across six pillars (reliability, security, cost, performance, automation, governance) for AWS and Azure environments, with Datadog as the operational foundation. **What clouds do you support? +** AWS and Azure. We do not currently support GCP. **What is your response time commitment? +** 15 minutes for SEV-1 and SEV-2 incidents; this is the contractual response time. The 60-minute recovery figure for SEV-1 is a target; recovery time depends on the nature of the incident. **What is the difference between Critical Support and Critical Support Lite? +** Critical Support is the full flagship: 24×7 coverage, 15-minute SEV-1/SEV-2 response, and 16-56 hours of improvement engineering per month. Critical Support Lite is designed for start-ups and smaller single-cloud environments; lighter coverage windows and fewer improvement hours, but the same SRE-driven, Datadog-native model. Lite customers can step up to Critical Support as they grow. **Do you use AI in your operations? +** Yes. Datadog's Bits AI SRE and Watchdog assist our engineers in triaging alerts and surfacing likely root causes. AI is advisory: humans approve all production, security, and cost changes. **Can I keep access to my Datadog environment? +** Always. Every Critical Support customer retains direct, full-fidelity access to their own Datadog environment. Nothing is hidden in a proprietary layer. This is one of the five partner principles we hold ourselves to. **How is the trust layer delivered? +** The trust layer for AI operations is delivered by Critical Support: 24×7 incident management plus monthly improvement engineering, with Datadog as the operational platform. We operate, secure, and govern the stack your AI runs on; we never touch your app, your model, or your business logic. How we operate the trust layer → ## Ready to move on from reactive firefighting? Tell us about your platform. We'll recommend the right plan and show you how Critical Support would work for your environment. Talk to us Critical Support Lite --- ## /cloud/critical-support-lite/ Source: https://criticalcloud.ai/cloud/critical-support-lite/ Cloud Managed Service for Start-ups, Powered by Datadog # Critical Support Lite Right-sized cloud cover that grows with you. Critical Support Lite is built for start-ups and smaller single-cloud environments that need professional incident cover and engineering improvement, without the overhead of the full managed service. Same SRE-driven, Datadog-native model. Right-sized cover that keeps an earlier-stage platform in control as it grows. Talk to us Need the full service? Scroll 01 Who it's for - -> Start-ups with a single AWS or Azure environment and no dedicated ops team - -> Scale-ups that need after-hours cover without a full 24×7 commitment - -> Product teams that want proactive improvement but not the cost of Advanced - -> Teams on a growth path to Critical Support as their platform matures 24×7 cover available as an add-on. Step up to Critical Support at any time. 15 min SEV-1 response time 8-24 hrs Improvement engineering/month AWS + Azure Single cloud per plan Powered by Datadog Full Datadog access always 02 Plans ## Three plans, Monitor + Fix, Engineer Assist, Partner Plus All plans cover a single cloud (AWS or Azure) with a single landing zone. Response times are contractual commitments. Recovery targets are targets. Talk to us for pricing. Feature Monitor + Fix Engineer Assist Partner Plus Coverage hours 09:00-17:00 Mon-Fri 17:00-09:00 Mon-Fri + 24×7 weekends 17:00-09:00 Mon-Fri + 24×7 weekends Severity covered SEV-1 SEV-1 & SEV-2 SEV-1 & SEV-2 Response time 15 min (SEV-1) SEV-1: 15 min · SEV-2: 30 min SEV-1: 15 min · SEV-2: 30 min Recovery target (SEV-1) 2-hr target 90-min target 90-min target Improvement hours/month 8 hrs 16 hrs 24 hrs Monitored endpoints 1 endpoint 2 endpoints 2 endpoints Out-of-hours callouts None 2 per month 2 per month RCA & reporting Basic RCA + monthly summary Basic RCA + SEV review Enhanced review + backlog follow-up Recovery times are targets, not contractual guarantees. Response times are the firm commitment. 24×7 cover is available as an add-on on all Lite plans. What's included ## The same operational model, right-sized - -> Datadog instrumentation , we set up and maintain your Datadog environment; you retain full access at all times. - -> Incident management , detection, triage, response, escalation, and recovery within contracted hours. Blameless RCA for SEV-1 incidents. - -> Improvement engineering , monthly hours spent on the improvement pillars most relevant to your environment: reliability, security, cost, and governance. - -> Monthly reporting , service review covering incidents, improvement work completed, and next priorities. - -> SRE-driven, not traditional MSP , engineers who understand modern cloud architecture, not ticket-handlers. - -> AI-augmented triage , Bits AI SRE and Watchdog assist our engineers. Humans approve all production changes. - -> ISO 27001 and Cyber Essentials Plus , our security posture applies to every customer environment we operate. - -> Customer retains admin control , IAM, access approvals, and material change decisions stay with you. 03 How we work ## Five principles, the same ones whether you're on Lite or Advanced **01 Transparency & Ownership You keep full access to your Datadog environment at all times. +** You keep full access to your Datadog environment at all times. When something breaks, we own it to resolution, not to the first handoff opportunity. **02 Collaboration & Enablement Shared backlog, visible work. +** Shared backlog, visible work. Runbooks and governance standards stay in your environment, you should be less dependent on us over time, not more. 04 Growing beyond Lite? ## Critical Support, the full 24×7 service Critical Support adds full 24x7 coverage, 15-minute response for both SEV-1 and SEV-2, more improvement engineering hours, and support for multi-account and hybrid environments. **Read the full context +** Critical Support adds full 24×7 coverage, 15-minute response for both SEV-1 and SEV-2, more improvement engineering hours, and support for multi-account and hybrid environments. The Datadog instrumentation, runbooks, and governance standards you built on Lite carry straight across. Explore Critical Support Talk to us about upgrading 05 ## FAQ **What is Critical Support Lite? +** Critical Support Lite is a right-sized cloud managed service for start-ups and smaller single-cloud environments. It combines incident cover (daytime or evening/weekend) with a smaller improvement engineering allocation, using the same SRE-driven, Datadog-native model as the flagship Critical Support. **Does Lite cover AWS and Azure? +** Each Lite plan covers a single cloud, either AWS or Azure. If you need multi-cloud cover, Critical Support Standard or Advanced is the right service. **Is 24x7 cover available on Lite? +** Yes: 24×7 coverage is available as an add-on on all Lite plans. Alternatively, stepping up to Critical Support gives you native 24×7 coverage plus more improvement engineering hours. **What are "recovery targets"? +** Recovery targets (e.g. 90-minute target for SEV-1) are our working objectives, we aim to achieve them, and they reflect our operational capability. They are not contractual guarantees, because recovery time depends on the nature of the incident. Response times (the time to first engineer contact) are the contractual commitment. ## Right-sized cloud cover for where you are today. Tell us about your AWS or Azure environment and we'll recommend the right Lite plan or tell you honestly if Critical Support is the better fit. Talk to us Critical Support --- ## /cloud/critical-response/ Source: https://criticalcloud.ai/cloud/critical-response/ Stay in control when things break: cloud incident response, powered by Datadog # Critical Response Rapid detection. Clear escalation. Fast recovery. Critical Response is incident-response-only cloud cover for AWS and Azure. We detect, triage, respond, and recover, within the coverage window you choose. No proactive engineering. Just reliable, SRE-driven incident management when you need it. Agents accelerate the analysis. A human owns the outcome of every incident. Talk to us Want proactive improvement too? Scroll 15 min SEV-1 response time Daytime · E&W · 24×7 Coverage options AWS + Azure Clouds covered Powered by Datadog Full observability always Who it's for - -> Teams that need after-hours or 24×7 cover but already have in-house day-to-day ops capability - -> Businesses that want to supplement their in-house on-call without replacing it - -> Scale-ups that need weekend and overnight cover as they grow but aren't ready for a full managed service Want proactive improvement too? See Critical Support or Critical Support Lite . How it works ## Five-stage incident lifecycle Every incident follows the same structured process, from first signal in Datadog to blameless postmortem. Stage 01 ### Monitoring **Detail +** Datadog telemetry, synthetic monitors, and alert rules watch your environment continuously. Bits AI SRE helps surface signal from noise. Stage 02 ### Triage **Detail +** On-call engineer classifies severity (SEV-1-4), assesses blast radius, and confirms ownership. Customer notified immediately for SEV-1. Stage 03 ### Response **Detail +** Runbooks executed, safe workarounds applied, rollback procedures followed as appropriate. All actions documented in real time in Datadog Incident Management. Stage 04 ### Escalation **Detail +** On-call routing, cloud-provider escalation, vendor coordination, and stakeholder communications, managed by our engineers so yours can focus on the fix. Stage 05 ### Recovery & Review **Detail +** Validated recovery, blameless RCA, and a written summary. Recovery time is a target (not a guarantee), SEV-1 60-120 min depending on plan. 01 Incidents covered ## What Critical Response handles - -> Service outages , complete or partial failures affecting end users or dependent services - -> Performance degradation , sustained latency spikes, elevated error rates, or throughput collapse - -> Security alerts , operational triage and containment only (not SOC/MDR/forensics) - -> Integration and API failures , broken upstream or downstream dependencies causing customer impact - -> Cloud provider incidents , AWS/Azure provider events that affect your environment, with response and workaround coordination 02 Plans ## Three plans, Daytime, Evenings & Weekends, 24×7 Choose the coverage window that fills your gap. All plans use the same 5-stage lifecycle and Datadog-native tooling. Response times are contractual. Recovery times are targets. Talk to us for pricing. Feature Daytime Evenings & Weekends 24×7 Coverage hours 09:00-17:00 Mon-Fri 17:00-09:00 Mon-Fri + 24×7 weekends Full 24×7×365 Severity covered SEV-1 SEV-1 & SEV-2 SEV-1 & SEV-2 Response time 15 min (SEV-1) SEV-1: 15 min · SEV-2: 30 min SEV-1: 15 min · SEV-2: 15 min Recovery target (SEV-1) 120-min target 90-min target 60-min target Incident management time/month 4 hrs 4 hrs 8 hrs Monitored services Up to 10 Up to 10 Up to 20 External endpoints monitored 1 @ 5-min interval 1 @ 2-min interval 5 @ 1-min interval Dashboards Standard Standard Standard + 1 custom Out-of-hours callouts None 2 per month 4 per month Runbooks & reports Standard + monthly summary Standard + monthly summary Customised + detailed RCA trends Recovery times are targets, not contractual guarantees. Response times (time to first engineer contact) are the contractual commitment. 03 Severity model ## Four severity levels, SEV-1 to SEV-4 Classification happens at triage. SEV-1 and SEV-2 trigger immediate response within contracted hours. SEV-1 · Critical Complete outage or material risk Total service unavailability, data loss risk, or severe breach of contractual obligations. Immediate response. 15-min response target on all plans. SEV-2 · High Significant degradation or partial outage Major feature failure, severe performance degradation, or partial loss of service affecting a significant number of users. Covered on E&W and 24×7 plans. SEV-3 · Moderate Limited impact, workaround available Non-critical issues with a viable workaround. Handled during business hours. Not covered under Critical Response out-of-hours plans. SEV-4 · Low Informational or minor Minor issues, informational alerts, or configuration questions. Handled in-hours. Not covered under Critical Response out-of-hours plans. 04 Shared responsibility ## Clear ownership during incidents Critical Cloud owns: detection, classification, response execution, escalation to vendors, stakeholder communication, and recovery validation. All documented in Datadog Incident Management. - -> We notify you at SEV-1 detection and at key recovery milestones. - -> Material changes (infrastructure, configuration) need your approval. You own: application code, business continuity decisions, customer communications, and access approvals. - -> You retain full IAM and admin control at all times. - -> You have full, real-time access to your Datadog environment throughout any incident. 05 Want proactive improvement too? ## Critical Support, incident management plus monthly engineering Critical Response covers you when things break. Critical Support also improves things so they break less often. **Read the full context +** Monthly improvement engineering across six pillars, reliability, security, cost, performance, automation, and governance. If you're spending engineering time on reactive firefighting, Critical Support is built to change that. Explore Critical Support Critical Support Lite ## FAQ **What is Critical Response? +** Critical Response is an incident-response-only service for AWS and Azure: detection, triage, response, escalation, and recovery within the coverage window you choose (Daytime, Evenings and Weekends, or 24×7). There is no proactive improvement engineering, for that, see Critical Support or Critical Support Lite . **Does Critical Response include proactive engineering? +** No. Critical Response is incident management only. Each plan includes a small allocation of incident management time for runbook maintenance and operational overhead, but no improvement engineering backlog. For proactive improvement, see Critical Support or Critical Support Lite . **What does "recovery target" mean? +** Recovery targets (e.g. 60-minute target for SEV-1 on 24×7) are working objectives we aim to meet. They reflect our operational capability and historic performance, but are not contractual guarantees, complex incidents take longer by their nature. Response time (time to first engineer contact) is the contractual commitment. **What happens after an incident? +** For all SEV-1 incidents: a blameless postmortem and written RCA summary, shared with you within the agreed timeframe. Findings can be fed into an improvement backlog if you're on Critical Support or Lite. ## Need reliable cover for the hours that matter? Tell us about your AWS or Azure environment and we'll recommend the right coverage window. Talk to us Critical Support --- ## /ai-operations/ Source: https://criticalcloud.ai/ai-operations/ The trust layer for AI operations # AI Runtime Operations. AI Runtime Operations is the AI expansion layer of Managed Runtime Assurance . We make AI workloads, agentic workflows and model-serving infrastructure observable, governed, cost-controlled and production-safe, while your team keeps ownership of the application, model and business logic. Critical Cloud delivers it as a Datadog AI operations partner: the world's first Powered by Datadog accredited MSP. Ship AI fast. Stay in control. We operate the stack. You own the product. Talk to us See how we operate Scroll World's first and only Powered by Datadog accredited MSP Advanced Datadog partner 60% Lower MTTR in production 24×7 UK-based operations 01 ## We operate the stack. You own the product. We operate, secure, and govern the stack your AI runs on. We never touch your app, your model, or your business logic. That boundary is what makes us a trustworthy, impartial layer: we have no agenda over your product, so we can stand behind whether your operations are sound. 02 The trust layer for AI operations ## Ship AI fast. Stay in control. Your product sits above the line. The reliability, security, and accountable governance that keep it in control sit below it. Explore the layers we own. Explore what we own You ship AI fast. We keep you in control of everything beneath the surface. Hover, tap, or focus a layer to see what each one does. Five layers, from the boundary down to the Datadog core. 03 ## Agents own the analysis. Humans own the outcome. Agents are good at the labour: root-cause analysis across telemetry no human can hold in their head, surfacing correlations, drafting fixes. What does not automate is ownership of the outcome. A human evaluates the plan, weighs the context the agent lacks, and stays accountable for accuracy, trust, and compliance. **01 In control of failures Datadog Bits AI detection and remediation, operated under our governance model. +** Datadog Bits AI detection and remediation, under our governance **02 In control of the attack surface AI Guard and runtime protection, operated by us on your behalf. +** AI Guard and runtime protection, operated by us **03 In control in production Agent Observability and the Agent Console, giving you full visibility in production. +** Agent Observability and the Agent Console **04 The foundation The full Datadog platform we are accredited to operate, end to end. +** The full Datadog platform we are accredited to operate 60% Lower MTTR in production 75% Lower than building an in-house team AWS + Azure Proven in production 24×7 Always-on coverage The 60% MTTR reduction is measured across the production environments we operate. Around 75% lower than building an in-house team. This is budget reallocation, not net new spend: the cost of recruiting, paying, tooling, and rota-ing an internal team for 24×7 cloud operations, redirected into a managed service that already runs at that standard. ## The foundation beneath the trust layer The trust layer is not a slogan. It is delivered by Critical Support, our 24×7 DevOps and SRE managed service, with Datadog as the operational platform. We are the world's first and only Powered by Datadog accredited MSP and a Datadog Advanced Partner, certified ISO 27001 and Cyber Essentials Plus. Explore Critical Support Our Datadog accreditation 04 ## AIOps, AI Ops and AI Operations: what is the difference? These terms are used interchangeably in the market. They mean different things. AIOps, short for Artificial Intelligence for IT Operations, uses machine learning and AI to improve how teams run technology infrastructure: detecting anomalies, correlating incidents, reducing alert noise and accelerating root-cause analysis. AI Operations is the broader operating model for teams running AI in production: reliability, security, observability, governance, cost control and accountable human oversight of autonomous systems. Critical Cloud sits at the intersection of both. We use AIOps tooling, including Datadog Watchdog and Bits AI, inside our operations service. And we operate the stack that AI systems depend on in production, with human governance of what agents do. We use AI inside operations, and we operate the stack AI runs on. **01 AIOps explained What AIOps is, how it reduces MTTR, and how Critical Cloud uses it inside managed operations. +** What AIOps is, how it reduces MTTR, and how Critical Cloud uses it inside managed operations **02 AI operations management The operating model for running AI systems safely in production: reliability, observability, security and governance. +** The operating model for running AI systems safely in production: reliability, observability, security and governance **03 Managed AI operations 24×7 managed service for the stack your AI depends on, delivered through Critical Support. +** 24×7 managed service for the stack your AI depends on, delivered through Critical Support **04 AI operations governance Human accountability for what autonomous systems do in production. +** Human accountability for what autonomous systems do in production **05 AI incident response 24×7 response to AI-specific failures: hallucination signals, prompt injection, cost runaway, agent misuse. +** 24×7 response to AI-specific failures: hallucination signals, prompt injection, cost runaway, agent misuse **06 Datadog for AI LLM observability, agent tracing, AI Guard and AI-powered operations on the Datadog platform. +** LLM observability, agent tracing, AI Guard and AI-powered operations on the Datadog platform ## FAQ **What is AI Operations? +** AI Operations is the operating model for running AI systems reliably and safely in production. It covers reliability, security, observability, governance, cost control, incident response and human accountability for the stack AI runs on. Critical Cloud delivers this as a managed service, 24×7, powered by Datadog. **What is AIOps? +** AIOps, short for Artificial Intelligence for IT Operations, uses machine learning and AI to improve how teams run technology infrastructure. Core capabilities include anomaly detection, event correlation, incident triage, root-cause analysis and alert noise reduction. We use Datadog Watchdog and Bits AI as AIOps tooling inside our managed operations service. **What is the difference between AIOps and AI Operations? +** AIOps is a technique: using AI to improve IT operations. AI Operations is the broader operating model for running AI systems in production. Critical Cloud does both. We use AIOps inside our operations, and we operate the infrastructure, observability, security and governance layer that AI systems need in production. **What is the trust layer for AI operations? +** The accountable layer that lets a company ship autonomous systems fast and stay in control of them in production. We operate, secure, and govern the stack the AI runs on, so the team can stay focused on the product. **What is an AI operations partner? +** A provider that takes ongoing operational responsibility for the stack a company's AI runs on: reliability, security, compliance, and accountable governance, 24×7. Critical Cloud delivers this through Critical Support, powered by Datadog. **Does Critical Cloud manage our AI model or application? +** No. We operate, secure, and govern the stack your AI runs on. We never touch your app, your model, or your business logic. That boundary is what makes us an impartial, accountable layer. **How does Datadog support AI Operations? +** Datadog provides LLM Observability for tracing model calls, Agent Tracing for monitoring autonomous agent behaviour, Watchdog for automatic anomaly detection, Bits AI for AI-assisted root-cause analysis, AI Guard for runtime security, and Cloud Cost Management for AI cost attribution. As the world's first and only Powered by Datadog accredited MSP, we operate these capabilities on behalf of customers. **Who is accountable when an AI agent proposes a production change? +** A human. Agents propose and draft; a Critical Cloud engineer evaluates the plan, weighs the context the agent lacks, and stays accountable for accuracy, trust, and compliance. Nothing changes in production without human ownership of the outcome. **Can Critical Cloud help with AI incident response? +** Yes. We respond 24×7 to AI-specific failure modes: latency spikes, token cost runaway, hallucination signals, prompt injection, tool misuse by agents, data exposure, and model or provider outages. A named engineer owns every incident. **Are you an AI consultancy or an AI security company? +** No. We are a managed service. We operate, secure, and govern your operational stack 24×7. Compliance and security are part of how we keep you in control, not separate advisory lines. ## Ship AI fast. Stay in control. Tell us what you are shipping. We will show you how the trust layer would work for your stack. Book a call --- ## /ai-operations/governance/ Source: https://criticalcloud.ai/ai-operations/governance/ AI operations governance # AI operations governance: agents own the analysis. Humans own the outcome. AI operations governance is the accountable layer around autonomous systems in production. As those systems do more, that layer matters more. We govern what your AI operations do: a human pair of eyes, accountable ownership of accuracy, and responsibility for compliance. Talk to us The trust layer Scroll 01 The governance spine ## What agents do well Agents are good at the labour: root-cause analysis across telemetry no human can hold in their head, surfacing correlations across services and time windows, drafting remediation plans, and doing it at any hour without fatigue. Used well, they collapse the time between a signal firing and a credible diagnosis existing. ## What does not automate Ownership of the outcome. A human evaluates the plan, weighs the context the agent lacks, the business context, the change freeze, the customer commitment, and stays accountable for accuracy, trust, and compliance. When something changes in production, a named engineer owns that change. 02 Governance in practice ## Governance in practice, on the Datadog platform The governance model is not abstract. Each part of staying in control maps to a capability we operate every day. In control of failures Datadog Bits AI detection and remediation, under our governance In control of the attack surface AI Guard and runtime protection, operated by us In control in production Agent Observability and the Agent Console The foundation The full Datadog platform we are accredited to operate 03 Governance vs automation ## Governance is the difference between AI-assisted operations and uncontrolled automation AI can analyse telemetry at a scale and speed no human team can match. It can correlate signals across thousands of services, surface root-cause hypotheses in seconds, and draft remediation plans before a human has finished reading the alert. That is genuinely useful. It is also genuinely risky if the output of that analysis goes directly to production without review. **Read the full explanation +** The difference between AI-assisted operations and uncontrolled automation is a human in the loop with real ownership. AI can analyse and propose. A named human owns what changes in production. Every agent-generated plan is evaluated against context the agent cannot see: the deployment state, the customer commitments, the regulatory obligations, the change freeze. Governance is not a brake on what AI can do. It is the thing that makes it safe to use AI at all. 04 The boundary ## We govern the operational outcome, not the product outcome We operate, secure, and govern the stack your AI runs on. We never touch your app, your model, or your business logic. Governance here means accountable ownership of what happens in production: that incidents are owned, that agent actions are evaluated by a human, and that the operational record stands up to scrutiny. Explore Critical Support Critical Response ## In this cluster AI operations overview AIOps explained AI operations management Managed AI operations AI incident response Datadog for AI ## FAQ **Who is accountable when an agent acts? +** A human, always. Agents propose, draft, and accelerate; a Critical Cloud engineer evaluates the plan, weighs the context the agent lacks, and stays accountable for accuracy, trust, and compliance. Nothing changes in production without accountable human ownership of the outcome. **What is the difference between AI-assisted operations and uncontrolled automation? +** AI-assisted operations uses agents to accelerate analysis and surface recommendations, with a human evaluating and owning every outcome. Uncontrolled automation lets agents act directly in production without human review. The governance model is the difference: a named engineer accountable for accuracy, trust, and compliance at every step. **Do you govern our model or our product? +** No, we govern the operational outcome, not the product outcome. We operate, secure, and govern the stack your AI runs on. Your application, your model, and your business logic stay yours. That boundary is what makes us an impartial, accountable layer. **How does governance work in practice on the Datadog platform? +** Every governance principle maps to a Datadog capability we operate: Bits AI for AI-assisted incident analysis under human review, AI Guard for runtime security and prompt injection protection, Agent Observability for monitoring what autonomous agents do, and the full Datadog telemetry stack for the audit trail that accountability requires. ## Ship AI fast. Stay in control. Tell us what your autonomous systems do. We will show you what accountable governance looks like for your operations. Talk to us --- ## /security/ Source: https://criticalcloud.ai/security/ Security & Compliance # Certified, audited, and operationally accountable. Critical Cloud holds ISO 27001 and Cyber Essentials Plus. Our operations, incident management, change control, observability, and access governance, are designed to help customers meet their own regulatory obligations, including DORA and NIS2. EU and UK data residency is available. Request documentation Partner credentials Scroll ISO 27001 Certified, ISO/IEC 27001:2022 CE Plus Cyber Essentials Plus certified EU + UK Data residency options DORA & NIS2 Operational support for regulated customers Our posture at a glance - -> ISO 27001:2022 certified , independent audit of our information security management system - -> Cyber Essentials Plus certified , UK government-backed scheme, technically tested - -> Least-privilege access , IAM controls, just-in-time access, no standing admin sessions - -> Structured incident management , SEV-based severity model, documented postmortems - -> Change control , CAB process, customer approval for material changes - -> Sub-processor transparency , documented, available on request 01 Certifications ## Independently audited and certified Both certifications are maintained and renewed, not a one-off assessment. Documentation is available on request for procurement and due diligence processes. Information security management ### ISO/IEC 27001:2022 ISO 27001 is the international standard for information security management systems (ISMS). Certification means our policies, controls, and processes for managing information security risk have been independently audited and verified against the standard's requirements. - -> Risk assessment and treatment across our operations - -> Access control, cryptography, and physical security controls - -> Incident management, business continuity, and supplier relations - -> Annual surveillance audits and three-year recertification cycle Certificate and scope statement available on request. UK government-backed cyber security ### Cyber Essentials Plus Cyber Essentials Plus is the UK government-backed certification scheme covering five core cyber security controls. The "Plus" level involves independent technical testing, not self-assessment, verifying that the controls are actually in place and functioning. - -> Firewalls and network boundary controls - -> Secure configuration of devices and services - -> User access control and privilege management - -> Malware protection and software patching, technically verified Certificate available on request for supplier onboarding and due diligence. 02 Regulatory readiness ## Supporting customers subject to DORA and NIS2 We don't provide legal advice, and regulations place obligations on you, not on us. What we do is run the operational infrastructure, observability, incident management, change control, postmortems, in a way that is designed to help you evidence and meet those obligations. ### DORA, Digital Operational Resilience Act DORA requires EU financial services firms to demonstrate operational resilience: ICT risk management, incident classification and reporting, digital operational resilience testing, and oversight of third-party ICT providers. How our operations support this: - -> ICT incident management: SEV-based incident classification, documented response timelines, and blameless postmortems, the audit trail DORA reporting requires - -> ICT risk management: ISO 27001 ISMS provides the risk assessment and treatment framework; Datadog observability provides continuous visibility into the operational risk posture - -> Third-party ICT oversight: we operate as a regulated third-party ICT provider; sub-processor documentation, SLA evidence, and audit rights are available on request - -> Resilience testing support: our alliance with Tarian Labs can provide penetration testing and resilience assessments alongside your DORA testing programme, see Continuous Runtime Security Validation ### NIS2, Network and Information Security Directive NIS2 extends security and incident-reporting obligations to a broader set of sectors, including digital infrastructure, managed service providers, and cloud services, across EU member states. It requires risk management measures, supply chain security, and timely incident notification. How our operations support this: - -> Risk management measures: ISO 27001 ISMS, change control, access governance, and the improvement engineering programme address the technical and organisational measures NIS2 requires - -> Incident reporting: our SEV-1 incident management process documents timeline, impact, and root cause, the raw material for NIS2 72-hour notification obligations - -> Supply chain security: sub-processor list and contractual security requirements available; we apply security obligations to our own suppliers - -> MSP scope: as a managed service provider, we operate within NIS2's scope and maintain the security measures the directive requires of providers in our category This is not legal advice. DORA and NIS2 impose obligations on your organisation as the regulated entity. How our operations map to your specific obligations depends on your sector, jurisdiction, and circumstances. We can provide operational documentation to support your compliance programme, talk to us or your legal advisers for guidance specific to your situation. 03 Data residency ## EU, UK, and data sovereignty options Where your observability and operational data is stored matters, for GDPR, for Swiss nDSG, and for customers in regulated sectors. Datadog offers both EU and US site options; we configure your environment to match your residency requirement. Region / requirement Datadog site Relevance EU data residency Datadog EU1 site (eu1.datadoghq.com), data stored in Frankfurt, Germany Meets GDPR cross-border transfer requirements for EU customers; supports DORA data-localisation preferences UK data residency Datadog US1 or EU1 (UK adequacy decision applies); Azure UK regions for cloud workloads UK GDPR-compatible; UK adequacy decision maintains equivalence for EU→UK transfers post-Brexit Switzerland (nDSG) Datadog EU1 (data stored in Germany / EU); Switzerland-EU SCCs where relevant Swiss nDSG aligns closely with GDPR; EU1 site and appropriate transfer mechanisms support Swiss data-residency expectations Ireland Datadog EU1 GDPR-native; EU site ensures all observability data remains within the EEA EMEA (other markets) EU1 or US1, configured per customer requirement We advise on the right Datadog site for each customer's jurisdiction and sector obligations Datadog's own sub-processors and data processing addendum are published by Datadog at datadoghq.com/legal/sub-processors . Critical Cloud's sub-processor list is available on request. Operating model ## How security is built into how we operate Security controls are operational, not documentary. They show up in how we manage access, run changes, and handle incidents, every day, for every customer. ### ISO 27001 IMS Our Information Management System defines how we assess and treat risk, manage policies, train staff, audit controls, and handle non-conformances. Annual surveillance audit; three-year recertification cycle. ### Structured incident management SEV-1 to SEV-4 severity classification, documented response timelines, and blameless postmortems for all SEV-1 events. Incident records include timeline, impact, root cause, and remediation actions. ### Change control & CAB All material changes to customer infrastructure go through our Change Advisory Board process. Customers approve changes that affect their environment. Emergency changes are documented post-event with full rationale. ### Least-privilege access Access to customer environments follows least-privilege principles. IAM controls and access governance reviews are part of the standard operating model. Customers retain admin and IAM control at all times. ### Observability and audit trail Datadog provides a continuous, tamper-evident audit trail of operational activity. Customers have full, real-time access to their own Datadog environment, visibility is not restricted to Critical Cloud engineers. ### Sub-processor handling We maintain a documented sub-processor list covering the tools and services used in delivering managed services. Available on request for procurement, DPA, and compliance purposes. 04 Runtime Evidence ## Runtime Evidence, not compliance theatre. Good operations should produce evidence. Critical Cloud operates and evidences the runtime controls that support customer, auditor and regulatory obligations: incident records, access governance, change context, recovery evidence, SLO reporting and assurance packs. We operate and evidence the runtime controls that support your compliance obligations. Learn more about Runtime Evidence as part of Managed Runtime Assurance. ## Need security documentation for procurement or compliance? ISO 27001 certificate, Cyber Essentials Plus certificate, sub-processor list, and DPA terms, available on request. Request documentation About Critical Cloud --- ## /about/ Source: https://criticalcloud.ai/about/ About Critical Cloud # We operate the runtime layer behind AI-era software. Critical Cloud is the trust layer for AI operations . We operate, secure, and govern the stack a company's AI runs on, so AI-first teams can ship fast and stay in control. We are the world's first and only Powered by Datadog accredited MSP, an independent designation awarded by Datadog after formal engineering review. We are an AI-native and Datadog-native business built for the next generation of cloud operations, where observability, AI, and operational discipline are the product. Talk to us Join the team Scroll World's first Powered by Datadog MSP AWS + Azure Multi-cloud managed operations Cardiff · London Dublin, UK & EMEA ISO 27001 Cyber Essentials Plus What we are - -> Datadog-native MSP , Powered by Datadog is the operational backbone, not a monitoring add-on. - -> AI-native operations , Bits AI SRE, Watchdog, and Datadog's AI capabilities built into how we run. Agents own the analysis. Humans own the outcome. - -> Founder-led , the same people who built and exited a managed service business before are doing it again, with more depth. - -> Customer data, customer control , we operate, secure, and govern the stack your AI runs on. We never touch your app, your model, or your business logic. Serving tech-led SMBs across the UK, Ireland, and EMEA. Our Datadog credentials → 01 Credentials ## Independently verified, not self-declared Every credential below is awarded or audited by the organisation named, not a badge we printed ourselves. 02 Leadership ## The team Founder-led from day one. The same people who sold services, built the technical platform, and managed the customer relationships before are doing it again, with the benefit of having done it once. JS James Smith Founder · CEO · Cardiff - -> Petrol head and start-up junkie. Started as an engineer at Dell, built Pizza Hut's first UK online ordering system, then founded and scaled a cloud services company to 100+ people over a decade before selling it. Then started Critical Cloud. Some people learn when to stop. - -> Director of the Year (SME) and Technology Leader of the Year. Happiest either underneath an engine or in front of a whiteboard. Both involve understanding how complex systems actually work. LinkedIn AP Andrew Phillips COO · Cardiff - -> The person who makes sure Critical Cloud actually delivers what it promises. 25+ years in managed services across digital agencies, cloud delivery, and technology consulting means Andrew has seen every way a service can go wrong and quietly makes sure it doesn't. He holds Azure, AWS, Agile and DevOps certifications because he actually uses them. - -> Golfer and lifelong Everton FC fan, which means exceptional patience, a high tolerance for adversity, and an unshakeable belief that things will eventually come good. LinkedIn CW Chris Webb CRO · London - -> 25+ years selling complex technical services to people who are very good at spotting nonsense. Has joined early-stage cloud and managed services businesses, built the commercial function, and seen them through to successful exits. Datadog Sales Specialist certified. - -> Card magician and pétanque world champion. If you've ever wanted to watch a CRO make a deal disappear and then reappear exactly where you needed it, you're in the right place. LinkedIn 03 Engineering ## The founding SRE team The team that keeps customers' platforms running. Collectively they have decades of cloud engineering experience operating real production environments across AWS and Azure, and they're all certified in Datadog, AWS and Microsoft Azure. This is the team you'd actually be working with. LR Linsey Rimmer Senior Site Reliability Engineer · Cardiff - -> Our Datadog whisperer. When a dashboard is noisy, an alert is misfiring, or Datadog needs to actually tell you something useful, Linsey's the one who makes it happen. - -> Triathlete off the clock, which explains a lot about the stamina and the refusal to leave a problem half-solved. - -> MSc Computer Science, Cardiff University. LinkedIn JW Jamie Ward Senior Site Reliability Engineer · Cardiff - -> Our AWS guru. Multi-account architectures, landing zones, the full Well-Architected stack - Jamie's the one you want when the platform needs to be done properly, not just done quickly. - -> Drummer when not debugging. The combination of precision and rhythm probably explains why his runbooks always make sense. - -> MSc Software Engineering + BSc Physics, Cardiff and Bristol. LinkedIn MD Matt Dibble Senior Site Reliability Engineer · Cardiff - -> Our Azure specialist. AKS, Azure DevOps, complex enterprise environments - Matt builds things on Azure that hold together under pressure. - -> Carpenter and welder outside work. There's a theme here: someone who builds things properly, whether it's infrastructure or furniture, and doesn't cut corners. LinkedIn DJ Dafydd Jones Senior Site Reliability Engineer · Cardiff - -> Can do it all. AWS, Azure, Datadog, infrastructure, security, whatever the customer needs - Dafydd's the person you put on the hardest problems because he'll figure it out. - -> Genuinely calm when things are on fire, which is the energy you want on-call at 3am. Native Welsh speaker, and arguably the most versatile engineer on the team. LinkedIn WS William Sawtell Associate AI & Data Engineer · Cardiff - -> Building the AI tooling that makes the rest of the team faster: applications that connect cloud infrastructure, automation and observability in ways that actually work in production. - -> Music lover with an ML and reinforcement learning background. Probably the only person on the team who thinks about algorithms and chord progressions with equal enthusiasm. - -> BSc Computer Science, Cardiff University. LinkedIn See yourself being part of our team? We're a small, tight-knit group and we hire carefully. If this is the kind of team you want to work with, take a look at what we're building. View open roles -> 04 Where we work ## Cardiff. London. Dublin. The people above are the founding team. Behind them is a growing group of engineers, delivery managers, commercial and operations people spread across three offices. Cardiff is the engineering and operations hub. London runs commercial and enterprise relationships. Dublin is where we deliver for Ireland and EMEA customers. Cardiff · HQ ### Engineering & Operations The SRE team, technical delivery, and platform operations are all based in Cardiff. This is where the on-call function lives and where the majority of hands-on engineering work happens. Cloud engineering, AI tooling, and infrastructure work all run from here. London ### Commercial & Partnerships Sales, client management, and strategic partnerships are run from London. This is where enterprise and financial services relationships are managed and where we engage with Datadog, AWS, and Microsoft at a commercial level. Delivery management for UK enterprise accounts also sits here. Dublin ### Ireland & EMEA Delivery Our Dublin presence covers Ireland-based customers and acts as the delivery hub for EMEA engagements. Client-facing account management, service delivery and commercial relationships for Irish and European customers are managed from here, close to where those customers operate. We're growing. The team is expanding across all three locations. Cloud engineering, delivery management, and commercial roles. If you want to build something serious, take a look at what we're hiring for . View roles How we operate ## Eight principles, one operating model These aren't values on a wall. They're the criteria we use to make decisions and the criteria we'd ask anyone joining the team to hold themselves to. 01 ### Customer First Every decision starts with what's right for the customer, not what's convenient internally. If the customer doesn't feel the benefit, it doesn't count. 02 ### Own the Problem We take problems to resolution. We don't hand them off and hope. When something breaks, one person owns it through to the postmortem. 03 ### Engineer Simplicity The best solution is the one that doesn't need maintaining. Build for the on-call engineer at 3am who has no context, not for the engineer who built it. 04 ### Stay Curious The people who grow fastest here ask why about every system they touch. Curiosity about infrastructure, customers, and the business is what turns operators into engineers. 05 ### Operate at Scale We run multiple customer environments simultaneously. Everything we build, runbooks, standards, automation, has to work without us in the room. 06 ### Move with Urgency Speed of communication is a service quality metric. Slow responses compound. Even "we're looking at it" is better than silence. 07 ### Be Resourceful Use what's available, improvise where needed, ask for help when stuck. The constraint that stops most teams doesn't have to stop us, it's usually a process problem, not a skills problem. 08 ### Earn Trust by Delivering Consistency is the only currency that matters. Every stable environment, resolved incident, and clean service review is how we earn the right to the next one. ## Work with us, or join us. Talk to the team about your cloud and Datadog challenges or read the job specs if you want to be part of building it. Talk to us Careers Datadog credentials ---