Azure Well-Architected Review

Before an Azure workload goes live or changes hands, establish what could stop it operating safely and what must be fixed first. Read-only, evidence-backed, and scoped to one workload.

Someone has to say whether it can go live

The date usually arrives before the certainty does. A launch is booked, operational ownership passes to another team, an acquisition hands the estate over, or a second incident turns the first one into a pattern. Someone senior then has to say whether the workload can carry production traffic, and say it on the record.

This review is the engineering evidence behind that answer. For one bounded workload it establishes what could stop it operating safely, what must be fixed before the date, and what can wait until after it.

You keep the decision. The review supplies the case for it, in a form you can hand to a board, an auditor or the team taking the workload on.

What the review hands you

All of it for one explicitly bounded workload, agreed in writing before the work starts.

  • The critical flows, and the operational requirements agreed as in scope
  • Release or handover blockers, each backed by evidence from your own estate
  • Evidence gaps named as gaps, including recovery claims nobody has tested
  • A prioritised backlog, every item with an owner, an effort estimate and an acceptance condition
  • An engineering readiness recommendation, and a readout that walks your team through it
  • The Well-Architected mapping behind the findings, all 59 checks with the evidence for each
  • A dated baseline, so a later review can show what actually moved

The review is sold separately from the work it recommends. The evidence and the backlog are yours to use with any delivery team, and remediation is quoted on its own. See the production delivery route →

Where the workload is not staying where it is, the move itself is a separate piece of work with its own sequencing, cutover and rollback. How an Azure migration runs →

Microsoft's framework, in full

The readiness decision is the product. Microsoft's Well-Architected Framework is how it is reached: five pillars, 59 checklist recommendations, and every one of them judged against your workload.

  • REReliability10
  • SESecurity12
  • COCost Optimization14
  • OEOperational Excellence11
  • PEPerformance Efficiency12
  • Total59

Reliability

Resilience to malfunction, and recovery to a fully functioning state after failure.

Two sturdy interlocked rings, blue and turquoise, holding firm over the line they bear on

Design principles

  • Design for business requirements
  • Design for resilience
  • Design for recovery
  • Design for operations
  • Keep it simple

Design review checklist · all 10, rated

  • RE:01Focus your workload design on simplicity and efficiency
  • RE:02Identify and rate user and system flows
  • RE:03Use failure mode analysis (FMA) to identify potential failures
  • RE:04Define reliability and recovery targets
  • RE:05Add redundancy at different levels, especially for critical flows
  • RE:06Implement a timely and reliable scaling strategy
  • RE:07Strengthen resiliency with self-preservation and self-healing measures
  • RE:08Test for resiliency and availability using chaos engineering
  • RE:09Implement structured, tested and documented disaster recovery (DR) plans
  • RE:10Continuously measure and track system health

Security

Protecting the confidentiality, integrity and availability of the workload and its data.

A closed chain link forming the shackle of an abstract padlock, sealed shut

Design principles

  • Plan your security readiness
  • Design to protect confidentiality
  • Design to protect integrity
  • Design to protect availability
  • Sustain and evolve your security posture

Design review checklist · all 12, rated

  • SE:01Establish a security baseline aligned to compliance requirements and standards
  • SE:02Align a secure development lifecycle (SDL) throughout the software development lifecycle
  • SE:03Classify and consistently apply sensitivity and information-type labels
  • SE:04Create intentional segmentation and perimeters
  • SE:05Implement strict, conditional and auditable identity and access management (IAM)
  • SE:06Isolate, filter and control network traffic across ingress and egress flows
  • SE:07Encrypt data using modern, industry-standard methods
  • SE:08Harden all workload components
  • SE:09Protect application secrets
  • SE:10Implement a holistic monitoring strategy with modern threat detection
  • SE:11Establish a comprehensive testing regimen
  • SE:12Define and test effective incident response procedures

Cost Optimization

The highest return per unit of spend - cost treated as a first-class design constraint.

A balanced beam with a stack of small blue tokens level against one larger turquoise token

Design principles

  • Develop cost-management discipline
  • Design with a cost-efficiency mindset
  • Design for usage optimization
  • Design for rate optimization
  • Monitor and optimize over time

Design review checklist · all 14, rated

  • CO:01Create a culture of financial responsibility
  • CO:02Create and maintain a cost model
  • CO:03Collect and review cost data
  • CO:04Set spending guardrails
  • CO:05Get the best rates from providers
  • CO:06Align usage to billing increments
  • CO:07Optimize component costs
  • CO:08Optimize environment costs
  • CO:09Optimize flow costs
  • CO:10Optimize data costs
  • CO:11Optimize code costs
  • CO:12Optimize scaling costs
  • CO:13Optimize personnel time
  • CO:14Consolidate resources and responsibility

Operational Excellence

DevOps culture, standardised process, observability, and safe, repeatable deployment.

Three chain links meshed in a row like a well-run production line, checks passing above each

Design principles

  • Embrace DevOps culture
  • Establish development standards
  • Evolve operations with observability
  • Automate for efficiency
  • Adopt safe deployment practices

Design review checklist · all 11, rated

  • OE:01Define standard practices to develop and operate the workload
  • OE:02Use standardization for routine, ad-hoc and emergency operations
  • OE:03Formalize processes across the full software development lifecycle
  • OE:04Enhance software development and quality assurance
  • OE:05Use a standardized infrastructure as code (IaC) approach
  • OE:06Build a workload supply chain with automated pipelines
  • OE:07Design a monitoring stack
  • OE:08Establish a clear, structured incident management process
  • OE:09Enhance the quality of your workload through testing
  • OE:10Design automation to be reliable, secure and maintainable
  • OE:11Clearly define safe deployment practices

Performance Efficiency

Meeting performance targets efficiently as demand and the system evolve.

A single chain link leaning forward in motion with speed lines trailing it

Design principles

  • Negotiate realistic performance targets
  • Design to meet capacity requirements
  • Achieve and sustain performance
  • Optimize for long-term improvement

Design review checklist · all 12, rated

  • PE:01Define performance targets
  • PE:02Conduct capacity planning
  • PE:03Select the right services
  • PE:04Establish consistent performance measurement
  • PE:05Optimize scaling and partitioning
  • PE:06Optimize performance by testing in a production-like environment
  • PE:07Optimize code and infrastructure
  • PE:08Optimize data usage
  • PE:09Prioritize the performance of critical flows
  • PE:10Optimize operational tasks
  • PE:11Respond to live performance issues
  • PE:12Continuously optimize performance

How the review runs

  1. Scope

    The one workload, its subscriptions, its critical flows, and the operational requirements it has to meet. The decision waiting on the review, and its date, are agreed here in writing.

  2. Baseline

    A read-only collection from your estate: resource inventory, Azure Advisor signal and Defender secure score.

  3. Assess

    All 59 checks judged against your estate - met, partial, gap or not applicable - with the evidence for each.

  4. Rate

    Every gap rated for severity and for the effort to fix it, then marked as a blocker for the date or as work that can follow it. Tradeoffs you made on purpose are recorded as decisions.

  5. Readout

    The readiness recommendation and the prioritised backlog, walked through with your team.

Two to three weeks, kickoff to readout · read-only access, nothing changed

Why a review, not a scan

The framework is public and anyone can read the 59 recommendations. What you pay for is a principal engineer applying them to your workload and standing behind the readiness call that comes out.

Judgement a scan cannot have
at least 40 of the 59
What a scan can read
at most 19 - configuration and telemetry

Findings from someone who has built it

The engineer who finds the gap is available to close it. The review itself is read-only and its output is evidence; remediation is quoted separately, as its own engagement or a Production Delivery Sprint.

All 59 rated, because DBHQ built the tooling to

All 59, across every subscription in scope, with nothing extrapolated from a favourite few. DBHQ looks at everything because the tooling makes looking cheap. The judging is not automated, and will not be.

Judgement a scan cannot have

About two-thirds of the checklist cannot be read off a machine: whether you have done failure-mode analysis, whether your data is classified, whether your segmentation was deliberate, whether your recovery targets match what the business can carry.

Your tradeoffs recorded as decisions

Microsoft's framework says plainly that tightening one pillar costs another. A scan marks it wrong. The review records why you chose it, and whether it still holds.

Evidence, severity and effort on every finding

Each gap is rated for what it costs you against what it costs to fix, and marked as a blocker for the date or as work that can follow it. Sequence it, or hand it to a team.

Independent

No licences to resell, no migration to sell, no partner funding behind the findings. A cloud vendor assesses you free because the assessment serves the sale that follows it.

Twenty-six years building where failure was expensive - defence, banking, insurance, energy, commodities, healthcare and identity.

DBHQ builds its own tools. Here is one, free

An open tool that scans your Azure estate and scores the five pillars from the platform's own signals. Deterministic, read-only, no sign-up, and nothing leaves your tenant. It runs in about a minute and shows roughly where you stand.

A machine sees configuration and telemetry. The judgement, and the readiness call resting on it, are what the review adds.

Get the free Well-Architected taster →

Questions, answered

What access do you need?
Read-only. The review works from resource inventory, configuration and the platform's own signals - Azure Advisor and Defender. Nothing in your live environment changes.
Do you make the go-live decision?
No. You do. The review gives you an engineering recommendation and the evidence behind it, in a form you can put in front of a board, an auditor or the team taking the workload on. The accountable owner still signs it.
Isn't the Well-Architected Review free from Microsoft?
Microsoft's own assessment is a self-service questionnaire - you grade your own homework. Partners and resellers will do more than that for nothing, including a technical review, workshops and a written report. This review uses the same framework. What differs is that a principal engineer does the judging, evidences every finding against your real resources, has no licence or migration riding on the answer, and produces a readiness decision for a named date rather than a coverage report.
How is this different from a free vendor assessment?
A cloud vendor or reseller funds the assessment because they then sell you the migration, the managed service or the licences. The findings serve that sale. This review has no product behind it and no licences to resell.
What does it cost?
A fixed fee for one agreed workload and readiness decision, quoted after a short call - never open-ended. The evidence and the remediation backlog are yours to use with any delivery team. Implementation is quoted separately.
What if we don't go ahead with the fixes?
The review, the evidence and the backlog are yours, and useful to any engineer. You are not obliged to use DBHQ for the remediation.
How long does it take?
Two to three weeks from kickoff to readout, depending on the size of the workload.

An independent review conducted against the Microsoft Azure Well-Architected Framework. DBHQ is not affiliated with, endorsed by, or certified by Microsoft.

Fifteen minutes will tell you if this fits

Bring the workload and the date it has to clear. You will get a straight answer on whether a readiness review is the right instrument, what it would cover and when it would report.

I reply within 24 hours