# Multi-Cloud Monitoring: How to Watch AWS, Azure and GCP Together

> Multi-cloud monitoring guide: why the same metric differs across AWS, Azure and GCP, and how to normalize it with OpenTelemetry into one view.

Source: https://last9.io/blog/multi-cloud-monitoring/

Multi-cloud monitoring means watching workloads on AWS, Azure and Google Cloud from one place, with the same names, labels and alerts for all of them.

The hard part is the data underneath the dashboards. Each cloud names, scales and samples its metrics differently, so CPU on AWS is a percent with 5-minute data points by default, while CPU on Google Cloud is a fraction between 0 and 1. Put them on one chart without fixing that and the chart lies.

This guide covers what changes when you monitor more than one cloud, the differences to normalize, how to do it with OpenTelemetry resource attributes and a Collector in each cloud, where egress costs come in, and how to alert across clouds. Metric definitions come from the AWS, Azure and Google Cloud documentation, and attribute names come from the [OpenTelemetry semantic conventions for cloud resources](https://opentelemetry.io/docs/specs/semconv/resource/cloud/).

## What is multi-cloud monitoring?

Multi-cloud monitoring is collecting metrics, logs and traces from services that run on two or more public clouds and querying them as one system. A checkout service on AWS that calls a pricing service on Google Cloud should show up as one request path, with one latency number and one alert, whichever cloud is at fault.

It sits next to two related terms:

- **Hybrid cloud monitoring** covers public cloud plus private infrastructure, such as an on-premises data center. The techniques below apply there too, since an on-premises host is one more environment with its own labels.
- **Cloud monitoring** usually means monitoring one provider, often with that provider's own tool. Our roundup of [cloud monitoring tools](https://last9.io/blog/best-cloud-monitoring-tools/) covers the per-cloud options.

Teams end up on several clouds for ordinary reasons: an acquisition brings a second cloud, a data team picks a different provider for analytics, or a customer contract requires a specific region. The monitoring setup rarely gets planned for that, so each cloud keeps its own tool and on-call engineers switch between consoles during an incident.

## Why is monitoring several clouds harder than monitoring one?

Each cloud's native tool works well inside its own cloud. CloudWatch, Azure Monitor and Google Cloud Monitoring all collect their own platform metrics with no setup. The trouble starts when you compare them, because the same idea arrives in three different shapes.

CPU utilization for a virtual machine is the simplest example:

|                    | AWS EC2                                        | Azure Virtual Machines                   | Google Compute Engine                                    |
| ------------------ | ---------------------------------------------- | ---------------------------------------- | -------------------------------------------------------- |
| Metric name        | `CPUUtilization`                               | `Percentage CPU`                         | `compute.googleapis.com/instance/cpu/utilization`        |
| Unit               | Percent                                        | Percent                                  | Fraction, typically 0.0 to 1.0                           |
| Default resolution | 5 minutes (1 minute with detailed monitoring)  | 1 minute                                 | Sampled every 60 seconds                                 |
| What it measures   | Physical CPU time EC2 uses to run the instance | Allocated compute units in use by the VM | Utilization of allocated CPU, reported by the hypervisor |

Three things in that table cause real incidents:

- **Units.** [Google Cloud documents](https://docs.cloud.google.com/monitoring/api/metrics_gcp_c) its VM CPU metric as "Fractional utilization of allocated CPU on this instance", with values typically between 0.0 and 1.0. Its own charts display it as a percentage, which hides the difference until you export the raw values. An alert written as `cpu > 80` for AWS and Azure stays silent on Google Cloud, because 0.95 is never greater than 80.
- **Resolution.** [AWS documentation](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html) states that each EC2 data point covers 5 minutes by default, or 1 minute with detailed monitoring. Azure's [Percentage CPU](https://learn.microsoft.com/en-us/azure/azure-monitor/reference/supported-metrics/microsoft-compute-virtualmachines-metrics) has a 1-minute time grain. A 2-minute spike can show on Azure and disappear into a 5-minute average on AWS.
- **Delay.** Google Cloud notes that after sampling, this metric's data "is not visible for up to 240 seconds." Every cloud has some ingestion delay, and alert windows shorter than that delay produce false resolves and late pages.

The same pattern repeats for memory, disk, network, load balancer latency and managed database metrics. Some clouds report memory for VMs only after you install an agent, some label regions as `us-east-1` and others as `eastus` or `us-central1`, and account identifiers mean different things (AWS account, Azure subscription, Google Cloud project).

## What should a multi-cloud monitoring setup standardize?

Before picking tools, decide what every signal must look like when it reaches your backend. Four things matter most.

**1. Where the data came from.** Every metric, log and span needs the same attributes for provider, region, account and resource. The OpenTelemetry semantic conventions already define them, so there is no need to invent your own:

| Attribute                 | Meaning                                                                                                                              | Example values                              |
| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------- |
| `cloud.provider`          | The cloud provider                                                                                                                   | `aws`, `azure`, `gcp`                       |
| `cloud.platform`          | The service running the workload                                                                                                     | `aws_ec2`, `azure.vm`, `gcp_compute_engine` |
| `cloud.region`            | The provider's region name                                                                                                           | `us-east-1`, `eastus`, `us-central1`        |
| `cloud.availability_zone` | The zone inside the region                                                                                                           | `us-east-1c`, `us-central1-a`               |
| `cloud.account.id`        | The account that owns the resource; the subscription ID on Azure                                                                     | `111111111111`                              |
| `cloud.resource_id`       | The provider's full ID for the resource: an ARN on AWS, a fully qualified resource ID on Azure, a full resource name on Google Cloud | `arn:aws:lambda:us-east-1:...`              |

With these in place, one query can group by `cloud.provider` or filter to one region on every signal, whatever cloud it came from.

**2. Your own labels.** Cloud attributes describe infrastructure. Questions during an incident are about services and owners, so add the same small set of labels everywhere: `service.name`, `deployment.environment`, and a team or owner label. Enforce them in the Collector rather than trusting every team's tagging policy in every cloud.

**3. Units and names for the metrics you alert on.** Pick one unit per kind of measurement (percent or fraction for utilization, seconds or milliseconds for latency) and convert at ingest. You don't need to rename every metric. Start with the twenty or so behind your dashboards and alerts.

**4. Resolution and freshness.** Know the default resolution and ingestion delay for each source, and set alert windows longer than the slowest one. Where a 1-minute resolution matters, turn it on at the source, for example with EC2 detailed monitoring.

## How do you normalize multi-cloud telemetry with OpenTelemetry?

The most practical pattern is to run an OpenTelemetry Collector in each cloud and have all of them export to one backend. Each Collector adds the cloud attributes, fixes units, and drops what you don't need before anything leaves the cloud. Our introduction to the [OpenTelemetry Collector](https://last9.io/blog/what-is-opentelemetry-collector/) explains receivers, processors and exporters if you are new to it.

### Add cloud attributes with resource detection

The Collector's [resource detection processor](https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/processor/resourcedetectionprocessor/README.md) reads metadata from the environment it runs in and adds it as resource attributes. It ships detectors for each major cloud, including `ec2`, `ecs`, `eks` and `lambda` on AWS, `azure` and `aks` on Azure, and `gcp` on Google Cloud.

```yaml
processors:
  resource_detection/cloud:
    detectors: [env, ec2, azure, gcp]
    timeout: 2s
    override: false
```

In current Collector releases the processor type is `resource_detection`; older releases call it `resourcedetection`. Setting `override: false` keeps attributes your SDKs already set, since the default is to overwrite them. Listing the detector for every cloud lets you ship one config file to all three; keeping only the local cloud's detector in each Collector means it queries only the metadata endpoint that exists where it runs.

### Convert units at ingest

The transform processor uses the OpenTelemetry Transformation Language (OTTL) to change data in flight. This statement converts Google Cloud's CPU fraction to a percent so it matches AWS and Azure:

```yaml
processors:
  transform/gcp_cpu:
    error_mode: ignore
    metric_statements:
      - set(datapoint.value_double, datapoint.value_double * 100) where IsMatch(metric.name, ".*instance/cpu/utilization")
      - set(metric.unit, "%") where IsMatch(metric.name, ".*instance/cpu/utilization")
```

The exact metric name depends on how you ingest Google Cloud metrics, so check the name in your backend before you match on it. Going the other way works too: OpenTelemetry's own system metrics express utilization as a fraction, so converting AWS and Azure to fractions is equally valid. What matters is picking one.

### Wire the pipeline

Each Collector then runs the same pipeline shape:

```yaml
service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors:
        [memory_limiter, resource_detection/cloud, transform/gcp_cpu, batch]
      exporters: [otlphttp]
```

Platform metrics that only the provider can see, such as load balancer or managed database metrics, come in through each cloud's export path: CloudWatch Metric Streams on AWS, the Azure Monitor REST API, and the Cloud Monitoring API on Google Cloud. Route those through the same Collector where you can, so they get the same labels and unit fixes as your application telemetry.

## Where should multi-cloud monitoring data live?

Somewhere outside at least two of the clouds, which means telemetry crosses the network. Data leaving a cloud provider's network is outbound data transfer, and AWS, Azure and Google Cloud all bill for it. High-volume logs and high-resolution metrics can turn that into a line item finance asks about.

Three choices keep it under control:

- **Filter before export.** A Collector gateway in each cloud can drop debug logs, unused metrics and noisy labels before they leave. Data you drop at the source costs nothing to move or store.
- **Aggregate where detail is not needed.** Per-pod metrics for a stateless service can often be rolled up to the service before export, keeping the detail only where you debug at that level.
- **Keep the gateway in-region.** Agents send to a gateway in the same region, and only the gateway crosses clouds. That keeps one exit point per region to watch and tune.

Cardinality grows in multi-cloud setups for the same reason: every resource carries provider, region, account and resource ID labels. That detail is what makes cross-cloud queries useful, so decide on purpose which labels belong on metrics and which belong only on logs and traces. Our guide to [managing high cardinality metrics](https://last9.io/blog/how-to-manage-high-cardinality-metrics-in-prometheus/) covers how to make that call.

## How should alerting work across clouds?

Alert on what users feel, and use cloud labels to route and explain, rather than writing a separate alert set for each cloud.

- **Service-level alerts first.** Error rate and latency for each service, measured at the edge users reach, work the same whatever cloud runs the service. One alert covers a service that runs in two clouds.
- **Infrastructure alerts on normalized metrics.** Once CPU, memory and disk use one unit, a single rule like "CPU above 90% for 15 minutes" applies to every VM, and `cloud.provider` and `cloud.region` appear in the notification.
- **Windows longer than the slowest source.** If AWS data points arrive every 5 minutes, an alert that needs three bad points in 5 minutes can never fire there.
- **Route by owner, label by cloud.** Send pages by `service.name` or team, and include the cloud and region in the message so the responder knows which console to open, if any.

During an incident, the first question is often "is this one cloud or all of them?" A dashboard that splits each service's golden signals by `cloud.provider` answers that in one look.

## Do you need a separate tool for multi-cloud monitoring?

Not always. If one cloud runs production and the others hold a few side workloads, the main cloud's native tool plus some exported metrics can be enough. A cloud-neutral backend starts to pay off when any of these are true:

- More than one cloud serves production traffic, so incidents can cross clouds.
- On-call engineers need access to every cloud's console to debug one service.
- The same alert has to be written and maintained in two or three alerting systems.
- Traces break at cloud boundaries because each cloud's tracing tool sees only its own spans.

When you compare tools, test them with your own data: the CPU unit problem above, a trace that crosses two clouds, and a query that groups by `cloud.provider`. Check that the tool accepts OpenTelemetry natively, keeps the cloud attributes as queryable labels, and handles the label count your fleet produces without sampling it away.

## How does Last9 fit with multi-cloud monitoring?

Last9 accepts OpenTelemetry from every cloud, so the Collector setup above exports to us without changes. For each provider's platform metrics we have documented paths:

- **AWS:** our [CloudWatch metric stream integration](https://last9.io/docs/integrations/observability/aws-cloudwatch-metrics/) delivers CloudWatch metrics through Amazon Data Firehose in OpenTelemetry format.
- **Azure:** our [Azure Monitor metrics integration](https://last9.io/docs/integrations/cloud-providers/azure-monitor-metrics/) uses the Collector's Azure Monitor receiver to poll the Azure Monitor REST API with a read-only service principal, with no Event Hubs or diagnostic settings changes, and adds resource attributes such as `cloud.provider`.
- **Google Cloud:** a [read-only service account](https://last9.io/docs/create-gcp-service-account-with-read-only-access/) lets Last9 read resource metadata and monitoring data from your projects.

All of it lands in one place and is queried with PromQL, so a single query can group by `cloud.provider`.

Last9 handles 20M series per metric per day by default, with no sampling, which leaves room for the provider, region and account labels that make cross-cloud queries work. Our [Control Plane](https://last9.io/control-plane/) lets you drop, remap, redact, forward and aggregate metrics, logs and traces at ingest, so label and unit fixes can also happen after the data arrives.

## Make every cloud answer the same question

Multi-cloud monitoring works when a question about a service gets one answer, whatever cloud it runs on. That depends less on dashboards than on the data underneath them: the same cloud attributes on every signal, one unit for each metric you alert on, and alert windows that respect each source's resolution and delay.

Start with one metric that you know differs, such as CPU, and fix it end to end: resource detection in each cloud's Collector, a unit conversion, and one alert rule that fires correctly on all three. Then extend the same pattern to the rest of your alerting metrics. When you want AWS, Azure and Google Cloud telemetry in one place, [Last9](https://last9.io/) takes OpenTelemetry from all three and queries it together with PromQL.

## FAQ

### What is multi-cloud monitoring?

Multi-cloud monitoring is collecting metrics, logs and traces from workloads running on more than one public cloud, such as AWS, Azure and Google Cloud, and viewing them in one place with shared names, labels and alerts. The goal is to answer questions about a service without first checking which cloud it runs on.

### What is the difference between multi-cloud and hybrid cloud monitoring?

Multi-cloud monitoring covers two or more public cloud providers. Hybrid cloud monitoring covers a mix of public cloud and private infrastructure, such as an on-premises data center or a private cloud. Most of the work is the same in both: collect telemetry from each environment, label where it came from, and send it to one backend.

### Can CloudWatch, Azure Monitor or Google Cloud Monitoring monitor other clouds?

Each native tool is built around its own cloud. They can accept custom metrics or agent data from elsewhere, but their built-in metrics, dashboards and alerts cover their own provider's services. Teams that run production on several clouds usually keep the native tools as data sources and send the data to one cloud-neutral backend.

### Why is CPU utilization different on AWS, Azure and GCP?

The three clouds name, scale and sample it differently. AWS EC2 reports CPUUtilization as a percent, with 5-minute data points unless detailed monitoring is on. Azure reports Percentage CPU as a percent at a 1-minute grain. Google Cloud reports compute.googleapis.com/instance/cpu/utilization as a fraction, typically between 0.0 and 1.0, sampled every 60 seconds. A threshold of 80 works on the first two and never fires on the third.

### How does OpenTelemetry help with multi-cloud monitoring?

OpenTelemetry gives every signal the same resource attributes for where it came from, including cloud.provider, cloud.region and cloud.account.id. The OpenTelemetry Collector's resource detection processor fills these in automatically on AWS, Azure and Google Cloud, and the same Collector can rename metrics and convert units before export, so data from all three clouds lands in one consistent shape.

### Does sending monitoring data between clouds cost money?

Usually yes. Telemetry that leaves a cloud provider's network is outbound data transfer, which the major providers bill for. Running a Collector gateway inside each cloud lets you drop, filter and aggregate data before it crosses the network, which cuts both the transfer volume and the amount stored in the backend.
