---
title: How to Define an Effective Cloud Monitoring Strategy
description: "Best practices for cloud monitoring encompass four major areas: latency, saturation, traffic and errors. NetApp Cloud Insights was built to address specifically to address such monitoring needs."
image: https://cloud.netapp.com/hubfs/Define%20an%20Effective%20Cloud%20Monitoring%20Strategy%20Blog.png
---

![NetApp-Blue-logo-sm1](https://bluexp.netapp.com/hubfs/NetApp-Blue-logo-sm1.svg "NetApp-Blue-logo-sm1")

![hamburger icon](https://bluexp.netapp.com/hubfs/NetApp-Blue/hamburger-icon.svg) ![close icon](https://bluexp.netapp.com/hubfs/NetApp-Blue/close-icon.svg)

- Product
  
  ![left nav icons-1](https://bluexp.netapp.com/hubfs/left%20nav%20icons-1.svg)Storage 
    - [NetApp On-premises](https://bluexp.netapp.com/netapp-on-premises)
    - [Cloud Volumes ONTAP](https://bluexp.netapp.com/ontap-cloud)
    - [Amazon FSx for NetApp ONTAP](https://bluexp.netapp.com/fsx-for-ontap)
    - [Azure NetApp Files](https://bluexp.netapp.com/azure-netapp-files)
    - [Google Cloud NetApp Volumes](https://bluexp.netapp.com/google-cloud-netapp-volumes)
  
  
  ![data mobility](https://bluexp.netapp.com/hubfs/NetApp-Blue/data-mobility.svg)Mobility 
    - [Copy and Sync](https://bluexp.netapp.com/cloud-sync-service)
    - [Tiering](https://bluexp.netapp.com/cloud-tiering)
  
  
  ![left nav icons-2](https://bluexp.netapp.com/hubfs/left%20nav%20icons-2.svg)Protection 
    - [Backup and recovery](https://bluexp.netapp.com/cloud-backup)
    - [Disaster recovery](https://bluexp.netapp.com/disaster-recovery)
    - [Replication](https://bluexp.netapp.com/replication)
    - [Ransomware protection](https://bluexp.netapp.com/ransomware-protection)
  
  
  ![left nav icons](https://bluexp.netapp.com/hubfs/left%20nav%20icons.svg)Analysis and control 
    - [Observability](https://bluexp.netapp.com/cloud-insights)
    - [Classification](https://bluexp.netapp.com/netapp-cloud-data-sense)
    - [AIOps & storage health ](https://bluexp.netapp.com/digital-advisor)
    - [Digital wallet](https://bluexp.netapp.com/digital-wallet)
- Resources
  
    - ![solutions database](https://bluexp.netapp.com/hubfs/NetApp-Blue/solutions-protection.svg)CALCULATORS
    - [TCO Azure](https://bluexp.netapp.com/azure-calculator)
    - [TCO AWS](https://bluexp.netapp.com/aws-calculator)
    - [TCO Google Cloud](https://bluexp.netapp.com/google-cloud-calculator)
    - [Cloud Volumes ONTAP Sizer](https://bluexp.netapp.com/cvo-sizer)
    - ![resources explore](https://bluexp.netapp.com/hubfs/NetApp-Blue/resources-explore.svg)EXPLORE
    - [Global Availability Map](https://bluexp.netapp.com/cloud-volumes-global-regions)
- [Help Center](https://bluexp.netapp.com/help-center)

[Get Started](https://cloudmanager.netapp.com/)

# How to Define an Effective Cloud Monitoring Strategy

### Share

[![Facebook](https://bluexp.netapp.com/hubfs/Facebook.svg)](https://www.facebook.com/sharer.php?u=https://bluexp.netapp.com/blog/how-to-define-an-effective-cloud-monitoring-strategy) [![Twitter](https://bluexp.netapp.com/hubfs/Twitter.svg)](https://twitter.com/share?url=https://bluexp.netapp.com/blog/how-to-define-an-effective-cloud-monitoring-strategy) [![LinkedIn](https://bluexp.netapp.com/hubfs/linkedin-1.svg)](https://www.linkedin.com/shareArticle?mini=true&url=https://bluexp.netapp.com/blog/how-to-define-an-effective-cloud-monitoring-strategy)

### Subscribe to our blog

Thanks for subscribing to the blog.

### April 4, 2019

### Topics: [Cloud Insights](https://bluexp.netapp.com/blog/topic/cloud-insights) [4 minute read](https://bluexp.netapp.com/blog/topic/4-minute-read)

****Cloud Infrastructure Monitoring Best Practices**

Monitoring is a skill, not a full-time job. In today’s world of cloud-based architectures implemented through DevOps projects, developers, site reliability engineers (SREs), and operations staff need to collectively define an effective cloud monitoring strategy. Such a strategy should focus on identifying when service level objectives (SLOs) are not being met and are likely to have a negative impact on the user’s experience.

![An example NetApp Cloud Insights dashboard showing key service level indicators (SLIs)](https://cloud.netapp.com/hubfs/image4%20(1).png)

*An example NetApp Cloud Insights dashboard showing key service level indicators (SLIs).*

Best practices for cloud monitoring encompass four major areas, or signals:

- Latency, the time it takes to service a request
- Saturation, how loaded a resource is
- Traffic, the demand for a resource
- Errors, the rate at which requests fail

It’s difficult to monitor these signals with single-element managers. Relying solely on element managers would involve manually identifying the relationship between individual resources and calculating the extent of their correlation, which would be likely to contribute to a service objective breach.  

To avoid such a tedious—and potentially noncompliant—process, you need to adopt a cloud monitoring tool that not only reports the metrics of a single resource, but also shows you how it’s connected to other resources in your overall infrastructure.

An effective cloud-monitoring strategy relies heavily on setting an alerting policy that recognizes and filters out false positives. For example, setting an alert to go off when the CPU reaches 90% saturation will probably set off a flood of alarm bells if your CPU regularly reaches 90% saturation as part of a normal load sequence, such as nightly backup. But if the alerting mechanism that's built into your cloud monitoring tool recognizes that your system reaches 90% saturation during those backups, the barrage of false alarms is prevented.

Further, when configuring your cloud monitoring service, you should set up an alerting mechanism that can detect an otherwise undetected breach. If left unchecked, such a breach can have a negative impact on the end user.

We built NetApp® Cloud Insights specifically to address such monitoring needs. Cloud Insights is a SaaS cloud-monitoring tool that gives you actionable knowledge of your infrastructure, including real-time data visualizations of the availability, performance, and usage of your entire IT infrastructure. Cloud Insights can also give you insight into public clouds—specifically, AWS, Azure, and Google Cloud—as well as on-premises multivendor resources. Cloud Insights supports more than 100 data collectors.

**Dashboards Create Opportunities for Granular SLI Monitoring**

With Cloud Insights, you can easily create custom dashboards to answer straightforward questions that are, nevertheless, difficult to answer. Questions like:

- Where is latency unacceptable?
- What resources are saturated?
- Where is traffic driving high latency, and on which systems?
  
  ![A dashboard sample list showing questions to be answered](https://cloud.netapp.com/hubfs/image2%20(1).png)

A key advantage of Cloud Insights is that it automatically discovers service paths, allowing you to see the relationships between resources and events to better understand cause and effect. This means that if you’re experiencing high latency, for example, you can easily see which resources are probably correlated and break them down into responsible resources and affected resources.

The following Expert View report is an example of a latency policy breach on VM win2K16serv35. It shows an 87% correlation to the host ocise-esx-…netapp.com. By selecting that host, you can overlay its CPU utilization graph; the overlay immediately exposes a close correlation. But you can also see a high likelihood that two other VMs are affected by this latency breach.

![Finding the root cause: a sample report showing top correlated and degraded resources](https://cloud.netapp.com/hubfs/image%203(3).png)

*Finding the root cause: a sample report showing top correlated and degraded resources.*

Getting these insights by using only VM and data storage element managers would have been difficult and time consuming.

**Setting Up Alerts**

Using Cloud Insights, you can create alerts that detect when a resource has exceeded a specific service level indicator. More importantly, you can create alerts based on the relationship between multiple indicators. For example, instead of setting alerts that ping when a single threshold is met (and suffering the consequent avalanche of alarms), you can specify the severity of an alert and when it’s triggered, in conjunction with other variables, like the amount of time that the threshold must be exceeded.

To create an alert, you simply specify the alert name, the object type, and any annotations. Here is an example alert creation dialog box.

![Adding a policy alert dialog box](https://cloud.netapp.com/hubfs/image%204(4).png)

*Adding a policy alert dialog box.*

Unique to the Cloud Insights alerting system is its ability to easily specify multiple thresholds. You can specify as many thresholds as needed for a given object type. You can set an alert to take effect if any of the thresholds are crossed, or you can specify that all thresholds must be met in order to trigger an alert.

In this video I recorded while at Re:Invent last year ,  I show examples of using Cloud Insights to monitor, troubleshoot and optimize your entire infrastructure.

To get started with Cloud Insights now, [watch the on-demand webinar](https://bluexp.netapp.com/define-effective-cloud-monitoring-strategy-webinar-ty), and visit [Cloud Central and register for our 14-day free trial.](https://bluexp.netapp.com/cloud-insights) To summarize, Cloud Insights helps you define an effective monitoring strategy for both your multivendor on-premises and cloud infrastructure. Cloud Insights goes beyond simple element managers that show you relationships between resources. It also allows you to set complex alerts that minimize false positives and maximize your ability to find problems before they affect your users.

[James Holden](https://bluexp.netapp.com/blog/author/james-holden-senior-product-manager-cloud-insights)

### Senior Product Manager, Cloud Insights

<https://bluexp.netapp.com/blog/author/james-holden-senior-product-manager-cloud-insights>

-

```json
{
  "@context" : "https://schema.org",
  "@type" : "TechArticle",
  "author" : {
    "@type" : "Person",
    "name" : "James Holden, Senior Product Manager, Cloud Insights"
  },
  "dateModified" : "Apr 4, 2019 2:27:55 PM",
  "datePublished" : "Apr 4, 2019 2:27:55 PM",
  "description" : "Best practices for cloud monitoring encompass four major areas: latency, saturation, traffic and errors. NetApp Cloud Insights was built to address specifically to address such monitoring needs. ",
  "headline" : " How to Define an Effective Cloud Monitoring Strategy ",
  "image" : [ "https://cloud.netapp.com/hubfs/Define%20an%20Effective%20Cloud%20Monitoring%20Strategy%20Blog.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://bluexp.netapp.com/blog",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://bluexp.netapp.com/hubfs/2016_dev/netapp-logo.png"
    },
    "name" : "Netapp"
  }
}
```