AI Cert Prep
Saisissez un mot-clé pour rechercher dans la documentation.

CloudOps Engineer – Associate

Monitoring and observability with CloudWatch

CloudWatch metrics, custom metrics, alarms, Logs, Logs Insights, dashboards, anomaly detection and composite alarms — the core of the SOA-C03 monitoring domain.

Operations notes for AWS Certified CloudOps Engineer – Associate (SOA-C03), the credential formerly called SysOps Administrator – Associate. This is the operator exam: monitoring, remediation, automation, cost control and governance, tested from the point of view of the person on call.

3 topics, 12 study points. Everything here is exam-oriented: each point is a fact or a distinction that SOA-C03 items are built on. Test yourself against the practice exam once you can explain a section without re-reading it.

1. Monitoring with CloudWatch

Amazon CloudWatch is the central observability service for AWS — it collects and tracks metrics, aggregates log files, and fires alarms based on configurable thresholds. CloudWatch covers an enormous range of AWS services out of the box: EC2 instances, Auto Scaling Groups, Elastic Load Balancers, Route 53 health checks, EBS volumes, Storage Gateways, CloudFront distributions, DynamoDB tables, ElastiCache clusters, RDS databases, EMR clusters, Redshift clusters, SNS topics, SQS queues, OpsWorks stacks, and Lambda functions. You can also publish custom metrics from your own applications and servers using the CloudWatch Agent or API.

By default, EC2 instance metrics are reported to CloudWatch at 5-minute intervals at no additional charge. Enabling detailed monitoring reduces this interval to 1 minute, which is required for fine-grained Auto Scaling reactions and CloudWatch dashboards that show data in sub-5-minute windows. The four default EC2 metric categories are CPU Utilization, Network In/Out, Disk Read/Write Bytes, and Status Check results. These are collected by the AWS hypervisor and do not require any agent installation on the instance.

RAM (memory) utilization is not a default CloudWatch metric because the AWS hypervisor cannot observe memory usage inside the guest operating system. To monitor memory, disk space usage, or other OS-level metrics, you must install the CloudWatch Agent on the instance and configure it to collect and publish those values as custom metrics. This is a critical distinction on the SysOps exam — always remember that RAM is a custom metric, not a default one.

EC2 Status Checks fall into two categories that require different remediation approaches. System Status Checks monitor the underlying physical host infrastructure — they fail due to loss of network connectivity, loss of power, hardware failures, or software issues on the AWS physical host. The fix for a failed system status check is to stop and start the instance (not reboot), which causes AWS to migrate the instance to a different healthy physical host. Instance Status Checks monitor the virtual machine itself — they fail due to misconfigured networking, exhausted memory, corrupted file systems, or incompatible kernel configurations. The fix here is to reboot the instance or modify the operating system configuration, because the problem is inside the VM rather than with the underlying host.

2. CloudWatch Metrics, Alarms & Logs

CloudWatch supports two metric resolutions. Standard resolution metrics are published at 1-minute granularity. High-resolution custom metrics can be published at 1-second granularity using the —storage-resolution 1 parameter in the PutMetricData API call. High-resolution metrics are useful for applications where rapid detection of anomalies matters — such as detecting a spike in error rates within seconds rather than waiting a full minute for the next data point.

CloudWatch Alarms continuously evaluate metric data against a threshold you define and take automated action when the threshold is breached. An alarm exists in one of three states: OK (the metric is within the defined threshold), ALARM (the metric has breached the threshold), or INSUFFICIENT_DATA (not enough data points have been collected yet to evaluate the alarm — common immediately after creating a new alarm or when an instance is first started). Alarm actions can include: sending a notification via SNS (which can then email, text, or invoke a Lambda function), triggering an EC2 action (stop, terminate, reboot, or recover the instance), or triggering an Auto Scaling policy to add or remove capacity.

CloudWatch Logs is the service for collecting, storing, and querying log data from any source — EC2 instances, Lambda functions, CloudTrail, VPC Flow Logs, Route 53, and custom applications. Log data is organized into Log Groups (one per application or service) and Log Streams (one per instance or invocation). CloudWatch Logs Insights provides an interactive query interface for analyzing log data using a purpose-built query language — you can filter log entries, extract fields, aggregate counts, and visualize results. Metric Filters extract numerical values from log entries and publish them as CloudWatch metrics, enabling you to alarm on patterns found in your logs (for example, counting occurrences of “ERROR” in application logs).

CloudWatch Events (now called Amazon EventBridge) delivers a near real-time stream of system events describing changes in your AWS resources. Events can be triggered by AWS service state changes (an EC2 instance entering a terminated state, an Auto Scaling event, a CodePipeline stage transition) or by scheduled expressions using cron syntax. Event rules match incoming events and route them to target services such as Lambda functions, SQS queues, SNS topics, or Step Functions state machines. Scheduled rules with cron expressions are a common operational pattern for triggering periodic maintenance tasks — for example, running a Lambda function every night at midnight to rotate old log files or generate a cost report.

3. Amazon CloudWatch (Advanced)

CloudWatch Metric Math enables you to perform mathematical operations across multiple metrics and display the result as a new time-series graph or alarm. For example, you can create an expression that calculates error rate as errors/requests * 100, or compute a sum of data received across multiple EC2 instances. Metric Math expressions use a simple language that supports arithmetic operators, statistical functions (sum, avg, min, max, percentile), and temporal functions like RATE (calculate the rate of change). The result of a Metric Math expression can be used as the basis for a CloudWatch Alarm, enabling you to alarm on derived metrics that are not published directly by AWS services.

The CloudWatch Agent collects system-level metrics and log files from EC2 instances and on-premises servers. It uses two protocols for receiving metrics from custom applications: StatsD, a widely-supported protocol available on both Linux and Windows, allows applications to push metric data to the agent using simple UDP datagrams; collectd is a Linux-only protocol used for system performance metrics collection through plugins. The agent configuration file (stored in Parameter Store for fleet-wide distribution) specifies which metrics to collect, at what granularity, and which log files to monitor.

CloudWatch Container Insights provides deep observability for containerized applications running on ECS, EKS, and Kubernetes. It collects, aggregates, and summarizes metrics and logs from containerized workloads — CPU, memory, disk, network at the cluster, node, pod, and task level — and surfaces them in pre-built dashboards. CloudWatch Application Insights automatically detects and configures monitoring for common application components (Java, .NET, SQL Server, IIS) on EC2 and sets up dashboards and alarms based on known failure patterns. CloudWatch Synthetics (Canaries) enables you to create scheduled monitoring scripts that simulate user behavior — loading a web page, making an API call, going through a login flow — and alert you if the synthetic transaction fails, giving you outside-in visibility into your application’s availability from a user perspective.

For multi-region monitoring in a single CloudWatch dashboard, each metric widget must specify the region it pulls data from. Standard (5-minute) metrics are sufficient for most dashboard widgets, but if you need 1-minute granularity on a dashboard widget for an EC2 instance, detailed monitoring must be enabled on that instance. Using CloudWatch cross-account observability, you can create a monitoring account that aggregates metrics, logs, and traces from multiple source accounts, giving operations teams a unified view of a multi-account environment without requiring IAM assume-role switching.


Where to go next

Dernière mise à jour le 18 sept. 2026