AI Cert Prep
Saisissez un mot-clé pour rechercher dans la documentation.

Cloud Practitioner

Analytics, big data and machine learning services

Athena, Glue, EMR, Kinesis, QuickSight and the AWS AI services a Cloud Practitioner should be able to place by purpose.

Foundational notes for AWS Certified Cloud Practitioner (CLF-C02). Read them in order the first time through, then use them as a revision sweep before you book.

10 study points. Everything here is exam-oriented: each point is a fact or a distinction that CLF-C02 items are built on. Test yourself against the practice exam once you can explain a section without re-reading it.

Big Data

Amazon EMR (Elastic MapReduce) is a managed big data platform that lets you process vast amounts of data using open-source frameworks like Apache Hadoop, Spark, Presto, and Hive. Instead of setting up and managing your own cluster of servers, EMR provisions the infrastructure, configures the software, and scales the cluster automatically. For example, a data science team can spin up a 100-node Spark cluster, run a machine learning job for two hours, and then terminate the cluster, paying only for the time it ran.

Amazon Kinesis is a suite of services for collecting, processing, and analyzing real-time streaming data at any scale. Kinesis Data Streams captures data from sources like IoT devices or clickstreams, Kinesis Data Firehose delivers streaming data to destinations like S3 or Redshift, and Kinesis Data Analytics lets you query streaming data using SQL. For example, a ride-sharing company can use Kinesis to process millions of GPS location updates per second to match drivers with riders in real time.

Amazon Athena is a serverless interactive query service that lets you analyze data directly in Amazon S3 using standard SQL. There is no infrastructure to set up or manage; you simply point Athena at your data in S3, define a schema, and start querying. You pay only for the amount of data scanned by each query. Athena is perfect for ad-hoc analysis, such as querying CloudTrail logs to investigate a security incident or analyzing web server access logs to understand traffic patterns.

Amazon QuickSight is AWS’s cloud-native business intelligence service that lets you create interactive dashboards and visualizations from your data. It connects to a wide range of data sources, including S3, RDS, Redshift, Athena, and even on-premises databases. QuickSight uses a machine learning engine called SPICE to deliver fast, interactive analysis. For example, a marketing team can build a dashboard that visualizes campaign performance metrics updated in near real time.

AWS Glue is a fully managed extract, transform, and load (ETL) service that makes it easy to prepare and load data for analytics. It automatically discovers and catalogs your data using the Glue Data Catalog, generates ETL code in Python or Scala, and runs the transformation jobs on a managed Apache Spark environment. For example, you might use Glue to take raw JSON log files from S3, clean and normalize them, and load the transformed data into Redshift for analysis.

Amazon Redshift Spectrum extends Redshift’s SQL capabilities to query data directly in S3 without loading it into Redshift tables first. This lets you run queries that join data stored in your Redshift data warehouse with data sitting in your S3 data lake. For example, you might keep recent sales data in Redshift for fast queries while archiving historical data in S3, using Spectrum to query across both when needed for long-range trend analysis.

A data lake is a centralized repository that stores all your data in its raw format, whether structured, semi-structured, or unstructured. AWS Lake Formation helps you set up a secure data lake in days instead of months. It automates data ingestion, cataloging, transformation, and security. Building a data lake on S3 gives you the flexibility to use different analytics tools like Athena, EMR, or Redshift Spectrum depending on the question you’re trying to answer.

Amazon OpenSearch Service (formerly Elasticsearch Service) is a managed service for search, log analytics, and real-time application monitoring. It’s commonly used to index and search large volumes of log data, build search functionality for websites, or create real-time dashboards. For example, a DevOps team can stream application logs from CloudWatch into OpenSearch and use Kibana dashboards to visualize error rates, latency, and request patterns in real time.

When building a big data architecture on AWS, a common pattern is to use S3 as the central data lake, Glue for ETL, Athena for ad-hoc queries, Redshift for structured analytics, EMR for complex processing, and QuickSight for visualization. Each service handles one part of the data pipeline, and together they form a scalable, cost-effective analytics platform. The key principle is to use the right tool for each job rather than trying to force all workloads through a single service.

AWS Data Pipeline is an orchestration service for scheduling and automating data movement and transformation across AWS services and on-premises data sources. While Glue focuses on ETL in a Spark environment, Data Pipeline is more general-purpose and can coordinate tasks like copying data from DynamoDB to S3, running EMR jobs on a schedule, or transferring data between on-premises databases and AWS. It provides retry logic and dependency tracking to ensure your data workflows run reliably.


Where to go next

Dernière mise à jour le 18 sept. 2026