AI Practitioner
Generative AI and foundation models
Foundation models and LLMs, inference parameters, prompt engineering, model customisation, retrieval-augmented generation and agents, at AIF-C01 depth.
Foundational notes for AWS Certified AI Practitioner (AIF-C01). Unlike the AI Business Strategist credential, this one does assess AWS AI services — so the vocabulary pages and the service pages carry equal weight.
6 topics, 21 study points. Everything here is exam-oriented: each point is a fact or a distinction that AIF-C01 items are built on. Test yourself against the practice exam once you can explain a section without re-reading it.
1. Foundation Models & Large Language Models
Foundation Models (FMs) are large AI models trained on enormous, diverse datasets using self-supervised learning. Unlike traditional ML models trained for a single narrow task, foundation models develop broad, general capabilities during pre-training that can be applied to many different downstream tasks — text generation, summarization, translation, code completion, question answering, and more. The word “foundation” reflects their role as a base that can be adapted and built upon. Notable examples include GPT-4, Claude, Llama 2, and Amazon Titan.
Large Language Models (LLMs) are a class of foundation models specifically focused on language understanding and generation. They are trained on text from the internet, books, code repositories, and other sources at a scale of hundreds of billions of parameters. LLMs excel at tasks that require understanding and generating natural language: summarizing documents, classifying text, answering open-ended questions, conducting conversations, extracting information, and generating creative content. Because they are trained on such broad corpora, LLMs exhibit emergent capabilities — abilities that were not explicitly trained for but arise naturally from the scale and diversity of training data.
The Transformer architecture is the neural network design that powers virtually all modern foundation models. Transformers use a mechanism called self-attention, which allows the model to weigh the importance of every other word in the input when encoding or generating a particular word — regardless of how far apart the words are in the sequence. This overcomes the limitation of earlier recurrent architectures (RNNs, LSTMs) that struggled to capture long-range dependencies. Positional encodings are added to input embeddings to provide the model with information about word order (since self-attention itself is order-agnostic). The encoder-decoder architecture enables tasks that transform one sequence into another (like translation), while decoder-only architectures (used by GPT-style models) are optimized for text generation.
The context window is the maximum amount of text a model can process at one time during inference. It is measured in tokens (not characters) — a token is roughly 3/4 of a word on average in English. The context window determines how much conversation history, document content, or instruction the model can “see” when generating a response. If your input exceeds the context window, earlier content is truncated and lost. Larger context windows (e.g., 100K+ tokens) enable processing of entire codebases, long legal documents, or extended conversations, but they also increase computational cost per inference. Context window size is a key differentiating factor between FM providers and model versions.
2. Inference Parameters
When you send a prompt to a generative AI model via Amazon Bedrock or another API, you can configure inference parameters that control how the model samples the next token at each generation step. Understanding these parameters is essential for building applications that produce the right balance of predictability and creativity for their specific use case.
Temperature is the most fundamental inference parameter — it controls the randomness of the model’s outputs by modifying the probability distribution over candidate tokens before sampling. A temperature of 0 makes the model fully deterministic: it always selects the highest-probability token (greedy decoding). As temperature increases toward 1 (and beyond), the distribution flattens, making lower-probability tokens more likely to be sampled and producing more creative, varied, and sometimes surprising output. Use low temperature (0–0.3) for factual tasks where consistency matters: summarization, data extraction, code generation. Use high temperature (0.7–1.0) for creative tasks: story writing, brainstorming, generating diverse variations.
Top K and Top P are vocabulary filtering parameters that limit which tokens are eligible for sampling at each step. Top K restricts the sampling pool to the K most probable next tokens, discarding all others — choosing Top K = 40 means the model only considers the 40 highest-probability candidates, regardless of how concentrated or spread out those probabilities are. Top P (nucleus sampling) restricts the pool to the smallest set of tokens whose cumulative probability exceeds P — for example, Top P = 0.9 means the model samples from the minimal set of tokens that together account for 90% of the probability mass. Lower Top P produces more focused, predictable output; higher Top P allows the model to consider less likely but potentially more interesting options. Top P is generally considered more principled than Top K because it adapts to the shape of the probability distribution.
Stop sequences are specific strings of characters that, when generated by the model, cause it to immediately stop generating further tokens. For example, specifying ”
” as a stop sequence causes the model to stop after completing a paragraph. Stop sequences are useful for controlling the structure and length of model outputs in applications that expect a specific format — preventing the model from rambling beyond the intended answer boundary.
3. Prompt Engineering
Prompt engineering is the practice of designing and refining the text instructions you send to a foundation model to elicit high-quality, accurate, and appropriately formatted responses. Because foundation models do not have traditional software APIs with typed parameters — they respond to natural language — the quality of your prompt is the primary lever for improving output quality without changing the model itself. Prompt engineering requires no training, no labeled data, and no changes to model weights, making it the fastest and cheapest technique for adapting FM behavior.
Zero-shot prompting asks the model to perform a task with no examples provided — only a description of what you want. This works well for tasks that fall within the model’s broad training distribution, like “Summarize the following text” or “Classify the sentiment of this review as positive or negative.” Few-shot prompting includes a small number of input-output examples in the prompt to demonstrate the desired behavior or format. This is particularly useful for novel tasks or specific output formats that the model might not infer correctly from a description alone.
Chain-of-thought (CoT) prompting is a technique for improving the model’s reasoning ability on complex, multi-step problems. Instead of asking the model to jump directly to an answer, you prompt it to reason step by step — either by providing examples of step-by-step reasoning (few-shot CoT) or by simply appending “Let’s think step by step” to the prompt (zero-shot CoT). By externalizing the reasoning process into the generated output, the model is forced to work through the problem sequentially rather than pattern-matching to a superficial answer, which dramatically improves accuracy on arithmetic, logical reasoning, and multi-step inference tasks. CoT prompting is a key technique for getting reliable, verifiable outputs from LLMs on tasks that require structured thinking.
4. Model Customization
When a general-purpose foundation model does not perform well enough out of the box for your specific business use case, you have three approaches for improving its performance, each with different costs, data requirements, and depth of adaptation. Understanding precisely when to use each — and critically, which ones change the model’s weights and which do not — is a key exam objective.
Prompt engineering does not change the model’s weights. You improve performance purely by refining the instructions, examples, and context in your prompt. It is the fastest and cheapest approach and should always be attempted first. RAG (Retrieval-Augmented Generation) also does not change the model’s weights. Instead, it dynamically provides the model with relevant information retrieved from an external knowledge base at inference time. RAG is the right choice when your application needs access to information that is not in the model’s training data (internal company documents, recent events, frequently changing data like pricing or inventory) without the expense and delay of retraining. Fine-tuning does change the model’s weights. It further trains the FM on a labeled dataset specific to your task, adjusting all or some parameters to specialize the model’s behavior for your domain. Fine-tuning is used when you need the model to reliably match a specific tone, format, or domain vocabulary that cannot be conveyed through prompting alone. It requires labeled examples and compute time but produces a model that is deeply adapted to your use case.
Continued pre-training also changes the model’s weights, but unlike fine-tuning it uses unlabeled data and expands the model’s general knowledge rather than teaching it a specific task. A healthcare company, for example, might continue pre-training a general LLM on medical journals, clinical notes, and research papers — not to teach it a specific task but to familiarize it with medical terminology, abbreviations, and reasoning patterns that are underrepresented in its original training data. Amazon Bedrock supports both fine-tuning and continued pre-training for compatible foundation models. A critical exam rule: fine-tuning uses labeled data and trains for a specific task; continued pre-training uses unlabeled data and improves general domain knowledge.
Feature engineering is a related concept in classical ML that deserves mention: it is the process of selecting, modifying, or creating new input features from raw data to improve model performance. Rather than changing the model, feature engineering improves the quality of information fed into the model. Good feature engineering can make a simple model perform far better than a complex model trained on poorly designed features. In the AWS ecosystem, Amazon SageMaker Data Wrangler is the primary tool for feature engineering at scale, offering 300+ built-in data transformations.
5. RAG (Retrieval-Augmented Generation)
Retrieval-Augmented Generation (RAG) is an architectural pattern that extends the capabilities of a foundation model by connecting it to an external knowledge base. The core problem RAG solves is knowledge staleness and scope: a foundation model’s knowledge is fixed at training time — it cannot know about events that occurred after its training cutoff, and it has no access to your organization’s private, proprietary documents. RAG solves this without retraining the model by dynamically retrieving relevant information and injecting it into the model’s context at inference time.
The RAG workflow has four steps. When a user submits a query, an embedding model converts the query into a dense vector. That vector is used to search a vector database containing embeddings of your knowledge base documents — the search retrieves the most semantically similar document chunks (not just keyword matches). The retrieved chunks are injected into the model’s prompt as additional context. The foundation model generates a response informed by both its parametric knowledge (from training) and the retrieved context (from your knowledge base). The result is a response that is grounded in current, authoritative, organization-specific information, with dramatically reduced hallucination risk. RAG is a passive approach: it retrieves and augments, but it does not take actions or modify external systems.
Amazon Bedrock Knowledge Bases is AWS’s managed RAG implementation. It handles the full RAG pipeline: ingesting documents from S3, automatically splitting them into chunks, generating embeddings using a specified model, storing embeddings in a vector store (Amazon OpenSearch Serverless, Aurora, or Pinecone), and orchestrating the retrieval and augmentation flow at query time. Amazon Q Business is an enterprise Q&A application built on RAG — it connects to your organization’s data sources (SharePoint, S3, Confluence, Salesforce) and enables employees to ask natural language questions answered from your internal knowledge base.
6. AI Agents
AI Agents represent a more powerful and autonomous paradigm than RAG. While RAG is a passive retrieve-and-respond pattern, an agent is an autonomous application that can reason about a goal, decide which actions to take to achieve it, execute those actions (such as calling APIs, querying databases, or running code), observe the results, and decide what to do next — repeating this cycle until the task is complete. Agents are appropriate when the task requires multi-step reasoning, interaction with multiple external systems, or decisions that cannot be made in a single pass.
The agent loop works as follows: the user makes a request; the foundation model at the agent’s core interprets the request and creates a plan; the agent executes actions using “tools” it has been given access to (APIs, search engines, databases, calculators, code interpreters); the results are fed back to the model, which decides whether the goal has been achieved or whether additional steps are needed; this reasoning-action-observation loop continues until the task is complete. Agents are therefore active — they take actions and affect the external world, not just generate text.
The practical difference between RAG and Agents maps cleanly to task complexity. RAG is ideal for information retrieval tasks: “What does our refund policy say about international orders?” — a single retrieval-and-answer cycle is sufficient. An agent is required for multi-step tasks: “Process this customer’s refund request — look up their order, verify eligibility, initiate the refund in the payments system, and send a confirmation email.” Agents for Amazon Bedrock is the AWS managed service for building these autonomous workflows. It provides the orchestration layer that connects a foundation model to tools (defined as Lambda functions or action groups), handles the reasoning loop, maintains conversation memory, and integrates with Bedrock Knowledge Bases for RAG-within-agent patterns.
Where to go next
- Back to the AWS Certified AI Practitioner overview.
- Look up any service you could not name in the AWS services glossary.
- Sit the 80-item practice exam once two or three note pages are solid.
Dernière mise à jour le 18 sept. 2026