AWS Cloud Engineering
Stand up a full production environment on AWS, from networking and compute to monitoring, and defend the monthly cost with confidence.
What you'll be able to do
Graduates can take a containerised application and stand up its entire production environment themselves.
- Design and build a segmented, multi-availability-zone network
- Deploy load-balanced, auto-scaling compute or a container service
- Operate a managed database with backups and a tested restore
- Apply least-privilege identity and externalised secrets management
- Define all infrastructure in Terraform with remote state and modules
- Deploy through a CI/CD pipeline using short-lived credentials
- Build dashboards and alarms that fire on real symptoms
- Produce an architecture diagram and defend the monthly cost
Certification Positioning
The module sequence aligns with the AWS Solutions Architect Associate objectives, and students are encouraged to sit the exam. The course is not exam preparation, however. The primary deliverable is two to three repositories of infrastructure code deployed to a real account, each with an architecture diagram and a monthly cost estimate. Cloud Practitioner is treated as optional.
Who it's for
- Developers who want to own their own deployments
- IT support staff and system administrators moving into cloud work
- Graduates targeting cloud support, junior DevOps and platform engineering roles
Prerequisites
- Linux command line comfort
- Basic networking concepts: IP, DNS, ports, HTTP
- Ability to read and modify code in at least one language
The coding prerequisite is screened at enrolment. Students without it stall in the infrastructure-as-code and container modules.
All students complete the two-session Engineering Onboarding module before Module 1.
Tools and technologies
Target roles
Course curriculum
- Concepts
- what cloud changes operationally and commercially; regions, availability zones and edge locations as design inputs; the shared responsibility model; IAM covering users, groups, roles, policies, policy evaluation, least privilege, assumed roles, instance roles and the risk of long-lived access keys; account structure basics; pricing models including on-demand, reserved, savings plans and spot; free tier limits; tagging for cost allocation; budgets and alerts; the AWS CLI and profiles; the Well-Architected pillars as a working checklist.
- Lab
- secure a fresh account with root lockdown, MFA and an administrative role; write and test three IAM policies, including tightening an over-permissive one; set and trigger a budget alarm; tag resources and produce a cost breakdown by tag.
- Project
- an account baseline document covering identity model, tagging standard, budget thresholds and a teardown checklist. Every subsequent lab is torn down or budget-alarmed.
- Concepts
- VPC design covering CIDR planning, public and private subnets across zones, route tables, internet and NAT gateways with their cost implications, security groups against network ACLs, and VPC endpoints; bastion access against Session Manager; EC2 instance families and sizing, AMIs, user data, EBS volumes and snapshots; application load balancing covering target groups, health checks and listener rules; auto scaling with launch templates and scaling policies; DNS with Route 53 and certificates with ACM; multi-zone availability as the default.
- Lab
- build a two-zone VPC manually once, so that the later infrastructure code is meaningful; deploy an application to private-subnet instances behind a load balancer; terminate an instance and observe replacement; misconfigure a security group, lose connectivity and diagnose it; attach a domain with HTTPS.
- Project
- a documented and diagrammed highly available network and compute environment.
- Concepts
- S3 covering buckets, storage classes, lifecycle policies, versioning, encryption at rest, bucket policies against IAM, presigned URLs and static hosting; CloudFront caching, invalidation and origin access; RDS for PostgreSQL covering sizing, parameter groups, multi-zone deployment against read replicas, automated backups, point-in-time recovery, snapshots, maintenance windows and connection pooling; DynamoDB partition and sort key design driven by access patterns, plus capacity modes; Secrets Manager and Parameter Store with rotation; KMS encryption; backup and recovery planning with stated recovery point and time objectives.
- Lab
- private bucket with presigned upload plus a CloudFront distribution for public assets; deploy RDS into private subnets and connect the application, then perform a point-in-time restore as a drill; move database credentials into Secrets Manager and remove them from configuration; design and query a DynamoDB table from stated access patterns.
- Project
- completed persistence layer with a written backup and recovery plan.
- Concepts
- ECR and image lifecycle; ECS on Fargate covering task definitions, services, desired count, rolling deployments, autoscaling and log delivery, plus the Fargate against EC2 launch type decision; Lambda covering the execution model, cold starts, memory and CPU coupling, timeouts, layers and event sources; API Gateway; SQS and SNS for decoupling and asynchronous work; EventBridge for scheduling and events; a decision framework for when serverless is cheaper and simpler and when it is not.
- Lab
- push an image to ECR and run it as a Fargate service behind the existing load balancer with autoscaling; perform a zero-downtime rolling deployment, then force a failed deployment and roll back; build a Lambda and API Gateway endpoint; decouple a slow task through a queue with a worker and dead-letter queue; compare the monthly cost of the same workload on Fargate against Lambda.
- Project
- Mini-project 1: a containerised application on ECS Fargate with private networking, RDS, S3, externalised secrets, HTTPS, autoscaling and logging, delivered with an architecture diagram and cost estimate.
- Concepts
- why manual console changes cannot be tracked, reviewed or reproduced; declarative provisioning and desired state; Terraform providers, resources, data sources, variables, outputs, locals and expressions; plan against apply; state management covering remote backends, S3 with DynamoDB locking, state as sensitive data, drift and import; modules for composition, reuse and versioning; workspaces and environment separation; dependency graphs and lifecycle rules; secrets handling in infrastructure code; reviewing infrastructure changes; CloudFormation and CDK as the AWS-native alternatives.
- Lab
- rebuild the entire Module 2 to 4 environment in Terraform in a clean account, then destroy and recreate it to prove reproducibility; refactor repeated resources into a module; make a console change, then detect and reconcile the drift; review a classmate's plan output and identify the destructive change before apply.
- Project
- Mini-project 2: the full environment as a modular Terraform repository with remote state and separate development and production variable sets.
- Concepts
- pipeline design covering build, test, plan, approve and apply; OIDC federation from GitHub Actions so no long-lived credentials exist; environment promotion and manual approval gates; immutable images and tagging strategy; database migrations in a pipeline; rollback, blue-green and canary strategies; observability covering CloudWatch metrics, logs, log queries, custom metrics, dashboards, alarms, X-Ray tracing and synthetic checks; deciding what to alert on against what to merely record; security review covering least privilege, public exposure audit, encryption coverage, GuardDuty and Security Hub, secrets scanning and WAF basics; cost engineering covering right-sizing, non-production scheduling, storage lifecycle, NAT and data-transfer charges, and Cost Explorer analysis; incident response and the postmortem format.
- Lab
- build the pipeline with plan output posted to the pull request and a gated apply; deploy and roll back through it; build a dashboard and three alarms, then break the service and confirm they fire; audit the account for security findings and remediate them; run a cost-optimisation pass and document the monthly saving.
- Project
- an operational runbook covering dashboards, alarms, escalation, deployment and rollback procedures.
- Concepts
- designing under constraints of availability target, budget ceiling and compliance requirement; reading and drawing architecture diagrams; documenting decisions with trade-offs; review of Solutions Architect Associate exam domains, including topics not covered in labs such as advanced networking, migration services and additional database options; specialisation paths across platform, security, data and reliability engineering.
- Lab
- timed architecture design exercises from written business requirements, presented and defended to the group; exam-style question review.
Capstone project
A production-grade AWS environment for a real application, defined entirely as code.
Requirements
- Multi-zone VPC with public and private segmentation
- Containerised application on ECS Fargate or EC2 with auto scaling, behind a load balancer with HTTPS
- RDS for PostgreSQL in private subnets with automated backups and a tested restore
- S3 and CloudFront for assets
- Secrets in Secrets Manager, with no credentials in code or configuration
- Least-privilege IAM and no long-lived keys anywhere, including CI
- All infrastructure in Terraform with remote state, modules and two environments
- CI/CD pipeline using OIDC, with plan on pull request and gated apply
- CloudWatch dashboard and alarms proven to fire
- Completed security audit with remediations
- Architecture diagram
- Monthly cost estimate with an optimisation section
- Runbook covering deployment, rollback and incident response
- README allowing a reviewer to destroy and rebuild the environment from scratch
Assessment
Out of scope
- Kubernetes and EKS
- Multi-account organisations and landing zones
- Hybrid networking
- Data engineering services beyond basics
- Migration methodology
Enquire about this course
Ask about the next cohort, schedule or prerequisites and our team will get back to you.



