Engineering
DevOps Engineering
Build and operate the delivery platform for a multi-service application, from containers and Kubernetes to incident response and postmortems.
8 Weeks4 Sessions/WeekIn PersonIntermediate
What you'll be able to do
Graduates can build and operate the delivery platform for a multi-service application.
- Diagnose a failing Linux server and trace a request across the network
- Author small, hardened, scanned container images and multi-service local environments
- Build CI/CD pipelines with test, security and quality gates
- Define infrastructure and configuration as code
- Deploy to Kubernetes with probes, resource limits and autoscaling, delivered through GitOps
- Instrument metrics, logs and traces, define service level objectives and write actionable alerts
- Lead an incident response and write a useful postmortem
- Produce security hardening and cost optimisation reports
Who it's for
- Graduates of the Cloud Engineering, Full-Stack or Python Backend courses
- Working developers moving toward platform work
- System administrators and IT operations staff
Prerequisites
- Linux command line
- Git and pull request workflow
- At least one programming or scripting language
- Basic cloud familiarity, including a personal cloud account and one deployed service
This is an intermediate course with prerequisites enforced by an entrance task.
Applicants without these should complete AWS Cloud Engineering or a development course first.
Tools and technologies
LinuxBashPythonDockerDocker ComposeTerraformAnsibleKubernetes with HelmArgoCDGitHub Actions (primary)GitLab CI (comparative labs)Jenkins (comparative labs)PrometheusGrafanaLokiOpenTelemetryTrivytfsecGitleaks
Target roles
DevOps EngineerPlatform EngineerSite Reliability EngineerBuild and Release EngineerCloud Infrastructure Engineer
Course curriculum
8 modules · 7-8 weeks
- Concepts
- the objectives of delivery engineering, expressed through lead time, change failure rate and restoration time; Linux processes, signals, systemd, permissions, file descriptors and scheduled tasks; resource inspection and diagnosis; package management; SSH key management and hardening; networking for operators covering DNS resolution, the TCP handshake, listening sockets, HTTP and TLS, proxies, load balancers and firewalls; diagnostic tooling; shell scripting for automation covering arguments, exit codes, strict modes, idempotency and logging; Python for operational tooling.
- Lab
- five timed diagnostic scenarios on deliberately broken servers, covering a full disk, incorrect permissions, a service failing to start, a port conflict and DNS misconfiguration; write an idempotent provisioning script with error handling; trace a request from DNS to response and document every hop.
- Project
- a personal operations toolkit repository with documented scripts.
- Concepts
- namespaces and control groups; images, layers and build caching; Dockerfile authoring covering multi-stage builds, minimal base images, non-root users, deterministic dependency installation, image size and startup time; build arguments against runtime configuration; volumes and persistence; container networking and service DNS; health checks; resource limits; Compose for multi-service local environments; registries, tagging strategy and image promotion; image security covering vulnerability scanning, base image currency and secrets in layers; debugging a running container; logging to standard output as the contract.
- Lab
- reduce a 1.2 GB image below 150 MB and document each technique's contribution; build a five-service Compose stack of API, worker, database, cache and reverse proxy that starts with one command; scan an image, triage real vulnerabilities and remediate; debug a container that exits immediately and one that starts but cannot reach its database.
- Project
- Mini-project 1: the reference multi-service application fully containerised, with a one-command local environment and a documented build and scan process.
- Concepts
- continuous integration as a working discipline; trunk-based development against long-lived branches; branch protection and required checks; pipeline design covering stages, caching, matrix builds, artifacts, parallelism, fail-fast behaviour and runtime as a tracked metric; test and lint gates; build once and promote the same artifact; semantic versioning and release notes; environment promotion with approvals; deployment strategies covering rolling, blue-green, canary and feature flags, each with a rollback plan; database migrations in pipelines, including backward-compatible migration patterns; pipeline secrets and OIDC; self-hosted against hosted runners; security gates covering static analysis, dependency scanning, infrastructure scanning and secret detection.
- Lab
- build a complete pipeline for the reference application covering lint, test, build, scan, push, staging deployment and gated production release; halve pipeline runtime through caching and parallelism and report the figures; implement blue-green deployment with an automated rollback trigger; execute a backward-compatible migration across two deployments; port the pipeline to GitLab CI and write a comparison.
- Project
- the delivery pipeline with a documented branching and release strategy.
- Concepts
- infrastructure as code principles and reproducibility; Terraform at working depth covering modules, remote state and locking, workspaces, drift, import, provider versioning, plan review as code review and testing infrastructure code; configuration management with Ansible covering inventories, playbooks, roles, idempotency and encrypted secrets; where configuration management fits alongside immutable images; image building with Packer; environment parity; secrets management architecture covering dynamic credentials and rotation; GitOps as an operating model.
- Lab
- provision a full environment in Terraform with modules and two environments; configure instances with Ansible and prove idempotency by running twice; destroy and rebuild everything from code; introduce and reconcile drift; review infrastructure pull requests against a checklist.
- Project
- the reference environment fully defined in code, destroyable and rebuildable, with a documented secrets architecture.
- Concepts
- when orchestration is warranted and the counter-argument for small teams; cluster architecture covering control plane, nodes, kubelet, scheduler and etcd; workload objects covering pods, deployments, stateful sets, daemon sets, jobs and scheduled jobs; services and their types; ingress and controllers; config maps and secrets; namespaces and role-based access control; resource requests and limits, and the behaviour of the cluster under pressure including out-of-memory termination, CPU throttling and eviction; liveness, readiness and startup probes; rolling updates and rollbacks; horizontal pod autoscaling; persistent volumes and storage classes; Helm charts, values and releases; managed against self-managed clusters and their costs; diagnosing common failures.
- Lab
- deploy the multi-service application to a cluster with services, ingress, config maps and secrets; set requests and limits, then trigger an out-of-memory termination and a throttle and observe both; break the deployment five ways and diagnose each from the command line; package the application as a Helm chart with environment values; configure ArgoCD so a commit deploys the change.
- Project
- Mini-project 2: the application running on Kubernetes through a Helm chart, delivered by GitOps, with probes, limits, autoscaling and ingress TLS.
- Concepts
- monitoring against observability; metrics, logs and traces and the purpose of each; structured logging and correlation identifiers; log aggregation, retention and cost; metric types, cardinality and query language basics; dashboards designed for use during an incident; distributed tracing for latency attribution; service level indicators, objectives and error budgets; alerting on symptoms rather than causes; alert fatigue and runbook attachment; on-call and escalation; incident response covering roles, communication, timeline and mitigation before diagnosis; the blameless postmortem and its action items; capacity planning; the cost of observability itself.
- Lab
- instrument the application with metrics and traces; build one dashboard answering whether the service is broken and where; define objectives and write alerts against error budget burn; participate in a live incident simulation in which the instructor breaks staging without warning, then diagnose, mitigate and write the postmortem; run one controlled failure experiment and verify system behaviour matches expectations.
- Project
- a full observability stack with objective definitions, alert rules and runbooks.
- Concepts
- embedding security within the pipeline rather than at the end of the cycle; supply chain security covering dependency provenance, lockfiles, software bill of materials and image signing; container and Kubernetes hardening covering non-root execution, read-only filesystems, security contexts, network policies, admission control and least-privilege access control; secrets management maturity; vulnerability management and prioritisation; compliance and audit trails; cost management covering tagging and allocation, showback, right-sizing, spot capacity, scheduling and the common sources of waste; backup, disaster recovery and a tested restore; documentation as an operational asset.
- Lab
- harden the cluster against a checklist and rescan to demonstrate improvement; implement network policies and verify that forbidden traffic fails; build a cost dashboard and write an optimisation proposal with figures; perform and time a full backup and restore drill.
- Project
- security and cost review documents for the reference platform.
Capstone project
A complete delivery platform for a multi-service application. The postmortem and runbooks carry the same assessment weight as the infrastructure.
Requirements
- All services containerised with hardened, scanned, minimal images
- Infrastructure in Terraform across two environments with remote state
- Configuration managed as code
- Kubernetes deployment through a Helm chart, delivered by GitOps, with probes, resource limits, autoscaling and ingress TLS
- CI/CD with test, lint, static analysis, dependency, infrastructure and secret scanning gates, build-once-promote-many, and a demonstrated blue-green deployment with rollback
- A backward-compatible database migration executed through the pipeline
- Secrets managed externally, with no credentials in version control
- Prometheus, Grafana and Loki observability with defined objectives and alerts proven to fire
- A live incident exercise conducted by the instructor on the student's platform, with a written postmortem
- Security hardening report
- Cost report with optimisation recommendations
- Runbooks and an onboarding document
Assessment
25%Weekly labs and mini-projects
20%Peer review of infrastructure and pipeline changes
35%Capstone
10%Incident exercise and postmortem
10%Demo Day presentation and technical questioning
Out of scope
- Service mesh
- Multi-cluster federation
- Custom operators
- Large-scale chaos engineering beyond one lab
- Internal developer platform theory
Enquire about this course
Ask about the next cohort, schedule or prerequisites and our team will get back to you.
Keep learning



