Model job description

Model Job Description: Cloud Engineer

This model job description sets out a cloud engineer’s practical responsibilities, work products, skills and decision interfaces. Employers can adapt its scope, tools and authority to their own environment; it is a learning template, not a live vacancy.

Explore the cloud engineering certificate
Resource
Model job description
Evidence
United States
Reviewed
October 7, 2026
Format
Reusable professional guide

Adapt an evidence-derived cloud engineer job description with responsibilities, outputs, skills, owner interfaces and decision boundaries.

Evidence scope: Evidence-derived role resource from a purposive review of 100 current U.S.-eligible cloud and platform engineering requisitions and a separate 12-source current-changes study through 7 October 2026. Vacancy mentions describe the reviewed sample, not national prevalence.

Model Job Description: Cloud Engineer

Professional Certificate in Cloud Engineering · U.S. role model · 7 October 2026

This is an evidence-derived, adaptable example of a role, not an open vacancy or a universal employer policy. It draws on a structured purposive sample of 100 U.S.-eligible cloud, infrastructure and platform engineering postings. Actual employers set the level, eligible location, work arrangement, systems, permissions, hours, review owner and requirements for their own position. The U.S. vacancy report explains the sample and its limits.

Role purpose

The Cloud Engineer builds, changes and operates the shared cloud or hybrid infrastructure that lets application teams deploy services safely and see whether those services are healthy. Depending on the employer, the work can emphasize infrastructure as code, a developer platform, deployment pipelines, observability, production reliability or a migration. Qrypt assigns AWS/Terraform infrastructure and customer-environment deployments; Rockstar Games emphasizes cloud-platform evaluation and support alongside on-premises systems. A title alone does not specify seniority or authority.

Local position fields to confirm: [U.S. work location or eligible remote states], [employer/team], [employment pattern and hours], [level and reporting manager], [platforms in use], [on-call expectations if any], [production-change approver], [security/access conditions if any], and [escalation route].

Responsibilities and intended work products

An employer selects the responsibilities that belong to its position. Each responsibility below includes a checkable intended work product or handoff; no posting assigns every bullet.

  • Define cloud infrastructure. Make reviewed, version-controlled resource changes and check the network, identity, security and cost effects. Qrypt assigns Terraform-managed AWS infrastructure; San Francisco describes Terraform/OpenTofu changes within established architecture. Intended output: an infrastructure definition and change record showing source, review, target environment, validation and owner.
  • Improve the delivery path when assigned. Build or maintain platform deployment workflows, automated checks, artifacts, verification and a rollback route. Dragos assigns release pipelines; OpenTeams assigns GitOps and CI/CD for reproducible, signed and scanned releases. Intended output: a reviewed workflow change, release evidence and a record of post-deployment checks.
  • Make service health visible and respond within the team's process. Implement or operate telemetry where owned, investigate assigned incidents and help restore service. OpenTeams assigns metrics, logs, traces, alerts and SLOs; Banyan assigns on-call response and restoration. Intended output: monitoring/alerting configuration, an owned incident record or a tested recovery result when specified.
  • Enable operational handoff. Cadwell assigns runbooks, escalation procedures and backup/recovery tests; First Due assigns documentation of operational processes. Intended output: a usable runbook, operating procedure or handoff guide when named, with owner, scope and update trigger.
  • Support a bounded migration or platform improvement. Heartflow describes migration patterns, Kubernetes infrastructure, benchmarks and post-migration operating documentation. Intended output: a migration/change plan, test and cutover evidence, open risks and a receiving owner.

The work products above are intended outputs. A vacancy does not prove that a deployed service met its objective, a recovery test passed or a runbook already exists. HP IQ places several IaC, deployment and on-call activities under “What You Might Do”; possible work must remain distinct from a definite assignment. The accepted analysis describes only what reviewed postings explicitly state.

Capabilities and tool status

Capabilities to specify as required only when essential for the actual position:

  • Reason about cloud infrastructure, service dependencies, networking and identity at the level the position owns.
  • Make or review a source-controlled infrastructure change and check deployment, service health and rollback evidence.
  • Troubleshoot and communicate an assigned operational issue; document the handoff and ask for a decision when authority is missing.

The employer must state which capabilities are required at entry, which receive guidance, and which are learned on the team. Amazon's Systems Development Engineer I works under senior guidance, while AlphaSense's staff role helps shape platform technical direction. The sampled roles do not establish one universal degree or years-of-experience minimum.

Preferred or role-specific capabilities may include:

  • Production Kubernetes operation, a named provider, specialized networking/security, or multi-region recovery.
  • Experience in restricted or regulated delivery, mentoring, or migration leadership where the position assigns it.

OpenTeams seeks production Kubernetes and at least one major cloud platform; Anduril asks for proficiency in at least one cloud provider and one of several programming languages. An “or” list offers alternatives; it does not require every option.

Tool categories may include:

  • Cloud platforms: AWS, Azure, GCP.
  • Infrastructure as code: Terraform, OpenTofu, CloudFormation, Pulumi.
  • Container platforms: Kubernetes, EKS, ECS.
  • Delivery: GitHub Actions, GitLab CI, Jenkins, Argo CD.
  • Telemetry: CloudWatch, Prometheus, Grafana, Datadog, OpenTelemetry.

The actual job description must mark each named product required, preferred, alternative, example or trained after hire according to that employer's wording. Qrypt makes AWS and Terraform hands-on requirements; other sources list several acceptable tools. A named product is not proof that every employer uses it or that its certification is mandatory.

Observable ways of working include testing a change, checking generated or AI-assisted work for correctness, explaining a trade-off, writing a handoff another engineer can use, and escalating a risk with the evidence already checked. Qrypt explicitly assigns accountability for AI-assisted output correctness; OpenTeams asks for independently usable runbooks and asynchronous collaboration. These are actions to observe, not personality labels.

Interfaces, authority and work rhythm

The engineer may work with application/product teams, platform or SRE colleagues, security/identity owners, support and—in some roles—customers or government stakeholders. Dragos collaborates with engineering on observability and security; OpenTeams works with security engineers and government stakeholders. Clarify which team owns application correctness, access, release approval, incident command, cost approval and compliance sign-off. “Owns infrastructure” is not itself a formal sign-off right. San Francisco's journey role operates within established architecture and senior guidance for complex decisions; Amazon's Engineer I explicitly escalates ambiguous problems to experienced teammates.

Work is often event-driven: a requested change triggers design, review, deployment and verification; an alert triggers triage and escalation; a migration or recovery exercise triggers a plan and evidence of the result. Heartflow describes a migration lifecycle, Cadwell names recovery testing, and Banyan names rotating on-call coverage. A daily operations duty, a shift, an on-call rotation and a project milestone are different timing facts. The employer states any daily, weekly or monthly routine and its coverage hours; the sample supplies no universal weekly or monthly schedule.

Some U.S. positions have additional applicant and access conditions. Defense Unicorns requires active Secret at start for one posting; OpenTeams distinguishes hybrid from U.S.-remote clearance conditions; Amazon's dedicated-cloud role requires active TS/SCI with polygraph. These are exact-requisition conditions, not default Cloud Engineer requirements. Work authorization, citizenship, U.S.-person/export access, eligibility, active clearance and later clearance steps must be read separately from the actual posting.

Local adaptation checklist

Local adaptation checklist

  • Confirm the eligible U.S. location, work arrangement, level, reporting line, hours and any on-call coverage in the actual posting.
  • Keep only assigned responsibilities; connect each one to an intended output, review check and named receiving owner.
  • Mark each capability and product as required, preferred, one-of, example or learned after hire; justify any degree or experience threshold for this specific position.
  • Name the actual source repository, deployment environments, observability systems, access permissions and change/incident decision owners.
  • State what the engineer may propose, implement, deploy, approve or escalate; verify any restricted-context applicant condition from the exact requisition.
  • Set daily, weekly, monthly and event triggers only where they exist locally. Remove unused model wording and leave unstated permission or timing unknown.

The public U.S. vacancy study, exact-version archived report and separate current-changes article provide the evidence and limits behind this model.

Quick reference

Use the resource in five moves

  1. Read the role purpose and expected outputs.
  2. Compare the model with the local role and authority boundaries.
  3. Select only statements supported by real evidence.
  4. Adapt the reusable fields without inventing experience or approvals.
  5. Review the result with the accountable person before operational use.