Corporate · onsite · online training worldwide
contact@DevOpsSchool.com· +91 99057 40781·
> Cloud Operations · DevOpsSchool Trainer

CloudOps Trainer

Private corporate batches, live online cohorts and 1-on-1 mentoring in running a cloud estate day two onward — guardrails, drift, observability, incident response, recovery and cost — taught by a practitioner who runs it in production.

20 years across DevOps, SRE and Security · 10,000+ engineers trained · Trained teams at JPMorgan Chase, Verizon, Nokia and the World Bank

DeliveryOnline · Onsite · Hybrid
FormatsCorporate · 1-on-1 · Cohort
AgendaCustomisable
Batch size8–30 engineers
Engineers we've trained work at
JPMorgan ChaseBank of AmericaWells FargoVerizonNokiaWorld BankGE HealthcareVMwareOracleQualcommMercedes-BenzAirbusDatadogSplunkDeloitteInfosysWiproCapgemini
# who teaches it

Your CloudOps trainer

Rajesh Kumar

Principal DevOps Engineer & Architect

Cloud architectureMulti-cloud estatesInfrastructure at scale20 years in productionPrincipal / architect roles10,000+ engineers trainedM.Tech BITS Pilani25+ certifications

Rajesh teaches CloudOps as estate management rather than console tours: multi-account and multi-subscription structure with preventative policy guardrails, infrastructure as code as the only sanctioned change path with drift detection behind it, observability designed around managed services and provider health events, incident response and runbook automation, rehearsed backup and disaster recovery against stated objectives, and cost attribution that traces a bill back to the architectural decision that caused it. Every module runs against a live estate, including deliberately introducing drift and failure and then operating through it.

Twenty years across DevOps, SRE and Security, in principal and architect roles at PayPay, SoftwareAG, ServiceNow, JDA Software, Intuit, Adobe and others. He has trained engineers at JPMorgan Chase, Verizon, Nokia, the World Bank, VMware, Oracle, Mercedes-Benz and Airbus — more than 10,000 people personally. He teaches what he runs, not what he reads.

One practitioner, not a bench

You are booked with a named engineer, and that is who turns up. Marketplaces and larger providers rotate whoever is free, so the person who sold you the agenda is rarely the person teaching it.

The same trainer is available for the next engagement, which matters when a team builds on what it learned last time.

18,000+certified learners
500+corporate batches delivered
50+countries served
100+certification programmes
# faculty

Who delivers CloudOps engagements

Your batch is assigned a named trainer before it starts, and that is who teaches it. See the full faculty.

How your CloudOps trainer is chosen

Engagements are matched on the tool, not the calendar. For CloudOps that means a trainer who has run it in production — running a cloud estate day two onward — guardrails, drift, observability, incident response, recovery and cost — rather than whoever is free that week. You are told who is teaching before you commit, and that person is on the discovery call that shapes the agenda.

Where a batch is large enough to need a second trainer, the pairing is declared up front. The lead trainer stays accountable for the syllabus and the assessment either way.

Rajesh Kumar

Principal DevOps Engineer & Architect

India20 yrsLead trainer

Twenty years across DevOps, SRE and Security in principal and architect roles at PayPay, SoftwareAG, ServiceNow, JDA Software, Intuit, Adobe, IBM/Emptoris, Ness, MindTree and Accenture. He has trained more than 10,000 engineers personally, at organisations including JPMorgan Chase, Verizon, Nokia, the World Bank, VMware, Oracle, Mercedes-Benz and Airbus. He teaches what he runs, not what he reads.

Durga Prasad

IndiaInstructorCoach

Gaurav Aggarwal

IndiaInstructorCoach

Harsh Mehta

IndiaInstructorCoach

Kapil Gupta

IndiaInstructorCoach

Kunal Jain

IndiaInstructorCoach

Nikhil Gupta

IndiaInstructorCoach

Pranab Kumar

IndiaInstructorCoach

Rohit Ghatol

IndiaInstructorCoach

Amit Agarwal

IndiaInstructorCoach

Anil Kumar

IndiaInstructorCoach

Balachandran Anbalagan

IndiaInstructorCoach

# how to engage

Four ways to work with this trainer

Private corporate batch

Teams of 8–30

Custom agenda, your timezone, onsite or online, NDA-friendly.

Request a quote

1-on-1 mentoring

Individual engineers

A private instructor and a curriculum built around your goal.

₹99,999

Live & Interactive cohort

Individuals who want peers

Scheduled batch, max 8 to 10 hours of live instruction.

₹34,999

Self-paced video

Self-starters

Full LMS access — 20+ courses and 50+ tools included.

₹833/mo
# private batches

Private CloudOps training for your team

A private batch starts with a discovery call. We look at the stack you actually run — the CI system, the cloud, the constraints — and map the agenda onto it, so examples use your topology rather than a generic one.

Delivery is onsite at your premises, live online, or hybrid, scheduled around your release calendar rather than ours. Batches run 8 to 30 engineers.

Every attendee leaves with recordings, slides, lab repositories and a completion certificate. You receive an attendance and assessment report. Invoicing supports PO and GST.

Talk to us about a private CloudOps batch

What you provide vs what we bring

  • You: the room or the call, and the engineers
  • Us: trainer, agenda, labs, assessment, certificates
  • Labs: we guide your team through provisioning their own free-tier cloud environment — the skill goes with them
# the technology

What is CloudOps?

CloudOps is the operating discipline that takes over once a cloud estate exists. Building an environment is a project with an end date; running it is a continuous job of keeping accounts structured, resources accounted for, workloads healthy, recovery rehearsed, access least-privilege and spend defensible — across an estate that changes every day because self-service provisioning is the whole point of the cloud.

What makes CloudOps distinct from traditional infrastructure operations is that almost nothing is manual and almost nothing is permanent. There is no rack to visit and no ticket queue standing between an engineer and a new database. The controls that matter are preventative rather than procedural: organisational policy that refuses a non-compliant resource, infrastructure as code as the only sanctioned way to change state, tagging enforced at creation so the bill can be attributed, and automated remediation for the drift that gets through anyway. Monitoring shifts too — you are watching managed services whose internals you cannot reach, so service-level signals and provider health events matter more than host metrics.

CloudOps overlaps with SRE and with FinOps without being either. SRE supplies the reliability practice — objectives, error budgets, incident review — and FinOps supplies the financial operating model. CloudOps is the estate-level work underneath both: which accounts exist, what runs in them, whether it is patched, backed up, tagged, monitored, recoverable and worth what it costs. In most organisations it is the function that decides whether the cloud stays an asset or quietly becomes an unmanaged liability.

Why this skill matters now

The first cloud project is usually a success and the second year is usually a mess. Accounts multiply, ownership blurs, half the estate is provisioned by hand, tags are missing on exactly the resources that cost the most, and nobody is certain which backups restore. That gap between a working migration and a governed estate is what CloudOps roles exist to close.

Demand has moved accordingly. Organisations that hired cloud engineers to build now hire for operation: someone who can define guardrails without blocking teams, detect and remediate drift, run an on-call rotation against managed services, evidence compliance to an auditor, and take twenty percent out of a bill without degrading a workload. Those are operational skills, and they are scarcer than build skills because they only develop on an estate that has been running long enough to go wrong.

There is a regulatory pull as well. Data residency, audit evidence and recovery obligations increasingly require demonstrable operational control rather than good intentions, and the ability to produce that evidence from the platform itself — rather than from a spreadsheet — is now a hiring criterion in regulated sectors.

CloudOps training
# outcomes

What your team can do afterwards

Structure an estate — accounts, subscriptions or projects — with preventative policy guardrails that hold as teams self-serve
Make infrastructure as code the only sanctioned change path, and detect and remediate the drift that appears anyway
Design observability for managed services, where host metrics are unavailable and service-level signals are what you have
Build alerting that pages a human only when a human is needed, and route everything else to a queue
Run an incident on cloud infrastructure — triage, provider health events, mitigation, communication and review
Rehearse backup and disaster recovery to a stated RTO and RPO, and produce the evidence an auditor will ask for
Operate a patching, upgrade and lifecycle programme across managed platforms without unplanned outages
Attribute cloud spend accurately and reduce it by changing architecture rather than by throttling teams
# curriculum

8 modules. Live demos in a real lab, not slides.

01What CloudOps is, and how the operating model changesLive & Interactive5 hrs · 2 assignments · 1 capstone

The shift from managing hardware to operating an API-driven estate. What day two actually contains, why traditional change control fails against self-service provisioning, and how CloudOps relates to SRE, FinOps, platform engineering and the security function without duplicating any of them.

Topics: Day one versus day two responsibilities · Why ticket-based change control breaks in a self-service estate · Preventative controls versus detective controls · CloudOps, SRE, FinOps and platform engineering — the actual boundaries · Operating models: central team, embedded, or platform-as-a-product · Defining ownership for a resource that anyone can create

  • Assignments: (1) Map who currently owns each operational responsibility in your estate; (2) Identify three controls that are procedural today and could be preventative
  • Capstone: Produce an operating model proposal with named responsibilities and escalation paths
02Estate structure, guardrails and policyLive & Interactive5 hrs · 2 assignments · 1 capstone

The controls that stop problems being created rather than finding them afterwards. Organisation hierarchy, account or subscription vending, policy that denies non-compliant resources at the API, tagging enforced at creation, and quota and region restrictions.

Topics: Organisation hierarchy, organisational units and management groups · Account, subscription and project vending · Preventative policy: service control policies, Azure Policy, organisation policy · Region and service allow-lists · Tagging strategy and enforcement at creation · Quotas, limits and requesting increases before you need them · Baseline resources every new account receives · Exception handling that does not become the default path

  • Assignments: (1) Write a policy that denies untagged or unencrypted resources and test it; (2) Vend a new account with a full compliant baseline
  • Capstone: Deliver a guardrail baseline that a new team can be handed an account under, unattended
03Infrastructure as code as the operational interfaceLive & Interactive5 hrs · 2 assignments · 1 capstone

Making the desired state authoritative. Pipelines for infrastructure change, review and approval, state and locking, environment promotion, then drift — how it happens even under discipline, how to detect it continuously and when to reconcile versus adopt.

Topics: Infrastructure pipelines: plan, review, approve, apply · State, locking and blast-radius partitioning · Environment promotion and parameterisation · Modules, versioning and an internal module registry · Drift detection and continuous conformance · Reconcile, adopt or delete — deciding what to do about drift · Emergency change and how to bring it back into code · Importing an existing estate into code

  • Assignments: (1) Introduce drift by hand, detect it, and reconcile it through the pipeline; (2) Import three unmanaged production resources into code without recreating them
  • Capstone: Bring a hand-built environment fully under code with a working drift alarm
04Observability for managed servicesLive & Interactive5 hrs · 2 assignments · 1 capstone

Monitoring things you cannot log into. Provider-native metric and log platforms versus open-source stacks, structured logging and retention cost, distributed tracing across managed components, synthetic checks, and the provider health events that explain outages you did not cause.

Topics: Provider-native monitoring versus Prometheus, Grafana and OpenTelemetry · Metrics that exist for managed services, and the ones that do not · Centralised logging, retention tiers and query cost · Distributed tracing across functions, queues and managed databases · Synthetic monitoring and real user signals · Provider health events and status feeds · Dashboards for operators versus dashboards for stakeholders · Instrumenting an estate consistently by default

  • Assignments: (1) Build a service dashboard using only signals a managed service exposes; (2) Reduce a log bill without losing the fields used in incidents
  • Capstone: Deliver a monitoring baseline applied automatically to every new workload
05Alerting, incident response and on-callLive & Interactive5 hrs · 2 assignments · 1 capstone

Turning signals into action. Alerting on symptoms rather than causes, routing and escalation, on-call rotations that are humane, the anatomy of a cloud incident including provider-side failures, and the review that turns an outage into a change.

Topics: Symptom-based alerting and alert fatigue · Severity definitions, routing and escalation policy · On-call rotations, handover and follow-the-sun · Incident command, roles and communication · Triage when the failure is on the provider's side · Runbooks that are executable rather than descriptive · Automated remediation and when to allow it · Blameless review and tracking the actions to completion

  • Assignments: (1) Rewrite five cause-based alerts as symptom-based ones; (2) Run a simulated incident with a defined commander and comms lead
  • Capstone: Deliver an on-call handbook with severity model, escalation, runbooks and review process
06Resilience, backup and disaster recoveryLive & Interactive5 hrs · 2 assignments · 1 capstone

Recovery you have actually performed. Deriving RTO and RPO from the business, backup coverage across every data store in the estate, cross-region and cross-account copies, restore testing on a schedule, and DR patterns priced honestly.

Topics: Deriving RTO and RPO, and costing each level · Backup coverage audit across every stateful service · Immutable and cross-account backup against ransomware and error · Restore testing as a scheduled, evidenced activity · DR patterns: backup-restore, pilot light, warm standby, active-active · Failover mechanics, DNS and data reconciliation · Zonal and regional failure modes, and correlated risk · Game days and chaos exercises against a real estate

  • Assignments: (1) Audit backup coverage and find the stateful resources nobody is protecting; (2) Restore a production-shaped database from backup and time it
  • Capstone: Run a documented DR exercise that meets a declared RTO, with evidence
07Security posture and compliance operationsLive & Interactive5 hrs · 2 assignments · 1 capstone

The continuous half of cloud security. Posture management and misconfiguration detection, identity hygiene as an ongoing task, secrets and key rotation, vulnerability and patch operations for images and managed platforms, and producing audit evidence from the platform rather than by hand.

Topics: Posture management and misconfiguration detection · Identity hygiene: unused credentials, stale roles, privilege creep · Secrets management, rotation and eliminating long-lived keys · Image, patch and managed-platform version lifecycle · Public exposure detection — storage, databases, endpoints · Audit logging, tamper resistance and retention · Continuous compliance and evidence generation · Handling a credential leak or exposed resource

  • Assignments: (1) Run a posture assessment and triage findings by real exploitability; (2) Rotate a long-lived key to a short-lived workload identity
  • Capstone: Produce a posture baseline plus an evidence pack an auditor could accept
08Cost, capacity and continuous optimisationLive & Interactive5 hrs · 2 assignments · 1 capstone

Operating the bill as a system signal. Attribution through tagging and account structure, unit economics, commitment and discount management, rightsizing from evidence, and the architectural decisions — storage class, instance family, egress, idle non-production — behind most waste.

Topics: Cost attribution: tags, accounts and shared-cost allocation · Unit economics — cost per customer, per transaction, per environment · Budgets, forecasts and anomaly detection · Commitment discounts, coverage and utilisation · Rightsizing from observed usage rather than requested size · Storage class and lifecycle optimisation · Egress, cross-zone and inter-service traffic cost · Non-production scheduling and idle resource reclamation · Showback, chargeback and making teams accountable without blocking them

  • Assignments: (1) Attribute one month of spend and quantify the unattributable remainder; (2) Deliver a rightsizing plan with a measured saving and a risk note
  • Capstone: Produce a cost operating loop: attribution, review cadence, backlog and measured savings

Need this mapped to your stack?

We rebuild the agenda around the tools you actually run.

Request a custom agenda
# hands-on

Labs and capstones your engineers actually build

LAB · GUARDRAILS

Account vending with policy that bites

Vend a new account with a compliant baseline, then prove the guardrails by attempting to create untagged, unencrypted and out-of-region resources.

policyvendingtagging
LAB · DRIFT

Break it by hand, fix it through the pipeline

Introduce configuration drift manually, detect it with continuous conformance checks, and reconcile it through the infrastructure pipeline without recreating the resource.

iacdriftconformance
LAB · OBSERVABILITY

Monitoring a service you cannot log into

Instrument a managed database and a serverless function using only exposed signals, build an operator dashboard, and cut log spend without losing incident-critical fields.

metricslogstracing
LAB · INCIDENT

Simulated outage with a real command structure

Run a full incident against a live environment — triage, provider health checks, mitigation, stakeholder comms — then complete a blameless review with tracked actions.

on-callincidentrunbooks
LAB · RECOVERY

Restore, timed and evidenced

Audit backup coverage across the estate, find the unprotected stateful resources, then restore a production-shaped dataset and record the real RTO.

backuprestoredr
CAPSTONE · COST

Twenty percent out of a bill

Attribute a month of spend, identify the architectural drivers behind the top lines, and deliver a rightsizing and lifecycle plan with quantified saving and risk.

finopsrightsizingattribution
# ecosystem

The tools CloudOps sits next to

AWS
Azure
Google Cloud
Terraform
Kubernetes
Ansible
Prometheus
Grafana
Datadog
OpenTelemetry
Vault
PagerDuty

Who this is for

  • Cloud and infrastructure operations engineers running an existing estate
  • DevOps and platform engineers inheriting day-two responsibility after a migration
  • SREs extending reliability practice across managed cloud services
  • Security and compliance engineers who need continuous posture rather than annual audits
  • Technical leads accountable for cloud availability, recovery and spend
  • System administrators moving from data-centre operations to cloud operations

Pre-requisites

  • Working familiarity with at least one public cloud console and CLI
  • Comfortable on a Linux command line
  • Basic networking: subnets, routing, DNS, TLS
  • Exposure to infrastructure as code, ideally Terraform
  • Access to a cloud account you can create and destroy resources in
# pricing

Straightforward pricing

Every plan includes 1 year of full LMS access — not just this course, the entire DevOpsSchool LMS: 20+ courses, 50+ tools, videos, quizzes, assignments and projects.

Self-paced video

₹833/mo

Billed yearly at ₹9,996

Enroll now

1-on-1 mentorship

₹99,999

Full program, private instructor

Enroll 1-on-1

Corporate / private batch

8–30 engineers · custom agenda · onsite or online · PO and GST invoicing

Get a custom quote

Refunds. If we cancel or postpone a cohort, you get a full refund within 15 days. There is no money-back guarantee otherwise.

Terms. Course material remains licensed to the attendee. Read the terms.

Your data. We don't share it with third parties. Privacy policy.

Every attendee gets a verifiable certificate

  • Issued per attendee on completion
  • Verifiable at devopsschool.com/certificates
  • Hard copy available on request
  • Corporate batches receive an attendance and assessment report
DevOpsSchool

CloudOps Training

Certificate of completion

# feedback

What engineers say

4.4 / 5 from 26 reviews on Trustpilot.

★★★★★
Great learning experience from a very knowledgeable instructor with well-prepared course notes. The lab exercises on AWS instance work well to learn the hands-on side of the course.
Ando Gg · Trustpilot
★★★★★
Good discussion, helped us to understand different tools in SRE.
Prashant Saxena · Trustpilot
★★★★★
Got good lab sessions which kept the new DevOps tool learnings to the point and it helped a lot in my career.
robin son · Trustpilot
★★★★★
I took Terraform training with the tutor named Mithilesh. I requested to tailor the course curriculum for my needs. He did an excellent job of showing me how to write the Terraform script per the instructions provided.
jason smith · Trustpilot
★★★★★
My experience with the AIOps training was positive. The course covered important topics in a structured way, and Rajesh Kumar explained the concepts patiently. I found the practical aspects particularly helpful because they made the technical content easier to understand.
AARTI KUMARI · Trustpilot
★★★★★
I was looking to improve my understanding of AIOps, and this training helped me achieve that goal. Rajesh Kumar explained the subject in a structured and practical manner. The sessions on different AIOps concepts were informative.
Sonali Tiwari · Trustpilot
# comparison

Why a named practitioner beats a marketplace listing

What mattersYouTube + blogsGeneric online courseFreelance marketplaceDevOpsSchool
Named practitionerNoRarelyVaries per bookingYes — same trainer each time
Production experienceUnknownUnknownUnverified20 years, named employers
Custom agendaNoNoSometimesBuilt from your stack
Onsite deliveryNoNoSometimesYes
Lab environmentNoneSandbox that expiresVariesYour own cloud — skill goes with you
AssessmentNoneQuizRarelyAssignments + capstone per module
Per-attendee certificatesNoSometimesRarelyYes
Corporate invoicingNoLimitedVariesPO and GST
Post-training supportNoneForum, time-limitedNoneLifetime forum access
# questions

Frequently asked

How is CloudOps different from SRE?
SRE supplies the reliability practice — service level objectives, error budgets, incident review. CloudOps is the estate-level work underneath it: account structure, guardrails, drift, patching, backup coverage, posture and cost. Most organisations need both, and module 1 draws the boundary explicitly.
How is this different from your cloud fundamentals course?
The fundamentals course teaches the model — service models, identity, networking, storage, resilience. This one assumes the estate already exists and teaches operating it: policy guardrails, drift, on-call, recovery drills, posture and cost attribution.
Which provider do you use for the labs?
Whichever you run. The practices are provider-neutral and the labs are built on AWS, Azure or Google Cloud to match your estate. For mixed estates we run the guardrail and observability modules across two providers deliberately.
Can the agenda be customised for our stack?
Yes — that is the normal case for a private batch. We start with a discovery call, look at your account structure, tooling and current operational pain, and rebuild the module list around them.
Our estate is largely hand-built. Is this still relevant?
It is the most relevant case. Module 3 covers importing an existing estate into code without recreating resources, and module 2 covers retrofitting guardrails onto accounts that already contain workloads.
What lab environment do we need?
Attendees provision their own environment — free-tier AWS, Azure or GCP — and we walk them through it. We deliberately do not hand out temporary sandboxes, because the environment they build is the one they keep.
How long does a private CloudOps batch take?
Typically four days. Three cover guardrails, infrastructure as code, observability and incident response; the fourth adds the recovery exercise, posture work and a cost workshop against your own billing data.
Do you deliver onsite?
Yes. Private batches run onsite at your premises, live online, or hybrid, scheduled around your change calendar.
What size are batches?
Private corporate batches run 8 to 30 engineers. Public Live & Interactive cohorts are capped at 10 so everyone gets time with the trainer.
Do attendees get a certificate?
Yes — every attendee receives a completion certificate, verifiable at devopsschool.com/certificates. Corporate batches also receive an attendance and assessment report.
What is your refund position?
If we cancel or postpone a cohort, you receive a full refund within 15 days. There is no general money-back guarantee, and GST and gateway fees are not refunded.

Still deciding?

Tell us the team, the stack and the timeline. You'll get a straight answer, not a sales sequence.

Talk to an advisor
# ready when you are

Book a CloudOps trainer — or ask a question first.

  • No spam, no drip sequence
  • Syllabus in 60 seconds
  • A human reply within one business day

Prefer to call or email?

More ways to reach us on the contact page.

Talk to an advisorRequest a quote