Corporate · onsite · online training worldwide
contact@DevOpsSchool.com· +91 99057 40781·
> IT Operations · DevOpsSchool Trainer

ITOps Trainer

Private corporate batches, live online cohorts and 1-on-1 mentoring in running and modernising the estate — builds, identity, patching, backup, capacity, monitoring and the queue — taught by a practitioner who runs it in production.

20 years across DevOps, SRE and Security · 10,000+ engineers trained · Trained teams at JPMorgan Chase, Verizon, Nokia and the World Bank

DeliveryOnline · Onsite · Hybrid
FormatsCorporate · 1-on-1 · Cohort
AgendaCustomisable
Batch size8–30 engineers
Engineers we've trained work at
JPMorgan ChaseBank of AmericaWells FargoVerizonNokiaWorld BankGE HealthcareVMwareOracleQualcommMercedes-BenzAirbusDatadogSplunkDeloitteInfosysWiproCapgemini
# who teaches it

Your ITOps trainer

Rajesh Kumar

Principal DevOps Engineer & Architect

DevOps transformationSRE adoptionTeam enablement20 years in productionPrincipal / architect roles10,000+ engineers trainedM.Tech BITS Pilani25+ certifications

Rajesh teaches ITOps from the estate outward: standard builds and drift control with Ansible, identity and certificate operations that quietly cause most outages, patch campaigns with a real rollback path, restore testing rather than backup reporting, and capacity work grounded in measured utilisation. The modernisation half is taught the same way — converting a written run book into automation, deflecting a ticket category with self-service, and moving alerting from CPU thresholds to service-level signals — so attendees leave with a sequenced plan rather than an argument for a rewrite.

Twenty years across DevOps, SRE and Security, in principal and architect roles at PayPay, SoftwareAG, ServiceNow, JDA Software, Intuit, Adobe and others. He has trained engineers at JPMorgan Chase, Verizon, Nokia, the World Bank, VMware, Oracle, Mercedes-Benz and Airbus — more than 10,000 people personally. He teaches what he runs, not what he reads.

One practitioner, not a bench

You are booked with a named engineer, and that is who turns up. Marketplaces and larger providers rotate whoever is free, so the person who sold you the agenda is rarely the person teaching it.

The same trainer is available for the next engagement, which matters when a team builds on what it learned last time.

18,000+certified learners
500+corporate batches delivered
50+countries served
100+certification programmes
# faculty

Who delivers ITOps engagements

Your batch is assigned a named trainer before it starts, and that is who teaches it. See the full faculty.

How your ITOps trainer is chosen

Engagements are matched on the tool, not the calendar. For ITOps that means a trainer who has run it in production — running and modernising the estate — builds, identity, patching, backup, capacity, monitoring and the queue — rather than whoever is free that week. You are told who is teaching before you commit, and that person is on the discovery call that shapes the agenda.

Where a batch is large enough to need a second trainer, the pairing is declared up front. The lead trainer stays accountable for the syllabus and the assessment either way.

Rajesh Kumar

Principal DevOps Engineer & Architect

India20 yrsLead trainer

Twenty years across DevOps, SRE and Security in principal and architect roles at PayPay, SoftwareAG, ServiceNow, JDA Software, Intuit, Adobe, IBM/Emptoris, Ness, MindTree and Accenture. He has trained more than 10,000 engineers personally, at organisations including JPMorgan Chase, Verizon, Nokia, the World Bank, VMware, Oracle, Mercedes-Benz and Airbus. He teaches what he runs, not what he reads.

Pranab Kumar

IndiaInstructorCoach

Rohit Ghatol

IndiaInstructorCoach

Amit Agarwal

IndiaInstructorCoach

Anil Kumar

IndiaInstructorCoach

Balachandran Anbalagan

IndiaInstructorCoach

Durga Prasad

IndiaInstructorCoach

Gaurav Aggarwal

IndiaInstructorCoach

Harsh Mehta

IndiaInstructorCoach

Kapil Gupta

IndiaInstructorCoach

Kunal Jain

IndiaInstructorCoach

Nikhil Gupta

IndiaInstructorCoach

# how to engage

Four ways to work with this trainer

Private corporate batch

Teams of 8–30

Custom agenda, your timezone, onsite or online, NDA-friendly.

Request a quote

1-on-1 mentoring

Individual engineers

A private instructor and a curriculum built around your goal.

₹99,999

Live & Interactive cohort

Individuals who want peers

Scheduled batch, max 8 to 10 hours of live instruction.

₹34,999

Self-paced video

Self-starters

Full LMS access — 20+ courses and 50+ tools included.

₹833/mo
# private batches

Private ITOps training for your team

A private batch starts with a discovery call. We look at the stack you actually run — the CI system, the cloud, the constraints — and map the agenda onto it, so examples use your topology rather than a generic one.

Delivery is onsite at your premises, live online, or hybrid, scheduled around your release calendar rather than ours. Batches run 8 to 30 engineers.

Every attendee leaves with recordings, slides, lab repositories and a completion certificate. You receive an attendance and assessment report. Invoicing supports PO and GST.

Talk to us about a private ITOps batch

What you provide vs what we bring

  • You: the room or the call, and the engineers
  • Us: trainer, agenda, labs, assessment, certificates
  • Labs: we guide your team through provisioning their own free-tier cloud environment — the skill goes with them
# the technology

What is ITOps?

ITOps is the function that keeps an organisation's existing IT services running. It owns the estate rather than a product: servers and hypervisors, directory and identity services, DNS, certificates, storage, backup, the network edge, the endpoint fleet, the monitoring stack and the queue of incidents and requests that arrives every morning. Where a product team ships features, ITOps is measured on availability, restore times, patch currency and how quickly the queue clears.

It is the least fashionable part of the field and the largest by headcount, and most estates are hybrid in a way conference talks rarely admit — a cloud footprint growing beside a Windows domain, a virtualisation platform, a decade of applications that were never designed to be redeployed, and a change process built around scheduled downtime. Teaching ITOps honestly means teaching that estate as it is: the operational fundamentals that keep it safe, and the modernisation path that gradually converts hand-run work into code.

That modernisation path is the second half of the subject. Standard builds replace snowflake servers, configuration management replaces the build document, run-book automation replaces the copy-and-paste procedure, self-service replaces a category of tickets entirely, and monitoring moves from CPU thresholds towards signals that indicate a service is actually degraded. Done well, ITOps ends up sharing most of its tooling with platform and site reliability engineering while still owning the parts of the estate nobody else will.

Why this skill matters now

The workload has grown while the operating model has not. Most organisations added cloud accounts, containers, SaaS integrations and remote endpoints without retiring anything, so the same team now maintains more surface with the same headcount and a ticket queue that never empties. The gap gets filled by automation or by unmanaged risk, and the second option has become expensive.

Security is the sharpest driver. Ransomware turned patch currency, privileged access, directory hardening and — above all — tested restores into board-level questions, and they are answered by operations rather than by a security policy. Cyber insurance renewals now ask for evidence of exactly the things ITOps has historically done informally: an accurate asset list, a patch cadence, immutable backups and a proven recovery time.

Careers move the same way. Infrastructure teams are being asked to become platform teams, which means the traditional skill set has to be paired with configuration management, infrastructure as code, pipelines and a service-oriented view of what they run. Engineers who can do both — hold the estate together and systematically automate it away — are the ones organisations promote rather than outsource.

ITOps training
# outcomes

What your team can do afterwards

Define and enforce a standard server build, then detect and remediate drift across the estate automatically
Operate the identity and name services that underpin everything else — directory, DNS, DHCP, certificates and their expiry
Run a patch campaign with maintenance windows, staged rings, verification and a rollback path that has been tested
Prove recoverability: design backup by RPO and RTO, defend it against ransomware, and run restore tests that are genuinely end to end
Build monitoring that indicates service degradation rather than machine metrics, and cut alert volume without losing coverage
Plan capacity from measured utilisation and growth rather than from vendor sizing guides
Convert manual run books into automation and self-service, and measure the tickets that stop arriving
Sequence a modernisation programme that improves operations while the estate stays in production
# curriculum

8 modules. Live demos in a real lab, not slides.

01What ITOps owns and how it is measuredLive & Interactive5 hrs · 2 assignments · 1 capstone

The shape of the function before the tooling. Services versus systems, the estate inventory, the daily queue, on-call and escalation, the relationship with service management and security, and metrics that describe outcomes rather than activity — availability, restore time, patch currency, change failure rate and backlog age.

Topics: Services, systems and the map between them · Asset and configuration data as the foundation of everything else · The operational day: queue, escalation, handover, on-call · Where ITOps meets service management, security and application teams · Outcome metrics: availability, MTTR, restore time, patch currency, change failure rate · Toil, interrupt load and why the backlog never clears · Run books, knowledge and the cost of undocumented work

  • Assignments: (1) Inventory one environment and identify every system with no named owner; (2) Measure a week of interrupt work and classify it by cause
  • Capstone: Produce a service map and operational scorecard for one part of the estate
02The estate — compute, virtualisation and storageLive & Interactive5 hrs · 2 assignments · 1 capstone

The platform the services run on. Hypervisor and cloud compute operations, resource allocation and overcommit, storage tiers and their failure characteristics, filesystem and volume management, and the lifecycle work — provisioning, decommissioning and the zombie systems nobody switches off.

Topics: Hypervisor operations, clusters, resource pools and overcommit · Cloud compute alongside on-premises: the hybrid reality · Storage tiers, IOPS, latency and what actually causes application slowness · Volumes, filesystems, snapshots and reclaiming space · Provisioning workflow and decommissioning discipline · Hardware and lifecycle management, warranty and end-of-support tracking · Finding and removing zombie systems

  • Assignments: (1) Audit an environment for orphaned VMs, unattached disks and end-of-support systems; (2) Diagnose a storage-latency-driven application problem from metrics alone
  • Capstone: Deliver a capacity and lifecycle report with a decommissioning plan and reclaimed capacity
03Standard builds and configuration managementLive & Interactive5 hrs · 2 assignments · 1 capstone

Ending the snowflake server. Golden images and image pipelines, the baseline every system must meet, configuration management to enforce it, drift detection and remediation, and the practical migration from hand-built systems to managed ones without a big-bang rebuild.

Topics: Golden images and image build pipelines · Baseline definition: packages, users, agents, hardening, logging · Configuration management with Ansible against Linux and Windows · Drift detection, check mode and safe remediation · Bringing existing hand-built systems under management · Secrets and credentials in build automation · Windows and Linux baselines side by side

  • Assignments: (1) Define a baseline and enforce it across three unmanaged hosts; (2) Introduce drift deliberately and prove your remediation catches it
  • Capstone: Ship a standard build plus enforcement that new systems inherit automatically
04Identity, name and certificate operationsLive & Interactive5 hrs · 2 assignments · 1 capstone

The services that cause outages out of proportion to the attention they get. Directory operations and privileged access, DNS and DHCP, and certificate lifecycle — where an expiry nobody tracked takes down a service that had nothing else wrong with it.

Topics: Directory services operations: Active Directory, Entra ID, LDAP · Privileged access, service accounts and tiered administration · Group policy, drift and the configuration that lives outside your tooling · DNS architecture, resolution failures and change safety · DHCP, addressing and the network dependencies operations owns · Certificate lifecycle: issuance, inventory, expiry monitoring, automated renewal · Time synchronisation and the failures it causes when wrong

  • Assignments: (1) Build a certificate inventory with expiry alerting across the estate; (2) Harden service account usage in one environment and document the access model
  • Capstone: Deliver an identity and name services runbook with monitoring for every expiry-driven failure mode
05Patching, vulnerability remediation and changeLive & Interactive5 hrs · 2 assignments · 1 capstone

Keeping the estate current without breaking it. Patch cadence and staged rings, maintenance windows and the negotiation around them, reboot orchestration for clustered services, emergency patching under a disclosed vulnerability, and the reporting that proves currency to security and audit.

Topics: Patch cadence, rings and pilot groups · Maintenance windows, freeze periods and negotiating both · Reboot orchestration for clusters and dependent services · Emergency patching and out-of-band change · Vulnerability scan output as an operations backlog · Exceptions, compensating controls and their expiry · Patch currency reporting that survives audit

  • Assignments: (1) Run a staged patch campaign with verification and a documented rollback; (2) Convert a vulnerability report into a prioritised, owned remediation plan
  • Capstone: Deliver a patch programme with rings, evidence and a measured currency figure
06Backup, restore and disaster recoveryLive & Interactive5 hrs · 2 assignments · 1 capstone

The area most estates fail an honest test on. Designing backup from RPO and RTO rather than from a schedule, immutability and isolation against ransomware, restore testing that includes the application rather than the disk, and DR that has been exercised recently enough for the run book to be true.

Topics: RPO and RTO as design inputs, agreed with the business · Backup topologies, retention and the 3-2-1 principle in a cloud estate · Immutability, air gaps and credential separation for backup systems · Application-consistent backup for databases and clustered services · Restore testing: file, system, application and full-service levels · DR strategy, failover, failback and dependency ordering · Documenting and exercising a recovery plan that stays true

  • Assignments: (1) Perform a full application restore into an isolated environment and time it; (2) Assess one backup design against a ransomware scenario and list the gaps
  • Capstone: Deliver a tested recovery plan with measured RTO for a named business service
07Monitoring, event management and capacityLive & Interactive5 hrs · 2 assignments · 1 capstone

Making the estate observable and the alerts worth reading. Coverage across infrastructure, services and endpoints; moving from threshold alerts on machines towards signals about service health; event correlation and noise reduction; and capacity planning built from measured utilisation and growth rather than guesswork.

Topics: Coverage: infrastructure, services, dependencies and synthetic checks · From CPU thresholds to service-level signals · Alert design: actionability, ownership, routing and escalation · Event correlation, deduplication and maintenance suppression · Log collection and retention for operations rather than security · Dashboards for on-call versus dashboards for reporting · Capacity planning and trend-based forecasting · Performance troubleshooting method under time pressure

  • Assignments: (1) Cut alert volume on one system by half without losing a real failure mode; (2) Build a capacity forecast for a growing service and state the trigger for adding capacity
  • Capstone: Deliver a monitoring rebuild for one service: coverage, alerts, escalation and capacity trend
08Automation, self-service and the modernisation pathLive & Interactive5 hrs · 2 assignments · 1 capstone

Turning the function into something that scales. Converting run books into automation, exposing safe operations as self-service so a ticket category disappears, infrastructure as code across cloud and on-premises, and sequencing a modernisation programme — including the honest transition towards platform and SRE practice and where it does not apply.

Topics: Choosing what to automate: frequency, risk and time saved · Run-book automation with Ansible and Rundeck-style execution · Self-service request fulfilment and ticket deflection · Infrastructure as code for cloud and virtualised on-premises estate · Version control, review and testing for operations code · Toil measurement and reinvesting the time saved · Migration and consolidation programmes without a freeze · The path from ITOps to platform and SRE practice, and its limits

  • Assignments: (1) Automate one recurring ticket type end to end and measure the tickets it removes; (2) Write a twelve-month modernisation sequence with checkpoints and rollback points
  • Capstone: Deliver an automation and modernisation plan with measured toil reduction on the first item

Need this mapped to your stack?

We rebuild the agenda around the tools you actually run.

Request a custom agenda
# hands-on

Labs and capstones your engineers actually build

LAB · BUILD

Baseline and drift control

Define a standard build, enforce it with configuration management across mixed Linux and Windows hosts, then introduce drift and prove detection and safe remediation.

ansiblebaselinedrift
LAB · IDENTITY

The expiry that takes you down

Build a certificate and service-account inventory with expiry alerting, then simulate an expired certificate and follow the failure through DNS, load balancer and application layers.

certificatesdirectorydns
LAB · PATCHING

Staged patch campaign with rollback

Run a patch campaign across pilot and production rings with reboot orchestration for a clustered service, verification checks and a rollback you actually execute.

patchingchangerollback
LAB · RECOVERY

Restore the application, not the disk

Recover a database-backed service into an isolated environment from backup, time the full restore, and record every undocumented step the run book was missing.

backuprestoredr
LAB · MONITORING

Halve the alerts, keep the coverage

Rebuild alerting for one service around service-level signals, remove threshold noise, and prove with a failure injection that the real failure mode still pages someone.

alertingcapacitynoise
CAPSTONE · MODERNISATION

Automate a ticket category out of existence

Take the highest-volume recurring request, automate it end to end with self-service intake and audit trail, and measure the tickets that stop arriving.

automationself-servicetoil
# ecosystem

The tools ITOps sits next to

Ansible
Active Directory
VMware vSphere
PowerShell
Zabbix
Nagios
Prometheus
Grafana
Rundeck
Terraform
Jira Service Management
Microsoft Configuration Manager

Who this is for

  • System and infrastructure administrators responsible for a production estate
  • IT operations engineers and NOC staff moving from manual work to automation
  • Operations team leads accountable for availability, patching and recovery commitments
  • Windows or Linux specialists broadening across the whole estate
  • Engineers being asked to turn an infrastructure team into a platform team
  • Service desk and second-line staff moving into infrastructure operations

Pre-requisites

  • Hands-on experience administering Linux or Windows servers
  • Understanding of core networking: IP, DNS, DHCP, routing and firewalls
  • Familiarity with virtualisation or a cloud provider's compute and storage services
  • Some scripting exposure — shell, PowerShell or Python
  • Access to two or three VMs, hosts or free-tier cloud instances for labs
# pricing

Straightforward pricing

Every plan includes 1 year of full LMS access — not just this course, the entire DevOpsSchool LMS: 20+ courses, 50+ tools, videos, quizzes, assignments and projects.

Self-paced video

₹833/mo

Billed yearly at ₹9,996

Enroll now

1-on-1 mentorship

₹99,999

Full program, private instructor

Enroll 1-on-1

Corporate / private batch

8–30 engineers · custom agenda · onsite or online · PO and GST invoicing

Get a custom quote

Refunds. If we cancel or postpone a cohort, you get a full refund within 15 days. There is no money-back guarantee otherwise.

Terms. Course material remains licensed to the attendee. Read the terms.

Your data. We don't share it with third parties. Privacy policy.

Every attendee gets a verifiable certificate

  • Issued per attendee on completion
  • Verifiable at devopsschool.com/certificates
  • Hard copy available on request
  • Corporate batches receive an attendance and assessment report
DevOpsSchool

ITOps Training

Certificate of completion

# feedback

What engineers say

4.4 / 5 from 26 reviews on Trustpilot.

★★★★☆
Helped to understand more on overall DevOps concepts.
Pankaj Malhotra · Trustpilot
★★★★★
Got good lab sessions which kept the new DevOps tool learnings to the point and it helped a lot in my career.
robin son · Trustpilot
★★★★★
Basics explanation was exemplary from Rajesh where he dealt with complicated topics to be simple. Great learning stuff personally for me.
Krishna Mohan Yelleti · Trustpilot
★★★★★
Very detailed explanation and has lots of patience in attending the questionnaire. Thanks again for your wonderful sessions.
Uttam Samudrala · Trustpilot
★★★★★
Good discussion, helped us to understand different tools in SRE.
Prashant Saxena · Trustpilot
★★★★★
I took Terraform training with the tutor named Mithilesh. I requested to tailor the course curriculum for my needs. He did an excellent job of showing me how to write the Terraform script per the instructions provided.
jason smith · Trustpilot
# comparison

Why a named practitioner beats a marketplace listing

What mattersYouTube + blogsGeneric online courseFreelance marketplaceDevOpsSchool
Named practitionerNoRarelyVaries per bookingYes — same trainer each time
Production experienceUnknownUnknownUnverified20 years, named employers
Custom agendaNoNoSometimesBuilt from your stack
Onsite deliveryNoNoSometimesYes
Lab environmentNoneSandbox that expiresVariesYour own cloud — skill goes with you
AssessmentNoneQuizRarelyAssignments + capstone per module
Per-attendee certificatesNoSometimesRarelyYes
Corporate invoicingNoLimitedVariesPO and GST
Post-training supportNoneForum, time-limitedNoneLifetime forum access
# questions

Frequently asked

Is ITOps just traditional sysadmin work with a new name?
It is the same responsibility, taught for the estate people actually run. The difference from a classic sysadmin course is that half the material is about systematically removing manual work — standard builds, configuration management, run-book automation and self-service — rather than only performing it more efficiently.
How does this differ from your SRE course?
SRE engineers reliability for a product using SLOs, error budgets and toil reduction, usually inside a development organisation. ITOps owns a heterogeneous estate of services other people built, including the ones nobody will rewrite. The practices converge — both end up automating and measuring — but the starting point and the constraints are different, and this course is built for the estate.
Do you cover Windows as well as Linux?
Yes. Standard builds, patching, monitoring and automation are taught for both, and the identity module is heavily Windows and directory oriented because that is where most estates carry their risk. For a private batch we weight the balance to match your environment.
We are migrating to cloud. Is this still relevant?
More so, because migrations run for years and the practices here are what keep the remaining estate safe while it happens. Patching, identity, backup, monitoring and capacity all apply to cloud instances too, and the modernisation module covers sequencing a migration without a freeze.
Can the agenda be customised for our stack?
Yes — that is the normal case for a private batch. We start with a discovery call, look at the hypervisor, directory, monitoring, backup and ticketing systems you actually run, and rebuild the module list around them. Examples then use your topology rather than a generic one.
Do you cover AIOps tooling?
Only in context. This course covers the data quality, event management and alert discipline that any AIOps platform depends on, because correlation over noisy, incomplete telemetry produces confident nonsense. If you want the analytics layer itself, the AIOps course is the right one and it assumes this material.
Do you deliver onsite?
Yes. Private batches run onsite at your premises, live online, or hybrid. You provide the room and the engineers; we bring the trainer, agenda, labs, assessment and certificates.
What lab environment do we need?
Attendees provision their own environment — free-tier AWS, Azure or GCP, or local VMs — and we walk them through it. We deliberately do not hand out temporary sandboxes, because the environment they build is the one they keep.
How long does a private ITOps batch take?
Typically four to five days. Estate fundamentals, builds, identity and patching fill three days; backup and recovery, monitoring and the automation programme take it to five.
What size are batches?
Private corporate batches run 8 to 30 engineers. Public Live & Interactive cohorts are capped at 10 so everyone gets time with the trainer.
Do attendees get a certificate?
Yes — every attendee receives a completion certificate, verifiable at devopsschool.com/certificates. Corporate batches also receive an attendance and assessment report.
What is your refund position?
If we cancel or postpone a cohort, you receive a full refund within 15 days. There is no general money-back guarantee, and GST and gateway fees are not refunded.

Still deciding?

Tell us the team, the stack and the timeline. You'll get a straight answer, not a sales sequence.

Talk to an advisor
# ready when you are

Book a ITOps trainer — or ask a question first.

  • No spam, no drip sequence
  • Syllabus in 60 seconds
  • A human reply within one business day

Prefer to call or email?

More ways to reach us on the contact page.

Talk to an advisorRequest a quote