kai.kzLab
DevOps simulator · built by an engineering academy

Get real practice — in the seat of a DevOps engineer

You join a tech company. Colleagues ping you in chat, production breaks, you take a night shift, and in November comes Black Friday. Nine months of work in two, an hour a day. At the end: a performance review and a level.

Get the offer5 tasks free · no card · in your browser
The KAI programme is taught at L. N. Gumilyov Eurasian National University
91 tickets9 on-call shifts · 44 alerts9 postmortems18 real outage deep dives9 core areas

#platform

HW

09:38hey — admin.maple.io has been down since this morning 🙏 I have a customer demo at 11

NB

09:40OPS-101 is yours. We migrated the admin panel last night — could be related.

SO

09:41Facts first, conclusions second. And no “let me just bounce it and see”.

critical02:40 · Alertmanager

PostgresReplicationLag · 184s

“Paid, but my account says order not found” — 38 support tickets in 10 minutes

🚨 Needs action🔕 Noise

Courses give you topics

Docker, Kubernetes, CI/CD — all covered. But at work nobody asks what a pod is. They ask why it is in CrashLoopBackOff and what you are going to do.

Work asks for decisions

On-call, rolling back a release, writing a postmortem, reading a Terraform plan that quietly destroys a database. Only the first year on the job teaches that — and its mistakes are expensive.

Here mistakes are free

Miss an incident, close a real problem as noise, force-push to a shared branch — and get the breakdown of what a mid-level engineer would have done instead.

How it works

01

You get an offer

You join Maple, a home goods marketplace in Berlin, as a Junior DevOps engineer. A team lead, an onboarding buddy, developers, a product manager and a CTO.

02

You close tickets

Colleagues write in chat. You investigate in the console, fix configs and answer from facts. Task first — then theory and how a mid-level engineer would have handled it.

03

You take on-call

An alert queue: what is burning and what is noise. Decisions are final and the clock runs — like on a real shift.

04

You write postmortems

Impact, root cause, timeline, action items. Then compare against a reference answer with a checklist.

05

You get a review

A level, a skills breakdown, strengths and gaps. Export the PDF for your résumé.

9 months = the 9 core areas of DevOps

Every month at Maple is one area turned into real work: networking, Linux, Docker, Git, Ansible, CI/CD, Kubernetes, monitoring and cloud.

Month 1free

Introduction to IT, networking and DevOps

«The first month»

Access, the first complaints from colleagues, and the first conclusion drawn from facts instead of guesses.

DNSHTTP/HTTPSportsTCP/IP
Month 2🔒

Linux

«Access to production»

Disks, permissions, systemd, backups — and a first night on call paired with Sam.

df/du/lsofsystemdpermissionsbashcron
Month 3🔒

Docker

«Everything into containers»

Moving catalog into containers: restart loops, heavy images, compose — and the first data loss.

Dockerfiledocker composevolumeshealthchecksecrets
Month 4🔒

Git

«The wild west in the repository»

A payment provider key in the commit history, a hotfix on production and a bad commit on a shared branch.

log/show/diffreverthotfix flowtagssecrets
Month 5🔒

Ansible

«No more configuring servers by hand»

Two new machines for catalog, a playbook from scratch, an unreachable host and a password in plain text.

playbooksinventoryhandlerstemplatesVault
Month 6🔒

CI/CD

«Deploy with one button»

A pipeline from scratch, a broken push to the registry, the first rollback through CI — and the night of the big release.

GitLab CIRunnerRegistrymulti-stage pipelinerollback
Month 7🔒

Kubernetes

«Moving to Kubernetes»

CrashLoopBackOff, a service with no endpoints, a production-ready manifest and rolling back a stuck release.

pods/deploymentsservicesprobesresourcesrollout
Month 8🔒

Monitoring

«Black Friday»

Alerts people trust, PromQL under load — and the biggest night of the year for a marketplace.

PrometheusPromQLAlertmanagerGrafanaLoki
Month 9🔒

Cloud and Terraform

«Infrastructure as code»

A security group open to the whole internet, a plan that quietly deletes the database, and a migration with no downtime.

Terraformsecurity groupsS3statemigrations

Every month: deep dives into real outages

How large companies broke production and what they learned. Written from the official postmortems, with questions so the lesson sticks. In an interview this is a ready answer to “tell me about a well-known outage”.

Meta · 2021 · month 1

How Facebook disappeared from the internet

Dyn · 2016 · month 1

Dyn: how an attack on a DNS provider took down Twitter, GitHub and Netflix

GitLab · 2017 · month 2

GitLab deleted its database — and not one backup worked

Linux / Reddit, Mozilla · 2012 · month 2

The leap second that pinned CPUs

Codecov · 2021 · month 3

Codecov: a leak through a Docker image build

Docker · 2020 · month 3

Docker Hub introduced rate limits — and CI went red worldwide

Uber · 2016 · month 4

Uber: a cloud key in a private repository

npm · 2016 · month 4

left-pad: 11 lines of code that broke builds worldwide

Amazon Web Services · 2017 · month 5

One typo — and half the internet lost S3

Microsoft Azure · 2012 · month 5

Azure and February 29: a certificate for a date that does not exist

Knight Capital · 2012 · month 6

Knight Capital: 45 minutes of a failed deployment

CrowdStrike · 2024 · month 6

CrowdStrike: one update, millions of blue screens

Reddit · 2023 · month 7

Reddit: a Kubernetes upgrade on Pi Day

Niantic / Google Cloud · 2016 · month 7

Pokémon GO: 50 times more traffic than planned

Cloudflare · 2019 · month 8

Cloudflare: one regular expression for the whole world

Datadog · 2023 · month 8

Datadog: a day without monitoring for thousands of companies

Atlassian · 2022 · month 9

Atlassian: a script deleted 883 customer sites

Capital One · 2019 · month 9

Capital One: how one cloud setting cost the data of 100 million people

Not sure about your level?

A free skills test: 18 real situations, 10–15 minutes, no sign-up. It shows your gaps across 9 areas and which month to start from.

Take the test

Answers for the tech interview

“Prod is down — walk me through it.” You will have dozens of situations you have actually worked through, not a memorized answer.

A performance review PDF

A level, skills across 9 areas, on-call shifts and postmortems. Attach it to your résumé.

Calm on your first day

The biggest fear of a junior is “can I handle this”. After a Black Friday night shift, that question is answered.

Built by KAI

KAI is an engineering academy in Kazakhstan: 13 years, 5,000+ graduates, a Cisco Networking Academy. The simulator is written by the engineers who teach there and run production themselves. Our graduates have been hired by:

The KAI programme is taught at L. N. Gumilyov Eurasian National University

KazakhtelecomFreedom HoldingJusan MobileBank CenterCreditBI GroupBeelineSamruk-KazynaForteBankEPAMKazakhaltynNITECMyBuh.kzSensata Group

For teams

Assess a DevOps candidate in an hour, not in a probation period

  • — Hiring assessment: the candidate takes an on-call shift and two tickets; you get a report with a level and decision times.
  • — Junior onboarding: their first months happen in the simulator, not in your production.
  • — Team check-up: a skills map across 9 areas for every engineer.

Scenarios built around your stack and your real incidents — on request.

Questions

+Are these real servers?

No, and that is deliberate. The console reproduces real command output and the scenarios come from real incidents. That way a task takes 5–15 minutes instead of an hour spent bringing up a lab, and you can do it from a phone.

+Do I need a course first?

No. The months follow the natural order of the job: networking, Linux, Docker, Git, Ansible, CI/CD, Kubernetes, monitoring and cloud. Students take a month alongside a topic, graduates consolidate, working engineers check themselves.

+Is Maple a real company?

No, it is fictional. But everything that happens there — a disk filled with logs, a key in Git history, OOMKilled pods, a Terraform plan that destroys a database — happens in every other tech company.

+How is this different from quizzes and video courses?

Those ask for definitions. This asks “production is down, what do you do?”. The answer is only found by hand: in the console, in the logs, in the config. Those are the questions asked in technical interviews.

+How do I pay?

One payment, no subscription. Leave a request and we email you a payment link and your personal access link. Access is permanent. If it is not a fit, we refund within 7 days, provided you have not gone past month 2, “Linux”. The first tasks are free: you only see the price after you finish them.

+Can I buy it for a team or as a gift?

Yes. For teams: access for any number of engineers, invoice payment and a skills report per engineer — write to us through the “For teams” section. A gift is arranged by a curator: you get a personal link to pass on to the recipient.

Get your offer at Maple

Your first ticket is 2 minutes away