You join a tech company. Colleagues ping you in chat, production breaks, you take a night shift, and in November comes Black Friday. Nine months of work in two, an hour a day. At the end: a performance review and a level.
#platform
09:38hey — admin.maple.io has been down since this morning 🙏 I have a customer demo at 11
09:40OPS-101 is yours. We migrated the admin panel last night — could be related.
09:41Facts first, conclusions second. And no “let me just bounce it and see”.
PostgresReplicationLag · 184s
“Paid, but my account says order not found” — 38 support tickets in 10 minutes
Docker, Kubernetes, CI/CD — all covered. But at work nobody asks what a pod is. They ask why it is in CrashLoopBackOff and what you are going to do.
On-call, rolling back a release, writing a postmortem, reading a Terraform plan that quietly destroys a database. Only the first year on the job teaches that — and its mistakes are expensive.
Miss an incident, close a real problem as noise, force-push to a shared branch — and get the breakdown of what a mid-level engineer would have done instead.
You join Maple, a home goods marketplace in Berlin, as a Junior DevOps engineer. A team lead, an onboarding buddy, developers, a product manager and a CTO.
Colleagues write in chat. You investigate in the console, fix configs and answer from facts. Task first — then theory and how a mid-level engineer would have handled it.
An alert queue: what is burning and what is noise. Decisions are final and the clock runs — like on a real shift.
Impact, root cause, timeline, action items. Then compare against a reference answer with a checklist.
A level, a skills breakdown, strengths and gaps. Export the PDF for your résumé.
Every month at Maple is one area turned into real work: networking, Linux, Docker, Git, Ansible, CI/CD, Kubernetes, monitoring and cloud.
«The first month»
Access, the first complaints from colleagues, and the first conclusion drawn from facts instead of guesses.
«Access to production»
Disks, permissions, systemd, backups — and a first night on call paired with Sam.
«Everything into containers»
Moving catalog into containers: restart loops, heavy images, compose — and the first data loss.
«The wild west in the repository»
A payment provider key in the commit history, a hotfix on production and a bad commit on a shared branch.
«No more configuring servers by hand»
Two new machines for catalog, a playbook from scratch, an unreachable host and a password in plain text.
«Deploy with one button»
A pipeline from scratch, a broken push to the registry, the first rollback through CI — and the night of the big release.
«Moving to Kubernetes»
CrashLoopBackOff, a service with no endpoints, a production-ready manifest and rolling back a stuck release.
«Black Friday»
Alerts people trust, PromQL under load — and the biggest night of the year for a marketplace.
«Infrastructure as code»
A security group open to the whole internet, a plan that quietly deletes the database, and a migration with no downtime.
How large companies broke production and what they learned. Written from the official postmortems, with questions so the lesson sticks. In an interview this is a ready answer to “tell me about a well-known outage”.
Meta · 2021 · month 1
How Facebook disappeared from the internet
Dyn · 2016 · month 1
Dyn: how an attack on a DNS provider took down Twitter, GitHub and Netflix
GitLab · 2017 · month 2
GitLab deleted its database — and not one backup worked
Linux / Reddit, Mozilla · 2012 · month 2
The leap second that pinned CPUs
Codecov · 2021 · month 3
Codecov: a leak through a Docker image build
Docker · 2020 · month 3
Docker Hub introduced rate limits — and CI went red worldwide
Uber · 2016 · month 4
Uber: a cloud key in a private repository
npm · 2016 · month 4
left-pad: 11 lines of code that broke builds worldwide
Amazon Web Services · 2017 · month 5
One typo — and half the internet lost S3
Microsoft Azure · 2012 · month 5
Azure and February 29: a certificate for a date that does not exist
Knight Capital · 2012 · month 6
Knight Capital: 45 minutes of a failed deployment
CrowdStrike · 2024 · month 6
CrowdStrike: one update, millions of blue screens
Reddit · 2023 · month 7
Reddit: a Kubernetes upgrade on Pi Day
Niantic / Google Cloud · 2016 · month 7
Pokémon GO: 50 times more traffic than planned
Cloudflare · 2019 · month 8
Cloudflare: one regular expression for the whole world
Datadog · 2023 · month 8
Datadog: a day without monitoring for thousands of companies
Atlassian · 2022 · month 9
Atlassian: a script deleted 883 customer sites
Capital One · 2019 · month 9
Capital One: how one cloud setting cost the data of 100 million people
A free skills test: 18 real situations, 10–15 minutes, no sign-up. It shows your gaps across 9 areas and which month to start from.
“Prod is down — walk me through it.” You will have dozens of situations you have actually worked through, not a memorized answer.
A level, skills across 9 areas, on-call shifts and postmortems. Attach it to your résumé.
The biggest fear of a junior is “can I handle this”. After a Black Friday night shift, that question is answered.
KAI is an engineering academy in Kazakhstan: 13 years, 5,000+ graduates, a Cisco Networking Academy. The simulator is written by the engineers who teach there and run production themselves. Our graduates have been hired by:
The KAI programme is taught at L. N. Gumilyov Eurasian National University






For teams
Scenarios built around your stack and your real incidents — on request.
No, and that is deliberate. The console reproduces real command output and the scenarios come from real incidents. That way a task takes 5–15 minutes instead of an hour spent bringing up a lab, and you can do it from a phone.
No. The months follow the natural order of the job: networking, Linux, Docker, Git, Ansible, CI/CD, Kubernetes, monitoring and cloud. Students take a month alongside a topic, graduates consolidate, working engineers check themselves.
No, it is fictional. But everything that happens there — a disk filled with logs, a key in Git history, OOMKilled pods, a Terraform plan that destroys a database — happens in every other tech company.
Those ask for definitions. This asks “production is down, what do you do?”. The answer is only found by hand: in the console, in the logs, in the config. Those are the questions asked in technical interviews.
One payment, no subscription. Leave a request and we email you a payment link and your personal access link. Access is permanent. If it is not a fit, we refund within 7 days, provided you have not gone past month 2, “Linux”. The first tasks are free: you only see the price after you finish them.
Yes. For teams: access for any number of engineers, invoice payment and a skills report per engineer — write to us through the “For teams” section. A gift is arranged by a curator: you get a personal link to pass on to the recipient.
Your first ticket is 2 minutes away