← CloudCédric MerlinFREN

Cloud galaxy Case 4 of 7

A routine maintenance takes the application down

The problem

After a hosting provider’s maintenance, the application shows “page not found” (404 error) on a whole series of servers.

The constraint

Everyone suspects the latest database migration. It is not the cause.

What I did

01

Too fast

Clue: the application restarts in 18 seconds instead of 10 minutes. Too fast to be healthy.

02

A card was renamed

The private network card changed name on reboot.

03

Three links fall

The route to the outside is gone, the secrets vault is unreachable, the application starts without its configuration.

04

Fixed for good

I fix each server, then the automation, so the next reboot breaks nothing.

Diagram: the chain of five links between the reboot and the 404The false lead, the database migration, is crossed out. A start in 18 seconds instead of 10 minutes is the clue. The private network card changes name, which takes down the route to the outside, access to the secrets vault, then the configuration. The fix, written into the automation, turns all five links green: the application answers 200.404200DB migration10 min18 sname Aname Brebootcard renamedroute lostsecrets unreachablestarts, no configsetting

The result

The cause as a chain of five links, and a written procedure.

5 links

between the reboot and the error

Stack

  • Linux
  • Private network
  • Secrets management
  • Automation