← CloudCédric MerlinFREN

Cloud galaxy Case 3 of 7

One crash a day, and no error at all

The problem

Once a day, the system abruptly killed the application. No error in the logs.

The constraint

The memory the application actually used was far below the machine’s capacity. The culprit was elsewhere.

What I did

01

A hidden reservation

The application reserved a huge share of memory in one go, in an area the usual tools do not show.

02

More than the machine

Together with the search engine and the other services, we went over the machine’s memory.

03

A setting that survives

I reduce that reservation, add swap space, and record the setting in the automation so it survives the next deployment.

04

The real culprit

I measure process by process: web traffic was flat, the spikes came from a batch job.

Diagram: a memory gauge that overflows, then the real causeA memory gauge. The application reserves a large share of it in one go, the search engine and the other services add to it and the gauge overflows: the system kills the application. The reservation is reduced, with swap space and a setting recorded in the automation. Finally, the per-process chart shows flat web traffic and spikes coming from a batch job.memoryswapreservesearchserviceskilled by the systemreducedsettingwebbatch

The result

Not a single crash since the fix, and a clear map of what to split out first.

1 crash, then 0 crash

a day since the fix

Stack

  • Linux
  • Memory analysis
  • Automation
  • Monitoring