SaaS helpdesk outage post mortem

General News

Summary

The errors were mostly cryptic and non-actionable: The timeout period elapsed while attempting to consume the pre-login handshake acknowledgement coming from the biggest database server in the cluster. I bolted back home after quickly checking Slack to ensure my team was in full panic mode (they were). After some more incantations through the AWS control panel, we finally managed to coax the server back to life, connect via ssh and discover the root drive is out of space (was it because of the error logs?). Were not sure if it was the errors that triggered the log-flood, which then filled all our disk space, leading to the file corruption when we gave the server a kick using the AWS console. Or was it the file corruption that set off the error explosion, which, in turn, caused the error-flood and filled the disk space?

Classifications

industries
Entertainment
applications
Customer Service & Support

AskAI Classifications

Labels
Help Desk Software SaaS Ticketing System

Linked Companies

Jitbit
$50M to $100M