Activity

Production Outage - RabbitMQ

Started
Lasted 2h 59min
CRITICALaffected
Machine Production (Prod 2)

RESOLVED

lasted 2h 59min

All sites have returned to normal production after restarting the problematic RabbitMQ nodes. We're continuing to monitor.

MONITORING

26min earlier

A fix has been implemented and we are monitoring the results. Job production is restoring.

IDENTIFIED

31min earlier

CloudAMQP restarting queues one at a time.

IDENTIFIED

1h 6min earlier

CloudAMQP identified there was a partial netsplit network failure which caused node queues to get in a bad state. We are continuing investigations.

IDENTIFIED

1h 11min earlier

We are now attempting a full reboot of the problematic node instance to conduct a full power cycle.

IDENTIFIED

1h 34min earlier

We've been able to determine that at least 1 node in our Rabbit cluster is having issues and have restarted the service on that node.

INVESTIGATING

2h 35min earlier

It appears most messages for work being produced are not being processed by the Production service correctly. This is resulting in some carton kick outs at Walmart sites. We've reached out to CloudAMQP.

INVESTIGATING

2h 57min earlier

We are investigating high exception count of RabbitMQ messages in Prod2.

INVESTIGATING

2min earlier

A transient network blip in Rabbit's platform managed by CloudAMQP occurred. This was later determined to be the root cause of this issue that interrupted production.