Some events such as presence and typing indicators are disabled.
Affected components:
- API
Timeline:
–
Incident updates:
-
Resolved
We have re-enabled the affected events as those events no longer appear to be a major source of issues.
-
Monitoring
Service has been stable for 24 hours. Events will continue to be disabled until we engineer an infra or software solution. Edit (4th Mar): Someone is assigned and actively working on a software solution.
-
Investigating
Aware some events are being dropped, currently adding event replicas to fix this. Update: Now being rolled out to cluster. Update: Rolled back event replicas, ended up causing more issues.
-
Monitoring
Just had a major face palm moment regarding the database, service should be a bit more stable now. More infra changes to follow. 🙂
-
Identified
Trying to ease up the burden on parts of the system, there will be rolling restarts of services which ), but they may increase load. This process should complete in about 15 minutes, up to 30 minutes at most. Update: It's dropping enough users in each batch that the database is tanking performance to zero, we'll be re-scaling this afterwards. Update: Preparing to re-scale the database.
-
Identified
Just saw a burst of users coming online, attempting to re-scale. Update: Scaled, monitoring performance, we may have to go Update: Trying to push throughput even further Update: We're at an architectural limit, going to temporarily disable some events (incl. typing indicators and user updates) to help ease congestion while an actual fix is being put together
-
Identified
I suspect we are hitting limits with our message pubsub, a solution is being put together. Update: Scaled vertically for now.
-
Monitoring
Production services are now scaled up. There is a possibility of hitting further bottlenecks but we should be okay for a moment. Will continue to monitor and improve the deployment pattern.
-
Identified
Ordered more servers, waiting for fulfillment. Service is generally stable right now, but more load is expected either today or tomorrow peak hours.
-
Identified
Single-node cluster deployment was successful, now scaling it up.
-
Identified
Deployment is taking a little longer than expected, but I would expect this to take less than an hour to resolve.
-
Identified
We are currently working on scaling up our services.