Starting at 9:50 AM Central Time on July 31st we received alerts around querying health and latency. The cause was eventually tracked down to a clash of two different factors. First, we have been introducing Lambda Managed Instances (LMI) into our query architecture, which behave similar to classic lambdas in many ways and had been introduced as part of work around querying performance improvements with the goal of a better and faster user experience. Unfortunately, a difference between classic lambdas and LMIs we discovered is the 15 minute hard stop of classic lambdas does not exist on LMIs, which means that certain assumptions built into our querying system no longer held. At the time of the incident, our newer Canvas feature was running batches of queries for a customer which should have a self imposed one minute time out. This timeout was not properly communicated back to the lambda, and while with classic lambdas the task would automatically cut off at 15 minutes, the new LMIs were not enforcing this sort of safety mechanism. As a result, some large queries were running for up to 30 minutes, driving up query latency across the board and resulting in errors for users trying to run fresh queries. The system recovered by 10:10 AM CT, 20 minutes later, when the large queries cleared out, but we continued to investigate the cause of the incident.
Once we identified this weakness in our querying architecture we began work to fix it and prevent the same incident from occurring again. We have shipped all associated incident follow ups.