
I never got around to having lunch as I was busy digging through things trying to work out of there was any way to undo everything easily. But although, in theory, it could be done, it would have been a lot of work. So I had to go to the infrastructure people to get the data restored from backup - which is a huge process in itself, due to everything involving loads of moving parts on AWS, and certain things needing to be done by them

But they were brilliant, with five or six of them jumping on a Slack call and assisting me through the whole process - first on our hotfix environment just to check there weren’t any nasty surprises lurking in the undergrowth, then in production. And the system allows for restoring to a point in time with a granularity of five minutes, so we didn’t lose any of the work people had done this morning before that, or very little. We got the whole thing resolved just before 17:00 and everything’s up and running again

We’ve got to have a full post mortem on Monday because it counted as a P1 incident due to needing to take the service offline for a while, and I gather they get reported upwards to senior management, only just below ministerial level. While there are potentially valid reasons for users being able to do what this one did, it really shouldn’t have been that easy - the system should have alerted them that something was unusual and given them a chance to correct it, rather than just charging ahead. So I’ll argue that even though we’re not supposed to be adding new features, we should at least take some time to strengthen the guard rails







Leave a comment: