Imagine sitting at your computer for 7 hours, pressing refresh every 5 minutes, hoping you could actually work.
That was my design team Monday.
It broke around 10am. We didn't think it was GitHub. We thought it was us. I opened an IT ticket, then we all just sat there, refreshing every 5 minutes, or typing "continue" into Copilot like it was a magic word. Like jiggling a doorknob you already know is locked.
Thank god we didn't have anything due that day. There was no way to get files out of the pipeline and back into Figma to actually work.
We babysat it straight through to 5, logged off having made nothing, and went home hoping tomorrow would just work. It did. Nobody ever told us why.
The part that stuck with me
Engineering has a whole discipline built around this moment. Runbooks. On-call rotations. A postmortem that gets read out loud in a meeting, with an actual answer for what broke and what changes because of it.
Design got a snow day. We sat there refreshing a page for seven hours and waited for the weather to pass.
Nobody built us a plan, because nobody ever thought of design as something that could go down. Design was the thing that happened on a whiteboard, independent of whatever engineering was running on. That stopped being true a while ago and nobody sent a memo.
We didn't even trust it was real
The instinct to blame ourselves first is the tell. A global outage, hitting error rates near 50% on GitHub's own numbers, and our first move was to open a ticket on ourselves. Not "the tool broke." "We must have done something wrong."
That's what it looks like when a team has fully absorbed a piece of infrastructure without anyone deciding to. You don't get a rollout announcement for that kind of dependency. You just find out the day it's gone, and your first thought is that you're the one who's broken.
Somebody wrote it down
GitHub has since said what happened, and it's worth reading. A critical component in their Central US data center didn't scale when traffic hit a new peak. A service mesh overloaded and set off a retry storm. A Copilot bug pushed authentication traffic to ten times normal. Neither this outage nor the one eleven days earlier came from a bad deploy. Both were capacity failures, on a platform whose monthly commits went from 1.4 billion to 2.9 billion since April.
So engineering got exactly what I described. A real postmortem, written down, with an answer in it and a list of what changes.
My team got a snow day and a shrug.
That gap is not GitHub's problem to fix, and it isn't IT's either. If your design org can't say what happens the next time its pipeline goes down, that's a design leadership failure, and it belongs to whoever runs design. Nobody else is going to write that runbook, because nobody else has noticed design turned into infrastructure.
Yours went down too. You just called it a bad day.