Waylight · what went wrong

Every incident of the last year.

Written by the person who fixed it, published whether or not anyone noticed. If this list were empty we would not expect you to believe it.

33
incidents, last 12 months
433433
minutes affected, total
  1. 14 Jan

    Boards slow to cut for 41 minutes

    degraded
    What happened
    A migration held a lock on the planning tables longer than we had tested. Boards still cut, but some took over a minute.
    What changed
    Migrations now run against a copy first, and the planner reports its own latency to the status page rather than to us.
  2. 2 Nov

    Gate view unavailable, 12 minutes

    outage
    What happened
    A bad config reached production because the deploy gate only checked the primary region. Two yards saw an error page at shift change, which is the worst possible minute.
    What changed
    Config changes now roll region by region with a five minute soak, and the gate view falls back to the last cut board offline.
  3. 8 Aug

    Exports delayed overnight

    degraded
    What happened
    A queue backed up behind one very large export and nothing shed load. Nightly exports landed at 06:40 instead of 23:00.
    What changed
    Large exports run on their own lane, and anything late now tells the customer before we notice.