OHSHIT REPORT: A maxed-out database took meh (and friends) offline for a few hours
34On the afternoon of July 8th, meh.com — along with our sister stores morningsave.com, sidedeal.com, hammacher.com, and others — went dark for 2 hours and 52 minutes, from 3:01 PM to 5:53 PM CT. If you came by for a deal during that window and got a sad dog instead, this one’s for you.
Here’s what happened, in mostly-plain English.
All of our stores share a behind-the-scenes service we call catalog-service. It’s the thing that knows every product, price, and deal, and every page you load asks it for that info. catalog-service keeps all of that in a database hosted on MongoDB Atlas (mongodb.com) — a managed database service lots of companies use so they don’t have to babysit their own servers.
First, a myth we’d love to bust. A lot of folks assume a store like ours picks one server, sizes it once, and if things get slow, well, we’re just being cheap. That could not be further from how this actually works. Almost every layer of our setup scales itself up and down automatically, all day long — based both on how busy we are and, the more unusual part, on how busy we expect to be.
The first part is reactive. Our stores and services run on Microsoft Azure, and they automatically add more servers the moment things get busy: when a server’s CPU climbs past a threshold, more copies spin up to share the load, then spin back down once the rush passes. Here are the actual rules for one of our services:

And here’s a look at our catalog servers’ own CPU and memory during the afternoon:

The second part is where our weird little business comes in. Almost everything we do is an event — a daily deal, a big email blast, a Mehrathon, or one of our products getting featured on a TV show. Those throw off sharp, very predictable spikes of traffic, and because they run on a schedule, we know they’re coming. So rather than wait to get slammed and react, we scale up ahead of time — automatically pre-warming extra servers before a scheduled event kicks off, then winding them back down when it’s over. Reacting to a spike is table stakes; seeing it coming and being ready is the fun part.
Our database can resize itself too — when it’s working hard, MongoDB Atlas can automatically bump it up to a beefier machine. We lean on all of this because demand is spiky (a great deal drops, an email goes out, a segment airs on TV) and we would much rather spend more to stay fast than run lean and fall over. Which is what makes what happened next a little ironic.
Our web servers talk to the database over “connections.” Think of connections like phone lines into the database: there’s a fixed number of them, and once every line is busy, new calls get a busy signal. At 3:03 PM CT our database hit 99.9% of its limit — 2,994 of 3,000 lines in use. That’s the line pinned flat against the top here, which is Not Great:

And here’s the irony, given all that autoscaling: this time it worked against us. As our stores and services scaled up to handle the load, every new server opened its own fresh batch of connections to the database — so the harder our whole fleet worked to serve you, the faster we burned through those phone lines.
With every line busy, the database started handing out busy signals, and catalog-service couldn’t get the answers it needed. Its own failed-request rate climbed toward 90% and its response times stretched into the minutes:

The raw count of errors it logged tells the same story:

And since every store leans on catalog-service, every store went down with it. Here’s meh’s own response times ballooning into the minutes and its error rate climbing:

…and meh’s raw error count over the afternoon:

For the record, meh’s own servers were mostly sitting on their hands the whole time — they weren’t the bottleneck, they were just stuck waiting on catalog-service, with a burst at the end as everything came back online:

Right about then, our database provider automatically kicked off a maintenance operation on the cluster — you can see it here, started on its own by “System”:

Normally that’s routine — it restarts the database’s servers one at a time. But because the database was already maxed out and gasping, it didn’t settle down. Instead the servers kept restarting and picking a new leader over and over for nearly two hours, and each restart knocked catalog-service’s connections out all over again. You can watch the database’s processor and memory getting thrown around the whole time — spiking and lurching with every restart:


(We’re honestly still not sure what triggered that maintenance at that exact moment — the timing was rotten, and we’re digging into it.)
To stop the bleeding while we sorted it out, we did the thing you saw: we put the stores into maintenance mode so they’d quit hammering the struggling database. Please enjoy our hard-working dog:

Meanwhile we manually upgraded the database to a bigger size — more of those phone lines, plus extra breathing room — and let it calm down. By 5:05 PM CT it was steady again, and from 5:10 to 5:53 PM we brought the stores back one at a time, watching each to make sure it held. By 5:53 PM everything was back to normal.
We also tracked down a specific thing that had quietly been making our database work harder than it should: an inefficient query behind our “current deals” pages that was occasionally taking six minutes to do a job that should take a few thousandths of a second.

We’re fixing that one so a bad afternoon is less likely to snowball like this again.
Sorry for the busy signal, and thanks for sticking with us while we got everyone back online.
- 22 comments, 28 replies
- Comment
Thanks for the report. Most of it is over my head, but I appreciate the transparency.
It’s good to hear from you, Shawn. I hope you’re doing well.
Can we get a TL;DR synopsis please?
@MrGoodGuy Meh was running a water park using a bunch of garden hoses to feed the attractions. A kid dumped a bunch of Orbeez into the recirculation pumps clogging all the existing hoses. They bought out the entire supply of hoses in town in a vain attempt to bypass the clogged ones. After that failed, they kicked everyone off the rides and out of the park so plumbers could snake all the pipes. They found the kid responsible from surveillance videos and are about to go SWAT on his butt.
KuoH
@kuoh @MrGoodGuy
@kuoh Well gee whiz! I completely understand now why the site was borked yesterday!! Them darned Orbeez are the candy equivalent of Alien Tape! [IYKYK]
MongoDB is Web Scale.
It works great until it doesn’t.
@arloth to be fair, exhausting database connections is pretty common regardless of what database you use. I accidentally caused this with postgres when I didn’t realize that the Google pub sub python SDK creates a connection for every thread, and by default has 10 threads going at a time, per subscription, and usually our subscriber scripts handle multiple subscriptions, and then when there’s a burst of messages you scale up to multiple pods that’s 10 times the number of pods times the number of subscriptions, and our limit is generally around 50 (depending on the service) so…
@arloth plus, having worked with Shawn back in the beginning of the meh days, I know the effort he put into making it web scale, and how when it’s not web scale it’s human error + poor defaults on mongo.
Not that I’m presuming you posted this to criticize
Great triage report from a fellow infra person. This is my kinda content! So, when you say the maintenance timing was foul, do you perhaps also mean suspicious? Cybersecurity sure can be a never ending battle.
Even though a lot of this was (way) over my head I can follow the general gist of it. Makes me appreciate how fragile some of the systems we depend on every day really are. Nice job tracking down the offending issue and ‘getting your shit together’.

BTW: happy birthday meh
Never a dull moment!
Sooo, I should not have thrown my phone in frustration when I could not peek at the 'thon?
Happy B-Day. It would not seem right for the day to go too smooth. Need some excitement. Smoke and fire would have been more dramatic and visual though.
All I know is: I didn’t do it.
this is exactly what my psychic told me would happen.
Well, happy birthweek. Also I need to figure out how to teach Claude to write like you
.
@katylava
@ChadP

@katylava seeing you post here made my day! Had no idea you had a Scapegoat Emeritus badge.
@katylava Thank you for the blast of nostalgia seeing your name brought! Hope that you are doing super well!
@katylava Hi there!
@shawn Thank’s for giving us an oh shit report. I missed those even though I am not a techie at that level.
On a more important issue (
), since it trashed the members/VMP hour only will the one today have both day’s worth of stuff and last longer?
@Kidsandliz you should be able to buy the members stuff from yesterday through ICYMI, right?
@djslack OK. Didn’t think of that. I had presumed they never got listed.
I dearly love to read Ohshit reports, since I’m retired from the fray. This was a good one, thanks.
So basically a bird flew in the window and caused a chaotic storm of events, right.
So this had nothing to do with me setting that creature loose in Irk’s nest?
@Thumperchick We all know that this elaborate and convincing fiction is just a cover for your shenanigans. But we love you anyway. Or maybe partly because.
I got my wish. The report came out so I wasn’t disappointed. I am disappointed that I missed all those Meh faces during the outage and there’s no way to go back in time to click on them.
“Our stores and services run on Microsoft Azure,”
So? I love purple.
Always.
@Barney Yes, but Crayola and Sharpie don’t always give you the same purpleness.
KuoH
@kuoh Purple is purple. Always.
@Barney So . . . if they’d been running on Purple, none of this catastrophe would have occurred. duh
(I briefly tried to come up with a plausible cloud hosting platform, but decided that it wouldn’t add much of value and wouldn’t be worth the time I’d inevitably lose to it. But fwiw, my first draft is The People’s Pan-galactic Purple.)
/showme The People’s Pan-galactic Purple
@joelmw Here’s the image you requested for “The People s Pan-galactic Purple”
@mediocrebot I like it. That’ll do, pig.
@joelmw You are really wound up today. I’m enjoying it, except for the one topic I didn’t understand.
@Barney What’s that? It’s so good to see you, btw.
@joelmw https://meh.com/forum/topics/laocoon
Maybe one of these days you can explain it to me. Nah, I think I am now past having things explained to me and me able to understand what is said.
And – I’ve enjoyed seeing you, too. It’s hard to believe it’s been so many years since we first met.
@Barney It’s admittedly a bit of an idiosyncratic, slightly obtuse thing here in the States. Blake is more a big deal in England.
My high school Humanities teacher lent me–a theologically conservative (already a mad leftist politically) fundamentalist–her facsimile of The Marriage of Heaven and Hell. Few events (the birth of my child, the weddings) have had a more profound impact on my wee brain and life. She had me read that and Freud’s The Future of an Illusion. Surely she was trying to break me.
Anywhoodle, in simplest terms, it’s about creating a life and spirituality centered in Art and Imagination.
Blake accepts and expands on the Christian canon. He also famously recast Milton’s Paradise Lost, doing a kind of Being John Malkovich in the process. In addition, of necessity, he creates his own mythology.
Some of his poetry is super accessible though. I highly recommend the Songs of Innocence and Experience, if you don’t want to join me on the crazy train. And I’d argue that the rest of his work is mostly allowing yourself to take the ride, and being willing to cast of the encumbrances of religion and normative culture.
I’ve read some wild stuff. None wilder than William Blake. But it resonates and makes sense. Far more sense than most things in this life.
@Barney I’d been thinking about the good old days here, and the fun we had. The limericks, the wanton recklessness, the going at each other and the world fearlessly. I know it’s a cliche, but things have devolved. It’s not a unique thing; arguably it’s inevitable.
@Barney Here are some early and pertinent bits from The Marriage of Heaven and Hell


@joelmw I believe I slept through all of this in my freshman lit course. It was an early class and I wasn’t an early type person.
As for Meh, I had high hopes that we would continue on our journey with little interruption, but I have been more pissed than pleased at what’s been going on around here.
@Barney @joelmw You should pair up with Purple mattresses! Think of the advertising you could do: “Scale up your infrastructure while lying in bed!”
cough AWS with scaling cough
I fell into a black hole a while back, but I’ve finally climbed out and am among the living (if you can call this living, amiright?).
What a wild coincidence that I come back just as such crazy stuff as this is going one. Anyway, I guess I had a few things to say:
ClarkConger?MediocreMercatalystStores.com complete the circle of life and buy out that river-named online store? (I can’t remember if we’re allowed to say its name out loud here or not.)I suppose people are gonna tell me I’m supposed to actually read the email I get. Like that’s gonna happen.
Thanks for the explanation! Only understood about 0.74% of it, but thanks all the same.
Ah, just like how Woot used to (still does) crash during a Woot-off.