At one point I got the something went terribly wrong page but instead of the irk video it had a picture of the rechargeable insect rackets from a few days ago…WTF?
I suspect the Meh AI power consumption has been increasing exponentially due to it possibly being on the verge of reaching sentience. Soon it will know what we want before we realize it and have already charged and shipped the order without us even lifting a finger.
This has been occurring for several days. Although it seems to have normalized at the moment. When it first started happening, I ran side by side CMD prompts. One continuously pinging Meh’s DNS entry and one pinging Google’s DNS server. Meh currently is about a 65-70ms response time, which isn’t bad and also seems to be very consistent at the moment. Google’s DNS server averages about 5ms. Because google.
Although over the last few days I have seen wild spikes in latency from Meh with blocks of dropped packets. I ping both Meh and Google so if Meh has issues I can look at Googles response and if it’s good I at least know it’s not on my end. I’m guessing they are running on Azure cloud and MS is pretty tight with their routing configurations, so perhaps Meh’s hosted VM(s) had a memory leak and they bounced it because it seems stable since this morning. Storage and IOPS shouldn’t be a problem for Azure.
Still having ‘issues’. Page doesn’t load/refresh when I post a comment or reply but it does show up if I close and reopen the window… which is a PITA to verify if it posted.
#1st world problems…
meh.com Experiencing Intermittent Connection Errors — Azure App Service
We’ve noticed a slight increase in error rates affecting meh.com over the past couple days. Nothing crazy alarming but frequent enough to be annoying. The Failed Transaction Rate (Average) increased to 0.5% and spiked a couple of times to 4.7%. Here’s what that looked like from our Application Performance Monitoring.
It was a confusing one because the root cause masqueraded as an occasional application bug when the real problem ended up being a network layer issue with our cloud provider. Specifically, we were battling SNAT (Source Network Address Translation) port exhaustion or what this Dynatrace blog calls “Azure’s most reliably invisible production failure”.
/showme SNAT
That’s a pretty good description because I wasted a lot of time combing over recent code changes to try and find something that slightly increased errors on meh.com but not for any other stores.com sites.
Turns out there was no code issue to solve. I proved that by taking the same code, deploying it to a different Azure App Service in a new subnet and lo and behold the error rate dropped back to 0.
@shawn Thanks for the update. I was almost in the ballpark with no troubleshooting Heh. I actually just went back to college (I’m old and this is a young man’s game) to take an AWS infrastructure class. We moved our Citrix VDI backend (DDS, Storefront, etc) to AWS but had a third party for our cloud space support, but because we were in a federally regulated environment, we couldn’t put client data in the cloud so it was kind of a hybrid environment that had its share of challenges. I took the class just because I wanted to understand more. Had a shitty professor, but I learned some stuff Hope things continue to run stable for you all. I haven’t worked with SNAT, but even NAT can be problematic if things go wonky
@f00l A whois lookup on the IP address and also a tracert to the server. edit - I’m a little bit of a tech nerd. And a cooking nerd. And a coffee nerd. Yeah… that’s me edit - I geek out about a lot of things
Thank you to the “powers that be” for parsing out this problem. As you say, things seem to be running more smoothly now. And even though the oh shit report went WAY over my head I appreciate the transparency.
It’s definitely you!






No but seriously. They’ve been a bit of a hot mess today in particular.
Yes, something went terribly wrong. A lot.
Worse than last week, not as bad as parts of last year, so no, it’s not just you.
I had a problem as soon as I clicked on the site
At one point I got the something went terribly wrong page but instead of the irk video it had a picture of the rechargeable insect rackets from a few days ago…WTF?
Wonder if we’ll get an oh shit report…
I suspect the Meh AI power consumption has been increasing exponentially due to it possibly being on the verge of reaching sentience. Soon it will know what we want before we realize it and have already charged and shipped the order without us even lifting a finger.
KuoH
@kuoh AI thAInk you AIre rAIght.
@kuoh
But it will still take forever to actually get to your house!
¯\_(ツ)_/¯
This has been occurring for several days. Although it seems to have normalized at the moment. When it first started happening, I ran side by side CMD prompts. One continuously pinging Meh’s DNS entry and one pinging Google’s DNS server. Meh currently is about a 65-70ms response time, which isn’t bad and also seems to be very consistent at the moment. Google’s DNS server averages about 5ms. Because google.
Although over the last few days I have seen wild spikes in latency from Meh with blocks of dropped packets. I ping both Meh and Google so if Meh has issues I can look at Googles response and if it’s good I at least know it’s not on my end. I’m guessing they are running on Azure cloud and MS is pretty tight with their routing configurations, so perhaps Meh’s hosted VM(s) had a memory leak and they bounced it because it seems stable since this morning. Storage and IOPS shouldn’t be a problem for Azure.
Shrug
@capnjb
Thanks… I think…
@capnjb
Ouch. That hurt!
Still having ‘issues’. Page doesn’t load/refresh when I post a comment or reply but it does show up if I close and reopen the window… which is a PITA to verify if it posted.
#1st world problems…
OHSHIT REPORT
meh.com Experiencing Intermittent Connection Errors — Azure App Service
We’ve noticed a slight increase in error rates affecting meh.com over the past couple days. Nothing crazy alarming but frequent enough to be annoying. The Failed Transaction Rate (Average) increased to 0.5% and spiked a couple of times to 4.7%. Here’s what that looked like from our Application Performance Monitoring.
It was a confusing one because the root cause masqueraded as an occasional application bug when the real problem ended up being a network layer issue with our cloud provider. Specifically, we were battling SNAT (Source Network Address Translation) port exhaustion or what this Dynatrace blog calls “Azure’s most reliably invisible production failure”.
/showme SNAT
That’s a pretty good description because I wasted a lot of time combing over recent code changes to try and find something that slightly increased errors on meh.com but not for any other stores.com sites.
Turns out there was no code issue to solve. I proved that by taking the same code, deploying it to a different Azure App Service in a new subnet and lo and behold the error rate dropped back to 0.
Apologies to @chienfou and others (@sillyheathen @pmarin @werehatrack @Star2236 @kuoh @cfg83 @capnjb @f00l) that reported here but we seem to have things running a bit smoother now.
@shawn Here’s the image you requested for “SNAT”
@shawn Thank you for promptly addressing this issue. It was a bit maddening, which is the last thing I need
@heartny @shawn Thanks for the SNOT, Succinct Network Outage Treatise.
KuoH
@shawn Thanks for the update. I was almost in the ballpark with no troubleshooting
Heh. I actually just went back to college (I’m old and this is a young man’s game) to take an AWS infrastructure class. We moved our Citrix VDI backend (DDS, Storefront, etc) to AWS but had a third party for our cloud space support, but because we were in a federally regulated environment, we couldn’t put client data in the cloud so it was kind of a hybrid environment that had its share of challenges. I took the class just because I wanted to understand more. Had a shitty professor, but I learned some stuff
Hope things continue to run stable for you all. I haven’t worked with SNAT, but even NAT can be problematic if things go wonky 
@capnjb
How did you know/guess it was Azure?
@f00l A whois lookup on the IP address and also a tracert to the server. edit - I’m a little bit of a tech nerd. And a cooking nerd. And a coffee nerd. Yeah… that’s me
edit - I geek out about a lot of things 
@shawn Ok… after a bit of reading I guess I’ve used SNAT quite a bit and because I’m old school and Y2K I’ve always used NAT in my brain
And a nerd joke for @shawn
I speak TCP. My wife speaks UDP. Usually when I’m on the other side of the house
@capnjb @shawn I’m just an IPv4 martian trying to creep around in the IPv6 world.
KuoH
@kuoh @shawn I mean… v6 is in use, but not as widespread as was expected when it launched. I’ll stick with my octets, thank you very much
@kuoh I had to disable IPv6 or I couldn’t get passed reCAPTCHA. I got tired of identifying crosswalks and fire hydrants over and over and over.
As a fellow robot, I understand your pain.
Thank you to the “powers that be” for parsing out this problem. As you say, things seem to be running more smoothly now. And even though the oh shit report went WAY over my head I appreciate the transparency.