Alright guys, I haven't had time to get on the forums and let everyone know what is going on with their hosting. As some client's have already stated, our Datacenter had a cooling failure. This is the official RFO I received
Summary:
On May 6th 2012 at approximately 1:30pm CDT Data Center DLLSTX03 experienced an interruption of utility power of approximately 15-20 seconds. UPS and generator systems operated and functioned as expected and provided uninterrupted power to critical systems on the Data Center floor. However the central cooling plant did not re-initialize properly after the utility power event. Our Network Operation Center dispatched our Facilities Engineering team. Facilities Engineers initiated our emergency response process. Facilities Engineer worked with system vendors to restore normal operation.
The system was fully brought back online by 05:30am. Temperatures are back to normal. We are monitoring the system closely. We are working with our system vendors and suppliers to fully investigate the outage and understand the root cause. An updated RFO will be released based on those investigations/analysis.
This cooling failure caused a great deal of problems for us as well as every other customer in that data center. We had a total of 16 Server(s) fully overheat and crash during the cooling outage. Of these 16 servers 9 of them were brought online within 2 hours of cooling being fully restored. The remaining 6 servers all of which were Core Nodes (Meaning they are very powerful and hold a lot of clients) had reached critical temperatures and were completely unresponsive . We soon found out that we had lost hard drives in every single one of the servers. We use RAIDs to protect your data, and so there was no data loss. In a perfect world a RAID will actually make it where the server can continue to function without a hard drive. However since the servers had overheated and were unresponsive we had no choice but to manually power cycle the servers. When we did power cycled the servers, the servers RAIDs became out of sync and so the servers refused to boot. The server(s) RAIDs would now all need to be completely rebuilt manually, and FSCK's run on all of the servers, in order to get them to boot normally. This task on a normal day would have been very easy to accomplish from our office in Tulsa, Oklahoma. Today however the data center was having such a support overload that our initial reboot requests took 1 hour to fully process, and that was with no less than 15 harassing phone calls from myself. I knew that I couldn't let our servers be down all day until the datacenter got their act together. I had to make a decision, and I had to get to Dallas. I loaded up all of my equipment in my car and headed off to Dallas. I called several emergency linux admins located in Dallas on my way there, I guess emergencies don't happen on Sundays. I was able to successfully bring two servers online while driving to Dallas with some cooperation from the Datacenter, we successfully brought two nodes online. I arrived in Dallas at 7:45 PM CST (It's a 4 hour drive). I quickly got to work, rebuilding the arrays, and running file system checks across the board. I was actually very happy with my speed, as I had brought 3 more of the nodes fully online by 9:00 PM CST. This left me with two nodes to go, I was sure I would have everyone online by days end. I consoled into node 7, and noticed it had much different boot errors than the other servers. In fact it didn't even get far enough to get a boot error, the kernel froze before it could even begin the boot cycle. The errors on the screen (which indicated memory addresses at which point the hang had occured) indicated a RAM issue, it is not uncommon for RAM to fail under high heat in fact this particular server has a setting that intentionally shuts itself down when the DRAM hits a certain temperature (I guess the thermometer is broken). We have spare parts of EVERYTHING literally we could pretty much have any part of a server fail, and replace it with fresh parts. I swapped out all of the RAM sticks from node 7, and attempted to boot it. The server started to boot and then hung at the same exact point. I knew the RAM I put in the server was good as we test all parts before even putting them in our datacenter to be used as spares (What's the point in replacing a broken part, with another broken part). I knew it must be a motherboard issue, that was casuing the DIMM's to be bad. I pulled one of our spare node's and filled it up with the spare RAM, I then placed the hard drives into the new server and fired her up. The server booted this time.....unsuccessfully, This time I had the same RAID error, I had on the first two servers, so I reassembled the RAID from scratch, and dumped the data onto brand new hard drives (If that server got hot enough to fry the board, I can no longer trust any part of it) . This process took quite a long time, and allowed for me to begin work on the only other down node at this time (VPS11, a Windows VPS) This node had the exact same issues as the first nodes, so I started the raid rebuild process,and went back to node 7. The RAID finally finished the rebuild process on node 7, and I was able to boot it far enough to require a FSCK. I started the FSCK which took around 4 hours to run on this server (2 Terabytes of Data). The FSCK finally finished repairing the file system, and I thought we were ready to go. I rebooted the server, to get a kernel hang error, however this was a good kernel error in that it simply was unable to properly mount the partitions. It actually took me quite some time to figure this one out, but the server I replaced it with had a newer bios than the other server, which had AHCI turned on by default, we did not originally install the OS with AHCI turned on, so we had to manually load the AHCI drivers, and rebuild the kernel all from linux rescue mode, Once this had finally been completed (Around 8:00 AM) node 7 finally came fully online. The only node left was VPS11, I had started the FSCK on this server at some point while I was figuring out why node 7 wouldn't boot, and so the fsck was fully completed, and the node booted successfully. The only problem was the logical volume was nowhere to be found since we had to rebuild the RAID all of the partition information (UUID's) were slightly different than before. I quickly pulled our LVM config for VPS11 off of our backup server, and brought the LVM online manually with a little help from backups

(The backup contained information on how the LVM was assembled there was 0 client data loss) Once the LVM was assembled, I was able to boot up every single VPS on node11, and Hostwinds was back online the way we should be.
I know that isn't your typical RFO, or response from a hosting provider, but I just wanted to tell you guys the truth everything that happened, and everything that has made up my weekend. I haven't slept since our service alarms woke me up yesterday at around 9:00 AM. I'm really sorry to everyone who was affected by this downtime, a few good things did come from this however:
Node 7 Clients (Shared3), you guys have faced some real problems lately, and this brand new hardware should eliminate all of them
The other nodes that went down got new RAM sticks where old ones did not pass with 100%, also brand new hard drives were deployed replacing the old hard drives just in case they got too hot. We ran CPU stress tests on the servers and our benchmarks indicated there was no CPU damage (Which is really lucky)
We learned a lot about how dependent on our data center we are, this is why we are now developing and working on a system, where we can have a set (5-10) of "cold servers" (servers that are not powered on but are 100% configured and ready to go) where we can easily swap out hard drives, or migrate data, and turn 12+ hours of downtime, into 10-20 minutes, even in the worst conditions.
I really appreciate how nice most of you have been to us, I know it's frustrating when your site goes offline, which is why I started Hostwinds in the first place, because I wanted to provide a better service than you would get anywhere else. I know we aren't perfect sometimes, however we try extremely hard to provide the best service we possibly can, all client's that had any downtime at all are receiving 1 month of free hosting, and all client's who had extended downtime will be receiving 2 months of free hosting. We know it can't possibly begin to make up for what has happened, but we absolutely want to try to make it up to you in the best way we can. I hope that you can find a way to forgive us for the downtime, and not let this incident mar your opinion of us as a company, we do the best that we can, and that's all we can do.
--Peter Holden
CEO Hostwinds.com