Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Alright guys, I haven't had time to get on the forums and let everyone know what is going on with their hosting. As some client's have already stated, our Datacenter had a cooling failure. This is the official RFO I received

Summary:

On May 6th 2012 at approximately 1:30pm CDT Data Center DLLSTX03 experienced an interruption of utility power of approximately 15-20 seconds. UPS and generator systems operated and functioned as expected and provided uninterrupted power to critical systems on the Data Center floor. However the central cooling plant did not re-initialize properly after the utility power event. Our Network Operation Center dispatched our Facilities Engineering team. Facilities Engineers initiated our emergency response process. Facilities Engineer worked with system vendors to restore normal operation.


The system was fully brought back online by 05:30am. Temperatures are back to normal. We are monitoring the system closely. We are working with our system vendors and suppliers to fully investigate the outage and understand the root cause. An updated RFO will be released based on those investigations/analysis.


This cooling failure caused a great deal of problems for us as well as every other customer in that data center. We had a total of 16 Server(s) fully overheat and crash during the cooling outage. Of these 16 servers 9 of them were brought online within 2 hours of cooling being fully restored. The remaining 6 servers all of which were Core Nodes (Meaning they are very powerful and hold a lot of clients) had reached critical temperatures and were completely unresponsive . We soon found out that we had lost hard drives in every single one of the servers. We use RAIDs to protect your data, and so there was no data loss. In a perfect world a RAID will actually make it where the server can continue to function without a hard drive. However since the servers had overheated and were unresponsive we had no choice but to manually power cycle the servers. When we did power cycled the servers, the servers RAIDs became out of sync and so the servers refused to boot. The server(s) RAIDs would now all need to be completely rebuilt manually, and FSCK's run on all of the servers, in order to get them to boot normally. This task on a normal day would have been very easy to accomplish from our office in Tulsa, Oklahoma. Today however the data center was having such a support overload that our initial reboot requests took 1 hour to fully process, and that was with no less than 15 harassing phone calls from myself. I knew that I couldn't let our servers be down all day until the datacenter got their act together. I had to make a decision, and I had to get to Dallas. I loaded up all of my equipment in my car and headed off to Dallas. I called several emergency linux admins located in Dallas on my way there, I guess emergencies don't happen on Sundays. I was able to successfully bring two servers online while driving to Dallas with some cooperation from the Datacenter, we successfully brought two nodes online. I arrived in Dallas at 7:45 PM CST (It's a 4 hour drive). I quickly got to work, rebuilding the arrays, and running file system checks across the board. I was actually very happy with my speed, as I had brought 3 more of the nodes fully online by 9:00 PM CST. This left me with two nodes to go, I was sure I would have everyone online by days end. I consoled into node 7, and noticed it had much different boot errors than the other servers. In fact it didn't even get far enough to get a boot error, the kernel froze before it could even begin the boot cycle. The errors on the screen (which indicated memory addresses at which point the hang had occured) indicated a RAM issue, it is not uncommon for RAM to fail under high heat in fact this particular server has a setting that intentionally shuts itself down when the DRAM hits a certain temperature (I guess the thermometer is broken). We have spare parts of EVERYTHING literally we could pretty much have any part of a server fail, and replace it with fresh parts. I swapped out all of the RAM sticks from node 7, and attempted to boot it. The server started to boot and then hung at the same exact point. I knew the RAM I put in the server was good as we test all parts before even putting them in our datacenter to be used as spares (What's the point in replacing a broken part, with another broken part). I knew it must be a motherboard issue, that was casuing the DIMM's to be bad. I pulled one of our spare node's and filled it up with the spare RAM, I then placed the hard drives into the new server and fired her up. The server booted this time.....unsuccessfully, This time I had the same RAID error, I had on the first two servers, so I reassembled the RAID from scratch, and dumped the data onto brand new hard drives (If that server got hot enough to fry the board, I can no longer trust any part of it) . This process took quite a long time, and allowed for me to begin work on the only other down node at this time (VPS11, a Windows VPS) This node had the exact same issues as the first nodes, so I started the raid rebuild process,and went back to node 7. The RAID finally finished the rebuild process on node 7, and I was able to boot it far enough to require a FSCK. I started the FSCK which took around 4 hours to run on this server (2 Terabytes of Data). The FSCK finally finished repairing the file system, and I thought we were ready to go. I rebooted the server, to get a kernel hang error, however this was a good kernel error in that it simply was unable to properly mount the partitions. It actually took me quite some time to figure this one out, but the server I replaced it with had a newer bios than the other server, which had AHCI turned on by default, we did not originally install the OS with AHCI turned on, so we had to manually load the AHCI drivers, and rebuild the kernel all from linux rescue mode, Once this had finally been completed (Around 8:00 AM) node 7 finally came fully online. The only node left was VPS11, I had started the FSCK on this server at some point while I was figuring out why node 7 wouldn't boot, and so the fsck was fully completed, and the node booted successfully. The only problem was the logical volume was nowhere to be found since we had to rebuild the RAID all of the partition information (UUID's) were slightly different than before. I quickly pulled our LVM config for VPS11 off of our backup server, and brought the LVM online manually with a little help from backups :) (The backup contained information on how the LVM was assembled there was 0 client data loss) Once the LVM was assembled, I was able to boot up every single VPS on node11, and Hostwinds was back online the way we should be.


I know that isn't your typical RFO, or response from a hosting provider, but I just wanted to tell you guys the truth everything that happened, and everything that has made up my weekend. I haven't slept since our service alarms woke me up yesterday at around 9:00 AM. I'm really sorry to everyone who was affected by this downtime, a few good things did come from this however:

Node 7 Clients (Shared3), you guys have faced some real problems lately, and this brand new hardware should eliminate all of them

The other nodes that went down got new RAM sticks where old ones did not pass with 100%, also brand new hard drives were deployed replacing the old hard drives just in case they got too hot. We ran CPU stress tests on the servers and our benchmarks indicated there was no CPU damage (Which is really lucky)


We learned a lot about how dependent on our data center we are, this is why we are now developing and working on a system, where we can have a set (5-10) of "cold servers" (servers that are not powered on but are 100% configured and ready to go) where we can easily swap out hard drives, or migrate data, and turn 12+ hours of downtime, into 10-20 minutes, even in the worst conditions.

I really appreciate how nice most of you have been to us, I know it's frustrating when your site goes offline, which is why I started Hostwinds in the first place, because I wanted to provide a better service than you would get anywhere else. I know we aren't perfect sometimes, however we try extremely hard to provide the best service we possibly can, all client's that had any downtime at all are receiving 1 month of free hosting, and all client's who had extended downtime will be receiving 2 months of free hosting. We know it can't possibly begin to make up for what has happened, but we absolutely want to try to make it up to you in the best way we can. I hope that you can find a way to forgive us for the downtime, and not let this incident mar your opinion of us as a company, we do the best that we can, and that's all we can do.


--Peter Holden
CEO Hostwinds.com
 
all client's that had any downtime at all are receiving 1 month of free hosting, and all client's who had extended downtime will be receiving 2 months of free hosting

That's awesome thanks Peter!

Appreciate you being honest & letting us know what's going on.

Shit happens... At least things are fixed now & a better future
with Hostwinds is in the works. I'm still sticking with Hostwinds
even after this ordeal.
 
so...after i got a downtime for a few days in a row, i receive a message that my account was terminated because it has expired (i had the hosting paid until 2013)

after i resolve that issue and my account got back, the sites are still down. the next day i ask again why my websites are not alive yet, and they say there are no sites in my account.
WTF ?! i log in the Cpanel account and i see i don't have any addon domains. so out of 5 websites i have 0 now.

after all this downtime and mistakes, deleting accounts etc, i receive the best news "Unfortunately we dont have backup of yours :("

i haven't done a backup for a long time so my websites are pretty much bye bye now, the last backup was only half of the posts each website had

guys, at least send a nice written "sorry you got f#@$ up" letter
 
Moved to croc web nice and pain free and cancelled my hostwinds account, didn't even recieve an apology just a one word line saying my account is now deleted.
 
LOL, greatest response ever after a few days full of mistakes (deleting accounts etc etc). here it is - "I have taken a look, and we did not delete or lose any of your data, the only way this could have happened is by someone logging into your website and deleting all of your files"

It's bad when you start treating your clients like this
 
There is no way that our servers could have "lost your data"

Peter Holden
You're telling fibs
[22/12/2011 03:30:25] Oxonbeef: 209.105.251.85
[22/12/2011 03:32:15] Oxonbeef: I just went to log into my VPS and couldn't I went to submit a ticket and I noticed you had moved me to a new VPS node. I logged in and none of my files, folders or programs are there.
[22/12/2011 03:46:49] Peter Holden: hmmmm
[22/12/2011 03:46:54] Peter Holden: oh no
[22/12/2011 03:47:10] Peter Holden: this is chris?
[22/12/2011 03:47:15] Oxonbeef: yes
[22/12/2011 03:47:42] Peter Holden: shit
[22/12/2011 03:47:48] Peter Holden: Also please note we will be taking the other server offline in 72 hours
[22/12/2011 03:48:04] Peter Holden: 209.105.251.85
[22/12/2011 03:48:06] Peter Holden: fuck fuck fuck
[22/12/2011 03:48:12] Peter Holden: lemme see if I can recover the files
[22/12/2011 03:48:16] Peter Holden: we were just about to format the drives
[22/12/2011 03:48:21] Peter Holden: this is gna tkae me like at least 8 hours
[22/12/2011 03:48:26] Peter Holden: hard drive recover is not fun
[22/12/2011 03:48:40] Oxonbeef: Please don't tell me you didn't back it up before transfreing?
[22/12/2011 03:50:40] Oxonbeef: http://www.cgsecurity.org/wiki/TestDisk best program I know.
[22/12/2011 19:03:39] Peter Holden: you there?
[22/12/2011 19:33:14] Oxonbeef: yes mate
[22/12/2011 22:45:29] Peter Holden: ok
[22/12/2011 22:45:35] Peter Holden: I dont think I can recover this data
[22/12/2011 22:45:42] Peter Holden: did you not get our email?
 
You're telling fibs
That was not our server, that was me deleting your VPS after you failed to read the ticket, even in you thread you said where I refrenced the original ticket

"[22/12/2011 03:47:48] Peter Holden: Also please note we will be taking the other server offline in 72 hours"


You were given 72 hours to transfer your files, and I removed your old server after those 72 hours were up, why are you posting in a thread about an issue that happened months ago that is in no way related to this at all, stop trying to stir up trouble.
 
Signed up last night see how it goes for a month if good for longer.
 
Down again...

Looks its online again, but was down like for few hours.
 
Last edited:
I subscribe Hostwinds for almost 1 year now..
I can remember experienced server down for 3 times including this month..

Personally still acceptable for me since it's still cheaper price..
 
Down again...

Looks its online again, but was down like for few hours.


After the data center problem, I started using 2 free site up-time checking services I found through Google to check every 30 minutes

I'm on shared business server 3 and up-time has been 100%

I'm sure that if you can provide Hostwinds with down time evidence (date, time and duration), they'll either fix your server, or move you to another server
 
We have recently upgraded our server status page so you guys can get up to the second information about your server and how it's doing

https://clients.hostwinds.com/serverstatus.php


We have also officially launched windows VPS's and Cloud Servers


After the data center problem, I started using 2 free site up-time checking services I found through Google to check every 30 minutes

I'm on shared business server 3 and up-time has been 100%

I'm sure that if you can provide Hostwinds with down time evidence (date, time and duration), they'll either fix your server, or move you to another server

Perhaps you could share with everyone which up-time checker you are using so that they can use it.
 
Try my luck with hostwind ~ Hope the coupon still work :P
 
Try my luck with hostwind ~ Hope the coupon still work :P

We will exceed your expectations I promise you, and yes all the coupons still work


We are working on coming up with some new ideas on discounts for our windows VPS platform, what kinds of discounts and coupons would you guys like to see here?
 
Shared2 down from yesterday, money lost. Thats it for me, I had alot downtimes in all this time and I am sure had much more then I saw.

Also all files was lost on earlier server?
 
Last edited:
I was 1 month in using hostwinds..
Got this email again today...why?
Thank you for your order from us! Your hosting account has now been setup and this email contains all the information you will need in order to begin using your account.
and also hostwinds down?
 
I was a customer with Hostwinds for a bout 1 year.. 1 fuckin year of shitty hosting that is.. lol.

I kept thinking they would get better and the price wasn't bad. There servers started going down at least once a month so Peter told me it was because I needed to upgrade to a VPS. So I did... I'm running a multi million dollar website here so I told them I couldn't have ANY downtime. A few weeks after upgrading, the same thing started happening. Peter said I needed to upgrade my VPS... so I did.. and this similar story goes on and on and would still be going on if I hadn't switched providers.

In the end, one of the tech support guys ended up fucking up on a mySQL query for my CRM and they didn't want to take any responsibility for it(Peter actually told me he could fix there fuck up if I payed him an hourly rate).. hahah.

Went and found another hosting provider and my websites haven't been down once in the past 3 months.

This is just my 2 cents. Some people may have positive experiences(at least i hope so lol) with them but I personally would absolutely NEVER recommend them to anyone.

:pat:
 
This is the worst host i have ever used.
Its down about 20% of the time and now all my files have just been deleted.
Add to that my account has been hacked 3 times and you have the perfect trifecta of a crappy host!
 
I still will stay with HW, server changed, maybe downtimes will be solved now. When I tryed to move to HG, the next day I got nonsense DMCA, this never happened with HW. Here is my thread

http://www.blackhatworld.com/blackhat-seo/blackhat-lounge/448217-warner-music-group-dmca.html

I dont understand how they can send these, when they even didnt checked is it really have any copyrighted material.


We appreciate you as a client, and we will continue to improve.

We are launching our optional backup system in about 10 days time, we are just waiting for all of the parts to arrive.
 
Status
Not open for further replies.
Back
Top