Thanks for your feedback and
NOT being a customer, you definitely would not fit into our family here.
Now that the dust has settled on the issues we were having on ONE of our shared servers which hosts a lot of BHW members, I wanted to give a full report of what happened. I'm doing this because we are 100% transparent, it's one of our values and to also let you know that ANYWHERE you host you're on hardware and software that can have issues from time to time. We, unfortunately, had issues with one of our shared servers that had a perfect storm of events. I'll detail the timeline below:
On Sunday August 1st 2019 I commissioned an outside System Administration team to work on one of our shared servers. The problem was that the server was provisioned wrong back in 2017. The / directory had only 20 GB and was consitantly filling up causing issues with logging into cPanel. To combat this, we tried multiple things such as clearing log files quickly (not advised) and running hacks such as clearing cache on the server every 5 minutes (not advised). Moving files around, etc. Nothing worked.
There were two options to solve this issue:
A) Make a complete backup of the server offsite / reformat the existing server to increase the / partition / install cPanel/WHM/ Customize the server back to its orginal performance/ then finally restore from the offsite backup (terabytes of customer data) and pray that everything worked. This would have resulted in 3-5 days of downtime for most customers as this server has 2 TB's worth of customer data and is highly customized for performance. This would have been a VERY risky thing to do, and I did not have much hope in it working very well since cPanel backups for some accounts were over 20GB which causes issues during packaging and restoring.
B) Find a way to increase the / partition without a complete reformat.
So after many weeks of trying to figure out with my team how we can keep the / directory from filling up. Setting up cron jobs to clear out archived logs, and some VERY risky commands we decided to consult outside help. The team we hired to come help us first decided it was the MySQL directory that was causing the issue (Databases). Without notifying us, they started to move the databases to another partition. The problem here was that the MySQL directory was not in / it was on its own partition. So it was not the cause of / filling up. We notified them of this AFTER they had started working with the databases.
When they got my message, they stopped and after many hours created a disk img on /home for cPanel files which were taking up much of the / directory. This freed up enough space (2gb) on / to have the server run as it is supposed to.
During this time however, MySQL went offline multiple times and http as they performed this maintenance.
On Monday 9/2/2019 I was informed the work was complete. However, all MySQL databases were not connecting. I relayed this information over with a couple of example sites, and they fixed them. So we thought everything was good, we did not check ALL sites as there are tons. But the major ones that I know about we did check and I sent to them to fix.
On Tuesday 9/3/2019 I was told the issue was fixed now. I then sent an email out to ALL users that are on the server to check their websites and let me know if they are seeing issues. We were flooded with clients saying their site had database issues. So obviously this was not fixed. My frustration with this company at this point was extremely high. Whatever they did during the initial MySQL moving, that didn't need to happen, crashed all database driven sites. I relayed to the company the impact that this was having and they worked over night to restore from a sandbox they had the databases.
As of this morning 9/4/2019 at 1am I have been notified that all sites are working and they sent over exactly what they did to fix it (restore the databases tables that were corrupt) which thankfully they had a backup of.
I tested all sites that were reported and have been testing even more all morning, everything seems to be functioning as it was before.
I can't tell you how upset this makes me as StressFreeHosting is supposed to be... stress free. This was not stress free for our users or for me. I haven't slept in over 48 hours while I worked with this company to resolve the issue they created.
I do apologize for this, it was not anything at all that was planned and I had no idea that they would execute commands that would end up having one of our shared servers down for so long. Very disappointing and lots of egg on our faces here at StressFreeHost.com.
Current Status:
All sites are functioning as they should
MySQL is repaired and databases are loading
/ is freed up (Which is what was the goal of the whole maintenance)
Speed/performance of the server is stable
Again my apologies for this and if you are still seeing issues with your site PLEASE contact us at
[email protected] (Do not PM we do not check that often on BHW).
Also, if your website was down during this time please email our support and we will credit you for a free month of service.
This ONLY affected one shared server (lots of customers on it though). It did not affect our VPS or Dedicated customers and customers on other shared servers.
-Josh