[METHODS] The Blackhat Guide to Social Network Botting - Re-Released Extended Cut + [PDF]

Status
Not open for further replies.
1. Did I miss something? If you can't log in using the same browser, then what are your options?
2. Also, do bots (not the imacros ones) count as these browsers?
I mean, let's say I want to use xtumblebot to bot tumblr or quickfire stumblr to bot stumbleupon, do bots like these act the same way as browsers do with regard to "my browser being unique"?
3. I'm inclined to think that these bots do not act the same way as conventional browsers do. However, what about social networks where I cannot find a decent bot around?

I asked because I was doing pretty ok with getting tumblr followers, reblogs and traffic from tumblr before I got the mass bans. I had been thinking if its worth the money. I also noticed that you have an affiliation with xtumblebot but judging by your posts and the white name, I trust that you will give an unbiased opinion.

I was thinking of getting xtumlebot once I have the money. If I ever decide upon it, I'll be sure to use your sig above.

Hey Black.Cat ... I was referring to using your accounts with a bot, and then logging in to them with your browser. Browser foot printing is possible these days and a handful of networks are using it. Our bots do not have unique footprints like your browser may. Foot prints can be as simple as useragents or as complex as analyzing your system fonts and browser plugins installed with a nasty little javascript.

For social networks without a decent bot, I recommend setting up different browser profiles in firefox and using noscript plugin + a user agent switcher and a flash blocker (for lso cookies)

xTumble bot is a good tool for intermediate to advanced botters, I do not recommend it to beginners. As a general judge of ability ... If most of the info in this guide is ground breaking then you are probably beginner to intermediate ... if 50% or more was familiar you are probably intermediate. If 100% is familiar + you are laughing at the simplicity of the guide then you are advanced, and you should add more info to this thread ;)
 
ZenoGlitch .. I think I love you.

Although yes, some of steps may be common sense but sometimes you go off track & completely miss out the common sense stages; a brilliant read - refreshed my mind completely.

Thank you for this exclusive BHW share! Shares like this make me feel very privileged to be apart of this community.
 
Last edited:
Thanks for the heartfelt bromance ;) Let me know if you have any specific questions on the content, i'm happy to clarify anything.

ZenoGlitch .. I think I love you.

Although yes, some of steps may be common sense but sometimes you go off track & completely miss out the common sense stages; a brilliant read - refreshed my mind completely.

Thank you for this exclusive BHW share! Shares like this make me feel very privileged to be apart of this community.
 
Nice thread, but i'd like to know what are some good bots that meet the requirements.

Check the BST section here on BHW, cross reference them with the requirements, and read the customer reviews in the threads is your best bet.
 
Thanks for the guide. Has lots of common sense info in it, but there's lots of people lacking it.
 
The best guide I've read about this topic to date, great work!
I've got a question about the Reblogging chapter, especially the following quote:
xTumble (and many other
bots) allow you to change the source URL for the reblogs, so that the click throughs go to
your landing pages

Is this still true? I manually tried to reblog other posts changing the click-through url but I couldn't manage to do it. The reblog url always remained the same.
Asking around told me that this were some recent changes made by tumblr and that it was possible before.
Has xTumbleBot a special feature so it can still change the click-through url for reblogs or is this out of question?

Greetings
 
The best guide I've read about this topic to date, great work!
I've got a question about the Reblogging chapter, especially the following quote:


Is this still true? I manually tried to reblog other posts changing the click-through url but I couldn't manage to do it. The reblog url always remained the same.
Asking around told me that this were some recent changes made by tumblr and that it was possible before.
Has xTumbleBot a special feature so it can still change the click-through url for reblogs or is this out of question?

Greetings

Now tumblr does show the "source url" but the image click through is not modified. xtumble is looking into a feature addition that allows users to set the re-blogger to work by reposting content as original instead of using the official reblog feature.
 
This should be a sticky. Best start-up guide to botting I have seen. I'd consider myself an intermediate botter and I got a few "ah ha moments" from it as well
 
BULLSHIT BULLSHIT BULLSHIT

First of all if, if you are going to talk about something in detail then atleast have the decency to educate yourself better on the subject instead of reinforcing the mob mentality BHW has towards social media filters when all your doing is spreading the same shit(ofcourse packaged in a nice format, but is still erroneous information like it or not) that is not true.Man, if I didn't call you out on this who would have done this? I guess no one.

Now I am going to tell you why in a very detailed constructive post with tons of pictures why your 3 types of filtering are all made up fairy tales bhw shit you just rehashed

There are 3 types of filtering networks have at there disposal. 1. Keyword filtering. 2. User based filtering (report spam buttons). 3. IP based filtering and browser filtering techniques.
MMMM... so you say that there are 3 types of filtering these networks have at their disposal, lol. Really? Who told you that? Did you learn all of this from bhw or did you actually do any research?Have you actually read any White Papers on the subject or does your knowledge on those supposely 3 types of filters come from reading hundreds of threads all saying the same thing from other uninformed bh'ers? Ever heard of the Facebook Immune System paper ?

Nah.... of course you didn't . Why would I even bother asking when the evidence is clear? How silly of me to think you did lol

If you?re account travels ?x? distance outside of these ?normal usage? statistics, a simple filter can be put in-place to either A. Send your account to a manual review, or B. Automatic ghosting, or account banning.
Oh wait... how could I forget, this thread also has one of those big famous bhw myth's , the "manual review". Geez, why do we need Random Forests, Naive Bayes, Regression ,Fuzzy Logic Matching Algorithms, and more Algorithms when there is manual review. I guess facebook is overpaying these engineers to use all kinds of Algorithms when you could just do the manual review route, so why build the Facebook Immune System to manage all that.

Funny how I never found any mentioning on their Immune System where there is "manual review" for accounts. I guess this "manual review" is really is an exclusive "BHW Urban Legend" lmao.

WARNING!!! LONG DETAILED CONSTRUCTIVE POST AHEAD WITH PICTURES

So before I start I would like to give credit to the first bhw member who posted http://www.blackhatworld.com/blackhat-seo/social-networking-sites/460856-must-read-8-page-report-facebook-spam-team-into-how-they-detect-fake-accounts.html almost a year ago but didn't get enough replies, however it did get 566 views but only 4 replies, and 9 thanks. From those 9 thanks one of them was from G-S-T who actually found this to be one of the best reads he had on that month, and even thanked the guy. Tacolypse thanked the guy too.That goes to show you the real NUGGETS are hidden somewhere in bhw and they don't have to be long to read(its all about the type of content that is in there not just the Quantity), in my honest opinion this is far from being a real NUGGET, but just a rehashed packed thread with infographics.

I guess sometimes quality is a lot better for the few members that appreciate Advanced Topics,even if its a simple link to a White Paper, than a long guide with a bunch of erroneous information with pictures from BHW that doesn't really bring anything new to these thirsty guys seeking for something new to read.For everyone else these guides are a wonderful ways to read your rehash information and for the rest it would make extra money rehashing this into a clickbank product.
Now with my constructive post....
So first lets go over some basics, as posted on the Research Paper.

This is the figure 1 - The adversarial cycle.
ximg.php

A quick explanation goes like this, the Attacker(can be a spammer, hacker, creeper) attacks the system, when the first Attack is detected the System does nothing to combat it. What it actually does is constructs a new training model , and collects as much data as it can in order to better classify this new threat to the Graph. So while the attacker might not know this, he is actually contributing to his own demise. By doing the same type of attack and not changing much, your training a new model where it can easily assemble a very good training model to defend itself once its ready. On the defense stage, this is where the attackers attacks become immune to the system. Either the attacker adapts and changes his tactics, before the Immune System mutates(this means it basically adapts before you to change its defense system before you even plan your next attack) or the cycle ends with the defenders winning the game.

Its a never-ending cat and mouse game played by Facebook and its Attackers(again, spammers, creepers, hackers etc you get the point) . Only the strongest survive this cycle for a long long time, the rest die off.

So if you can get something out of this, just get this. This is a self-learning system, that actually uses data and statistical models to adapt to the environment without much human input. In other words, it can easily have a defense for thousands of different spammers based on their signature of attack, or a big botnet like koobface that uses other compromised computers to do its dirty work, or just annoying fuckers who send those chain letters.It can fight all of those.


Ghosting and Mutate
ximg.php


The phases the attacker controls are the Attacks, and the Detection phases. If you fail to detect before the system mutates, that makes your Detection phase shorter and useless which can lead to your Attacks shorter too. That is the ultimate goal of this system, prevent your detection rate, and your attacks and make their defend and mutate phases longer. That is why they have Ghosting, as a self-defense mechanism that outputs obscure messages, and mutate features like requiring a different IP(or restricting activities per IP basis) which makes it more expensive to keep playing this game for the Attackers.


Actual Decision Making - Not manual review
ximg.php


However, in times of emergency they do have to do damage control. Because it takes time to train a new model and build up the defenses while the attack is happening. It is sometimes better to focus more on saving 98% of 100k users than saving 99% of 1k users, so is not viable to wait longer until a more accurate classifier is made in order to combat the attack. That is where human intervention occurs, especially when you have those viral pages going on. Those are manual reviews that are done in order to eliminate the threat when it becomes viral.


Its time for the fun part, the 4 Main Components of the Facebook Immune System

4 Main Components of Facebook Immune System
ximg.php


Now pay attention to where it says
"
In addition to responding quickly, it is important to target features
that are difficult for the attacker to detect (Defense) and
change (Mutate). This differs from traditional machine-learning
where the features are chosen solely on how strongly they improve
the accuracy of the classifier.
"

This is key to understanding why there IP, Cookies, IP:ACCT ratio are what you shouldn't just solely focus on. It clearly says there that it is important to target shit that you ain't looking at. So to make it a lot easier for all of you to understand, if you only thing that facebook can track you by your proxies, regular and flash cookies, comments(you know the main shit people say when you get detected)then you are very misinformed because no one is talking about the other 495 other feature datasets that we ain't tracking, unless that is you are part of the .0001% that tracks more than just the main 4-5 feature datasets and make your own models.

Just understand this IP and browser user-agents filters, Comment filters, and User based filtering are not FILTERS, they are Feature DataSets, WHICH IS A VERY BIG DIFFERENCE OF WHAT ZENOGLITCH IS CALLING THEM


NOTE: IP filtering is not a main filter like zenoGlitch claims, but one of the many 500 Feature sets

500 Feature Sets
ximg.php


This is a picture of the whole Immune System with all of its components.
Figure 3 - High-Level design of the Immune System
ximg.php


Policies are what you would call your filters, these filters require data which comes from the feature datasets, they use classifiers to do their algorithm modeling for them providing values that are sent back to the policy manager to do its decision. Inside the Feature Data providers 3 loops keeps the data up-to-date with counters there are 3 loops, inner, middle, and outer loop(More on that later).

Policy Layer
ximg.php

Policy Layer-2
ximg.php


There are 2 types of policies, business and logic . The main difference is that business doesn't require any trained data and doesn't mutate. This basically means that its a hard-coded rule, the logic one uses trained models and is always-mutating. It is important to know this, because every single featuredata set can actually trigger another set of policy rules. In other words, lets say you do 1 comment on 1 fb acct and you do have logged into that acct from that ip.

If you are only tracking your comments per acct, and you get banned then you would automatically think it was because you did x comments. That is the wrong way to think about it, because it could have been from other feature datasets that triggered that banned. Which means that IP, Comments to IP ratio aren't as important as you think because those are only part of certain policies that use those feature datasets as params to evaluate the outcome.Remember this exert from the previous quote"where the features are chosen solely on how strongly they improve
the accuracy of the classifier." well this means that basically where IP, IP:Acct , Comment:ACCT ratio were very important at first, now they matter less because other feature sets are given more weight than those feature datasets when you are using normal rates for those feature datasets. Just remember again there are over 500 Feature Datasets


Design - ClassifyScore not TrustScore
ximg.php


This is actually called the ClassifyScore , and it is not for just Accounts. This is used for every classifier
It?s not always as simple as just figuring out the what and the when. Sometimes gaming the popular feeds have more complex algorithms that take into consideration the ?trust score? of the account, the relevent content on the account, the keywords used, and even manual reviews for popular content in rare cases.

Classify Score - Not Trust Score
ximg.php



Classifier in real-time
ximg.php

Classifier in real-time-2
ximg.php

From those 2 paragraphs what you need to get out of this is that there are mainly 2 types of classifiers, real-time and not real-time. The real-time are classifiers that models and classify feature data sets in real-time. For example, for the facebook chat if you are sending the same message over and over, this is an example of real-time where its instant classification of a spam message. The non-live require time to read more data and understand what is going before making a decision. That is why it takes time for you to get banned, because it completely analyzes all of the feature datasets before making a hard decision. YOu ever notice why you don't get banned right away after doing x actions, is because evaluation is in process through these modeling is not instant.

Just so you know they compile test cases of the entire world(all of it facebook) to build new models and test certain certain models. Ask yourself are you doing that?

Alongside of keyword detections there will be string analyzers that search for duplicate match strings over ?x? volume. So if exact match string exceed acceptable duplicatent comment rate in ?y? ammount of time then flag/ban/delete/ send this account to a queue for manual review.
They use Fuzzy Logic Matching Algorithms they don't need to do what you say they do.
Ever heard of
-N-gram
- Q-gram-based Algorithms
- Levenstein

Yea well... it clearly says they use Fuzzy logic Matching Algorithms

Fuzzy Logic Matching
ximg.php


Now for creating the rules you use FXL, now it is important to know this. The expressions use extensive use of subtrees and memoization techniques, this is not some "if this x then y" bullshit like OP is saying. Ever heard of Ensemble learning?
Expressions as trees - memoization, subtrees. This is mostly just a scripting language they created to create the policies and the classifiers.
ximg.php

Feature Data Providers


ximg.php

Feature Loops
ximg.php

Loops - Inner_Middle_Outer
ximg.php


Now the important stuff to remember is that those feature datasets need to be updated with counters, that is why they have 3 counters.

The inner loop keeps track of simple counters on the feature datasets(like how many times user has done, comments, posts,likes, etc), the outer does a little more computing which computes expressions for matching certain ip's and urls. The outer loop is a way more complicated which does something different, it basically keeps track of the number data from across the network. So in other words, it actually can keep track of the number of times x url has been posted across the whole network, or it can match any message from all the users on facebook to see how many times certain comments are similar to each other.

So remember, innerloop is to keep track of the feature datasets like simple counters, number of likes, comments, posts, etc
Middleloop keeps track of expressions using these feature data sets, like "if x comment = bla bla bla" but using fuzzy logic matching algorithms, especially N-grams.
Outerloop, keeps track of the overall featuredata sets across the whole network to spot similarities across the network . This is to prevent mostly phishing scams, viral pages, malware infections that take over the network virally.


There you have it folks, a very simple breakdown of how the Facebook Immune System works, not just some fairy tale bullshit that the OP is rehashing and packaging for you. You see where there is no mention of "manual review quee", there is no 3 filters based on IP, commenting, and user-based filters.

Remember, the biggest advantage Facebook has over you is user-feedback and Global Data that it uses for Data Modeling, most people don't do that.

If you have read this whole thing up to now. Here is a trick I will show you as a token of appreciation for reading this whole damn long post,this again was posted on bhw over 2 years ago. I can't find the thread but I can tell you it got buried down under so many threads like this.


If you want to create an account on facebook that is captcha free after acct creation(without being pva) then when you go and create a facebook account, CHOOSE AN EMAIL THAT IS FROM A UNIVERSITY, even if you don't own it. Here is how it works.

say you put in the email something like this [email protected], when you finish creating that account. Change your email to the one you were going to use, and resend the verification email. Guess what? You will be captcha free on that account. You see how changing the email can make a difference. That is how Facebook works, they check a lot of your data and I can tell you a university Email is more valuable than a yahoo email when doing an acct creation. Trust me they do look into the type of email you are using. So if they are looking at the email used and they even classify different email providers as spammy, and good then don't you think they would use other Feature DataSets as a way to determine if you are legit or a spammer besides jsut IP, cookies, and x comment,likes, post ratio. Think about it.
 
Last edited:
Wow that's interesting "behind the scenes" analysis.
And no need to be pushy :).
I think then ...dawg and ...glitch are right (at least in some parts) but on a different level. Side note: this immune system is much more interesting then the whole social part of FB :).
...glitch is giving some practical advice how not to trigger all this n-gram magic on the FB side, while ...dawg is telling you that:
1.it's not how it works internally
2.it could and will change and your X likes rule won't work anymore.
Funny thing is that with all these algorithms edu email & change advice ...dawg has given could after some time also trigger the spam alarm.
Thanks ...dawg it was a long and d@mn interesting read although machine learning was my hobby some time ago.
Take care,
D.
 
Some meaty stuff there Dawg!

I'm also gonna check out the studies/references in that white paper.

Thanks.
 
Thanks for the post youfeelmedawg, but you are missing one key point this is not a facebook specific writeup or specific to any network at all, it is a general reference that is supported by my first hand experience inside the operations of a trust and saftey team of a large start-up (not facebook) so it's not just coming out of my ass. Filters and queues do exist and are not "fairy tales" I can verify this first hand for two large networks and with other associates that have worked within the T&S teams with those networks.

Further more I clearly stated this is a beginner to intermediate write up. But thanks for attempting to elevate the conversation, I can look past the ignorant and disrespectful stance you took by assuming I was referring to facebook solely. However... applying facebooks immune system white papers across all social networks is extremely in-accurate.

If the advice I provided is taken into consideration you will run a more stable botting operation on average on MOST networks.
 
Last edited:
I also want to apologize for coming too strong and pushy, i over stepped my boundries there. I was also a lil bit too pumped up on 2 20 oz cups of coffee so i was ranting off too much at the beginning.

My whole point on this thread was to open awareness to the skills that are needed in order to fight back the algorithms set in place with these social networks.

Just how we would need programmers to make the bots for automating these social networks, we would also need people highly skilled in some statistical program, something like R . These would be the people that would be able to take data from different observations and give conclusive data based on statistical models, how the social spam filters are reacting to certain botting activity.

With that in mind, its also good to know that in todays economy you can get very good PhD guys that are underemployed at some university getting paid a very low amount of money for what they perform, or even the PhD guys who are not even working on their fields of study that can be highly used in a system where they provide the models based on the data the bots give.

What data? data like ip:acct ratio, how many comments has an acct made, likes, posts, events, delay per each action, all of this data and more can be used to know what variables could be changed to get a higher success ratio.
However, there is one thing I do agree with zenoGlitch on his thread, commenting is by far the most profitable thing you can do on any of these networks, commenting on instagram is highly profitable, same as top comments on youtube(I know this from first hand).

So 250k seems like a low balling amount if you were to almost semi-automatically be able to bypass the anti-spam algorithms whenever you encounter them, or even better, being able to change yoru tactics automatically before you get a trained model built after your botting.I would think that if there were statisticians employed with a system that can deploy automatically tests based on the tests the statistician would want for his/her modeling , I believe its possible to reach a 7 figure yearly straight from commenting.

In that case, there would have to be complete anonymous accts and eliminate the trail cash and become anonymous, because at that level am pretty sure you would have the social networks try to come after you instead of trying to play a cat-n-mouse game.
 
No apology needed, you added some good facebook related info, and I wish there were more of that in the thread from our more edumahcated members. Carry on.
 
Status
Not open for further replies.
Back
Top