[Method] Finding Articles Directories for AMR without Footprints

ct2272

Registered Member
Joined
Nov 13, 2010
Messages
56
Reaction score
44
One of the problems I have always had with Scrapebox footprints was for me they can take forever to produce a result using public proxies or will often burn out my bandwidth quota if I am using private proxies. The last time I used footprints for article directories it took me a week to generate the list and the end result was a slight incremental increae in the amount of directories that were valid after importing them into AMR.

With that said, I have come up with a way to create import lists faster and more efficiently. In a nutshell it goes on the premise of finding AA URLs except for article directories. My typical AMR blast before I figured this out was 274 submitted/100 auto approved. After 1 pass of importing URLs into AMR I was able to increase my result to 700 submitted/300 auto approved.

What you will need:
1. An AA list for an AMR article you have submitted.
2. Scrapebox Link Extractor
3. Scrapebox Backlink Checker
4. AMR

Steps:
1. Take the AA list for the article you have and import it into Scrapebox.
1a. Trim the list to the root domain.
1b. Remove duplicate domains.

This will leave you with a site list that you will be able to load into the Link Extractor add-on.

2. Load the Link Extractor add-on and set to Internal links.
2a. Import the list from the Harvestor.
2b. Hit 'Start'.

The Link Extractor will go to each url and attempt to scrape the entire site. You are ultimately searching for articles because within those articles are the backlinks to the author's site. Once the Link Extractor has worked, Export the list to somewhere you can import it again into the Harvestor. I have noticed with the this add-on that it only seems to go down 1 level into the site. So it is necessary to repeat steps 2 - 3 a few times.

3. Remove Duplicate URLs
3a. Remove URL Containing the Word, author
3b. Remove URL Containing the Word, category
3c. Remove URL Containing the Word, tag

Step 3 is not really necessary but my end goal is to build up a list where the url contains "domain.name/article.name" because that is where you will find the author's anchor text. For efficiency I just removed what I have deemed "non essential" URLs.

4. Repeat Step 2 - 3 until you have a large list of "domain.name/article.name" URLs.
5. Change the Link Extractor to now be External.
6. Import the last list into the Link Extractor that you were able to generate from Steps 2 - 3.
7. Start

You're going to end up with all of the External links from these pages which will contain the author's anchor text URL (the goal).

8. Load the Scrapebox Backlink Checker.
8a. Take the last list you had and import it into the Backlink Checker.
8b. Start
8c. Download Backlinks

The result after step 8 will be a list of article directories (and every other backlink method) from these domains.

9. Trim this list to the root domain.
10. Import into AMR

My result from this test was a list of about 125,000 to import. I ended up with 20% that were accepted by AMR. After signing up and removing the errors I ended up with a list of about 4,000 with a submitted result of 15% and an auto approve result of around 9%.
 
This sounds like a classic case of "no pain no gain"! How long did this take to accomplish?
 
Working out the method took about 20 minutes on a treadmill.

The actual implementation and getting a list together to import is just a few hours. For me it cut down the time about 5 days. It took me that long to scrape footprints.
 
Hi,

it works like a charm and it brought many new sources.

Thousand thanks for the hint.


Randolph
 
props for thinking out of the box but you're making it way too complicated. you either don't have enough proxies, you don't know how to scrape or both. it is much more reliable to use footprints to scrape but there's a catch. most people don't know how to properly scrape. they use too few seed keywords. I use about 60,000 keywords to scrape. every 10,000 keywords might give me just several hundred URLs more but in the end I have a list of about XXX,XXX distinct domains scraped.
 
My result from this test was a list of about 125,000 to import. I ended up with 20% that were accepted by AMR. After signing up and removing the errors I ended up with a list of about 4,000 with a submitted result of 15% and an auto approve result of around 9%.

Thanks for sharing your think outside the box method, appreciated.
However, the results dont seem to be appealing at all; In the end, we are seeing only 54 auto-approved directories?

15% of 4000 = 600
9% of 600 = 54...
 
Very good method, if we keep this method to scrape AA sites everyday, we'll get enough AA sites.
 
I have never scraped. I downloaded Sick Scraper and will probably buy Scrape Box. What is the AA list for AMR? I have AMR need to get better directories since the included directors are over used.

Thanks!
 
props for thinking out of the box but you're making it way too complicated. you either don't have enough proxies, you don't know how to scrape or both. it is much more reliable to use footprints to scrape but there's a catch. most people don't know how to properly scrape. they use too few seed keywords. I use about 60,000 keywords to scrape. every 10,000 keywords might give me just several hundred URLs more but in the end I have a list of about XXX,XXX distinct domains scraped.

You're right and that was the way I was doing it. I was taking every footprint I could find and then crossing it with every word in the dictionary and it was taking me weeks to create lists. Public proxies only though because I only have 10 private proxies I use for bookmarking and so forth.
 
8. Load the Scrapebox Backlink Checker.
8a. Take the last list you had and import it into the Backlink Checker.
8b. Start
8c. Download Backlinks

The result after step 8 will be a list of article directories (and every other backlink method) from these domains.


What is the last list? The links that were external? What websites do you check these against?
 
My result from this test was a list of about 125,000 to import. I ended up with 20% that were accepted by AMR.

Those 125K urls obviously include various and sundry blog platforms as well (wordpress, vbulletin, etc) - why didn't you mention posting to those (non-AMR) sites using scrapebox?
 
I never had good results with scraping with keywords or footprints in Scrapebox so I'm going to try your method. Thanks for sharing.
 
Superb method mate. This can be repeated for other types of backlinks as well. Hint: pligg sites :D
 
I have been using a similar method since it's getting tougher to scrape long lists using scrapebox and it comes down to the quality of those URLs, right?

thanks for confirming I'm on the right path, cheers
 
Hmm...it sounds good!i will do a scrapebox scraping using your methods & let you guys know if it really works or not.
 
Now this is gold! If you have tool like ultimate demon you can add that final list and get pretty good new sites in your tool ;)
 
Well, I've been following the instructions of your guide step by step. I got to the backlink checker part, but forgot about it not working anymore! Thought I would try again anyway but low and behold, 0 backlinks found for every site. :( Any easy step to get the backlinks short of searching each and every site one at a time?
 
Back
Top