Scrapejet Videos - How Tos and more

thanks loopline you're the man gonna check them out asap
 
How do you test this?
I load a profile with anchor1, edit profile, change to anchor2, click "update profile",assign a test list,run,check manually anchor text by loading 3-4 webpages in firefox.
After I check I see anchor1, not anchor2.

Note:after I "update profile" if I edit it again it shows anchor2.
If I remove profile, load it again and run, the anchor will be anchor2.

I sent you a PM, pls. check.
 
I started to play with ScrapeJet and I feel like a pimp :)
I'm sending ScrapeJet babe out there and she brings home backlinks. How cool is this? It's even more amazing because this test is using public proxies.

f3xfb


I want to start using bad words list. How is this working?
It search bad words in urls or in page content?
 
I started to play with ScrapeJet and I feel like a pimp :)
I'm sending ScrapeJet babe out there and she brings home backlinks. How cool is this? It's even more amazing because this test is using public proxies.

f3xfb


I want to start using bad words list. How is this working?
It search bad words in urls or in page content?

The bad word list is applied to the content, not to the url.
 
Thank you for answering my question loopline.

I have one more when assigning a list of urls to post to does it post using all of the def files or only the original ones that came with sj?
 
Thank you for answering my question loopline.

I have one more when assigning a list of urls to post to does it post using all of the def files or only the original ones that came with sj?

It will use the def files you have enabled on the settings screen.
It will try to match a def file to the blog type.
 
I am thinking to use a keywords file with different footprints that contain quotation marks and ":" like:

Code:
"keyword1"
"keyword1" "leave a reply"
"keyword2" "leave reply"
link:website.com

Is this possbile?
Do I have to use http://www.scrapejet.com/encode.html to encode this keyword file?
Thanks!
 
I started to play with ScrapeJet and I feel like a pimp :)
I'm sending ScrapeJet babe out there and she brings home backlinks. How cool is this? It's even more amazing because this test is using public proxies.

f3xfb


I want to start using bad words list. How is this working?
It search bad words in urls or in page content?

Haha, Beautiful.


I am thinking to use a keywords file with different footprints that contain quotation marks and ":" like:

Code:
"keyword1"
"keyword1" "leave a reply"
"keyword2" "leave reply"
link:website.com
Is this possbile?
Do I have to use http://www.scrapejet.com/encode.html to encode this keyword file?
Thanks!

You only need to encode footprints, not keywords, SJ does that automatically as needed.

You could use that as keywords, but you have to bear in mind that whatever you put in the keyword list is combined with whatever is in the footprint field of the def file that it is using to harvest with at any given time.

So if you put in

link:website.com

and then your footprint file is like

"powered by wordpress" +"leave a comment" -"comments closed"

SJ is going to be searching for

"powered by wordpress" +"leave a comment" -"comments closed" link:website.com

Which might not give you the desired result.

Your thinking outside the box though, so thats good. Have you wanted my video no expanding your list? It deals with the concept of what you are wanting to do, but just with site:website.com but same concept basically.

http://www.youtube.com/watch?v=OEbL5sZaa8I
 
Last edited:
Have you wanted my video no expanding your list? It deals with the concept of what you are wanting to do, but just with site:website.com but same concept basically.

Using footprints I'm trying to find new domains, because I'm finding again and again same AA domains.

Expanding AA list is good if used wise. I wasn't and I lost 120k out of my 150k aa list.
What happened. Last 2 months I builded the 150k list: posting/link checker using scrapebox.
When I tried to use the list the approval rate was down... so I tried to be a little Sherlock Holmes. I tried manually some links and many have installed captcha meanwhile and I saw messages like "your request was queued and webpage will show soon...".

Now I'm thinking web hosting companies are trying to limit this, because posting very quick on many pages can crush their servers. Most blogs use shared hosting. There are hundreds/thousands shared accounts on a server.

Back to the expanding. I'm posting now on a row only 1 link per domain.
What do you think?
 
I only use private proxies and think its pointless not to. Why would I try to use public proxies, if you set your connections right you can scrape from the engines forever with private proxies. I have scraped over 50 Million urls in a run I have going on right now from google, with only 38 private proxies.

proxiesscraping.jpg



Not using private proxies for scraping is a common misconception, you just have to do it right. Keep your connections at 20% or less of your proxies and you can pretty much scrape forever. Its faster then dealing with checking public proxies and then dealing with all the failure rates and low speed.

Even if you get your private proxies blocked, in a matter of hours they will be unblocked again. You can't really "burn" them persay, when you go posting google never sees the IP that you post from only the blog admin sees it. So there is no connection.

I would go as far as to say that if you have private proxies, your shorting yourself by not using them to scrape, you already pay for them, get their use out of them. Its been many months since I have used public proxies for anything. I fail to see the point in not using your private proxies.

Just my 2 cents.


Also there is no token use, such as %Blogtitle% in scrapejet. Sweetfunny emailed me back.

Dude this is an awesome thread. Ive been looking for this kinda thing since i bought scrapejet. Out of curiosity, do you pay for your own private proxies and if so, where from?
 
Don't think too small mate, the sky is the limit. Back when Yahoo had their API, I scraped 180 million urls just from yahoo with just 1 instance of scrapebox in about a week. Here is the latest on that scrape, as you can see its just climbing, with now over 100 million results, all with 38 private proxies. :)

how do you save all those scraped links if SB limit is 1 million?
 
i don't know about scrapejet features , but it's work with non english sites ?
because scrapebox not works .

Of course it works. I use scrapebox in Spanish all the time.

Scrapejet is more restricted about that because it does not allow you to add other countries search engines. For instance you can't search from google.pt while in scrapebox you can.

I hope they work on that soon.
 
Of course it works. I use scrapebox in Spanish all the time.

Scrapejet is more restricted about that because it does not allow you to add other countries search engines. For instance you can't search from google.pt while in scrapebox you can.

I hope they work on that soon.

This feature was already requested and is on our to-do list. It will be implemented with the next update, stay tuned :)
 
softtouch is captcha sniper integration coming soon too?
 
This feature was already requested and is on our to-do list. It will be implemented with the next update, stay tuned :)

Sweet, thanks softtouch, we appreciate your hard work!
 
Captcha sniper will be supported too in future, its on the to-do list, but further down.

Thanks softtouch. Adding the captcha platforms that captcha-sniper supports to ScrapeJet would be AWESOME :)
 
Back
Top