Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
How to save article scraped with Article Scraper for GSA ? When im scraping article becomes like that:


"TITLE NAME "<FILED>"BODY ARTICLE"

I Dont want like that. There written TAG or Something there " <FILED>" . Tow To separate title and body each other? I Want to scrape article like that :


"TITLE NAME "

"BODY ARTICLE"
 
How to save article scraped with Article Scraper for GSA ? When im scraping article becomes like that:


"TITLE NAME "<FILED>"BODY ARTICLE"

I Dont want like that. There written TAG or Something there " <FILED>" . Tow To separate title and body each other? I Want to scrape article like that :


"TITLE NAME "

"BODY ARTICLE"

On the export dialog, there is a field "Separate fields with:", which has as default value <FIELD>. When you hover with the mouse over it, you see that you can also use <CR>, <LF> and <TAB> as field separator.
 
I would change up the queries a bit to be honest. With those you plan to use... you will get a lot of unrelated results.

do something like this:

Code:
san francisco attractions intitle:blog

I wouldn't use the "" either, as this is requesting results to be in that exact order.

To do this properly you shouldn't use public proxies, your results will be abysmal at best.

I would recommend either stormproxies for something like this. But for other projects I also use buyproxies.org .
Disloyal, thank you for your reply. Could you clarify on "I wouldn't use the "" either, as this is requesting results to be in that exact order." Don't understand that part.
 
Please tell me how to put "TITLE " and "BODY Article" separately . The example is shown below.


Scrapebox.png
 
Disloyal, thank you for your reply. Could you clarify on "I wouldn't use the "" either, as this is requesting results to be in that exact order." Don't understand that part.

By inserting the keyword in the middle of quotation marks like: "KEYWORD IS USED HERE"

This is only looking for an exact match.

You are requesting the search engines to only show you results that have the keyword in that exact order. So let's say you were also interested in seeing results that included a variant as well, you wouldn't be shown that. Because you told the search engine that the results MUST have the keyword in that EXACT order or else don't show it to me.

So you potentially are losing out on various targets who didn't use the keyword exactly how you wrote it. Test out your queries manually and you will see what I mean.

Whatever results you see when you search manually will be what Scrapebox will see(or pretty close to that).

There's obviously a time and place to use the exact match modifier, I just feel this wouldn't be one of those cases. At least in my opinion.

But just like @loopline said, your settings will have to be set up properly and have some good proxies since you are using advance operators.
 
Guys hi,

If i want to scrape google results showing to USA users should i use USA proxies?
 
Last edited:
Guys hi again,

is it possible to scrape websites / webpages that link to specific url with scrapebox?

For example let say i need to know what pages/websites link to this url:
Code:
hXXps://www.website2.com/en/22559/tours/San-Francisco/Must-See-sf

Is it possible with scrapebox?
 
Guys hi again,

is it possible to scrape websites / webpages that link to specific url with scrapebox?

For example let say i need to know what pages/websites link to this url:
Code:
hXXps://www.website2.com/en/22559/tours/San-Francisco/Must-See-sf

Is it possible with scrapebox?
I mean you can try and put

"hXXps://www.website2.com/en/22559/tours/San-Francisco/Must-See-sf" -site:website2.com

and scrape google, but thats about all I can think of. The back link checker uses moz and moz bases it on the domain not the exact url. So you can get the data from majestic for example, with paid account, but aside from paid accounts the engines don't give away near the info they used to.
 
Sorry for all the questions lately, the more i use scrapebox the more questions i have :)

1 When IP is banned from google for advanced operator use but ok for regular searches, do those bans ever revoked? or is it permanent ban?

2 Is it possibly to make setting to harvest only 1 url per domain per search (keyword)?

I know you can remove/filter them later, but it's not the same. I am trying to collect 20-30 results per KW. Often it works good, but sometimes there might be website that occupies most of the results and that way potential to gather other url is lost.

Example for some KWs it can be:
website1.com
website2.com/page1
website2.com/page2
website2.com/page3
website2.com/page4
.........................
website3.com

so this 2nd website is taking space from others.
 
Last edited:
Sorry for all the questions lately, the more i use scrapebox the more questions i have :)

1 When IP is banned from google for advanced operator use but ok for regular searches, do those bans ever revoked? or is it permanent ban?

2 Is it possibly to make setting to harvest only 1 url per domain per search (keyword)?

I know you can remove/filter them later, but it's not the same. I am trying to collect 20-30 results per KW. Often it works good, but sometimes there might be website that occupies most of the results and that way potential to gather other url is lost.

Example for some KWs it can be:
website1.com
website2.com/page1
website2.com/page2
website2.com/page3
website2.com/page4
.........................
website3.com

so this 2nd website is taking space from others.

1 - The bans are revoked, typically in 48 hours or less.

2 - no, scrapebox only gets only what google gives, it doesn't alter it along the way.


But if your using normal google engine, its pulling from the 100 results pages. So it doesn't matter if you are keeping 1 results or 100, scrapebox is still getting back the top 100 results from google. So you could do like 70 results instead of 20 to 30 and then remove duplicates when your done. That would leave you with what you want and still not add any additional requests to google (keeping your proxies still at minimum use)
 
Has anyone had any difficulty with the Scrapebox "Google Meta Scraper Addon" ? I cannot get it to load results. I've tried with Proxies enabled, with proxies turned off. It returns 0 results every time, no matter how many keywords I have.
 
Has anyone had any difficulty with the Scrapebox "Google Meta Scraper Addon" ? I cannot get it to load results. I've tried with Proxies enabled, with proxies turned off. It returns 0 results every time, no matter how many keywords I have.

Please check the Meta Scraper Update 1.0.0.7
 
I did a quick google search and the average web page is 3MB (https://speedcurve.com/blog/web-performance-page-bloat/), plenty are larger of course. So 1000 threads X 3MB (which is probably conservative) is 3000MB or 3GB of data at 1 time.

So try going into windows explorer and opening 1000 3MB text files at 1 time... Id bet your pc will choke and go unresponsive.

Scrapebox is massively efficient, but your asking it to deal with a huge amount of data at one time, process it, all sorts of file types, deal with memory, purge the memory and handle the load of downloading thousands of components and files of websites, all at the same time for a sustained amount of time.

I would call 1000 exceedingly high. The fact it can even run is an accomplishment to a grand degree in and of its self.

To be blunt, I think your expectations of what processing 1000 simultaneous websites should use, resource wise, is unrealistic. Loads of other programs are not close to this resource efficient.

My 2 cents.



:D

Is Scrapebox able to take advantage of dual xeon cpu setups if someone needed to use a few thousand threads at once?
 
Is Scrapebox able to take advantage of dual xeon cpu setups if someone needed to use a few thousand threads at once?
Yes. I mean scrapebox is built so it can use all avaialble cpu cores, however assigning the cores isn't something scrapebox can control, so Ive seen cases where windows will leave cores idle when they would have really benefited scrapebox.

Also windows may not like a "few" thousand threads. Typically it tends to start to choke at 1500 to 2000 simultaneous connections. Also I would advise spreading connections across multiple instance of scrapebox as windows likes many instances running less threads better then 1 instance running a ton of threads.
 
Status
Not open for further replies.
Back
Top