Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
He probably wants to harvest from custom sites. Scrapebox maybe could harvest from wordpress search results.

Though harvesting a small site with 3500 connections would be a D/DOS attack.

Can you clarify? I don't understand.
 
Can you clarify? I don't understand.

He probably wants to harvest from custom sites. Scrapebox maybe could harvest from wordpress search results.

Though harvesting a small site with 3500 connections would be a D/DOS attack.

S1.PNG
S2.PNG
Yes wordpress search results. Plenty of fish there. Usually wordpess blog has a search box and an xml? sitemap.
Combine these two means make scraping a breeze.
 
I won't help. All you'll do is crash webservers.

There's already a sitemap scraper on scrapebox. It's on site.com/sitemap.xml usually. Or harvest the outbound links from the sitemap if there are sub-sitemaps, just don't fucking query wordpress sites.

Especially not with 3500 threads.

View attachment 73768
View attachment 73769
Yes wordpress search results. Plenty of fish there. Usually wordpess blog has a search box and an xml? sitemap.
Combine these two means make scraping a breeze.
 
Remove subdomain from url: remove subdomain of blogspot not work.

~~ ~~ ~~ ~~ ~~ ~~
on Vanity Name Checker
L.PNG

V.PNG



If find a available but not keyword rich, is it worthy of be registered?
 
The Vanity Name Checker's potential platforms---are they always are some very very popular websites? For example always within Alexa Rank 100k even within 10k? I have method find/filter possible vanity platforms--if very low Alexa Rank can't have the chance to be a vanity platform--then they can be filtered rapidly.
 
Remove subdomain from url: remove subdomain of blogspot not work.

~~ ~~ ~~ ~~ ~~ ~~
on Vanity Name Checker
View attachment 73776

View attachment 73777



If find a available but not keyword rich, is it worthy of be registered?

With the new ability to register generic tlds it has gotten a bit confusing. So Scrapebox uses a database to remove subdomains from urls.
However sometimes you will have a list and most of it seems to work, but you will be left with stuff like
something.blogspot.com
and wonder why that subdomain wasn't removed. The answer is that is not a subdomain thats a domain because blogspot.com is an actual tld. So it would seem.com is the tld, but its not its blogspot.com So
car.something.blogspot.com
is a subdomain but
something.blogspot.com is a regular domain just like car.com is a regular domain. You can view the complete list here:
https://publicsuffix.org/list/effective_tld_names.dat




The Vanity Name Checker's potential platforms---are they always are some very very popular websites? For example always within Alexa Rank 100k even within 10k? I have method find/filter possible vanity platforms--if very low Alexa Rank can't have the chance to be a vanity platform--then they can be filtered rapidly.

You can add your own sites to the vanity checker, I have a video here:
https://www.youtube.com/watch?v=-F2nr_ltCRo
 
What would be a good paid private proxie service to use for scrapebox?
 
With the new ability to register generic tlds it has gotten a bit confusing. So Scrapebox uses a database to remove subdomains from urls.
/QUOTE]
Jesus for years and years I think something before blogspot is a subdomain.
 
R.pngNU.PNG
It sound a little boring but "Remove any non-url from harvester grid"
.co removed;
co. can't be removed.
I stuff too many strange stuff in the grid.:)

Problem solved. "Remove Entries Which Are Not Urls"
 
Last edited:
Just paid. But nothing delivered to my inbox yet. Shall await for arrival of this software and give me review.
 
What would be a good paid private proxie service to use for scrapebox?

These are the places scrapebox recommends
http://www.scrapebox.com/proxy

Just paid. But nothing delivered to my inbox yet. Shall await for arrival of this software and give me review.

Check your spam folder, it takes all of 5 seconds to get the auto reply mail. It goes to the email address you entered (if you paid with a credit card) or the paypal primary email address (which is also the email address you need to use when activating scrapebox).

Else its

http://www.scrapebox.com/payment-received
 
SS.PNG
Sitemap Scraper out of memory during scrape 167K domains, a 500 connections of Custom Data Grabber gabbing url title is running simultaneously. Sitemap result have been saved in session in real time.
4G RAM VPS.
 
@loopline you are the man as always. I am one of your loyal subscriber to your list for GSA SER. Thanks for your great service to the community.

These are the places scrapebox recommends
http://www.scrapebox.com/proxy



Check your spam folder, it takes all of 5 seconds to get the auto reply mail. It goes to the email address you entered (if you paid with a credit card) or the paypal primary email address (which is also the email address you need to use when activating scrapebox).

Else its

http://www.scrapebox.com/payment-received
 
Matt,

do you know if it's possible to post images with tumblr?

How would you do that?

Tumblr's default editor is the rich text editor not your standard html editor. Not sure if using html img src tags in your tumblr posts would work.

 
mat have u noticed drop in g-passed proxies last 2 days.?
ive gone from 5-600 on scanned
and 6-800 on scraped
both all way down to 100 or so
 
View attachment 73850
Sitemap Scraper out of memory during scrape 167K domains, a 500 connections of Custom Data Grabber gabbing url title is running simultaneously. Sitemap result have been saved in session in real time.
4G RAM VPS.

Well out of memory means you have to work with a smaller list or you can install more memory. Really thats all there is to it.

@loopline you are the man as always. I am one of your loyal subscriber to your list for GSA SER. Thanks for your great service to the community.

Thanks mate!

Matt,

do you know if it's possible to post images with tumblr?

How would you do that?

Tumblr's default editor is the rich text editor not your standard html editor. Not sure if using html img src tags in your tumblr posts would work.

yes I think so. Tumblr posting supports Markdown

https://www.tumblr.com/docs/en/posting

here is a cheat sheet which covers images
https://github.com/adam-p/markdown-here/wiki/Markdown-Cheatsheet#images

mat have u noticed drop in g-passed proxies last 2 days.?
ive gone from 5-600 on scanned
and 6-800 on scraped
both all way down to 100 or so

I haven't been scanning proxies the past few days, so don't know. I know they are ever getting tighter and banning proxies faster, but that probably won't make much difference on port scanned proxies, unless others are of course scanning the same ranges and getting blocked before you have a chance to test them.
 
is it possible sbox tester needs a tweek? due to poth scanned or scrapped
pass rate only dropping in last 2 days from +500 to 100
all testers need a tweak from time to time
 
The only tweak is users and scrapebox develop team add more and more source to it and remove items that not work anymore or bad. What more google passed proxies are not so important these day especially after V2 released.~~Premium Artice Scraper got updated but what is the details?
 
Status
Not open for further replies.
Back
Top