How to find specific websites info using tools like scrapebox or GSA SE ranker?

dooku

Registered Member
Joined
Jan 30, 2014
Messages
54
Reaction score
12
My goal: I want to find a list of websites based on their niche and if they use technologies like Disqus, Jetpack, CommentLuv....etc..etc....
So I need to be able to search using probably keywords that are "intitle" and common pointers in htm that disclose the use of tools like Discus, Jetpack.....

If there are better tools or methods available then Scrapebox or GSA SEranker then please let me know.
Because I have none of these tools and need to buy one based on your input.

And last but not least, if someone can post instructions for any of those tools like which function to select and what to enter to accomplish the above that would be great!
 
You can use GSA SER to identify urls based on patterns in the page's html using the script language. I use it for this purpose myself. But it's not really it's intended use, has a high learning curve and might not really be worth it for you.

Scripting manual is here: https://docu.gsa-online.de/search_engine_ranker/script_manual
You will be wanting to look at the 'page must have' command along with the data extraction section.

But really a programming language coupled with knowledge of web requests and some regex/pattern matching will probably be a better way.
 
That is what I was afraid of :-( I can have a developer create a Python tool that will do this......as I usually do with tasks like this.
However I was hoping if there was a more simple tool that can do relatively small tasks like this with very little time investment.
 
That is what I was afraid of :-( I can have a developer create a Python tool that will do this......as I usually do with tasks like this.
However I was hoping if there was a more simple tool that can do relatively small tasks like this with very little time investment.
There is not a tool like you were wishing for.
 
ThreadMoved.jpg
Thread moved to the BH SEO Tools section. :)
 
@OP: The best bet would be to buy data packs from builtwith or you can find a BST which I remember was offering any data packs for $50.
 
@loopline would be able to suggest if Scrapebox has the potential to add that functionality.
 
@loopline would be able to suggest if Scrapebox has the potential to add that functionality.
Actually scrapebox has this feature to search through keywords and footprints (disqus, commentluv, etc..)

But what I suggested to OP is a cheaper method to directly go with data bundles instead of going through Proxy headaches, google search bans, filtering the list, etc..

In builtwith, you can get the data bundles which are using specific plugin, platform... so it's much cheaper and time saving solution.

I've seen a BST here a few days ago, which was selling the data packs for $50. If you would do the scraping by yourself, then only the search scraping proxies would cost much more.
 
@JasonS, from who do I buy these data packs?
 
The only issue with scrapebox I’m having when it comes to analyzing source code for footprints is it gets blocked by WAFs frequently since it doesn’t use a browser (could also be proxies). More than half of the websites I’ve tried to analyze for whatever footprint I’m looking for block scrapebox when I was using it this year to hunt down my competitors PBNs.

SER is in the same boat. It also doesn’t use a browser. Or even a headless browser.

The best tool that comes to mind for this which is noob-to-programming friendly is zennoposter. It uses a real browser with modern antidetect features. Zennoposter is what I wound up using in conjunction with scrapebox (for scraping google first). And I did find all their (indexed) PBN links.

But it will require at least a few hours of time and debugging. Granted that you already know how to use it. And an understanding of HTML (and css/xpath/regex is even better).
 
Last edited by a moderator:
Thank you all for the responses. Purchasing a datapack seems to be the easiest solution for this particular request, else I need to have my usual developer create a python scrape script.
 
I just came across the tool Paigham Bot .......and it seems almost exactly what I need, except for the part of checking the html code afterward (but I can use other solution for that).
Is anyone using Paigham Bot currently for web scraping and posting to contact forms and can tell me how well this works?

I can set up my own vps, domains, proxies, email addresses etc...etc... for the posting part using Paigham bot, but ANY good pointers and tips regarding using Paigham Bot are welcome!
 
Back
Top