Scrapebox: Is there a way to scrape websites for contact forms?

regensteiner

Regular Member
Joined
May 9, 2012
Messages
244
Reaction score
82
I know Scrapebox has the ability to find websites with contact forms, but it only seems to include results that have my keywords on the same page as the contact form.

I'm trying to find a way in Scrapebox to do the following:
  1. Find all websites with my keywords
  2. Instead of searching for emails, search for contact forms
This seems like something Scrapebox should be able to do but I haven't figured out a way to do so yet.
 
I know Scrapebox has the ability to find websites with contact forms, but it only seems to include results that have my keywords on the same page as the contact form.

I'm trying to find a way in Scrapebox to do the following:
  1. Find all websites with my keywords
  2. Instead of searching for emails, search for contact forms
This seems like something Scrapebox should be able to do but I haven't figured out a way to do so yet.
You can find the urls for your keywords then run the link extractor 1 time and extract internal urls.

Then filter out all the urls that don't have the word "contact" in the url. You can put together a list of words in the url. If you get a good list my data shows that using the word list in the urls gets about 85% of the potential urls that have contact forms.

The remaining 15% ish, give or take a little mostly have contact forms on the home page, as they are the single page type sites.
 
You can find the urls for your keywords then run the link extractor 1 time and extract internal urls.

Then filter out all the urls that don't have the word "contact" in the url. You can put together a list of words in the url. If you get a good list my data shows that using the word list in the urls gets about 85% of the potential urls that have contact forms.

The remaining 15% ish, give or take a little mostly have contact forms on the home page, as they are the single page type sites.
sorry for necroposting but, for every doubt i have on scrapebox, i find a useful post/video from you answering it. thanks for everything, you are the man.
 
Could use builtwith and look up any sites using contact form. Be your best bet imo. There's a little pricey though, but would get you the results you want
 
I am a complete newbie with SB and find it very confusing so my apologies if I'm asking something very dumb.

But in Google, I can do a search for something like:
san jose dentist "contact"

This will, for the most part, bring up contact pages of the various dentist sites. Can something like this be done in SB to have it scrape Google for the contact pages, then filter it even more to make sure it's really only the contact pages? Seems like it would be more efficient than scraping sites, then all their URLs, then filtering down. Or is that just the way it is by design?

Wish there was a newbie non-techie tutorial I could follow. All the videos on YT are too confusing.
 
I am a complete newbie with SB and find it very confusing so my apologies if I'm asking something very dumb.

But in Google, I can do a search for something like:
san jose dentist "contact"

This will, for the most part, bring up contact pages of the various dentist sites. Can something like this be done in SB to have it scrape Google for the contact pages, then filter it even more to make sure it's really only the contact pages? Seems like it would be more efficient than scraping sites, then all their URLs, then filtering down. Or is that just the way it is by design?

Wish there was a newbie non-techie tutorial I could follow. All the videos on YT are too confusing.
Scrapebox already has built in footprints to do this, and these footprints return the pages that are compatible with posting using scrapebox.

So all you have to do is enter in your keyword like

san jose dentist

and then click the platforms you want and scrape and it will return (for the most part) the exact contact form pages. so then just load and post.
 
Scrapebox already has built in footprints to do this, and these footprints return the pages that are compatible with posting using scrapebox.

So all you have to do is enter in your keyword like

san jose dentist

and then click the platforms you want and scrape and it will return (for the most part) the exact contact form pages. so then just load and post.
Thanks for the reply.
Great to know that what I'm looking for is already in Scrapebox. I'll have to do some work to unpack the steps though as I really have no idea where to click the platforms (and have vague ideas of what platforms are to begin with sadly), and what to do when I get results that are (for the most part) the contact pages.

But knowing the stuff is there means I just have to figure it out rather than trying to put a clumsy band-aid solution myself.
 
Thanks for the reply.
Great to know that what I'm looking for is already in Scrapebox. I'll have to do some work to unpack the steps though as I really have no idea where to click the platforms (and have vague ideas of what platforms are to begin with sadly), and what to do when I get results that are (for the most part) the contact pages.

But knowing the stuff is there means I just have to figure it out rather than trying to put a clumsy band-aid solution myself.
Platforms button is in the upper left hand quadrant of the main scrapebox window. Click the platforms radio button signifying that you want to use those footprints. Then click the platforms rectangle button and select the ones you want, there is a section for contact forms.

Then just add your keywords to the keywords box and scrape and your off to the races, as the saying goes.
 
Platforms button is in the upper left hand quadrant of the main scrapebox window. Click the platforms radio button signifying that you want to use those footprints. Then click the platforms rectangle button and select the ones you want, there is a section for contact forms.

Then just add your keywords to the keywords box and scrape and your off to the races, as the saying goes.
Think I got it.
Thanks again for the explanation. Appreciate it greatly.
 
Just gave this a try and getting some strange results.

If I run "web design CITY" I seem to get the same results as I do in the browser. I usually stop it after 99 to make sure I don't get my IP banned, but assume it will keep going and scrape more.

When I run "web design CITY" with the contact forms selected in the platform, the search stops on its own rather quickly and brings me back 20ish results. And most of them are not in the city I stated. Am I supposed to do something differently in the search string?
 
Just gave this a try and getting some strange results.

If I run "web design CITY" I seem to get the same results as I do in the browser. I usually stop it after 99 to make sure I don't get my IP banned, but assume it will keep going and scrape more.

When I run "web design CITY" with the contact forms selected in the platform, the search stops on its own rather quickly and brings me back 20ish results. And most of them are not in the city I stated. Am I supposed to do something differently in the search string?
Scrapebox is adding a long footprint to the query to return only results that are compatible with the scrapebox poster (at least its asking google for this, google does not always give accurate results) and then its only also returning the exact contact form url. So many businesses do not actually have a contact form so those are out, and some are not compatible so those are out, so that lowers the number

As for from other cities, google may be returning other results it thinks are relevant if there are not enough results that are actually relevant for your city query.

Export the footprints when its done, it pops a box where you can see the queries when its done, just export it and put some of the footprints in google and see what you get.
 
If I get a list of root domains, is there an newbie easy way to have SB scrape those sites only for contact pages?
 
Scrapebox is adding a long footprint to the query to return only results that are compatible with the scrapebox poster (at least its asking google for this, google does not always give accurate results) and then its only also returning the exact contact form url. So many businesses do not actually have a contact form so those are out, and some are not compatible so those are out, so that lowers the number

As for from other cities, google may be returning other results it thinks are relevant if there are not enough results that are actually relevant for your city query.

Export the footprints when its done, it pops a box where you can see the queries when its done, just export it and put some of the footprints in google and see what you get.
By the way, I get what you're saying. Had to re-read your answer multiple times and go back and forth with the Google and SB, but I think I finally am getting it through my head.
Lol.
Took long enough.

Thanks.
 
I tried running the Link Extractor. And the plan was to search for or filter URLs that didn't have the word "contact".
But as dumb as it sounds, I don't know what to do with the Excel file I saved after running the Link Extractor. It shows how many internal pages it found for each domain, but I don't know how to load that back into the SB to filter it as the Excel file doesn't seem to show the actual internal page URLs.

Feels like I'm missing something "simple".

Any help in either this or another simpler way maybe would be greatly appreciated.
Thanks.
 
I tried running the Link Extractor. And the plan was to search for or filter URLs that didn't have the word "contact".
But as dumb as it sounds, I don't know what to do with the Excel file I saved after running the Link Extractor. It shows how many internal pages it found for each domain, but I don't know how to load that back into the SB to filter it as the Excel file doesn't seem to show the actual internal page URLs.

Feels like I'm missing something "simple".

Any help in either this or another simpler way maybe would be greatly appreciated.
Thanks.
You don't want the excel file, thats just a count for informational/analytical purposes.

You want the actual output, which is auto saved in a folder. In the link extractor addon there is a Data Folder button, click that and it will open the data folder where the actual urls were auto output. Thats the file you want to reload into scrapebox.
 
Back
Top