Scrapebox how to find urls for contact forms only

Pannath

Newbie
Joined
Jul 2, 2021
Messages
2
Reaction score
0
When I have scrapebox searching for URLs by keywords, I want to know if there is a way go get the urls for the contact forms. I know there is a way to configure scrapebox afterwards to search the existing urls for a contact form then go post on it. But what I want is a list of sites specifically that have contact forms, and the url for the contact form. Since some sites don't have contact forms.

Any one have any ideas on how to do this?
 
Maybe just search for the term "Contact us" in the url itself?
 
Maybe just search for the term "Contact us" in the url itself?
I've tried that, but sometimes the contact page doesn't have a contact form at all, it just has their contact details.
 
scrape keywords as much as possible
filter unique domains
then search for contact us inner pages
 
When I have scrapebox searching for URLs by keywords, I want to know if there is a way go get the urls for the contact forms. I know there is a way to configure scrapebox afterwards to search the existing urls for a contact form then go post on it. But what I want is a list of sites specifically that have contact forms, and the url for the contact form. Since some sites don't have contact forms.

Any one have any ideas on how to do this?
Your question is about footprints. So you need a better foot print. First you want to understand that past the first several results google returns irrelevant results all the time, because 99% of people stay on page 1. So your not going to get even close to 100% accuracy unless your scraping like the first 3 results for each query.

so that out of the way scrapebox already has some built in footprints for contact forms. So first you can try those. Second you can build your own footprint and I will put the video on that below. But realistically if your looking for mass results your going to have lower accuracy. The lower of the results you take, the higher the accuracy, its just how it works. So you will have to find that sweet spot that you are comfortable with.


 
Another way is to use Scrapebox to harvest domains for your niche.
Then run them through the link extractor plugin with some filter rules.
For example, filter any URLs not containing, /contact/ or /contact-us/ etc
This will get Scrapebox to go through your target websites looking for the contact pages.
 
I have always found their Youtube tutorial videos useful. Check their channel on youtube anytime you need to know something.
 
Get a keyword list.
Use footprints: inurl:contact inurl:message inurl:talk - or whatever words you think will be present in the contact us URL. Merge your KWs and footprints.
Run your scrape.
Remove duplicate urls. Perhaps trim the urls to first folder or trim in such a way to get rid of any garbage at the end of the url.
Remove urls containing: wiki, wordpress, blogspot, yelp, yellow, and any other of these big sites. If you don't remove them it will be a pain later on.
Use the scrapebox page scanner to scan the pages for notorious contact elements like send button, submit button etc.
Once you have scanned the pages, remove any duplicate urls again.
Then send your contact form poster out to send the messages.

The heavy lifting here is done by the page scanner. You're using the harvester to grad a bunch of urls, but you will need to filter them with the page scanner to get the best results. This is a multi-step process and it will take some time to get right. I know because I spent a lot of time several years ago doing this exact thing. I haven't tried to do it for a long time, but I listed my basic process above.

Good luck!
 
Get a keyword list.
Use footprints: inurl:contact inurl:message inurl:talk - or whatever words you think will be present in the contact us URL. Merge your KWs and footprints.
So if I understand you correctly, the foot print list will be these exact words and one in a line in the txt file?

I use footprints like these but I don't add inurl.
I end up with millions of results, but after cleaning dupes and removing wikis and all that plus keeping only URLs with the contact/message/etc (gotten from GSA Contact pages in different languages), I still get very low results remaining
 
Old can you explain the SB built in footprints for contact forms?

Thanks
what do you mean explain? There are footprints that scrapebox has developed some time ago, these footprints target the platforms that scrapebox supports. In the upper left hand quadrant of the main scrapebox window, there is a platforms button.

If you click the button and then tick off the ones you want, then click the platforms radio button right next to the platforms rectangle button.

Then when you harvest, just put your keywords in the keywords box and the footprints will be automatically appended to your keywords when scraping. Test it out with a single keyword and like 20 results so you can see how it works.
 
Ok, I will do as you said.
Thing is that I don't care the platform, I just want to grab the contract page URLs. When Indo, I transfer them to GSA contract to post. Whichever success I get is ok, but first I need only the exact pages, platform not important
 
Back
Top