Scrapebox - Email Scraper - Custom Crawler Markers

Mich

Newbie
Joined
Jun 11, 2016
Messages
27
Reaction score
5
The markers seem to sometimes work but sometimes don't - looking for guidance on this. For example, how far back do you go when creating a marker - some pages you can have a before marker that's hundreds of characters long..?

Here's an example of one I'm struggling with. I'm creating a definition to crawl a bunch of pages that are coded almost identically on the same site.

There is a list like this in the source code:

address:'123 High St',
city:'New York',
state:'NY',
zip:'90210',
name:'John Doe',
phone:'123-456-7890',
email:'[email protected]',

I've used the markers:
  • Before = email:'
  • After = ',
This will successfully pull the email.

Same with the name, if I simply replace the word email with name in the before marker.

But for some reason, the address doesn't work.

The list appears as above in the source code - so the address is formatted almost exactly the same as the name and email - would there be any reason as to why this marker wouldn't work but the others would?
 
Your using the custom data scraper in the email scraper plugin yes?

Its important to note that it goes in order down the page

so lets say your columns in the email scraper plugin are in this order left to right

Name Email Address

like that

Its going to go down the page html looking for the name marker, then when it finds it it will scrape it. Then it will proceed from the after marker for the name, down the html looking for email. Then when it finds email it will proceed from the after marker of the email down the page looking for the address. However in your example the address would appear in the page html ABOVE the email so scrapebox would not see it.

In other words it does not scrape the first column and then start back over at the top of the html looking for the next column (like email) it proceeds down from where it left off to find the first column.

So your columns must be in order that the html appears on the page. Then if you want them in a different order you can rearange them in excel after you export.

Does that sort it?

If not Id need to see a specific url example.
 
Your using the custom data scraper in the email scraper plugin yes?

Its important to note that it goes in order down the page

so lets say your columns in the email scraper plugin are in this order left to right

Name Email Address

like that

Its going to go down the page html looking for the name marker, then when it finds it it will scrape it. Then it will proceed from the after marker for the name, down the html looking for email. Then when it finds email it will proceed from the after marker of the email down the page looking for the address. However in your example the address would appear in the page html ABOVE the email so scrapebox would not see it.

In other words it does not scrape the first column and then start back over at the top of the html looking for the next column (like email) it proceeds down from where it left off to find the first column.

So your columns must be in order that the html appears on the page. Then if you want them in a different order you can rearange them in excel after you export.

Does that sort it?

If not Id need to see a specific url example.

Oh man I've spent so many hours on this today - such a simple solution! That has indeed fixed it. Thanks!
 
I am stuck on this one though. Let's say I'm trying to scrape this page using the custom scraper in the email scraper plugin: (It's a home listing on Redfin - forum won't let me post it)

I want the property address.

I've tried before = "street-address" title=" and after =">

Didn't work.

So I tried regex (?<="street-address" title=").*.*(?=">)

Also didn't work :(
 
One more post and I can post the link
 
Perfect - https://www.redfin.com/AZ/Buckeye/941-S-200th-Ln-85326/home/28319952
 

"PostalAddress","streetAddress":"|"

or

class="street-address" title="|</span>

or

<title>||

Any of those should work, although one may be better then the other depending on if you want state, zip etc..

That last one, I did not try, but they are using the pipe key to separate the title listing and scrapebox uses the pipe key for separate things so it may cause an error or it may be fine.

I see you tried the middle one but had no luck, its possible that site is changing the html and blocking you or something due to the user agent or the ip etc..

Worst case you could try the requst in the scrapebox socket tool addon and see if you get the html back or a block page or what.

Cloudflare for example blocks some user agents and not others and various other things.

Glad you got the other working. :)
 
Yeh it works on some pages but not all for some reason :(

Like only 20% of pages do these work on...
 
Yeh it works on some pages but not all for some reason :(

Like only 20% of pages do these work on...
so that means either

1 - those html markers do not exist on the other 80% and you need to add markers for those pages

2 - Somethign is altering the html before it gets to scrapebox or its not being able to fully load the page due to proxy failure, network congestion etc.

So try putting connections at 1 for a test and see if that helps and then go up from there.

Also try without proxies, if you using proxies.

Lastly of course whitelist in all security software to make something isn't altering the html.

Another possibility I guess is that the end site is redirecting you to a block page due to ip blocks and thus its failing.
 
so that means either

1 - those html markers do not exist on the other 80% and you need to add markers for those pages

2 - Somethign is altering the html before it gets to scrapebox or its not being able to fully load the page due to proxy failure, network congestion etc.

So try putting connections at 1 for a test and see if that helps and then go up from there.

Also try without proxies, if you using proxies.

Lastly of course whitelist in all security software to make something isn't altering the html.

Another possibility I guess is that the end site is redirecting you to a block page due to ip blocks and thus its failing.

It would be 100% fine if you had a video to link to on your channel if you want to offer that.

The staff knows you are here to contribute and not consume.
 
It would be 100% fine if you had a video to link to on your channel if you want to offer that.

The staff knows you are here to contribute and not consume.

Thanks!

Ill keep that in mind in general.

For this one this video is relevant, however I feel I need more info on the exact urls that are not working to really give an accurate response.

@Mich can you post like 1 or 2 or a handful of urls that are not working with the markers I suggested?

Also this video may help:


Hi loopline. Can i use scrapebox to scrape email from gg map?

Unfortunately no, because scrapebox uses raw sockets and threads, these do not support javascript by convention. A couple of years ago google removed the no javascript results for google maps. So if you visit google maps with javascript turned off in you browser it will just redirect you to google.com

Thus scrapebox can no longer scrape it. Same applies to bing maps.
 
Thanks!

Ill keep that in mind in general.

For this one this video is relevant, however I feel I need more info on the exact urls that are not working to really give an accurate response.

@Mich can you post like 1 or 2 or a handful of urls that are not working with the markers I suggested?

Absolutely - I'm still working on this and your help is much appreciated!

So I tried your suggestion (only using 1 thread) and that didn't work either.

Using
before = "street-address" title="
after = ">

I can get 3/10 of the following URLs:

https://www.redfin.com/NV/Las-Vegas/10171-Kermode-Ct-89178/home/29490514
https://www.redfin.com/CA/Richmond/36-Lakeshore-Ct-94804/home/12128781
https://www.redfin.com/CA/Citrus-Heights/8519-Pronghorn-Ct-95621/home/19245794
https://www.redfin.com/FL/Pembroke-Pines/9219-NW-16th-St-33024/home/144494362
https://www.redfin.com/NV/Las-Vegas/Undisclosed-address-89147/home/29564592
https://www.redfin.com/CA/Walnut-Creek/16-Christmas-Tree-Ct-94596/home/1471943
https://www.redfin.com/NJ/West-Long-Branch/50-Wall-St-07764/home/37630705
https://www.redfin.com/NV/Las-Vegas/6312-Stonegate-Way-89146/home/29519912
https://www.redfin.com/CA/Lafayette/867-Acalanes-Rd-94549/home/1845863
https://www.redfin.com/NJ/Manalapan-Township/4-Mercer-Ln-07726/home/37060562

The 3 that worked are:
https://www.redfin.com/CA/Richmond/36-Lakeshore-Ct-94804/home/12128781
https://www.redfin.com/FL/Pembroke-Pines/9219-NW-16th-St-33024/home/144494362
https://www.redfin.com/NV/Las-Vegas/Undisclosed-address-89147/home/29564592

(The last one being an undisclosed address)

If then add in:
before = listingAgentName\":\"
after = \"

I can also scrape the Agent name - but only from those 3 URLs as well.
 
Update: Right after I posted my comment above, I went back through to check any other of your suggestions that I may have missed.

One of them was to try this without proxies - so I did and all 10 scraped perfectly! What's going on with that??
 
Update: Right after I posted my comment above, I went back through to check any other of your suggestions that I may have missed.

One of them was to try this without proxies - so I did and all 10 scraped perfectly! What's going on with that??
What kind of proxies are you using?

My immediate guess is 1 of 2 things

1 - your using rotating proxies and your landing on redfin with ips from around the world and some of those countries/ips redfin is redirecting to what it thinks is relevant for those ips/countries and thus the formatting is different and/or its just redirecting to a page saying that country is blocked/not supported etc.

2 - some of the proxy ips are flat our banned by redfin.

You can try the urls with the proxies in the scrapebox socket tool addon to see the exact html that scrapebox is seeing

But at the end of it we know it works without proxies and not with proxies. So just at a basic level you could try running it slow without proxies or get some proxies from the country your ip is from or the country redfin is from or at least the country your trying to scrape addresses for as thats likely the target market.

Another program I use outside of scrapebox is http debugger pro. It has a full featured free trial and you can run it and then just run scrapebox and then go in and look and see what scrapebox sees.
 
What kind of proxies are you using?

My immediate guess is 1 of 2 things

1 - your using rotating proxies and your landing on redfin with ips from around the world and some of those countries/ips redfin is redirecting to what it thinks is relevant for those ips/countries and thus the formatting is different and/or its just redirecting to a page saying that country is blocked/not supported etc.

2 - some of the proxy ips are flat our banned by redfin.

You can try the urls with the proxies in the scrapebox socket tool addon to see the exact html that scrapebox is seeing

But at the end of it we know it works without proxies and not with proxies. So just at a basic level you could try running it slow without proxies or get some proxies from the country your ip is from or the country redfin is from or at least the country your trying to scrape addresses for as thats likely the target market.

Another program I use outside of scrapebox is http debugger pro. It has a full featured free trial and you can run it and then just run scrapebox and then go in and look and see what scrapebox sees.

Hmmm they were supposed to be U.S.A semi dedicated proxies from buyproxies.org - this is my go to proxy site, as I always find they are the highest quality. Disappointing that they might be the issue...

I'm looking to scrape around 1,000-2,000 URLs from this site a day and they'll start using captchas and other tools (I believe) at a certain point. Even visiting the site today I got a captcha, since I had been scraping these URLs from my own IP. So I'll need to find a proxy solution that works...
 
Then I would try the socket tool addon or http debugger. Im not certain that buyproxies.org is an issue. I have hundreds of proxies from them and they have been good for years as a company. But even great proxies can get blocked, I mean your own ip got blocked.

But here is the thing, did you tell buyproxies you want to scrape redfin? Because you may be linked up with some other person that is also scraping redfin or maybe redfin has a really long memory and the proxies you are using were used by someone else on refin a year ago and they are still blocked. Its all guesses without data, thats why the socket tool or debugger pro gives the data you need to be clear.
 
Back
Top