category websites checker

mohamedabdo

Newbie
Joined
Jun 28, 2020
Messages
39
Reaction score
11
Hello all

anyone has idea's in order to get category from the website by Scarpebox
regular expression OR foot footprint

Categories List​

 
Do you have an example or footprint or regular expression to extract category for URL websites
 
Do you have an example or footprint or regular expression to extract category for URL websites
No, because its near impossible to build because there can be so many false positives. the list of potential categories is in the eye of the beholder as well, because each person is going to define different things.

Means there is no 1 concrete answer, so you would have to build your own and it till take 1 footprint at least for each category you want.
 
basically your asking for something that is a bit abstract. So it will take some trial and error on your part.
 
i have url list and i would like to find industry name as below list
Information Technology and Services
Hospital & Health Care
Construction
Retail
Education Management
Financial Services
Accounting
Computer Software
Higher Education
Automotive
Government Administration
Marketing and Advertising
Banking
Health, Wellness and Fitness
Real Estate
Food & Beverages
Telecommunications
Oil & Energy
Hospitality
Mechanical or Industrial Engineering
Primary/Secondary Education
Internet
Electrical/Electronic Manufacturing
Insurance
Medical Practice
Human Resources
Consumer Services
Transportation/Trucking/Railroad
Pharmaceuticals
Restaurants
Management Consulting
Civil Engineering
Research
Design
Logistics and Supply Chain
Law Practice
Architecture & Planning
Apparel & Fashion
Consumer Goods
Facilities Services
Food Production
Non-Profit Organization Management
Machinery
Entertainment
Chemicals
Wholesale
Arts and Crafts
Utilities
Farming
Legal Services
Mining & Metals
Airlines/Aviation
Leisure, Travel & Tourism
Sports
Building Materials
Environmental Services
Professional Training & Coaching
Medical Devices
Music
Individual & Family Services
Cosmetics
Mental Health Care
Industrial Automation
Security and Investigations
Staffing and Recruiting
Aviation & Aerospace
Graphic Design
Biotechnology
Textiles
Import and Export
Consumer Electronics
Public Relations and Communications
Broadcast Media
Business Supplies and Equipment
Writing and Editing
Military
Media Production
Computer Networking
International Trade and Development
Renewables & Environment
Events Services
Civic & Social Organization
Photography
Computer Hardware
Defense & Space
Furniture
Computer & Network Security
Printing
Fine Art
Investment Management
E-Learning
Outsourcing/Offshoring
Warehousing
Law Enforcement
Publishing
Religious Institutions
Maritime
Information Services
Supermarkets
Executive Office
Animation
Government Relations
Semiconductors
Program Development
Plastics
Online Media
Public Safety
Packaging and Containers
Judiciary
Alternative Medicine
Performing Arts
Commercial Real Estate
Motion Pictures and Film
Veterinary
Computer Games
Luxury Goods & Jewelry
International Affairs
Investment Banking
Market Research
Wine and Spirits
Package/Freight Delivery
Newspapers
Translation and Localization
Recreational Facilities and Services
Sporting Goods
Public Policy
Capital Markets
Paper & Forest Products
Libraries
Wireless
Venture Capital & Private Equity
Gambling & Casinos
Ranching
Glass, Ceramics & Concrete
Philanthropy
Dairy
Museums and Institutions
Shipbuilding
Think Tanks
Political Organization
Fishery
Fund-Raising
Tobacco
Railroad Manufacture
Alternative Dispute Resolution
Nanotechnology
Legislative Office
Mobile Games
 
No, because its near impossible to build because there can be so many false positives. the list of potential categories is in the eye of the beholder as well, because each person is going to define different things.

Means there is no 1 concrete answer, so you would have to build your own and it till take 1 footprint at least for each category you want.
can give me 1 sample of 1 footprint at least for each category and I will create it
 
Parse all of your content urls for the cats.

https:/
https:/ null/
https:/ null/ www.
https:/ null/ www. string /
https:/ null/ www. string / string /
https:/ null/ www. string / string / string
https:/ null/ www. string / string / string.filetype

(the actual category is often one of those strings with the others holding a close relationship to the content)...

be as granular as you want, I just split by folder delimiter /, but you can go down to the word or even the char if you wish, however your datasets in the following will get huge quick...

Push all of the post parsing url/string info (string:from:url) into a csv (structured text) (one triple per line) and then imp[ort the csv files onto a string database... like NEO4j ( graph bd )

This will give you both the possible semantic strings used for cats, but also the string relationship to its parent urls out of the box...

A high density of nodes with the same string in thier URLs is usually a good category candidate or a share URL footprint for the future . ;) :)

Re-export this info into a structured form again (preferably postgres db for next step)

Push all content text at URLs into a doc db. (like Solr)

Form Queries on Docs (ElasticSearch).

Bring all of the data together visually in OmniDB for easier db project management and planning.

Run the companion ML toolbox https://www.2ndquadrant.com/en/resources/2uda/ to Omnidb against the dataset.

Attaching the blob texts as a node like you would an image for classification and using the url split strings (and optionally manually supervise against the vocab list you have above) as your "correct" classification and with a little luck you will have the basis for a content classifier.

I am sure there are missing pieces and the above gist may look more like I am pissing in the snow than writing an etl / discovery recipe... but it might at least get the ideas flowing.
 
Parse all of your content urls for the cats.

https:/
https:/ null/
https:/ null/ www.
https:/ null/ www. string /
https:/ null/ www. string / string /
https:/ null/ www. string / string / string
https:/ null/ www. string / string / string.filetype

(the actual category is often one of those strings with the others holding a close relationship to the content)...

be as granular as you want, I just split by folder delimiter /, but you can go down to the word or even the char if you wish, however your datasets in the following will get huge quick...

Push all of the post parsing url/string info (string:from:url) into a csv (structured text) (one triple per line) and then imp[ort the csv files onto a string database... like NEO4j ( graph bd )

This will give you both the possible semantic strings used for cats, but also the string relationship to its parent urls out of the box...

A high density of nodes with the same string in thier URLs is usually a good category candidate or a share URL footprint for the future . ;) :)

Re-export this info into a structured form again (preferably postgres db for next step)

Push all content text at URLs into a doc db. (like Solr)

Form Queries on Docs (ElasticSearch).

Bring all of the data together visually in OmniDB for easier db project management and planning.

Run the companion ML toolbox https://www.2ndquadrant.com/en/resources/2uda/ to Omnidb against the dataset.

Attaching the blob texts as a node like you would an image for classification and using the url split strings (and optionally manually supervise against the vocab list you have above) as your "correct" classification and with a little luck you will have the basis for a content classifier.

I am sure there are missing pieces and the above gist may look more like I am pissing in the snow than writing an etl / discovery recipe... but it might at least get the ideas flowing.
really thanks for your support but i an not understand well what you want to tell me
 
can give me 1 sample of 1 footprint at least for each category and I will create it
Sorry I cant spend hours and hours to build a footprint for you for each category. You may be better off to hire someone to do it.

to be clear I do Not do work for hire, so I can not help you with this.
 
Sorry I cant spend hours and hours to build a footprint for you for each category. You may be better off to hire someone to do it.

to be clear I do Not do work for hire, so I can not help you with this.
give me sample for 1 category footprint only pls
 
give me sample for 1 category footprint only pls
Its a near impossible thing to do by scanning the content of the page alone, thats why I don't do it. The point being the false positives will be thru the roof high.

The simplest way though if you want to start is use the word as the filter So for

Libraries

use

Libraries

its that simple. Scrapebox does not offer any heuristics. If you want higher accuracy you need machine learning/AI/heuristics to look at the page and categorize it. Or you need human eyeballs.

It just depends on how much accuracy you need really.
 
Back
Top