[Journey] Reverse Engineering Google with AI (Fine-tuning only. Advanced level)

following and supporting, i understand what you trying to achieve (maybe like 50%), but clueless with all what u did, but i know its something BIG, i read through all until page 5 though. keep going mate
 
Wow, what a journey!! I'm just following this now, but I will now try to replicate it for my knowledge as well!! Thank you for this!
 
Will be updating soon. Just working on some stuff for my main business then I can get back to the AI.
 
I've now got around 7000 examples from gpt4 of websites with their class, like "blog cat", "top 10 reviews", "best page", "news", "info", "tutorial" and so on.

I gave gpt4 EXTRA info to make sure it gave me an accurate reply, but it has a context window size of 8k.

For the actual fine tuning I only have 2k, but that's ok. The goal is to be able to do this with smaller outlines and not need a ton of info from the page to classify it. Ie, a practical goal for this is that it's cheaper. gpt4 with up to 8k used is expensive.

curie, with 1.5k used is much cheaper.

First though I'll try fine tuning babbage JUST TO SEE if we can get an accurate fine tuned model.

Well, I'll train babbage AND curie, then test them both and see how they perform.

I won't even bother with davinci for this. NO point, and it's more expensive than gpt4. It's complete overkill anyway. You do not need davinci for such a simple fine tuning task. Davinci is for way more complex stuff.
 
Wberre
I've now got around 7000 examples from gpt4 of websites with their class, like "blog cat", "top 10 reviews", "best page", "news", "info", "tutorial" and so on.

I gave gpt4 EXTRA info to make sure it gave me an accurate reply, but it has a context window size of 8k.

For the actual fine tuning I only have 2k, but that's ok. The goal is to be able to do this with smaller outlines and not need a ton of info from the page to classify it. Ie, a practical goal for this is that it's cheaper. gpt4 with up to 8k used is expensive.

curie, with 1.5k used is much cheaper.

First though I'll try fine tuning babbage JUST TO SEE if we can get an accurate fine tuned model.

Well, I'll train babbage AND curie, then test them both and see how they perform.

I won't even bother with davinci for this. NO point, and it's more expensive than gpt4. It's complete overkill anyway. You do not need davinci for such a simple fine tuning task. Davinci is for way more complex stuff.
Where did you get the examples? Manual or scraped?
 
Wberre

Where did you get the examples? Manual or scraped?

I built a program to generate advanced outlines/footprints for a page. I shared it above. The full code.

I then pass that outline/footprint to gpt4 and give it a list of classes, and ask it to choose one and give a confidence level. I save those with a confidence of 7+.

To get the actual sites to check I start off on chatgpt4 and ask it :-

Give me search keywords that are most likely to result in company service pages. Give me 25 keywords

Then I say

"Another 25" if I want more.

Or "Now give me ones that will result in blog category pages"

Then I can just say

"now company service pages", "now ecommerce collection pages" and so on. The search keywords it gives are excellent. I then just store them in search_keywords.txt and my program loops through them, searches, pulls the top 10 results, scrapes them and builds the outline/footprint, passes that to gpt4 via the api with the text

classification = get_classification_from_chatgpt(model, url, outline, p)

where outline is the outline generated from my function, and p is samples words from the first paragraphs. The prompt begins with

PROMPT = '''
I will give you the outline of a webpage.
I want them classified into one of the following categories:
info - informational style article
tutorial - guide/tutorial. A more in depth article that teaches something rather than just informs.
ecom cat - ecommerce category page
ecom product - single ecommerce product page
best X - a best X type page
reviews top 10 - X reviews, top 10 X. Different than best which is more best 2-4 products
single product review - a review of just 1 product
news - news article. Reporting on current events in the world.
faq - a faq. Generally in the sense of the old school faqs as opposed to a people also ask style PAA page.
forum - a forum post
service - a service being sold. It can be physical or digital, but it's a service, not a product.
recipe - a cooking recipe
blog cat - a blog category/silo/tag page
directory - a directory page/list of links
profile - a profile link, business or person
gallery - A page with only images, or 80-90% images. A gallery
contact - A contact page
legal - legal documents like privacy policy etc
paa - People also ask page
Also write out your confidence score out of 10 that scores how confident you are that the category is correct. If you are ABSOLUTELY CERTAIN, then score 10, if you are certain, and there's a tiny chance you might be wrong, score 9, if you are confident it's correct, but there's a slight chance you're wrong, score 7 or 8. If you are fairly confident, but there's a not insignificant chance you are wrong, then score it 5 to 6. If you are not quite sure and making a guess you feel is a good guess, score it 3 to 4. If you have no confidence in your guess and feel it's essentially like rolling a dice, then score it 1 to 2.
Don't explain, just give a category and confidence. Do not make up your own categories, only use the ones I have given you. Examples of results(ALWAYS give the result in this format):
info:8
blog cat:9
best X:10
news:8
Here's the outline:

'''


I'm not going to manually check these as I don't NEED 100% accuracy, and I wouldn't get that anyway, but the few I check here and there gpt4 is definitely getting it right.

Sometimes it'll label a normal article as a blog category page, but that's when the article is structured like a big list of links type post. And that's probably because that type of article is a "listacle", and I didn't include that class.

I'll actually add that now.

If anyone can think of any other page types I've missed, let me know before I fine tune the final model.
 
Finished preparing the data!
Damn that was no joke. This is one of the hardest fine tunes.


(! 2010)-> ls -hl page_classification_data.json
-rw-rw-r-- 1 tom tom 11M Jul 31 16:14 page_classification_data.json

11 MEG json :)

The real training file contains 6000 entries, but here's a sample so you can see the format

[{"prompt":"page:https:\/\/www.bonappetit.com\/gallery\/healthy-breakfast-recipes\nexternals:50\ninternals:166\na:220\nem:3\nform:2\nh2:61\nh3:0\nh4:0\nh5:0\nh6:0\niframe:1\nimg:69\ninput:2\nli:107\np:72\nsvg:24\nul:32\npicture:69\nfigure:62\nfigcaption:61\ng:7\nsample:To revisit this article, select My Account, then View saved stories To revisit this article, visit My Profile, then View saved stories By Anya Tchoupakov We take our sweet and savory healthy breakfast ideas very seriously. Why? Well, you know how, when you work out first thing in the morning, you tend to eat healthier throughout the day and drink more water and overall just feel better? We apply that idea to the first meal of the day, like setting an intention: Make it a healthy and delicious breakfast\u2014maybe a smoothie bowl or steel-cut oats or Mexican-style scrambled eggs with tortillas\u2014and it will be a healthy and delicious day. (And breakfast tends to be the easiest time to get in the good stuff too. Sugary snacks may show up in the office kitchen and decadent dinner invitations may come your way, but breakfast is fully under your control.) You with us? Let\u2019s go!\ntitle:61 Healthy Breakfast Ideas to Start the Day Right | Bon App\u00e9tit\nh1:61 Healthy Breakfast Ideas to Start the Day Right\nh2:Baked Eggs and Greens in Harissa Tomato Sauce\nh2:Whole Wheat\u2013Oat Waffles\nh2:Herby Dutch Baby With Smoked Salmon\nh2:Cashew Yogurt\nh2:Morning Glory Baked Oatmeal\nh2:Gluten-Free Apple and Oat Muffins\nh2:Salad for Breakfast\nh2:Dashi Oats With Crunchy Veg\nh2:Breakfast Blondies\nh2:Kimchi Toast\nh2:Blueberry, Lime, and Cashew Smoothies\nh2:Nut Butter Granola Bars\nh2:No-Nut Granola\nh2:Eggs in Purgatory\nh2:Almond Butter and Jam Quick Bread\nh2:Coconut Curry Greens with Runny Eggs\nh2:Acai Power House Bowl\nh2:Power Butter\nh2:Healthyish Breakfast Sandwiches\nh2:Smoked Fish and Rice Breakfast Bowl\nh2:Sweet and Salty Fig Toast\nh2:Hazelnut Granola and Chia Pudding Bowls\nh2:Quinoa-Banana Muffins\nh2:Spiced Chickpeas and Greens Frittata\nh2:Flatbread with Smoked Trout, Radishes, and Herbs\nh2:Chocolate-Cashew Butter\nh2:Roasted Strawberry and Tahini Buttermilk Shake\nh2:Roasted Red Pepper Frittata\nh2:Overnight Oats With Cashews, Seeds, and Turmeric\nh2:CBD Mango Smoothie\nh2:Simple Quiche With Sweet Potato Crust\nh2:Grain-Free Tahini Granola\nh2:Overnight Oats With Soft-Cooked Egg and Miso-Braised Kale\nh2:Coconut-Date Power Breakfast Bars\nh2:Gluten-Free Oat and Buckwheat Pancakes\nh2:Muesli Toast with Labneh, Hazelnuts, and Honey\nh2:Sumac Fried Eggs With Red Chile and Garlic\nh2:Breakfast Salad With Smoked Trout and Quinoa\nh2:Overnight Oats With Banana, Maple Syrup, and Tahini\nh2:Morning Glory Breakfast Cookies\nh2:5-Grain Porridge With Bee Pollen, Apples, and Coconut\nh2:Egg White Chalupa\nh2:Smoked Salmon Breakfast Salad With Crispbread\nh2:Chile-and-Olive-Oil-Fried Egg With Avocado and Sprouts\nh2:Blueberry-Chia Smoothie\nh2:Mango Toast With Hazelnut-Pepita Butter\nh2:Turmeric Fried Eggs With Kale, Yogurt, and Bacon\nh2:Breakfast Rice Bowls With Smoked Fish\nh2:Gluten-Free Chocolate Buckwheat Waffles\nh2:Hemp Milk Chai\nh2:Mixed Grain and Coconut Porridge\nh2:Fried Egg Tacos With Chile Jam\nh2:Fried Brown Rice With Kale and Turmeric\nh2:Tropical Carrot, Ginger, and Turmeric Smoothie\nh2:Oat and Apple Pancakes\nh2:Chia Pudding With Dried Apricots and Pineapple\nh2:Green Shakshuka\nh2:Tropical Energy Bars\nh2:Seeded Whole-Wheat Banana Bread\nh2:The Greenest Smoothie\nh2:Overnight Oats With Coconut, Dates, Almonds, and Honey\n\n\n\n###\n\n","completion":" gallery\n\n###"},{"prompt":"page:https:\/\/www.allrecipes.com\/gallery\/best-authentic-mexican-recipes\/\nexternals:52\ninternals:200\na:257\nem:1\nform:3\nh2:22\nh3:0\nh4:0\nh5:0\nh6:0\niframe:2\nimg:82\ninput:3\nli:214\np:24\nsvg:142\nul:34\npicture:0\nfigure:23\nfigcaption:23\ng:1\nsample:Carl Hanson is a Senior Editor at Allrecipes who has been writing about food and wine for nearly 20 years. He enjoys creating content that informs, entertains, and assists busy home cooks get nourishing meals on the table for their families. What makes a recipe an authentic Mexican recipe? Certainly, traditional ingredients and time-tested preparations make a recipe authentic. Of course, as time passes and cultures collide, cuisines evolve, transforming into fusion foods, modern takes on the traditional, which we also love (looking at you, Tex-Mex). But with this collection of Mexican recipes, we pay tribute to the originals, the tried and true, comforting, authentic Mexican food recipes that were built to last. These are just some of our favorites; for more, check out our collection of Authentic Mexican Recipes.\ntitle:Authentic Mexican Recipes\nh1:Our 21 Best Authentic Mexican Recipes\nh2:Homemade Mexican Chorizo\nh2:Authentic Enchiladas Verdes\nh2:Carne en su Jugo (Meat in its Juices)\nh2:Authentic Mole Sauce\nh2:Menudo Rojo (Red Menudo)\nh2:Birria de Res Tacos (Beef Birria Tacos)\nh2:Churros\nh2:Carne Asada al Cilantro\nh2:Mexican Mango and White Fish Ceviche\nh2:Sweet Orange Tamales\nh2:Mexican Enchiladas Suizas\nh2:Migas\nh2:Guacamole with Corn\nh2:Tamales Oaxaque\u00f1os (Oaxacan-Style Tamales)\nh2:Authentic Mexican Chili Rellenos\nh2:Cochinita Pibil (Mexican Pulled Pork in Annatto Sauce)\nh2:Chiles en Nogada (Mexican Stuffed Poblano Peppers in Walnut Sauce)\nh2:Authentic Mexican Hot Chocolate with Chile\nh2:Green Rice with Cheese\nh2:Pescado en Achiote (Mexican Fish in Annatto Sauce)\nh2:Authentic Tacos al Pastor\nh2:More Mexican Recipes and Inspiraton\n\n\n\n###\n\n","completion":" gallery\n\n###"},{"prompt":"page:https:\/\/www.epicurious.com\/recipes-menus\/19-ice-creams-of-your-dreams-gallery\nexternals:43\ninternals:133\na:178\nem:2\nform:0\nh2:47\nh3:0\nh4:0\nh5:0\nh6:0\niframe:1\nimg:57\ninput:0\nli:83\np:61\nsvg:19\nul:27\npicture:57\nfigure:48\nfigcaption:47\ng:7\nsample:To revisit this recipe, visit My Account, then View saved recipes. To revisit this recipe, visit My Account, then View saved recipes By Lisa Elbert and The Editors of Epicurious Start your summer strong by whipping up some of our best homemade ice cream recipes. No ice cream maker? No problem. Below you\u2019ll find no-churn ice cream recipes, ice pop recipes, and ice cream cake recipes you can make using your favorite store-bought pints. Avoiding dairy? We have plenty of vegan recipes too. Sorbet, gelato, sherbet, semifreddo, you name it\u2014it\u2019s all here. There are recipes for ice cream connoisseurs and novices alike. For the creatures of habit, there are recipes for chocolate ice cream and vanilla ice cream\u2014and flavors like prune-Armagnac and coffee-cardamom for the thrill seekers. Welcome to the vast and wonderful world of homemade ice cream. Aren\u2019t you glad you\u2019re here?\ntitle:47 Ice Cream Recipes To Make Right Now | Epicurious\nh1:47 Homemade Ice Cream Recipes to Make Right Now\nh2:Tin Roof Ice Cream\nh2:Pista Kesar Kulfi (Pistachio and Saffron Kulfi)\nh2:Blackberry and Chocolate Ice Cream Icebox Cake\nh2:Bourbon Butterscotch Ice Cream\nh2:No-Churn Almond and Raspberry Swirl Ice Cream\nh2:Nutterbuddy Ice Cream\nh2:Boozy Pi\u00f1a Colada Ice Cream\nh2:Prune-Armagnac Ice Cream\nh2:Vegan Banana Ice Cream\nh2:Double Ripple Ice Cream Cake\nh2:Buttermilk Ice Cream\nh2:No-Churn Fresh Mint and Chocolate Ice Cream\nh2:Raspberry and Pistachio Ice Cream Icebox Cake\nh2:Concord Grape Sorbet With Rosemary And Black Pepper\nh2:No-Churn Labneh and Lime Ice Cream With Granola\nh2:Strawberry Honey Balsamic With Black Pepper Ice Cream\nh2:Pomegranate-Yogurt Ice Pops\nh2:Peach and Butter Pecan Ice Cream Icebox Cake\nh2:Dark Chocolate and Cardamom Ice Cream\nh2:Peanut Butter, Banana, and Jelly \"Ice Cream\"\nh2:Sweet and Sour Strawberry Semifreddo With Black Sesame\nh2:True Vanilla Ice Cream\nh2:No-Churn Salted Caramel Ice Cream\nh2:Eggnog Ice Cream\nh2:White Peach and Bourbon Vanilla Ice Cream\nh2:Roasted Banana Vegan Ice Cream\nh2:Roasted Strawberry\u2013Buttermilk Sherbet\nh2:Lemon Ice\nh2:Earl Grey Tea Ice Cream\nh2:Apple Crumble Ice Cream With Calvados and Cr\u00e9me Fra\u00eeche\nh2:Yogurt-Peach Semifreddo\nh2:Date Ice Cream (Buza \u2018Ala-Tamr)\nh2:Mint Chip Ice Cream\nh2:Cheesecake Ice Cream With Strawberry Sauce\nh2:Peanut Butter Ice Cream With a Hard Chocolate Shell\nh2:Chocolate Malted Ice Cream\nh2:Coffee-Cardamom Ice Cream with Figs\nh2:Goat Cheese Ice Cream With Roasted Red Cherries\nh2:Malted Vanilla Ice Cream With Peanut Brittle and Milk Chocolate Pieces\nh2:Gianduia Gelato\nh2:Brown Sugar Ice Cream With a Ginger-Caramel Swirl\nh2:Grapefruit \"Creamsicle\"\nh2:Brownie Ice-Cream Sandwich\nh2:Matcha Coconut Ice Cream\nh2:Lime, Ginger, and Lemongrass Sorbet\nh2:Peanut Butter Ice Cream\nh2:Sweet Corn Ice Cream With Butterscotch\n\n\n\n###\n\n","completion":" gallery\n\n###"}]


Now to train on babbage, then curie.

After that I'll do a test on 30 fresh live pages with

1) babbage fine tuned
2) curie fine tuned
3) gpt4
4) human

And compare the results.

I'll also test 200-300 without human. I don't fancy classing 300 pages myself manually ;-)

I'll update with the results once it's trained. Let's see how this pans out.
 
This is probably going to take a while.

I forgot this is a multi classification fine tune.

This means I can split my training data into training and validation sets and get accuracy stats at the end.

This is better because it means I can just start with ada, then do multiple fine tunes, changing the number of epocsh, the batch size and the learning rate multiplier and seeing which is best.

I'll then test the top 2 in babbage and make sure those hyperparameters perform the same on babbage, then again with curie.

If anyone has any experience with hyper parameters and fine tuning on openai feel free to chip in. I doubt anyone will though. It's such new territory. You can't even find any info online for this stuff. People are fine tuning, but no one is sharing data.
 
This is probably going to take a while.

I forgot this is a multi classification fine tune.

This means I can split my training data into training and validation sets and get accuracy stats at the end.

This is better because it means I can just start with ada, then do multiple fine tunes, changing the number of epocsh, the batch size and the learning rate multiplier and seeing which is best.

I'll then test the top 2 in babbage and make sure those hyperparameters perform the same on babbage, then again with curie.

If anyone has any experience with hyper parameters and fine tuning on openai feel free to chip in. I doubt anyone will though. It's such new territory. You can't even find any info online for this stuff. People are fine tuning, but no one is sharing data.
Hahaha, don't laugh at me; I'm still confused about what we are talking about. Is it about NLP entities? Can you please explain a few words about what you are trying to focus on with GPT-4 and its benefits?

BTW, I really appreciate your hard work, explanations, and guidance to this community. I don't know why people like me appreciate you. It's very sad.
 
Hahaha, don't laugh at me; I'm still confused about what we are talking about. Is it about NLP entities? Can you please explain a few words about what you are trying to focus on with GPT-4 and its benefits?

BTW, I really appreciate your hard work, explanations, and guidance to this community. I don't know why people like me appreciate you. It's very sad.

It's not gpt4 I'm working with.

I'm fine tuning the base gpt3 models.

ada, babbage, curie and davinci.

Just now I'm fine tuning them to classify a web page. Ie, is it an info article, a tutorial, a blog category page, an ecommerce product, a top 10 review, a single product review, an ecom collections page.

This is fundamental to everything that comes next, since if you can't classify a webpage, you can't do much else, since google ranks specific types of pages for specific keywords.
 
Learning a lot from this journey, thanks for continuing to share updates :)

On a related note, I tried out the new LLAMA2 foundation model today, it's quite impressive, IMO it's definitely better than the 7b models from a few months back...more coherent, I suppose. I was pretty hesitant about using quantized versions but the q8 quantization (GGML) seems pretty indistinguishable from FP16 and even q4 is usable.
 
Ada fine tune is done

It's pretty much 100% accurate on everything except about 30-40% of the time it thinks ecom product pages are ecom category pages. That in its self isn't a huge problem. We'll see how babbage does, but we can always fine-tune with more product examples, or absolute worst case every time I get ecom cat page, I can just run it through gpt4 to confirm. Still a HUGE cost saving since gpt4 costs around 2 cents per page to check and ada costs 2 cents to check 25-30 pages. So ada is about 1300-1500 pages per $1 and gpt4 is 50 pages per $1. HUGE difference :-) But if we only have to check ecom cat pages, then that doesn't add much. We wouldn't have to check ecom product pages, since it's not mistaking category for product, just product for category.

(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2054)-> python check_class.py https://www.gq.com/story/how-to-buy-jeans-online
getting page: https://www.gq.com/story/how-to-buy-jeans-online
https://www.gq.com/story/how-to-buy-jeans-onlinePage class is tutorial - Token cost = $0.0007472000000000001
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2055)-> python check_class.py "https://www.meundies.com/collections/print-collections"
getting page: https://www.meundies.com/collections/print-collections
https://www.meundies.com/collections/print-collectionsPage class is ecom cat - Token cost = $0.0003456
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2056)-> python check_class.py "https://www.meundies.com/products/boxer-brief?pc=TEN"
getting page: https://www.meundies.com/products/boxer-brief?pc=TEN
https://www.meundies.com/products/boxer-brief?pc=TENPage class is ecom cat - Token cost = $0.0008096
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2057)-> python check_class.py "https://www.houzz.com/professionals...w-york-city-ny-us-probr0-bo~t_11817~r_5128581"
getting page: https://www.houzz.com/professionals...w-york-city-ny-us-probr0-bo~t_11817~r_5128581
https://www.houzz.com/professionals...w-york-city-ny-us-probr0-bo~t_11817~r_5128581Page class is directory - Token cost = $0.0006352
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2058)-> python check_class.py "https://riteplumbingnyc.com/"
getting page: https://riteplumbingnyc.com/
https://riteplumbingnyc.com/Page class is service - Token cost = $0.0008112
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2059)-> python check_class.py "https://www.theguardian.com/world/2...otesters-storm-eritrean-festival-in-stockholm"
getting page: https://www.theguardian.com/world/2...otesters-storm-eritrean-festival-in-stockholm
https://www.theguardian.com/world/2...otesters-storm-eritrean-festival-in-stockholmPage class is news - Token cost = $0.0007216000000000001
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2060)-> python check_class.py https://wolfofblogstreet.com/privacy-policy/
getting page: https://wolfofblogstreet.com/privacy-policy/
https://wolfofblogstreet.com/privacy-policy/Page class is legal - Token cost = $0.0007344000000000001
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2061)-> python check_class.py http://brainsreport.com/review/top-10-toasters/
getting page: http://brainsreport.com/review/top-10-toasters/
http://brainsreport.com/review/top-10-toasters/Page class is reviews top 10 - Token cost = $0.0009104
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2061)->
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2061)-> python check_class.py https://nymag.com/strategist/article/best-toasters.html

getting page: https://nymag.com/strategist/article/best-toasters.html
https://nymag.com/strategist/article/best-toasters.htmlPage class is best X - Token cost = $0.0013008000000000002
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2062)->
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2062)-> python check_class.py https://www.lawhelp.org/dc/resource/frequently-asked-questions-about-student-loan

getting page: https://www.lawhelp.org/dc/resource/frequently-asked-questions-about-student-loan
https://www.lawhelp.org/dc/resource/frequently-asked-questions-about-student-loanPage class is faq - Token cost = $0.0007136
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2063)->
(projects-env)
(tom@localhost)-(jobs:0)-(~/projects/models/page-classifier/app)
(! 2063)-> python check_class.py https://www.vangoghmuseum.nl/en/collection
getting page: https://www.vangoghmuseum.nl/en/collection
https://www.vangoghmuseum.nl/en/collectionPage class is gallery - Token cost = $0.0003264
(projects-env)

I'm fine tuning babbage now, then I'll fine tune curie. Curie is too expensive for this and I want to be able to do this with ada/babbage, but I want to see if curie picks up the ecom product pages without more training examples to compare them. After that I'll create another 300 examples for ecom product pages and fine tune the fine tuned ada/babbage with those then re-test.

But for everything else it's pretty much 100% which is insanely good.

It's detecting info pages, news articles and tutorials differently.


It also seems to be struggling with listacles and blog category pages.

On further review my data for blog category pages is BAD. I'm going to have to just manually edit it.

Listacles is also hard to find. I have only 19 data samples which is pointless.

I'm also looking at my data for ecommerce single product and it's equally bad.

The data for the other classes is good, so the takeaway here is even Ada is extremely powerful, IF, you train it on good data.

I'm going to manually clean up the single ecom product, blog category and listacle data and then fine tune again and re-test.

But we pretty much almost have gpt4 level quality at ada pricing. Now imagine what you can do with fine-tuning davinci, if ada fine tuned on a downstream task is as powerful as gpt4. ada is something like 1.3b params. It's tiny. :-)
 
I've also switched to using scrapefly for scraping, because a lot of the ecom sites aren't scrapable without residential proxies/headless browsing with js etc.

I'm also continuing to clean up the data for the single ecom product pages, then I'll add another 100 of them and manually clean those up too.

It's a lot of work for this, but it's a good solid base fine-tune to have since when doing anything in SEO you need to figure out the page type.

Ie, if you're writing content, when the AI is researching like a human, it needs to be able to determine page type to learn the correct outline for what it wants to rank for.

Once that's done I'll probably just directly do a babbage fine tune.

I think openai limits you to 10 fine tunes per month, so unfortunately I can't really do multiple fine tunes to test on this which is unfortunate. If I do I won't be able to fine tune anything else until next month!

One of the most exciting fine tunes coming up is my content creation fine tune.

I'm going to gather articles that rank as well as some other tricks and fine tune davinci to write articles that won't be detected as AI and will have built-in rankability. Ie, like combining surfer seo with an AI writer, except much more advanced than surfer since this is using real ranking articles.
 
Wow, I learned more in this 3 hours reading all the OP posts than studying 3 months the subject. What an incredible injection of first hand knowledge. Priceless.
Thanks a thousand times !
 
Wow, I learned more in this 3 hours reading all the OP posts than studying 3 months the subject. What an incredible injection of first hand knowledge. Priceless.
Thanks a thousand times !

I will be getting back to this soon. Been too busy with the last couple months of wife's pregnancy, but my 2nd daughter was born 10 days ago now so things are getting back to normal.
 
Back
Top