Wrong answer. Maybe six months ago this may of been the case.
Sure. OpenAI has first mover advantage. But both Anthropic's Claude is fantastic IMO. But who wants to be a sucker and let these fuckers suck up your key, business critical for data.
Because with a former NSA chief administrator now at OpenAI .. its not in our interest.
Meta's llama 3.1 is FANTASTIC. I have been using uncensored forks with extremely impressive results .
The key aspect is to place your data into vector database. I would start with small text arguments. It's key to note that "tokens" .. are akin to segments of words (like n-grams in linguistics.. which compromise parts of words) .. vRAM is CRITICAL for performance. And the baseline model is held in vRAM along then the vectored data AND then the results.
This is why responses are limited and the workflow concept of RAG is very popular.
If you are going to pour cash into other people's tools when you can buy a GPU (or rent bare metal by the hour.
In closing, anyone serious about using open models on small scale runs (20 to 30 4090 GPUs) .. hit me up.