while most of the training is relying on public web data, big tech companies are building huge private datasets manually created by experts to tune the models to answer requests and do agentic stuff, and that’s becoming very valuable for companies and AI training in general. Are leaks of those data or weights already a thing? that would allow to get great open source models, without all that export drama… And probabily there are people out there that would pay for that stuff…