LLM training data is actually something I've been thinking about a lot lately. The idea is solid - if your content gets picked up into AI training sets, you essentially get cited by ChatGPT, Perplexity, Claude without any traditional SEO work.
The problem is nobody really knows what gets scraped and when. Common Crawl is the main dataset most LLMs train on, and they crawl everything but there's no timeline or guarantee your stuff makes the next training run.
What actually increases your chances - get your content on platforms that are definitely in training data already. GitHub, Reddit, Hacker News, Wikipedia citations, Stack Overflow. A page that gets referenced from those has way higher odds of ending up in the next dataset than a random blog post.
Also structured data helps. LLMs love clean, factual, well-structured content. If your page looks like something an AI would want to quote, it probably will get quoted.
3-4 months is optimistic honestly, training runs don't happen that frequently. But it's worth doing alongside normal SEO, not instead of it.