I tested several articles, all handmade, from a lifetime ago before AI existed. They were all articles written by journalists for major organizations, consultants, political staff and some of my own personal stuff. Many were flagged.
The only thing I can guess at (until one of these platforms state how detection occurs, which probably won't happen) AI was trained on published / publicly available content (probably some restricted content as well that it probably can't say it was trained on) and if the writing style closely resembles that, then the content may be falsely flagged. If you guys ever read business articles from major news sources, white papers, academic papers.... hell, even MBB or Big Four stuff, they all pretty much have a very similar tone and style.
There are footprints that AI content does leave behind (some very obvious), so sometimes that's what's being caught. There's also certain writing styles / phrases that seems to be more prone to being flagged as well. I don't test often enough to make an accurate assessment, and honestly, don't care too much so take with a grain of salt. Unless there's a huge shift in search results, and while actual handwritten content gets falsely flagged, it's going to be hard to disseminate what's truly AI or not. More so, a lot large corporations have already started using AI as well - so where does the line get drawn? OK for the big guys to do, but not smaller sites?