Looks like you answered your own question there
Fair enough, so I used your service to run a few tests, till I ran out of free credits.
Test 1: I used two different GPT models (GPTChat 3.5-turbo & Cohere) to write a short article of around 140 words. GPTZero and Originality.ai both detected it was AI generated. I had your tool rewrite it and the rewrite passed GPTZero's check but Originality.ai flagged it as 71% probability it was AI generated.
Test 2 & 3: I generated a longer article with GPTChat 3.5-turbo and ran a small section of around 120 words and a longer section of around 420 words and had your tool rewrite both. Both passed GPTZero's check. The shorter section just barely passed Originality.ai's check (51% human/49% AI). The longer section also passed GPTZero's check, however, Originality.ai flagged it as a 97% chance it was AI generated.
This is obviously a limited set of tests but it failed 2 out of 3 and only passed the other by a hair (and it seems that's because the content was so short). And this is using a detection system made by a very small company, with essentially no resources compared to what Google can throw at detecting ML spun text.
I've gone down the road of mixing up different GPT text generator APIs and using various methods to mashup or tumble generated content through them in different ways to try and lose the ML fingerprint but it just doesn't work well. You'd need a fairly large model, well optimized and trained, with good human feedback, designed to do nothing but take the entire premise of a body of text and rewrite it as human-like as possible. Stacking different mediocre models not designed to do specifically this, with some model off Huggingface or Github that you tweaked, I just don't think it will cut it.
But stick with it and prove me wrong.