The problem is the detection tools are terrible. They just feed an AI the text and ask it to predict each next part based on the last repeatedly. The higher accuracy it can predict what the writer will write next the higher the 'score' is deemed. This obviously leads to a lot of false positives.