Simulated AI-generated-text detection AUC across paraphrase-attack strength for 2400 configurations spanning four detection methods (GLTR, hard watermark, semantic-robust watermark, DetectGPT), five application contexts, source-model quality, and text length, grounded in the watermarking and detection-evasion literature. Forward-looking task predicts detection defeat under heavy paraphrase attack from cheap no-attack baseline performance.
Simulates detection-AUC degradation under paraphrase attack. 2400 configs across five contexts (academic integrity, news provenance, misinformation screening, job-application screening, content moderation), four methods, source-model quality, text length. Ten sources (2019–2024), including the foundational watermarking paper, DetectGPT, DIPPER's paraphrase-attack paper, and Sadasivan et al's fundamental-limits argument.
Lets A Platform Choose A Detection Method Estimate Attack Robustness From The Cheap No-attack Baseline Every Method Reports, Before Running A Real Adversarial Red-team Evaluation.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.