Point it at a live landing page. A deterministic rubric — readability, jargon density, reader focus, CTA strength, specificity, passive voice, seven measures with published weights — scores the copy with no model involved, and that scorecard becomes the brief the model rewrites against. Every variant is re-scored by the same code, so the improvement is measured rather than claimed, and what a rewrite traded away is reported alongside what it gained. Promote a variant to the results ledger, enter what the traffic did, and a two-proportion z-test calls the winner — which lets the ledger do the thing tools like this never do and grade the rubric: how often did the higher-scoring arm actually win? A competitor diff scores several pages against the same measures with no model at all, and a stability check runs one prompt repeatedly to show how much the model disagrees with itself. Ad variants, SEO briefs, and brand-voice enforcement round out the lab.
Anyone can generate landing copy now, so generating it is not the product. The rubric is: seven measures with published weights, run identically on the original and on the model's rewrite, in code the model never sees. It reports what each variant traded away as readily as what it gained, prices the test, and then — the part that costs something to admit — records what really happened and scores the rubric against it. If the higher-scoring arm keeps losing, that is the rubric's problem and this says so.