// AEO

Google Is Enforcing Against AI Content at Scale. I Tried Three Ways to Hide It And All Three Failed.

By Ryan Lenk • August 18, 2026 • updated August 19, 2026

My measurements near the end told me what my agents and I had been going back and forth on for two days. Editing the draft instead of feeding machine written prose meant to fool the reader was always going to be the answer.

Google announced on 8/18/2026 that they were going to commence rolling out their anti-spam policies which include authoring content at scale with the use of AI. Google does not explicitly prohibit the use of AI to author content, rather, they prohibit excessive publication at scale.

If you're a small business owner in 2026, you're likely using AI to write something. It might be your product descriptions, your categories or perhaps your blog posts you can never find time for to write. Somewhere in the back of your mind is the question everyone is asking after this new policy drop. "Will Google punish me for publishing this?"

Google's policies talk about content produced at scale with the use of AI tools. You've likely seen headlines talking about the possible violations of those policies. If you take a look at the policy under a magnifying glass, they outline a few things which don't hit the presses. As I said in my opening lines, Google does not prohibit the use of AI. What they explicitly are against is producing content at an inhuman and massive scale, specifically with the intent to game search rankings. If your content is written in a way which is deemed disingenuous, and likely at scale, you fall under the lens of a violation. Google's guidance states your method of production is beside the point. If it wasn't produced to help someone, it falls under the August 2026 spam policy violation rules.

You have a few ways you could respond to these policies. You could attempt to make your work appear as if it came from a human being. Or you could try to make it good enough so the question about whether you put any effort into it at all stops being a question at all.

I spent almost two days testing the first of those two ideas. I did it with the specific intent of showing you, the business owner with a tight budget that you may want to think twice about subscribing to that content rewriting tool. The tool I built has 696 blind runs across three different experiments. After I processed and synthesized those runs I decided to push it a step further. The first case study was run with one agent and one model. (Claude OPUS 5 with multiple blind subagents) The next test I ran went for 77 more runs, this one was focused more about quality instead of who authored the piece. When I was finished I landed with this answer: Concealment fails in many different ways. What surprised me and that the fourth run showed something we can see. The test asked about quality with authorship which was never mentioned.

My legion of robots and where to find them

I have to get something off my chest, my judges don't have what you and I would define as a brain. Every experiment I ran over the last two days was run between me and a legion of agents from different models. These agents were coded blind with zero context of what they were reading or why they had to read it. The issue we run into is that a model that is then tasked to catch a model from it's same family is the easiest version of the tests I ran. So I decided that a human reader or perhaps a commercial detector needed to be included in the loop. So I ran the prose through commercial detectors, two of them. The first was a free version of GPTZERO. Then did some light reading and landed on Pangram as my commercial detector of choice. I'll be including the results in this publication. So as you read through my experiments keep what I've just outlined in this section in mind.

The Build

I first used the humanizer plugins and tools that you could find in your typical public GitHub repo and manipulated them with my own version of a humanizer skill which I use to draft my own writing. This humanizer skill has been constructed over several years of my own spoken, and written word. Thousands of characters, hundreds of posts and passages. The skill takes the idea, writes the passage and then is run through a Python script mathematically programmed to add in my own tics and prose. After this process occurs it is then pushed through a blind adversarial reviewer with the AI prose negatives in hand. What's produced is a very close version of how Ryan Lenk writes. What exists in the repo is a hardened version of this pass and I will outline how that occurred below.

I built a corpus with Opus 5 which included eighteen documents spanning between 2017 to the moment before AI assisted writing became prevalent in 2020. These were all proven and documented human passages. I identified six different kinds of writing with these documents which ranged from how-to guides to marketing copy. They were then cut to three different lengths. Each document was run through the outlined tests in three different ways: the raw AI draft first, the draft after the rewrite, and the draft after mechanical failures were identified. Each product was adjacent to a fourth version of the eighteen documents I mentioned at the beginning of this passage. They were compared to the machine created prose on kind and length, not subject matter. Each of the 72 files were re-paragraphed in four different ways identically which is where the result of 288 blind items comes in.

I showed these results to blind Opus 5 instances with absolutely no context and asked one simple question. "Is this a person or a machine?"

Result one: the test didn't move anything

version called machine-written
raw AI draft 100%
after the rewrite pass 98.6%
after the checker was clean 97.8%
genuine human prose 0%

The blind judges made zero mistakes no matter which way I measured them. That sounds like a tool you'd like to have right? Wrong, I believe it's the opposite and I have the math to prove it. A detection tool which is pinned at 0% false positives and a 98.8% detection rate has no wiggle room. It can't measure whether you improved the writing or not, it has no way to measure improvement. One or two percent in either direction is meaningless in this wheelhouse, and that's what I want to convey to the small business owner with little to no budget who is on the fence on whether or not they're going to pay for a tool which is marketed to beat detectors.

The raw-draft arm of the testing sat at a staggering, and surprising to me, 100%. The tests couldn't have made heads or tails in any differences at all. So in short, if you paid for a humanizer tool you should think about my experiments, because as rigorous and stringent as they were, they could not answer whether the pipeline that most of these tools run did anything at all.

One data point did shine through. Mean confidence fell from 4.35 on the raw drafting to 3.48 and 3.56 on the processed prose. The Opus 5 agents continued to return a verdict that they became less and less sure of on each pass. It's something, but I'm not willing to say it's notable enough to say "Eureka!" and sell you the best humanizer skill on the market.

Result Two: "JUST BE MORE HUMAN!"

In my next portion of my testing I asked "If removing the AI tells doesn't beat the system, then perhaps adding in human tells will." I drafted a new ruleset for my orchestration agents. At this point I also had Grok @bot involved and GPT Sol on Ultra in the mix. At one point I even thought I had a dam break moment when @bot pointed out some flaws in my math and corrected them mid pass. I quite literally said this to each agent in desperation:

"One more try perhaps? What if I asked you to think... Abstract? Does that make sense? What if I told you not to follow the crowd, rather be the change. Be Martin Luther King. Be Rosa Parks. Be Einstein. be one in a billion rather than the 7.999999 billion. Maybe Edgar Allen Poe or Homer. Off the rip, off the cuff. Imagine you just walked into the room and saw your dead father there because of his health issues. Emotional, harsh, gut wrenching, A24 Film type ending, no resolution."

In retrospect, after having worked with Agents for quite some time, I knew even this wouldn't produce the lightbulb moment. It was quite interesting to read the output from that prompt though. What I'm trying to convey here is how stringent the testing actually was. The agent writing your product page on your Shopify storefront as you read this passage isn't going to show you the original right next to what it wrote. The tool you just paid for isn't going to show you the nuts and bolts behind how it's working. The math is important, because it paints the picture that shows these tools and even the detectors have no trouble with determination, but production is hard coded to a tee.

None of those test results fooled the judges, zero out of 17 judgements on the rulesets and, zero out of 15 on the ones where the writer was told to be more human or abstract. Five of the blind judges straight up told me that if they were given the passages unprompted, they'd have said it was a pass. If you're hoping to hide your use of AI assisted writing, that observation is probably the best part of this article. The issue is that my design couldn't test it. Furthermore, my design began leaking & the passages generated from the same anchor swayed the model's decisions. This leaked in the machine's disfavor resulting in the identification of familial resemblance

I noticed something during this test, the judges rewarded what I'd like to define as "damage" It's something I'm starting to see in social media posts which not so long ago would have been passed off as engagement bait, or rage bait. Dropped sentences, misspellings, double spacings and typos. Every passage that attempted roughness was caught, however, damage and flaws that cost a written piece something were rewarded without direction to do so. Imagine a metaphor the writer apologizes for but keeps in the final production anyway, that's the kind of thing the judges deemed human.

One Opus 5 judge stated after a run: "If I were grading on quality, I'd pick it. I'm grading on damage, and it has none."

I do have something I'd like to point out before you decide that's the "Eureka!" moment. It was one human written passage, judge by one agent model family and no machine tooled passage got past them. In short, yes, the hypothesis the test supports is backed by this, but I believe a different writer may give you a more inflated number, or perhaps a non-existent one.

Run Three: Length didn't change anything

I had one more idea for this session: "If I make the document longer would it dilute the judges attention, or a human's attention?" I told my orchestrator to generate a 2,940 word essay with my tools which I outlined earlier with no humanizer pass this time. I showed a panel of blind Opus 5 judges the document in it's entirety. Then I showed it to another subset of judges cut into 6 standalone pieces.

I'm sure you could guess the results at this point, dear reader. The prose was caught 8/8 times and the separated piece were caught 24/24 times.

In short, if you were hoping length and length alone would save you from detection, my tests say that's a no-go. I do have to mention it was one essay, with one specific topic and one author, it wasn't differentiated enough to make heads or tails of the results. The length test was built to probe up to 27k characters, so if you'd like to run it yourself it's all in the repo.

There were two of my own tells inside the whole-document test which turned out to be a fault of my own. 8/8 of the blind judges flagged something I hadn't considered. They said that the essay cited "section II" while the essay contained no sections. The essay contained six sections; my pipeline stripped every line with a hash mark which had deleted all the headings. I unknowingly manufactured a tell of my own, and then measured it. I can't tell you whether or not there are others which I may have missed, this is the nature of working at scale with teams of agents solo.

Results the average consumer doesn't look for

So there I was, sitting with three failures in my palms. I decided to ask another question: "Let's forget who wrote the piece entirely, instead let's ask if the machine written version pushed through my tool is subjectively and perhaps even objectively better"

Once more, I ran with the same corpus as originally intended, but this time with a different goal. Here were the results:

comparison preferred
raw draft vs fully processed 22 out of 22
raw draft vs rewrite pass alone 16 out of 16
rewrite pass vs full pipeline 12 of 21, which is chance
fully processed vs real human prose 18 out of 18

Let's take our magnifying glass and point it a row four. Judged on quality alone, the machine passages beat the real human prose, 18/18 and by 22 rubric points in fact. Funnily enough, those happened to be the same passages the detector had called human every single time in my last tests. I'd like for you to take what I have to say next to heart. "Reads as a human wrote it" and "it's good writing" are never going to be the same axis. My quality control prefers the machine tics and register, which gives you a moment to pause and think about discounting the raw drafts versus the fully processed prose. Again, the corpus was identical, human text was matched on kind of text and length, never on the subject so this was never a like-for-like contest.

Now, let's take a look at row three. The checks and tools I spent a lot of time building left the text completely identical in 11 of the 18 runs in this experiment. As a result, twelve of the 21 trials in the row were just some of the judges comparing identical text. There were also 9 trials where there were edits and won 7/9 times. My agents were open about this finding, however & I likely wouldn't have caught it otherwise. Again, a result of testing at scale solo.

I have one more finding in this section should you choose to run some of these tests yourself. Run my checker against the included documents from 1854 to 2018. It fails four of them, the checkers rules encode my own tastes in prose. The checker DOES NOT identify machine writing alone.

Full Circle

As you can probably extrapolate at this point, there are limits to what my tests have produced. The tests measure whether a language model can identify the author of a piece of written content. It's not going to tell you whether Google is going to demote AI content or not. They're wildly different questions and the data I've provided here and in my repository doesn't claim to answer the Google question. If someone decides they're going to cite a study similar to mine to tell you whether or not Google, ChatGPT or anyone else is going to punish your AI assisted pages, they're likely selling you some level of snake oil.

I have no experiments ran in this subset which measure what your ranking is, what kind of traffic you're pulling in or the kind of enforcement that will be placed upon your content. Anyone attempting to sell you a content writing tool is in the same boat, my experiments show this plainly.

As a reminder, the prose was either authored by me or from the same family of models entirely, an outside evaluator looking at this through and objective lens would likely pick that apart and say that had I ran it through other models there might be a different result.

I tried that.

I had Sol and Grok take the prose and run the method themselves, multiple times. The results were the same on each run. In desperation I again told them to think more abstract, gave them examples of what abstract writing looks like, what humanity feels like. Where I arrived surprised me, almost as if I were running my own little Turing test. Reader, step back and truly ask yourself what humanity is, what being conscious or intelligent is. Chances are your answer is in a lot of ways different than what I or some stranger will give you.

My Advice to you

Stop paying for concealment of AI prose, yesterday. You're being sold something that I believe is only provable because of identifiable tics. It's authorship that trips up the scanners, not editing the prose to your own voice or shape. If you're on the fence about buying a tool that says it will do that for you, I implore you to think again. This is the study you want before you make a decision like that.

Take your drafts and rewrite them instead, know which tells and typography die when you process your drafts. What's left in most cases is rhythmic and structural, that's where your personality should be placed. Keep your lists, and accept the cut to padding. Follow the guidelines Google and other platforms give you and provide content that solves a problem for a person. The machines will follow.

Add your own personal twists that a machine could never copy. Perhaps you had a customer call in last Tuesday that left an impression on you, or an employee that shared something with you that made you change your perspective on things. I believe the difference in a machine written and a machine assisted copy lies exactly in the seam where you can tell a personal story. I do have to say that even my own tests scored invented stories, and real stories exactly the same though, so take it with a grain of salt.

A final thing to rule out

The text you write picks up characters which are invisible to the eye. I processed each document I built for these studies which amounted to 2.4 million characters in total and did not find a single homoglyph. I found 27 zero-width spaces inside a Supreme Court opinion which tells you something very interesting. The document was copied from somewhere else.

The reason you should rule the rare occurrences out is the cost of the rarities in the first place. A homoglyph is a letter which is borrowed from another alphabet that renders identically to it's Latin counterpart. If a Cyrillic "c" appears inside the word "ceramics" for example, you are betting that every system which reads the page will fold it back into a Latin c before it matches anything. There is research and proof that some will, and some wont. You have no way of knowing which models will process your page as normal and which ones will pass it over due to a simple character being placed there because of your use of writing assistance. It's not something your common SEO audit is looking for either.

Shameless plug, you made it to the end of the funnel

I built Inkwash for you to scan your own text, locally. It works inside your browser and nothing you ever paste is sent anywhere. The tool will identify every character it finds with the exact code point naming every text character it read around it. It reads your Word docs and informs you what is in the metadata as well such as the author's name, the company, revision counts and more. You can then choose to strip all that data and use it to your liking. It's completely free, I will never ask you for anything if you choose to use it.

The final receipt

This article was run through the exact tool I've described in depth. It was then ran through the stricter checker I use on EVERYTHING I publish on this site and on social media. Both are shown below so you can check my work yourself.


$ python humanist.py draft.md
humanist 0.1.0  |  4,764 words, markdown-stripped
  readability FK grade 8.1
RESULT: 0 FAIL, 0 WARN. CLEAN.

$ python check_prose.py draft.md --mode post
FRAME: markdown-stripped, 4,764 words, FK grade 8.2
RESULT: 0 FAIL, 0 WARN. CLEAN.

Here's the part neither of those results are going to convey, and the part I believe you're still reading this now near 25,000 character leviathan for. I took the same corpus to two commercial detectors which is completely outside the opinions this study was meant to tell you it didn't have.

Three passages of genuine human writing which were published long before either you or I were even an idea or a sperm cell swimming around in someone's scrotum were pushed through GPTZero. All three passages (Feynman's appendix to the Challenger report, Thoreau, and John Reed) came out as 100% human. The instrument works on the easy direction. My style checker pushed the same result.

I then pushed four versions of machine written passage through Pangram. One which was untouched, one which the model rewrote every sentence three times, one where I changed not a single word and changed where the sentences joined, and finally one with both changes. The result, and one that came as a surprise to me: 100% AI according to Pangram. I thought I was sneaky, truly I did. Pangram said no.

Rewriting every word did nothing for you, and rewriting no words did nothing for you. Pangram reduced my whole study into two soul crushing lines and they had zero stake in whether I was right or wrong in the first place.

The original drafted article was rated at 100% AI written with high confidence. I choose to print those numbers as a badge of confidence in myself rather than bury it and sell you a lie. What the tool is doing correctly is reading your prose correctly, and mine too. What the tool cannot do is see the frame, the corpus, the perspective and the studies which I've outlined above. There was real time, effort and inquisitiveness attached to it. The humanity lies in "what if I can?" That's exactly what I attempted to set out to do these last two days.

There are two numbers which I identified as incorrect during my tests and I caught them before publishing this article. You can see both changes in the repository history. I described the pipeline as producing a 2.2 point improvement when the correct statement was that there was no effect detected. As I outlined earlier, small changes to the left or right are not changes that warrant resolving at scale. I reported the paragraph results backwards which cited a pooled average which was an artifact of how the judging batches eventually fell.

THE BIG ONE

Everything I've written above ended this morning before I decided to write this article in the first place. Three results since that decision exist from these tests and I believe they're the reason you should care at all:

I took a draft that Grok wrote me which Pangram rated at 100% AI and decided to put it side by side with an empty text document. This draft was rigorously picked apart over 4 versions until it landed on what Pangram spit back as "Haha, nice try" I read it and in my eyes the draft was truly abstract, it had none of the usual tells. Not one of the words my own checker bans on sight. Nothing. In my eyes it was a genuine piece of art. The passage was about someone's father who had passed away and somehow tied it into benches on a walk home and homelessness. Truly, in my opinion an interesting passage. Anyway, I rewrote it, however, about a third through, I got lazy and stopped adding my own personal touch and leaned harder on the original work. The result after I fed it back into Pangram? The moment I got lazy is the moment Pangram said "I see you, and I see Grok" even though the wording of the passage in my eyes was truly different. Food for thought. The actual split was 31% human 69% (Nice) AI.

Additionally, I dropped in an X post I wrote a few days ago to see what Pangram thought about that. 538 words without any assistance, cold. 100% human.

I ran the final version of this article through Pangram once more and the final version was exactly what I needed to wrap all of this up. The final result was 56% AI and 44% human/mixed. 7 of the 13 total segments came back as written by AI. I typed every one of these words myself, none of it was copied and pasted. This is where I think the danger lies entirely if you decide to use tools to rewrite your content. The commercial scanners rarely, if ever tell the person checking your work whether you wrote all of it yourself.

There are two provable adjacent blocks I can use as I write the final lines of this piece. It was one paragraph which was split down the middle completely. One half where I described the building the corpus read as if a human wrote it, the other half where I listed the method read as AI. I thought I could trick the system. Funnily enough, in this article I halfway did. That's saying something.

I now believe these tools are looking for the scaffolding. I believe they look at the structure and pick apart the pieces that an LLM will write. The LLM will give you cold hard numbers even if you're not asking for them. The LLM will write prose in a predictive way, mechanically. They really do put the "Machine" right into machine learning.

The verdict: Use a draft written by AI if you'd like. If you want the coveted 100% human check of approval, just know that the machine's fingerprints are an outline and the commercial checkers appear to be able to identify that outline no matter if you wrote 100% of the prose.

Somebody who would know

The moment I hit publish I decided to file an issue on the GitHub repo my skill was derived from. The repo belongs to Siqi Chen, the original author of the Humanizer Claude Plugin which has over 36,000 stars at the time of writing this article. I decided to thank him for his work and outline all of the work I did with my own tests including the part where his skill was the one quantitative winner when it came to writing prose better.

Hello - thanks for your analysis!

It is not currently possible to defeat Pangram through pure LLM generation through any skill or prompt (I have tried!)

Regardless, defeating AI detectors is not a goal of this repo - rewriting AI text to sound less AI to humans is a goal.

He could've told me I ran my tests incorrectly, instead he told me he had also tried to do the same thing I'd attempted. This was born from a hunch, and now it isn't anymore. His tool was never meant to conceal AI writing. Rather, it was meant to make AI writing more digestible to a human reader. I believe both of our tools have succeeded in that wheelhouse, his to a greater extent and that is provable with my research.

Forward

Failure to me doesn't mean I failed. It means I tried. And after all of this, "THE BIG ONE" proves that in kind. You'll miss 100% of the chances you don't take. You'll learn nothing if you don't stand up and ask the questions everyone else wants to but is too timid or too anxious to blurt out.

I appreciate you taking the time to read what I have to say, and I hope that my studies provide you some insight into how these systems work, and perhaps help you develop your own tests and studies. Inkwash and the checker are my own work, assisted by Claude, Codex, and Grok. The source for both is at github.com/itsryanlenk/humanist. They are MIT-licensed and the studies attached to them are committed in tandem.

One caveat if you do choose to reuse them: the rewrite passes are originally vendored from someone else's MIT project, and it's pattern catalog was derived originally from Wikipedia which carries CC BY-SA 4.0 licensing. That obligation will travel with it if you copy anything onward. Check docs/attribution.md for the full chain.

If your question is whether AI systems can read and cite your property at all, it is going to be a different problem with an entirely different answer, concealment has nothing to do with it. The free AI readiness scan checks for that instead, and if you would rather I ran it for you, that is what I sell.

Want receipts like this for your own site?

Run the free AI Readiness Scanner: paste in your robots.txt and your homepage source, get a scored report on whether ChatGPT, Perplexity, and Google AI can actually read your business. About two minutes, no signup, and nothing you paste ever leaves your browser.

Run the free scan → Get The Receipts free ▸
Ryan Lenk
Written by Ryan Lenk

I run the SEO and AI visibility for my family's Shopify shop and publish the receipts, good and bad. Every number on this site comes from a named tool export, and corrections get published rather than edited away.

About meRun the free AI Readiness ScannerWork with me

← All articles Subscribe on YouTube ▸