AdSense Rejected My Site for Low Value Content
TL;DR
An automated blog had published thousands of posts built from a shared library of paragraphs. An audit showed 95% of posts were 90 to 100% copied word for word from other posts. I paused the generator, kept the five hand-written posts listed and indexed, and marked the rest noindex, unlisted and ad-free without deleting them.
Key takeaways
- AdSense judges the whole site, not just the pages that show ads.
- Template-assembled posts count as scaled content no matter how many there are or how well they rank.
- A short script that counts repeated paragraphs shows the problem in numbers.
- noindex plus removing pages from listings and the sitemap is a reversible way to take them out of Google's view.
Google AdSense reviewed this site and said no: "We found some policy violations. Low value content." The message asked for a site that "provides authentic, high-quality information", "exhibits ongoing curation and structural maintenance" and "generates and sustains genuine user interest".
I suspected the cause. This post is about measuring it properly and fixing it without destroying anything.
The suspect: an automated blog
For several months, this site ran an automated publisher. Every 30 minutes it assembled a new post: it took a title from a queue, picked sections from a library of pre-written paragraphs for that topic, filled in the keyword, and added FAQs, a summary and a closing call to action. It published 48 posts a day. By the time AdSense reviewed the site, there were 6,346 posts.
Each post looked fine on its own. The problem only shows up when you compare them.
Measuring the duplication
I wrote a short Python script that splits every post into paragraphs, normalises whitespace and case, and counts how many posts contain each paragraph. Here is the core of it:
import collections, glob, hashlib, re
posts = {}
for path in glob.glob("src/content/blog/*.md"):
body = open(path, encoding="utf-8").read().split("\n---\n", 1)[-1]
paras = [re.sub(r"\s+", " ", p).strip().lower() for p in re.split(r"\n\s*\n", body)]
posts[path] = [p for p in paras if len(p) >= 80]
h = lambda s: hashlib.md5(s.encode()).hexdigest()
seen = collections.Counter(x for ps in posts.values() for x in {h(p) for p in ps})
for path, ps in posts.items():
total = sum(map(len, ps)) or 1
copied = sum(len(p) for p in ps if seen[h(p)] > 1)
print(f"{copied / total:.0%} {path}")
My full version also replaced each post's own title and keyword with a placeholder, so templated sentences matched even when the inserted keyword differed. The results, grouped:
| Share of a post's text copied word-for-word from other posts | Posts |
|---|---|
| 90–100% | 6,013 |
| 70–90% | 297 |
| 30–70% | 31 |
| 0% | 5 |
Three boilerplate paragraphs each appeared in 6,341 posts. Only five posts had no copied text at all, and they were the ones written by hand.
Google's spam policies describe this almost word for word. Scaled content abuse includes "stitching or combining content from different web pages without adding value". It does not matter that the paragraphs were originally written for this site, or that some of the posts got search traffic.
What I changed
I wanted a fix that removed the problem from Google's view but could be reversed, and that did not break any existing links.
1. Paused the automated publisher. Every new post would have added to the problem. The cron line is commented out rather than deleted, so it can be restored in seconds.
2. Added a curated-only mode. One switch in the code decides which posts are part of the site:
export const CURATED_ONLY = true;
export function isListed(post: Post): boolean {
return !CURATED_ONLY || post.curated === true || CURATED_POST_SLUGS.has(post.slug);
}
Only curated posts appear on the blog page, the homepage, the RSS feed, the sitemap and related-post links, and only they load ads. A post becomes curated by adding curated: true to its frontmatter.
3. Kept every other post, but out of search. The template posts still open at their URLs, so no old link breaks. They are served with noindex, follow, carry no ads, and their "keep reading" links now point to the curated posts.
4. Removed thin listing pages. Category pages now need at least three listed posts to exist. The hundreds of category and pagination pages that listed template posts now return 404.
The sitemap went from 6,924 URLs to 25. I submitted all 6,899 changed URLs through IndexNow so Bing would recrawl them, and Google will see the noindex tags as it recrawls.
What I did not do yet
I did not request another review straight away. Google's index still held about 5,200 of the template pages, and they will take weeks to drop out as Google recrawls them. Five original articles is also not much "substantial unique value".
The right next step is the slow one: publish original articles, written from real experience, on a steady schedule. This article is one of them.
Lessons
- AdSense reviews the whole site. Pages without ads still count.
- Volume is not value. Thousands of templated pages can make a site look worse, not bigger.
- Measure before you argue. A 30-line script turned a vague rejection into a precise, fixable number.
- Prefer reversible fixes.
noindexplus delisting removes pages from Google's view while keeping every URL working.
Frequently Asked Questions
What does 'low value content' mean in AdSense?
It means Google judged that the site does not offer enough original, useful content to support advertising. Thin, duplicated, auto-generated or templated pages are the most common causes.
Should I delete low-quality pages or noindex them?
Both remove pages from search. noindex keeps the URLs working for anyone who follows an old link and can be reversed easily. Deleting is cleaner but permanent. Either way, remove them from your navigation and sitemap.
How long should I wait before requesting another AdSense review?
Wait until search engines have processed your changes, which you can watch in Search Console's indexing report, and until the site has enough original content to stand on its own. Requesting review too early usually produces the same result.
Sandeep Kumar Chaudhary
Full Stack Software Developer· Nepal's SEO, AEO, GEO & AIO expert and share-market educator. More about me
