July 6th, Reddit published an article in which they stated that they're upgrading their filters to curb AI web spam. Using large language models, the platform now very accurately detects fake hype, coordinated behaviour, and manipulation, which earlier filters used to miss. 23 million spam views are getting blocked, and 25k spammy posts and comments are getting filtered out daily, as per the numbers disclosed.
Only a week prior, Google's Spam update rolled out, targeting scaled and slop content published by AI content mills.
A Redditor and a gaming nerd for 7 years says, “Believe me, when I read something on sites like Reddit, Quora, Substack, Medium, LinkedIn, you name it, half of the time it's like I'm reading the same thing again and again and again, regurgitated.”
Wanna know what's scary here?
If some writing or an idea or a thought was bad before, you could catch it instantly, but not now. Now that same flimsy idea, basic fluff, is getting beautified thanks to AI assistance and made into sales pitches, presentations, websites, blogs. Grammatically immaculate, with beautiful structure but hollow in depth and information.
No information is real information thanks to AI
AI content today looks beautifully immaculate but in reality is hollow in depth and information
I wouldn't exaggerate it because good pieces of content still exist on the web, but half of what you're reading right now is probably GPT-generated.
Now tell me: if GPT's answering your query, but you actually needed the answer to come from someone with the actual knowledge and lived experience, is that fair for you as the seeker?
This is now the reality of all forums and reading platforms. Earlier, people used to go to Reddit cause they didn't believe Google's beautified search results. Or they went to Quora and published their question hoping for in-depth answers. But now, what matters is satisfying query fan-outs via structured, short, and direct snippets of shallow information. This is what deep research agents and AI bots crawling the web love.
Open source/commercial chatbots have this affinity to retrieve whatever they can easily scrape at the shell level. Too many 4-5 star reviews, all neutral positive seedings and snippets. In reality, most of it's now inorganic, paid slop getting seeded day in and day out.
Sure, there are platform guardrails regarding repeated content, spam, manipulation, and self-promotion. But coherence and fluency merely can’t mimic real knowledge. A recent habit that I've adopted now is to check the author or poster’s activity history. Their past answers, comments, and posts, and how much they adhere to the platform's regulations before actually believing that something's worth reading.
Some newer platforms like Medium and Substack have also taken strict measures against AI-generated writing. Medium limits exposure of author-generated work which feels like AI, and quality articles stay behind paywalls. Sure, this reduces amplification of slop, but then again, whatever gets indexed on Google is primarily what’s freely readable and accessible.
Recently, on 30th July, even LinkedIn introduced a “Seems like AI Slop” reporting feature to upgrade their classifiers. The company said to mark slop writing by tapping on the button. The irony is, there’s also an “enhance your post” pop-up that comes when you try to write a post, basically AI-assisted proofreading to add flair and fluency to the poster’s broken words and thoughts.
Traditional search simply depends on ranking and relevance. Retrieval Augmented Generation (RAG) doesn't. RAG synthesises multiple answers, pages, and passages, and picks up bits and pieces from each and merges them into an ingredient for the next step.
The next step, which is vector search, involves numerical tagging of queries and passages and the coordination between their semantic proximity, or in simple words, how closely the answer resembles the ask of the question. Basically, whatever seems the most relevant becomes of higher proximity (but that's not the truth). Blogs, landing pages, Reddit posts, and Quora answers, which are written to best match popular query intents get retrieved before others.
The final output hence comes from a collective search results, which typically hold high domain and page authority. More enriched with semantic LSIs, but can be far, far away from the actual truth.
How are vector embeddings in RAG manipulated by AI search spam?
How vector images embeddings in RAG manipulated by AI search spam
Let me explain these via an example. Imagine an online medicine delivery app called X. They have a mix of good and bad reviews online and many other competitors, some of them a little better in terms of support and delivery speed than they are. Now this company's SEO team is smart.
They move to Reddit and, using a cluster of surrogate profiles or bots, so to speak, start creating posts where the titles either exactly mimic mid-tail or long-tail queries or are basically fan-outs of the original query. A few of the other profiles start astroturfing below those posts, writing neutral, positive in-favour responses for X. Mimic that for a couple of months; some removed, some accounts banned. But even if one or two of those posts sustain, they'll shoot up in Google rankings next time someone searches anything about X along those same queries.
Hallucinations are the basis of this, and this is totally opposite to what Google publicly discloses as their metrics for ranking and sourcing good-quality content. Google defines scaled content abuse as an activity of generating lots and lots of low-value fluff content, primarily to jam up the search engine so that the crawler's forced to retrieve these pages. It's the polar opposite of their people-first guidance, which has always emphasised detailed reporting, lived experience, and analysis of the issue and viable solution. That's useful, adaptive, and deep.
Yet again appears the paradox of depending on algorithmic retrieval to display the search results. See, algorithms love whatever's relevant and whatever's trending. Post something they love, and the chances of your content also climbing to the top increase. YouTube, TikTok, Instagram, everything's riding on this train.
But with AI being associated nowadays with real-time search operations, I think to myself: if the models are getting trained on ranked fluff, won't they recommend even more fluff?
Ways to sort out slop from good information
How to sort slop from good information
If you've been chronically online like me, then manual ways would be the best for you. Look for telltale AI sounds like plasticky content reading, having multiple em dashes, not x but y formats, and so on. They sound very formulaic, neutral, and lack the burst of uncanny emotions and subtle abrupt sentences as we humans write.
The writing industry now has AI detectors to check watermarked AI texts and these kinds of formulaic outputs. But then again, they also hallucinate, can’t comprehend languages other than English and oftentimes class original writings as AI-generated text.
Mike Caulfield’s SIFT method is a reasonable way to segregate quality from slop. It involves:
Stop: Basically, stopping and checking whether it's a genuine finding-based read or has a CTA attached at the end, asking for sign-ups, credentials, etc.
Investigate the source: Backtracking the author's experience and domain of expertise is necessary to understand where the knowledge's coming from.
Find better coverage: Most data on the web is a secondary source. But only a few of these cite and reference the original primary sources. That should be your main point of information.
Tracing claims: Read, understand, take your time, check the data for its statistics and methodology, and see if there are any limitations faced by the author. That completes your sourcing arc.
There are a few other public digital provenance solutions which can help battle misinformation, for example, the C2PA manifest. Simply explained, it's a cryptographic signing protocol which is attached to the digital asset, kind of like invisible watermarks that trace back to the original owner.
Other checkmarks involve the DIN SPEC 33461 SEO methodology (to check content for AI fluff) and adding author schemas in published pages.
Trust's now a two-way street
The dead internet theory always referred to the greater part of the internet to be inorganic. We're living that reality now.
You can't put your faith in what you're reading online. Neither can someone who's writing assure their content is getting viewed by real humans. It's the platforms and editors who need to put up the necessary guardrails, measures, and governance.
Added to this, disclosure of AI use can also be a signal of genuineness. Whether that's for the images used, content synthesised, or simply your thoughts structured. A chain of authenticity starting from the writer to the editor and finally to the platform that's gonna publish would ultimately help keep AI spam in check.