How AI Is Erasing the Internet's Memory — And What We Can Do About It
AI search summaries, link rot, and archive deletions are erasing the web's collective knowledge. Here's what's happening and practical steps to fight back.
The Slow Death of the Internet's Memory
For decades, we treated the internet as a permanent record. Whatever you published, whatever you read, whatever you referenced — it would always be there. That assumption is breaking down in real time, and AI is both the accelerant and the match.
A recent piece from The Walrus titled "Google Search Is Dying. What Comes Next Is Worse" sparked a massive discussion on Hacker News, hitting nearly 300 points and almost 300 comments. The article lays out something many of us have felt intuitively but struggled to articulate: the infrastructure that once stored and served humanity's collective knowledge is crumbling — and AI is making it worse, not better.
Google Search: From Library to Slot Machine
Google built its empire on returning the right answer. That reputation is now actively eroding. The company's AI Overviews have told people the sunset already happened when it hadn't. They've suggested eating rocks, putting glue on pizza, and other hallucinations that would be funny if they weren't appearing where trusted information used to live.
The problem isn't just bad answers. It's that AI summaries sit between you and the original source. Even when the underlying page still exists, it becomes practically undiscoverable. Google interposes an error-prone layer of synthetic text that displaces the very pages it claims to summarize. You no longer find the source — you find Google's summary of the source, complete with all the fidelity of a game of telephone.
The result is a feedback loop: fewer clicks to original sources means less incentive for publishers to maintain those sources. Less maintenance means more link rot. More link rot means more gaps for AI to fill with hallucinations. And so on.
AI Slop Is Polluting the Upstream
It gets worse. According to reporting from 404 Media, companies are now deliberately planting content on Reddit and other platforms to influence the answers that AI search engines generate. They're not gaming search rankings anymore — they're gaming the training data and the real-time retrieval pipelines that feed AI summaries.
This is contamination at the source. If AI search summaries are built on top of Reddit threads, and those threads are being strategically manipulated, then the knowledge you get from an AI search isn't just incomplete — it's actively shaped by whoever is willing to pay to plant it there. The public record is being edited before you ever see it.
The Archives Are Under Attack
If the front door to knowledge is broken, what about the back door? That's where the Internet Archive comes in. The Wayback Machine has been the web's closest thing to a fail-safe backup — hundreds of billions of snapshots capturing pages that have since disappeared. Journalists, researchers, lawyers, and historians rely on it daily.
But the Internet Archive is buckling. It faces engineering strain from indexing and storing an ever-growing repository, cyberattacks targeting its infrastructure, costly litigation from publishers who sued over digital lending programs, and news organizations blocking Wayback Machine crawlers out of fear that archived pages can provide AI companies with an indirect source of copyrighted material.
Each restriction limits the archive's ability to serve as a comprehensive backstop. The irony is bitter: organizations are blocking the Internet Archive to prevent AI from training on their content, but in doing so, they're destroying the only reliable record that their content ever existed.
Wikipedia: Infrastructure of Its Own Demise
Wikipedia may be the most consequential case study. For years, search engines sent billions of viewers to Wikipedia pages. That traffic sustained the encyclopedia through donations, volunteer engagement, and the network effects that make a community-run project viable.
Now, AI systems scrape and ingest Wikipedia's content directly, presenting it in their summaries without sending users to the site. Wikipedia has become the infrastructure of its own demise. Traffic drops. Donations drop. Volunteer engagement drops. The encyclopedia that trained the AI that now displaces it slowly starves.
FiveThirtyEight and the Deletion Problem
Consider FiveThirtyEight. The data journalism site, owned by Disney, was a go-to source for election analysis, sports predictions, and economic commentary. When Disney decided it was no longer an active asset, they didn't just shut it down — they deleted nearly the entire archive. Years of methodology, analysis, and data visualization simply vanished.
This is the pattern. When a platform decides content is no longer profitable, it doesn't just stop updating — it erases what was already there. A banned book can still be found. A scrubbed webpage can disappear so completely that few people would ever realize it existed.
What Can Be Done: Practical Steps
The situation is dire but not hopeless. Here are practical steps that individuals and organizations can take:
- Support the Internet Archive financially and politically — it's the web's only large-scale memory backup
- Self-host critical content rather than relying on platforms that can delete it without notice
- Use independent search engines like Qwant or Kagi that don't replace sources with AI summaries
- Back up important web pages locally using tools like SingleFile or archiving them through the Wayback Machine
- Donate to Wikipedia — the infrastructure that supports AI also needs support from the humans who benefit from it
- Be skeptical of AI-generated search summaries and always click through to original sources when accuracy matters
The Broader Lesson
The internet was never designed to be a permanent archive. It was designed to be a network — a living, changing system. But we built our collective memory on top of it anyway, assuming it would hold. Now that assumption is failing, and the tools we're adding to the web — AI search, AI content generation, AI summarization — are accelerating the loss rather than preserving it.
The question isn't whether AI is useful. It clearly is. The question is whether we're willing to lose the web's memory in exchange for convenience. Because that's the trade that's happening right now, and most people don't even know it's on the table.
If we want a future where knowledge persists, we need to treat digital preservation as infrastructure — not as an afterthought, not as a nice-to-have, but as something as essential as the power grid. The alternative is a world where the past is whatever an AI says it was, and nobody can check.
Related Posts
Why Your Local LLM Feels Dumber Than It Is: The Hidden Quality Gap
Your local LLM is not broken. Quantization, weak system prompts, and basic inference engines silently degrade quality. Here is what to fix.
AI Blindness: When Your Brain Learns to Stop Reading AI-Generated Content
A growing number of people report their brains automatically filtering out AI-generated text, like banner blindness for LLM output. This phenomenon reveals something deeper about trust, attention, and the future of human-AI interaction.
When GitHub Went Dark: Inside the 7-Hour Outage That Paralyzed the World's Code
GitHub's August 17 outage lasted nearly 8 hours and took down the entire platform — including Copilot. The root cause wasn't code: it was capacity. Here's what happened and what it means for every platform team.