When AI Eats the Web: Who Protects Our Collective Memory?
AI is changing the way people reach information online. Search engines now generate instant overviews. Chatbots answer questions in full sentences. Publishers, forums, and archives say their work is being scraped, summarized, and reused at scale. Recent online debates have described this as AI “eating” the web: machines pull knowledge from millions of pages, then increasingly stand between readers and the original source.
This matters because the web is more than a pool of facts. It is a record of reporting, argument, local history, specialist knowledge, and everyday language. The main tension is not hard to see. AI tools can make information faster to find and easier to translate. But if they reduce traffic, revenue, and visibility for the people who publish the original material, they can weaken the system that keeps public memory alive. For Arabic and bilingual readers, that risk is even sharper, because the online record is already uneven.
What we know
No single study proves that AI is making the internet’s collective memory disappear. The web has always been fragile. Links break. Small sites shut down. Newsrooms close. Old forums vanish. Archives go behind paywalls or into apps. So this is not a problem created by AI alone.
But AI adds a new layer of pressure. It does two things at once. First, it consumes huge amounts of online material for training and retrieval. Second, it often presents the result in a way that lets users stay on the platform instead of visiting the source. That changes the economics of publishing. If fewer people click through, fewer publishers can afford to keep writing, updating, and preserving their work.
That is not just a complaint from media companies. The same pattern affects hobbyist forums, local history blogs, research communities, niche wikis, and volunteer archives. Much of the web’s memory has never lived inside large institutions. It lives in small pages that are easy to overlook and easy to lose.
Access is not preservation. A quick summary may be useful, but it is not the same thing as a durable, visible, credited source.
Why summaries are not enough
A summary can tell you what happened. It often cannot show you how we know. That difference matters.
Original pages carry context. They show the date, the author, the source list, the edits, the argument, and sometimes the disagreement. A machine-written answer usually compresses that into a few lines. Even when the summary is accurate, the compression removes texture. You may get the conclusion without the evidence, the quote without the speaker, or the event without the local setting.
Think about a long forum thread where engineers solved a rare hardware problem over several years. Or a local newspaper archive that tracked a town’s land dispute. Or a blog post in Arabic explaining a legal change in plain language for residents. These pages are not just containers for answers. They are records of how knowledge was built and contested.
Once that material is turned into a clean response box, something important can be lost: provenance. Readers need to know where a claim came from, whether it is old or new, and whether other sources disagree. That is how public memory stays accountable.
The risk is higher for Arabic and bilingual knowledge
The web has never treated all languages equally. English dominates online publishing, digital archiving, search indexing, and training data. Arabic is large and rich, but its online record is fragmented. Some material is in Modern Standard Arabic, some in dialect, some in French or English, and some in mixed scripts or transliteration. Much of it sits on small media sites, aging blogs, regional forums, or pages with weak technical preservation.
When AI systems become major gateways to knowledge, they tend to favor what is easiest to index, scrape, rank, and summarize. That often means bigger English sources, cleaner formats, and already visible institutions. The result can be a narrower memory of the Arab world, filtered through the sources that are easiest for machines to process.
A bilingual reader may see this in small ways. Search for a regional event, a labor dispute, a court case, or a cultural debate, and the machine may lean on international coverage rather than local reporting. Ask for a historical explanation, and it may summarize the best-known English references while missing Arabic commentary, archives, or oral-history projects. This is not always bias in the dramatic sense. Often it is simply a structural preference for the material that is most available to the system.
But structural preference shapes memory. If some knowledge is rarely cited, rarely surfaced, and rarely linked back to, it becomes harder to find, harder to fund, and easier to forget.
The case for AI is real
There is a serious argument on the other side, and it should be acknowledged. The web is full of spam, duplication, junk SEO pages, and low-quality summaries written to catch search traffic. Many users do not want to open ten tabs just to get one straightforward answer. AI tools can save time. They can also help with translation, accessibility, and discovery.
For non-native English speakers, this matters a lot. A good summary can lower the barrier to entry. A translated explanation can open material that would otherwise remain inaccessible. For people on slow connections or small screens, faster answers are not a luxury. They are a practical benefit.
There is also a fair point that not every lost click means a lost reader. Some users would never have visited the source in the first place. In some cases, summaries may even send new audiences to original work they would not have found on their own.
These are real gains. It would be a mistake to pretend the old web was working perfectly or that every original page deserves equal attention.
Why the current balance is still wrong
The problem is not that AI makes information easier to use. The problem is that the value chain is becoming lopsided.
Original reporting, research, archiving, and community moderation cost time and money. Training models, generating answers, and keeping users inside a platform can be highly profitable. If the second layer captures the value while the first layer absorbs the cost, the system becomes extractive. Over time, fewer people will invest in the hard work of making and maintaining trustworthy source material.
This is the core of the debate. AI companies often describe themselves as organizing information more efficiently. Critics say they are also free-riding on the institutions and communities that produced that information. Both views contain some truth. But if efficiency destroys the incentive to create and preserve the underlying record, then the efficiency will not last.
There is another issue: correction. When an article is wrong, it can be updated. When a source is disputed, readers can compare it with others. When a summary is detached from its source, correction becomes harder. Errors can spread in a cleaner, faster form than the original material ever did.
What protection should look like
My position is simple: AI should be built on top of a living web, not at the expense of it. That means the systems that benefit from public knowledge should help preserve, credit, and support the sources that produce it.
- Stronger source visibility. AI answers should make citations prominent, not decorative. Users should be able to see where a claim came from, how recent it is, and what other sources say.
- Fair licensing and compensation. Large-scale commercial use of archives, journalism, and databases should not depend on vague assumptions that everything public is free to absorb. Publishers and archives need workable licensing models, especially for high-value material.
- Better controls for publishers. Site owners need clearer technical tools to say what can be indexed, summarized, retrieved, or used for training. Today, the choices are often too blunt: allow everything or block everything.
- Public investment in preservation. Libraries, universities, public broadcasters, and national archives should treat born-digital material as part of the historical record. That includes local news, forums, and language communities that the market alone will not protect.
- Multilingual accountability. AI products should be measured not only for accuracy, but also for language coverage and source diversity. If a system answers questions about Arabic-speaking societies mostly through English-language material, users should know that.
- Design that encourages click-through when it matters. Not every query needs a visit. But on topics like health, law, history, politics, and breaking news, platforms should push users toward original reporting and primary documents.
This is also a cultural question
Collective memory is not just a storage problem. It is a visibility problem. A page can still exist and effectively disappear if no one sees it, cites it, or can afford to maintain it. That is why this debate goes beyond copyright and product design.
Who gets remembered online has always depended on power: which institutions publish, which languages get indexed, which stories get linked, which archives survive. AI does not create that imbalance from scratch. But it can harden it by turning existing visibility gaps into default answers.
For Arabic and bilingual communities, this should be taken seriously now, not later. Once a thinner record becomes the standard machine-readable record, rebuilding what was missed will be difficult and expensive.
Protect the source, not just the shortcut
The web does not preserve itself. It survives because people report, write, edit, host, fund, and archive. AI can help readers navigate that work. It can widen access. It can reduce noise. But it should not be allowed to consume the public record while making the original record less visible and less sustainable.
If machines are becoming the main readers of the internet, then society needs rules that keep original knowledge worth creating. The test is practical: does the new system send value back to the sources, or does it only extract value from them? If we fail that test, we will not just lose traffic. We will lose memory.