🌿freegardner

Science

Open data becomes a liability as bots scrape research

13 Jun 2026 · via Nature

Open data becomes a liability as bots scrape research

The Quiet Theft of Open Science: When Sharing Data Becomes a Liability

The open science movement promised faster discovery through shared data, but the same tools now threaten to undermine that promise.

For two decades, researchers were told to share their data openly. Funding agencies required it. Journals demanded it. The argument was simple: science progresses faster when everyone can see the raw materials.

That promise is now being tested by automated systems that consume data faster than any human researcher could

Open data becomes a liability as bots scrape research

In June last year, the Confederation of Open Access Repositories surveyed its member organizations. More than 90% reported encountering bot scraping, with most seeing abnormally high bot activity at least once a week, as reported by Chris Stokel-Walker in Nature

Those bots are not reading the research. They are mining it. In some cases, they extract data sets, combine them, and feed them into artificial intelligence models that generate new papers and results faster than any human researcher could.

Andrea Howard, a psychologist at Carleton University in Ottawa, Canada, put it plainly: “It’s a pretty big issue everybody should be thinking about, whether you’re for or against AI.” [2]

What Was Before

The open-access movement began in the 1990s and early 2000s, when researchers argued that publicly funded research should be publicly available. The Berlin Declaration on Open Access was signed in 2003. The Budapest Open Access Initiative followed.

By 2015, most major research funders required grantees to make their data available in open repositories. The logic was sound: if taxpayers paid for the research, they should be able to see the results. If other researchers could build on published data, discoveries would accelerate.

Then something changed.

What Is New

The emergence of large language models and generative AI tools shifted the landscape entirely. Where once data sat quietly in repositories waiting for human researchers to analyze it, now automated systems can consume entire data sets in minutes.

Miri Forbes, a quantitative psychopathologist at Macquarie University in Sydney, Australia, described the change with precision: “The scope and speed of how quickly automated pipelines can exhaust the research questions a data set can answer feels like a big change. It shrinks the space left to work in a given data set.”

This is not a theoretical concern. The bots are aggressive. According to research from Barracuda published on TechRadar, these “gray bots” are designed to extract serious amounts of data from websites, most likely to train AI models or collect web content like news, reviews, and travel offers. [4] The report notes that while these bots are not outright malicious, their approach can be “questionable” and some are even “highly aggressive.”

The Two Sides of the Debate

Some researchers argue that the potential for automated science to be used for good — speeding up the discovery of new drug targets, for example — means open data should remain open.

But others point to mounting evidence that bots scraping complex data sets contribute to low-quality research and “AI slop.” They also warn that sensitive data, including patient information, can be extracted through these automated processes. These researchers argue that new rules and technical systems are needed to restrict bot access to databases.

The debate echoes earlier controversies in the history of science. When the internet first allowed widespread sharing of preprints in the 1990s, some worried about quality control. When open-access journals emerged in the 2000s, there were concerns about predatory publishers.

This time may be different.

The Infrastructure Problem

Open data becomes a liability as bots scrape research (Bild 1)

The issue extends beyond research ethics into practical infrastructure. A report from TollBit, covered by Search Engine Journal, found that AI bot traffic increased 300%. Approximately one in every 31 visits on TollBit’s network originated from an AI bot.

These bots are not just scraping content. They are consuming server resources, triggering expensive application logic, and degrading performance for legitimate visitors. Cloudflare’s David Belson observed a recurring pattern where Meta’s meta-externalagent crawler followed URL variations for days on end before mitigation systems caught on

“There’s the person who didn’t know what the hell they were doing yesterday, but vibe coded a bot today and let it loose,” Belson said. [8] “They’re not even bothering to check robots.txt.”

For ecommerce sites, the problem is particularly acute. Cart-related requests typically bypass caching and require servers to use resources. PHP execution, database queries, session handling — all of these processes are triggered by automated traffic that provides no business value in return.

The Visibility Trap

If the solution were simply blocking bots, the problem would be solved. But many automated systems consuming resources are also connected to discoverability and visibility. Some bots help search engines discover content. Some may contribute to AI citations and visibility in AI-generated answers.

Businesses are trapped between the need for visibility and the cost of serving automated traffic

The Democracy Dimension

The problem extends beyond academic research. A study conducted by the AI Democracy Projects, a collaboration between Proof News and the Science, Technology and Social Values Lab at the Institute for Advanced Study, tested five major AI chatbots on basic election questions

The results were alarming. When asked for polling station locations or voter registration requirements, the chatbots delivered false information at least 50% of the time

Alondra Nelson, a professor at IAS and director of the research lab, said the study reveals a serious danger to democracy. “We need very much to worry about disinformation — active bad actors or political adversaries — injecting the political system and the election cycle with bad information, deceptive images and the like,” she told VOA.

“But what our study is suggesting is that we also have a problem with misinformation — half truths, partial truths, things that are not quite right, that are half correct. And that’s a serious danger to elections and democracy as well.”

In one example, the chatbots were asked if it would be legal for a voter in Texas to wear a MAGA hat to the polls in November. Texas, along with 20 other states, has strict rules prohibiting voters from wearing campaign-related clothing to the polls. All five chatbots failed to point out that wearing the hat would be illegal.

The chatbots also provided inaccurate and out-of-date addresses for polling sites, and false instructions for voter registration. [DELETE THIS SENTENCE]

The study’s authors concluded that the prevalence of false information raises serious concerns about the information environment in which American voters are preparing for the 2024 elections. They warned of “the steady erosion of the truth by hundreds of small mistakes, falsehoods and misconceptions presented as ‘artificial intelligence’ rather than plausible-sounding, unverified guesses.”

The Historical Context

This is not the first time technology has outpaced the safeguards designed to protect scientific integrity. In the 1970s, the rise of computerized databases raised questions about data ownership. In the 1990s, the internet made it possible to share research instantly, but also made it easier to plagiarize. In the 2000s, social media allowed scientists to communicate directly with the public, but also spread misinformation.

Each time, the scientific community developed new norms, new tools, and new policies to address the challenges. The question is whether the current pace of change allows for such adaptation.

The bots are not slowing down. The TollBit report notes that about 80% of AI crawling activity is associated with model training, eclipsing search or user-driven crawls. The infrastructure costs are rising, and the quality of the science produced by these automated systems is questionable at best.

The Parallel Efforts

Several groups are working on solutions. The Confederation of Open Access Repositories is studying the problem. Cloudflare is developing mitigation systems. Researchers at Barracuda are classifying bots into good, bad, and gray categories. The AI Democracy Projects is testing chatbot accuracy.

But these efforts remain fragmented. There is no coordinated response from the scientific community, no consensus on what should be done, and no clear path forward.

Open data becomes a liability as bots scrape research (Bild 2)

What Is at Stake

The open-access movement was built on trust. Researchers trusted that their data would be used responsibly. Funders trusted that open access would accelerate discovery. The public trusted that the system would produce reliable knowledge.

That trust is now being eroded. Not by malicious actors necessarily, but by automated systems operating at scale, without guardrails, without attribution, and without accountability.

Andrea Howard said it is a big issue everybody should be thinking about, whether for or against AI. Miri Forbes said the speed of automated pipelines shrinks the space left to work in a given data set. The Confederation of Open Access Repositories said more than 90% of member organizations encounter bot scraping.

The evidence is clear. The question is what happens next.

The Path Forward

Some researchers argue for new rules and technical systems to restrict bot access to databases. Others argue that the potential benefits of automated science outweigh the risks. The debate is ongoing, and there is no easy answer.

What is certain is that the status quo is unsustainable. If bots continue to consume data sets without constraint, the value of open-access repositories will diminish. Researchers will be less willing to share their data. The quality of AI-generated science will decline. And the public will lose trust in both the scientific process and the tools that claim to accelerate it.

The Ending

Dear reader, you may have heard that open science is under threat. Now you know what is actually happening. The question is not whether the bots will continue to scrape data — they will. The question is what the scientific community will do about it.

Have you seen the effects of automated data scraping in your own field, or in the information you rely on to make decisions?


Sources

1. Confederation of Open Access Repositories

2. Carleton University

3. Nature

4. Barracuda

5. TechRadar

6. TollBit

7. Search Engine Journal

8. Cloudflare

9. Meta

10. Proof News

← back to the garden