Science Should Be Open, For LLMs Too
Few principles unite researchers as strongly as openness in science. It has become nearly tautological, right up there with “politicians against corruption” or “companies for innovation.” Yet here, too, support is often stronger in principle than in practice. Researchers like the idea of openness in science more than the reality, which involves handling too many data requests or having people replicate your studies only to find you reported something wrong.
Still, one part of the open-science agenda struck me as largely uncontroversial until recently: that scientific knowledge should be freely accessible and widely disseminated. Maybe that’s why I found it curious that so many of my peers opposed ACM’s proposal of giving LLMs access to its digital library. In this essay, I argue that restricting LLM access to scientific content is a remarkably ineffective way to address the legitimate concerns people have about AI. On the contrary, it is likely to make several of those problems worse.
Responses to ACM’s proposal ranged from “of course” to “this is absurd,” illustrating how polarized discussion of LLMs has become, particularly within computer science. As a scientist who studies LLM use and an arXiv moderator who spends a couple of hours a week sifting through (too often sloppy) submissions, I find that nuance is often the first casualty in debates around the impacts of AI. LLMs are not the solution to all our problems; indeed, they create plenty of new ones. But neither are they useless. They are, in fact, a little bit magical, as transformative technologies often are. More importantly, beyond evil and good, it is increasingly difficult to ignore that LLMs are already deeply embedded in the workflows of many scientists, including myself.
There are really two ways of thinking about ACM’s proposal. One view is that scientific institutions should resist the adoption of LLMs as much as they can. Another is that these systems are already part of scientific practice, and that our job is to shape how they are used rather than wish them away. I am skeptical that wholesale rejection of generative AI is either feasible or desirable. The incentives to use AI for productivity increase are too strong, and blanket prohibitions are difficult to enforce and likely to push use underground. But perhaps most importantly, such a stance treats a broad class of tools as inherently illegitimate rather than distinguishing between better and worse uses. With only a shovel, you are far more likely to succeed by redirecting a river than by trying to dam it.
So, from here on, let us assume that LLMs are part of science and that the relevant policy task is not to decide whether they should exist, but to balance their potential benefits against their potential harms. In that context, rather than making a general case for LLMs, I want to organize the rest of this essay around the strongest potential harms associated with their use. As I see it, these are threefold: LLMs may 1) hallucinate and can enable bad scholarship; 2) alter the incentives for producing information in a bad way; and 3) concentrate power in a small number of AI companies. These are real concerns, but my argument is that restricting LLM access to credible scientific literature does little to solve most of them, and typically makes them worse!
Problem 1: LLMs Hallucinate and Can Lead to Bad Scholarship
The most immediate objection is linked to one of the big threats of language models: that they hallucinate. There’s no guarantee that the content they produce is faithful to the source material they are citing. LLMs can fabricate citations and misrepresent results, which runs counter to the scientific ideals of precision and proper attribution. Thus, the proliferation of seemingly correct yet false AI-generated content is a serious problem stemming from the widespread adoption of language models. However, I’d also argue that providing LLMs with access to authoritative sources helps mitigate this issue. Many recent advances in reducing hallucinations rely on grounding model outputs in retrieved external sources rather than on models generating answers solely from their internal parameters (see here, here, and here).
It is also a mistake, in my view, to treat the fact that models hallucinate (and perhaps always will) as an impediment to their use. The world of engineering is replete with reliable systems built from imperfect components. Rather than demanding infallibility, reliability emerges from a continuous effort to understand and mitigate errors. Even then, systems are never perfect, and disasters are averted by adapting their use to their expected failure modes. Airplanes have redundant systems because individual components may fail; medical tests are interpreted in light of their sensitivity and specificity; computer networks are designed to account for failures at all levels of their protocol stack. The scientific endeavor is itself built around fallible instruments, fallible processes, and fallible people. Focusing narrowly on the fallibility of these models distracts us from more interesting questions: Can their mistakes be characterized, reduced, and caught? Do these tools improve scientific work when used within their limits?
A reasonable counterargument here takes a step back and looks at the bigger picture: providing access to the library may legitimize their use in contexts where they ought not be trusted. It may be that citation-grounded systems become sufficiently convincing such that researchers become overreliant on them. While I partially agree with this take, I’m not particularly convinced due to how unobservable and hard to deter LLM use is. The situation is somewhat analogous to abstinence-only approaches to STIs: when a behavior is already common and hard to eliminate, refusing to make it “safer” may simply increase the harms associated with it.
The analogy is imperfect, however, because some uses of LLMs are potentially beneficial to science. Most relevant to ACMs’ call for commentaries, I strongly believe that LLMs can be extremely helpful for searching large amounts of related literature and finding very interesting connections that are not easily surfaced by semantic search (I should write more about my adventures with Claude and Zotero in a subsequent post). But if you want to check for an example in another discipline, consider Mathematics. Many early groundbreaking math results obtained by the models came from literature discovery. Terence Tao has described the literature review as among the most productive near-term uses of AI in mathematics, noting that six Erdős problems had been reclassified from open to solved through such AI-assisted investigations.
Over the longer run, as models and harnesses improve, I expect the benefits described above to grow while the risks associated with hallucination decline, even though they will likely never disappear.
Problem 2: LLMs Change the Incentive Structure for Producing Information
Another real problem with the proliferation of LLMs is that they drastically alter the incentives of our information ecosystem. This does not happen in a monolithic way, as different parts of our knowledge ecosystem rely on different incentive structures. Journalists are rightly concerned that if people increasingly consume news through ChatGPT and newspapers receive neither traffic nor revenue, then financing news reporting will become progressively harder. On Wikipedia, the challenge is almost the opposite: the cost of producing content has decreased dramatically, while the cost of verifying it has increased (hence initiatives like the AI Cleanup work group). Community-driven Q&A websites like StackOverflow face yet another challenge: if programmers increasingly ask LLMs instead of asking questions, there will be fewer and fewer people participating in their Q&A community. And this has already happened: Stack Overflow is dying.
I believe that studying the way LLMs reshape our information ecosystem (and how we can design around their flaws) is one of the key questions of our age. But at the same time, increasing LLMs’ access to reliable scientific information is a force for good in this story, not for ill. Unlike the journalism scenario, scientists are already the suckers. We conduct and review the research, and often edit journals almost entirely for “free” (or more precisely, with taxpayers’ money). Yet, publishers routinely charge authors, universities, and readers exorbitant amounts to make all that research openly accessible, all while operating one of the highest-margin businesses in the world.
Thus, restricting LLMs’ access to scientific repositories does little to improve the incentives that actually motivate scientists. On the contrary, researchers want recognition: citations, impact, and influence. If scientific work is increasingly conducted with the aid of LLMs, then reducing the visibility of research to those systems will lead researchers to devalue venues that do not allow their work to be discovered and used through language models. Many of us already post our work on preprint servers such as arXiv because rapid dissemination is valuable. If AI-assisted scientific workflows continue to grow, restricting LLM access to peer-reviewed repositories may similarly diminish the value of publishing in those venues.
And here I’m going to waste some ink making a distinction between a broad claim that I am not making and the narrower claim that I am. I am not arguing that LLMs are an unambiguous good for science. They do introduce new problems and distort incentives. My claim, instead, is a conditional one: if LLMs are becoming central to scientific work (and I believe they already are), then making them better able to interact with the scientific literature is preferable to making them less knowledgeable.
Problem 3: AI Companies Have Too Much Power
Beyond objections to LLMs as technology, scholars (and the public) have increasingly criticized what my Princeton colleague Janet Vertesi and her co-authors described as “the political project of AI.” Their argument is that AI is a world-building project through which organizations are drastically expanding their power. And, indeed, this project has generated substantial public resistance. Creative workers face intense scrutiny when incorporating generative AI into their work; communities are mobilizing against the construction of energy-hungry data centers; and educators are increasingly questioning who benefits from the deployment of these systems.
I share these concerns. But I believe an important mistake occurs when opposition to this political project becomes indistinguishable from denial of the technology’s usefulness (as I have written previously). The worth and potential of AI as a technology is not wholly determined by whether you think the project of AI was misguided. Indeed, science is filled with transformative technology created in the context of shitty political projects: from rocketry to disease surveillance campaigns. It is fine that I believe the social hygiene movement was atrocious and at the same time believe that disease surveillance can be extraordinarily valuable. Condemning the political project in which a technology developed does not require denying the value of the technology itself. Indeed, if it did, we would have to discard much of the scientific and technical inheritance of human history.
And, to be clear, my point here is not that political origins do not matter but that technologies can be detached from the projects that first organized them and placed in the service of other values. Disease surveillance did not have to remain an instrument of eugenics. Public health could instead be constructed around universal welfare and health as a human right. But these transformations are not automatic; they instead happen because people build alternative institutions.
In my view, then, opposition to the political project of AI cannot stop at wholesale rejection. It must involve the construction of alternative technological projects, like open-weight models. And this is where access to information, like the one ACM is debating about providing, becomes crucial. OpenAI, Anthropic, and Google already have access to every damn PDF in ACM’s digital library. At this rate, they may already have access to every damn PDF ever uploaded on the Web. But you know who wouldn’t have access? You! If you were developing your own agent, your own model, your own harness.
Restricting machine access to science, thus, risks producing exactly the political outcome that critics of AI should fear: a world in which only a handful of corporations can build systems grounded in scientific records.
The history of science is, in some way, a history of making knowledge easier to build upon. Private correspondence gave way to journals. Libraries replaced personal collections. The internet made scientific work accessible beyond the institutions in which it was produced. If LLMs are becoming part of the infrastructure through which science is conducted, then they, too, should be able to engage with the scientific record. Denying them such ability makes AI systems less reliable, weakens the visibility of scientific work, and further entrenches the handful of organizations that already possess enormous proprietary corpora.

