Artificial intelligence and Open Science
Section outline
-
The problem for Open Science from AI
A paper from COAR (the Confederation of Open Access Repositories) articulates the problem very nicely in Navigating the Uneasy Interdependence of AI and Open Science:
Artificial intelligence (AI) is proving highly disruptive to open science as it becomes both widely used in the analysis and production of scholarly literature and deeply embedded in the public information commons. On the one hand, researchers across all domains are harnessing the power of AI and machine learning to do things previously unimaginable – such as rapidly processing massive datasets or synthesising large corpora in many different languages – greatly accelerating scientific progress and leading to new discoveries. On the other hand, AI is seriously challenging some of the fundamental assumptions on which open science rests, putting the open science ecosystem at risk in ways that demand urgent attention.
Open Science and AI aren’t merely complementary, they are structurally interdependent. Open resources are the raw materials for training AI models and for their application; while well-functioning AI tools are increasingly critical for conducting groundbreaking research. Yet AI, as a major consumer of open science outputs, also brings with it attribution problems, the potential for information contamination, and aggressive automated traffic that strains the very infrastructure on which it depends. Left unaddressed, these pressures threaten to reverse much of what the open science movement has achieved.
Generative AI allows the creation of content that looks like a real research paper, but is fabricated. Journals can be misled, and false research published. As noted by the Center for Open Science:
The is relevant to licencing of open publications. We discuss licences which grant permission for the use of copyrighted openly published material elsewhere in this course, but this system falls down where the AI models just hoover up whatever they can find without the attribution which is the usual component of the licence. Should the AI models be able to access copyrighted material, and should they attribute the work which their models then produce?
AI and scholarly publication
Stephen Downes, in a provocative paper The Present and Future Landscape of AI in Scholarly Publishing makes the point that the arrival of AI will have a substantial impact on scholarly publishing. While AI has benefits, reducing operational costs and democratising access to academic resources, there are a number of concerns including accuracy, research integrity, and ethical issues.
AI should even make us re-think the role of journal publications.
Should AI itself be 'open'?
Within the AI sector there is a vigorous debate about whether the underlying program code should be open source to allow developers and users to see and modify it. While much of the argument is commercial to allow the big private companies to maintain ownership and profit from their models, there are other arguments. The proponents of open quote democratisation of access, transparency, the ability to innovate, and ability to maintain privacy, while the opponents quote safety, misuse, and the fear of loss of control.