By Ernesto Spinak
Introduction

Image: Nathan Dumlao via Unsplash
The traditional model of the scientific article, conceived centuries ago to accelerate the exchange of results, appears to have come to an end. We are currently in its survival phase, in which the validity of the immutable PDF document as a “black box” (which some consider a “fossil”) is being questioned, since its rigidity facilitates fraud by allowing manipulated data or inserted metadata to be concealed.
A possible transition points toward science as a continuous and constantly evolving process, where knowledge is not condensed into a single product but rather into open research notebooks (such as Jupyter Noteboks). Here, code and data coexist in an auditable manner, ensuring transparency by documenting who made a decision, when it was made, and what information was used.
Current status: integrity, ghost citations, hidden references
One of the threats is the manipulation of metadata through so-called “sneaked references”. These are citations that appear neither in the body of the article nor in its visible bibliography.
The first study on the topic, published on arXiv last year, Detection of metadata manipulations: Finding sneaked references in the scholarly literature 1,demonstrates that the global scientific communication infrastructure has a vulnerability whereby “it is possible to artificially manipulate scientific impact by inserting nonexistent citations into an article’s metadata without altering the published document, thereby compromising the reliability of international citation-based evaluation systems. The results are extraordinarily revealing. An analysis of 4,077 documents identified more than 80,000 artificially inserted fraudulent citations, distributed across 2,787 manipulated articles.”
The main purpose of these phantom citations is to artificially inflate citation-based performance indicators, such as the h-index, journal Impact Factors, and Field-Weighted Citation Impact (FWCI). Since these metrics form the basis of prestigious global rankings (such as the Shanghai Ranking or Clarivate’s lists of highly cited researchers), fraud ultimately undermines the reliability of the entire academic evaluation system.
Ghost citations are inserted directly into the metadata of articles registered with agencies such as Crossref, causing them to spread across major scientometrics platforms. As a result, these frauds have forced infrastructure institutions to take drastic measures, such as permanently revoking Crossref membership and, more importantly, triggering a “confidence crisis”, s pointed out in an editorial published by ACS Nano in early 2026, Peer Review and AI: Your (Human) Opinion Is What Matters 2,which details how AI-generated evaluations are homogenizing science, destroying critical thinking, and creating “false efficiencies” that overwhelm human editors).
This phenomenon exposes an “ontological crisis” in science, in which the “publish or perish” incentive system leads publishers to prioritize volume and impact at the expense of the soundness of the discovery. (See the editorial entitled More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review 3, which analyzes the behavior of scientific journals following the mass adoption of LLMs and concludes that AI, combined with traditional incentives from funding agencies and transformative agreements, is driving academia to produce more volume rather than better research, thereby eroding standards of readability and conceptual depth.)
Consequently, the saturation of the system through the use of AI to generate fraudulent articles and citations threatens to create a “closed loop” in which some computers write and others review, stripping the scientific validation ecosystem of empirical rigor and human intervention. This makes it difficult for human researchers to distinguish between legitimate science and “sophisticated junk” based on fabricated data.
What if we moved toward the complete automation of science?
The risk of full automation in science is not merely a technological shift but is described in the sources as a response to the ontological crisis affecting the very essence of knowledge and its validation. The main risks mentioned are as follows:
- Loss of rigor: Full automation could lead to the approval of documents lacking actual research, integrity, or minimum scientific standards.
- Lack of human intervention: By eliminating the expert intervention of human researchers, the scientific validation ecosystem is deprived of empirical content.
- Atrophy of critical thinking: The act of writing a paper is not a mere administrative formality, but a way of thinking. The struggle to articulate a discovery forces scientists to detect methodological inconsistencies and gaps in their reasoning.
- System saturation: So-called “paper mills” use AI to generate flawless synthetic images, graphics, and data, creating coherent texts at nearly zero cost that are difficult for human reviewers to identify as fraud. In Kobak et al., Delving into LLM-assisted writing in biomedical publications through excess vocabulary 4, published in Science Advances, the study analyzed more than 15 million abstracts in PubMed through the end of 2024, finding that at least 13.5% of biomedical articles exhibit unmistakable linguistic traits of having been processed or written by LLMs.
- The automation of the scientific process tends to standardize science, eliminating the diversity of epistemological concepts that enrich it.
- Complacency (sycophancy): There is a risk that evaluation algorithms will tend toward complacency, validating only ideas that fit pre-established statistical patterns and marginalizing disruptive or counterintuitive ideas.
A Radical Shift in the Global Narrative Toward Minimum Verifiable Units
The present may be paving the way for the future by proposing a radical shift in how science is evaluated. Instead of reviewing an article as an indivisible narrative block, there are platforms such as OpenEval (which propose breaking down each article into “minimum verifiable units” or concrete claims).
- Detailed evaluation: AI can automatically extract these claims and classify them according to the evidence supporting them (experimental data, previous references, or statistical inferences).
- Efficiency: In tests conducted with the journal eLife, the AI system was able to evaluate 93% of the claims in a manuscript, surpassing the 68% coverage achieved, on average, by human reviewers, as noted in the article The next unit of science: Is the scientific paper due to be replaced?.5
The Homogenization of Science by AI
While there are advantages, there is an existential risk associated with total automation. We are witnessing the emergence of a “closed loop” in which AI writes and AI reviews.
- System saturation: Tools like Google’s Paper Orchestra [5] can transform lab notes into a complete manuscript in 40 minutes. This has led to as many as 21% of peer reviews at computer science conferences being generated entirely by LLMs, as pointed out in the preprint PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing 6
- The human author is rendered invisible: The widespread use of these tools tends to homogenize language, creating a world where AI speaks “for us,” but not necessarily “like us,” eroding creativity and epistemic diversity.
- Techno-scientific reductionism: By attempting to automate evaluation through fragmentation into “minimum verifiable units,” there is a risk of adopting a reductionist approach. While useful for verifying isolated data, this method may fail to capture complex synthesis, holistic theorization, and epistemic discussion, especially in fields such as the humanities and social sciences.
One reminder—which I hope is not merely nostalgic—is the adage that “to write is to think.” The difficulty of drafting an article forces the researcher to confront gaps in their reasoning and detect inconsistencies they did not notice during the experiment. By automating the writing process, there is a risk of eliminating this “work of reflection,” which is essential for deep understanding.
OpenEval for evaluating documents and other forms of forensic analysis
There are several editorial tools available to detect and prevent many of the types of fraud mentioned. In this section, we provide details about OpenEval and will include technical details on other procedures in the Annex at the end of this post.
OpenEval proposes that scientific journals separate two functions that are currently intertwined in articles:
- Reporting of results: these should be published in a structured, machine-readable format.
- Communication of ideas: the traditional narrative (the PDF) would become simply an interpretive layer over this structured, searchable database.
Proven Effectiveness
In a large-scale test using the eLife journal corpus (approximately 16,000 articles), the system delivered remarkable results.
- Mass extraction: It identified nearly 2 million individual statements.
- Accuracy: The AI and human reviewers agreed in 81% of cases.
- Coverage: OpenEval assessed 93% of the statements in the manuscripts, significantly surpassing the 68% coverage achieved, on average, by human reviewers due to time constraints and the complexity of the texts.
“OpenEval”-type platforms propose a revolution in scientific evaluation by replacing the overall review of an article with an automated analysis of each individual statement contained in the document, demonstrating that artificial intelligence can actively participate in the validation of scientific knowledge and challenging the traditional peer-review model that has dominated science for centuries, as demonstrated in the paper Science should be machine-readable7.
The Role of Open Access and Preprint Servers
Open Access emerges in the literature as one of the most powerful tools of structural resistance for dismantling ghost citation fraud and other problems introduced by AI, as well as the oligopoly of journal publishers with article processing charges (APCs) included. The role of Open Access is not merely a form of distribution, but also one of auditing and shifting incentives.
In the traditional model, an article is reviewed by two or three reviewers, usually in a process that lacks transparency. In contrast, Open Access allows thousands of scientists to review the document simultaneously, which makes it possible to:
- Anomaly detection: Platforms such as PubPeer allow the community to publicly comment on and report fraudulent patterns, such as unexplained excessive citations or inconsistent metadata, which has proven to be far more effective than the traditional closed-review process.
- Post-publication analysis: By making the document openly available, the chances of detecting statistical manipulation or inserted references increase exponentially.
By breaking away from the “black box” model and the concept of final, immutable products—such as commercially published PDFs—the Open Access model drives the transition toward “Process Science,” in which not only the text but also the research context, code, and unprocessed data are published. Furthermore, if the process is open and auditable in real time, the point of fabricating citations or manipulating metadata disappears, since any AI algorithm or researcher can verify the traceability of the evidence.
The preprint server system for open-access articles changes the economic model and incentives, working against “paper mills”, since commercial journals that charge high article processing charges (APCs) currently have a financial incentive to accept a large volume of articles, which loosens quality filters and opens the door to industrialized fraud.
True open access (in which neither the author nor the reader pays—the Diamond model) eliminates commercial pressure for volume, allowing for stricter ethical standards and reducing the appeal of “citation gaming” tactics used to inflate market metrics.
We might also add that the traditional model rarely publishes studies with unsuccessful results, as they do not generate citations. Open Access facilitates the publication of these results, which is vital for training AI models with honest, unbiased data, thereby avoiding the need for researchers to “fudge” or inflate citations to make their findings appear more successful than they actually are.
Open Access promotes interoperability by publishing in formats such as XML-JATS, which transform the article into a computational object. This allows machine-readable knowledge networks to automatically collect and verify references, immediately detecting whether a citation in the metadata lacks a corresponding reference in the actual text.
Finally, we can observe that Open Access facilitates automated technical auditing. For forensic tools such as GROBID or cross-referencing algorithms to function, access to the full text of articles is essential.
Validation cannot be a one-time “approved/rejected” event prior to publication. Preprint platforms point the way: peer review should be a public, post-publication, community-driven, and evolving dialogue. AI can act as a front-line assistant (checking statistical consistency, formatting, and verifying the bibliography with forensic tools), but the assessment of heuristic and conceptual value must be carried out by humans and the community of stakeholders.
Conclusion
The future of the scientific article lies not only in improving fraud detection algorithms, but in a paradigm shift:
Moving from “producing text to be indexed” to “producing verifiable knowledge to be shared”
The architecture that replaces the current model must strike a balance between algorithmic efficiency and the human reflection that has defined science for centuries.
In short, the total automation of the scientific process threatens to transform the scientific article from a vehicle for discovery into a mere bureaucratic bargaining chip, where the volume of publications takes precedence over the soundness and veracity of the findings.
Finally, preprints and Open Access decentralize epistemic authority, stripping publishers of the power to decide what constitutes “valid science” based on opaque metrics, returning validation to the global scientific community, and transforming it into an exercise in real-time transparency—one capable of withstanding the pace of automation and incorporating it into Science in Progress.
Annex – Some forensic review methods
The editorial tools currently used to detect and prevent fraud are:
- XML/PDF analyzers
- Automated cross-referencing
- Performing an intersection operation between sets of citations
$$\text{Anomalies} = \text{References}_{\text{Crossref}} \setminus\text{References}_{\text{PDF}}$$
$$\text{References}_{\text{Crossref}}
it is the set of all references that actually appear at the end of the text in the article’s PDF file.
$$\text{References}_{\text{PDF}}
It is the set of all the references that actually appear at the end of the text in the article’s PDF file.
Operation ($\setminus)
The difference operator means: “Take everything in the first set and subtract what also appears in the second.” In other words, only the elements that belong to the first set but not to the second remain.
Therefore, the set $\text{Anomalies}$ yields the following result: All references that the publisher registered with Crossref but that do not physically exist in the PDF text.
In the context of scientific integrity analysis or metadata auditing, this result reveals serious discrepancies, such as:
- Ghost citations: references that have been officially indexed (which inflates the citation counts of other authors in the databases), but which the reader will never see in the actual article.
- Production errors: mistakes made by the publisher when exporting XML metadata to Crossref, or last-minute cuts to the PDF text that were not updated in the official record.
- Extraction biases: If the list on the right was obtained using text extraction software (such as GROBID or ParsCit), a result here could also indicate that the extractor failed to read certain pages of the PDF.
- If the resulting set is not empty, this means that the publisher has inserted metadata that is invisible to the reader of the article.
- Network Analysis and Statistical Anomalies (Graphs)
To understand how fraud involving infiltrated references is detected, the process is divided into two technological phases:
- extraction and structuring of the actual content of the PDF (notably using GROBID: GeneRation Of Bibliographic Data),
- application of forensic detection algorithms that cross-reference this data with official metadata from agencies such as Crossref.
- Extraction and Automated Classification
The system uses advanced natural language processing to automatically extract each statement from the text. Once extracted, it classifies them according to the type of evidence supporting them
- Direct experimental data.
- Previous bibliographic references.
- Statistical inferences.
- Speculative interpretations.
- Evidence Assessment
Once classified, artificial intelligence assesses whether the evidence presented actually supports the claim made. This makes it possible to identify hidden connections; for example, the system was able to detect studies that reached complementary conclusions but did not cite one another because they belonged to different subfields.
Notes
1. BESANÇON, L.; CABANAC, G.; LABBÉ, C.; MAGAZINOV, A.; DI SCALA, J.; TKACZYK, D.; WEBER-BOER, K. Detection of metadata manipulations: Finding sneaked references in the scholarly literature. arXiv [online]. 2025 [viewed 20 July 2026]. https://doi.org/10.48550/arXiv.2501.03771. Available from: https://arxiv.org/abs/2501.03771↩
2. BURIAK, J. M. et al. Peer Review and AI: Your (Human) Opinion Is What Matters. ACS Nano [online]. 2026, vol. 20, no. 4, [viewed 20 July 2026]. Available from: https://pubs.acs.org/doi/10.1021/acsnano.6c00490↩
3. GARTENBERG, C. et al. More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review. Organization Science [online]. 2026, vol. 37, no. 3, pp. 795–812, ISSN: 1047-7039 [viewed 20 July 2026]. https://doi.org/10.1287/orsc.2026.ed.v37.n3. Available from: https://pubsonline.informs.org/doi/10.1287/orsc.2026.ed.v37.n3↩
4. KOBAK, D. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances [online]. 2025, vol. 11, no. 27, pp. eadt3813, ISSN: 2375-2548 [viewed 20 July 2026]. https://doi.org/10.1126/sciadv.adt3813. Available from: https://www.science.org/doi/10.1126/sciadv.adt3813↩
5. The next unit of science: Is the scientific paper due to be replaced? [online]. The Transmitter: Neuroscience News and Perspectives. 2026 [viewed 20 July 2026]. Available from: https://www.thetransmitter.org/from-bench-to-bot/the-next-unit-of-science-is-the-scientific-paper-due-to-be-replaced/↩
6. SONG, Y. PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing. arXiv [online]. 2026. [viewed 20 July 2026]. https://doi.org/10.48550/arXiv.2604.05018. Available from: https://arxiv.org/abs/2604.05018↩
7. BOOESHAGHI, A. S. Science should be machine-readable. bioRxiv [online]. 2026. [viewed 20 July 2026]. https://doi.org/10.64898/2026.01.30.702911. Available from: https://www.biorxiv.org/content/10.64898/2026.01.30.702911v1.full↩
References
BESANÇON, L.; CABANAC, G.; LABBÉ, C.; MAGAZINOV, A.; DI SCALA, J.; TKACZYK, D.; WEBER-BOER, K. Detection of metadata manipulations: Finding sneaked references in the scholarly literature. arXiv [online]. 2025 [viewed 20 July 2026]. https://doi.org/10.48550/arXiv.2501.03771. Available from: https://arxiv.org/abs/2501.03771
BOOESHAGHI, A. S. Science should be machine-readable. bioRxiv [online]. 2026. [viewed 20 July 2026]. https://doi.org/10.64898/2026.01.30.702911. Available from: https://www.biorxiv.org/content/10.64898/2026.01.30.702911v1.full
BURIAK, J. M. et al. Peer Review and AI: Your (Human) Opinion Is What Matters. ACS Nano [online]. 2026, vol. 20, no. 4, [viewed 20 July 2026].Available from: https://pubs.acs.org/doi/10.1021/acsnano.6c00490
GARTENBERG, C. et al. More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review. Organization Science [online]. 2026, vol. 37, no. 3, pp. 795–812, ISSN: 1047-7039 [viewed 20 July 2026]. https://doi.org/10.1287/orsc.2026.ed.v37.n3. Available from: https://pubsonline.informs.org/doi/10.1287/orsc.2026.ed.v37.n3
KOBAK, D. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances [online]. 2025, vol. 11, no. 27, pp. eadt3813, ISSN: 2375-2548 [viewed 20 July 2026]. https://doi.org/10.1126/sciadv.adt3813. Available from: https://www.science.org/doi/10.1126/sciadv.adt3813
SONG, Y. PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing. arXiv [online]. 2026. [viewed 20 July 2026]. https://doi.org/10.48550/arXiv.2604.05018. Available from: https://arxiv.org/abs/2604.05018
The next unit of science: Is the scientific paper due to be replaced? [online]. The Transmitter: Neuroscience News and Perspectives. 2026 [viewed 20 July 2026]. Available from: https://www.thetransmitter.org/from-bench-to-bot/the-next-unit-of-science-is-the-scientific-paper-due-to-be-replaced/
This post was prepared with the assistance of Gemini 3.1, Perplixity, and NotebokLM
About Ernesto Spinak
SciELO collaborator, Systems Engineer and Librarian, holding a Diploma of Advanced Studies and a Master’s degree in Information Society from the Universitat Oberta de Catalunya, Barcelona, Spain. He currently runs a consulting company that serves 14 government institutions and universities in Uruguay with information projects.
External Links
Translated from the original Spanish by Lilian Caló.
Como citar este post [ISO 690/2010]:















Recent Comments