Anantam IASPost · 25 July 2026

Research Integrity: Plagiarism, Paper Mills and the Ethics of AI-Written Authorship (UPSC Ethics — GS IV)

Study Notes · Education · Ethics, Integrity & Aptitude · General Studies · Governance · GS IV · Science & Tech

Research misconduct in its textbook form — fabrication, falsification, plagiarism — is rare and detectable. The practices that corrupt the scientific record at scale are legal, common, career-rewarded, and invisible to a similarity checker.

Research integrity is usually taught as a prohibition on cheating, which makes it sound like a discipline for the weak-willed and lets everyone else stop reading. The obligations are wider than that, and they run in four directions at once: to the record, which every later researcher will rely on; to co-authors, whose names carry the risk of one’s own carelessness; to the human beings who supplied data or bodies; and to whoever paid, which for most Indian research is the public.

That framing matters because the textbook offences are the rare ones. Outright fabrication is uncommon and, when found, career-ending. The practices that actually degrade the published record are legal, widely taught by example, rewarded in every promotion cycle, and entirely invisible to a similarity-detection tool. A department can hold a perfect plagiarism record and still be publishing conclusions that nobody will ever reproduce.

What the Obligations Actually Are

To the record. A published paper is an instruction to other people about where to spend their next three years. A wrong result that survives in the literature does not merely fail to help; it actively consumes the time of everyone who builds on it. This is the obligation that most cleanly separates research ethics from ordinary honesty, and it is why an uncorrected error is a live harm rather than a closed episode.

To co-authors. Authorship transfers risk. Every name on a paper is warranting something about work most of them did not perform, which is why the accountability criterion in authorship standards is not a formality.

To subjects. People who consented to be studied consented to a specific use. Data collected under one protocol and mined for another has been taken rather than given.

To the public. Publicly funded work carries a duty to publish, to publish where the finding can be reached, and to report what was found rather than what was hoped for. The three duties are frequently in tension with the incentives an academic actually faces, which is the theme running through most of what follows and through the wider account of integrity as a settled disposition rather than a set of rules.

The Taxonomy, and the Grey Zone That Does More Damage

The standard definition of research misconduct is narrow and deliberately so: fabrication (inventing data or results), falsification (manipulating materials, equipment, processes, or altering or omitting data so the record misrepresents the research) and plagiarism (appropriating another’s ideas, processes, results or words without credit). Honest error and differences of interpretation are excluded, which is correct — a definition that captured mistakes would make research impossible.

The larger problem sits just outside. Questionable research practices are not misconduct under any code and do more cumulative damage than all three offences together.

P-hacking is trying analyses until one reaches significance and reporting only that one. HARKing — hypothesising after the results are known — is presenting a hypothesis constructed from the data as though it had preceded them, which converts a fishing expedition into a confirmed prediction. Selective reporting drops the measures, conditions or subgroups that did not work. Salami slicing splits one study into the maximum number of publishable units. Honorary authorship adds names that contributed nothing; ghost authorship omits names that did, including industry writers. Citation manipulation covers coercive citation by reviewers and editors, and reciprocal arrangements between journals. Text recycling reuses one’s own prose, which is defensible in a methods section and dishonest in an introduction that implies novelty.

None of these requires a bad person. Each is what a rational academic does when evaluation counts outputs.

Table separating fabrication, falsification and plagiarism from questionable research practices, with what each does to the record and why detection fails
Misconduct is narrow and detectable; the grey zone is wide, legal and rewarded
Flow diagram of four questions to ask before submitting AI-assisted research writing, covering reasoning, citations, disclosure and responsibility
The four questions that separate a language tool from a substitute for thinking

What the Reproducibility Crisis Exposed

The evidence accumulated from several directions. John Ioannidis argued in 2005 that most published research findings are false, on the statistics of small studies, small effects and flexible analysis. Industrial laboratories reported that they could confirm only a small minority of published preclinical cancer results — one widely cited 2012 account put it at six of fifty-three landmark papers. The Open Science Collaboration’s 2015 replication of a hundred psychology experiments reproduced roughly a third of the original significant findings. A large 2016 survey of researchers found that most had failed at some point to reproduce another group’s work, and many had failed to reproduce their own.

The finding underneath all of this is not that scientists are dishonest. It is that the incentive structure selects for unreliable results. Journals prefer novel, positive, clean findings. Careers are built on publication counts, on journal prestige and on indices derived from citation, all of which reward volume and surprise over reliability. A researcher who runs a large, well-powered, careful study that finds nothing has produced better science and a worse curriculum vitae.

Two institutional responses are worth naming: the San Francisco Declaration on Research Assessment of 2012, which asks that journal-level metrics not be used to judge individual researchers, and the Leiden Manifesto of 2015, which sets principles for the use of quantitative indicators. Both are declarations rather than rules, and their adoption has been patchy.

At the bottom of the market sit predatory journals — outlets that charge a publication fee and provide no meaningful review. A 2013 sting in which a deliberately flawed paper was submitted to more than three hundred open-access journals saw it accepted by roughly half. Their existence is a direct consequence of count-based evaluation: where a line on a form is worth money, someone will sell lines.

Paper Mills

A paper mill is a business that sells authorship. It produces manuscripts to order — sometimes wholly fabricated, sometimes assembled by recombining figures and text across templates — and sells author slots to academics who need publications. It also sells its way through peer review, by supplying reviewer identities and email addresses so that the manuscript is refereed by the vendor.

What the exposure of these operations demonstrated is scale. A major publisher retracted several thousand papers from a special-issue programme in 2023 after systematic mill activity was identified, and detection has repeatedly come from volunteer sleuths spotting recycled images and templated phrasing rather than from journals’ own systems. Two conclusions follow. The first is that peer review, which is the whole basis for treating a published claim as credible, can be captured wholesale. The second is that no similarity checker catches this, because a mill-written paper is original text.

The cost of the real cases is not abstract. The retracted 1998 Lancet paper linking measles vaccination to autism led to a fall in vaccination coverage in several countries and, for its lead author, removal from the United Kingdom’s medical register in 2010. A physicist at a major American laboratory had a series of celebrated papers retracted in 2002 after an inquiry found data fabricated across them, and his doctorate was later revoked. A Dutch social psychologist resigned in 2011 with dozens of papers eventually retracted. In India, a chemistry professor at a state university had dozens of papers withdrawn or retracted around 2008 following findings of duplicate publication and unreliable data. In each case the collateral damage fell on students and junior co-authors who had done nothing wrong.

Who Counts as an Author

Authorship standards converge on a small set of conditions: a substantial contribution to conception, design, or the acquisition, analysis or interpretation of data; involvement in drafting or critically revising the work; approval of the final version; and agreement to be accountable for it. The fourth is the one that does the work. Accountability is the price of the credit.

The CRediT approach solves a different problem, by replacing the single undifferentiated label “author” with a list of named contributions — conceptualisation, methodology, software, formal analysis, data curation, writing, supervision, funding acquisition. It makes gift authorship harder because a gifted author has to be assigned a role they visibly did not perform.

The Indian version of the problem is specific and rarely stated plainly: the supervisor as automatic co-author. A research scholar depends on their guide for the thesis, the viva, the fellowship and the reference letter, and in that relationship there is no realistic way for a student to contest an authorship claim, or the ordering of names, or the addition of a departmental colleague. Coercive authorship of this kind is not covered by any misconduct definition and is not detectable by any tool. It is a straightforward abuse of an asymmetry of power, and the person paying for it has no route of complaint that does not end their own career — which is where this subject touches accountability and responsibility in its most concrete form.

India’s Institutional Response

The principal instrument is the University Grants Commission’s 2018 regulations on promotion of academic integrity and prevention of plagiarism in higher educational institutions. Their architecture is worth knowing precisely.

They mandate a similarity check for every thesis, dissertation and submitted manuscript, and require an undertaking from the student that the work is original. They establish a two-tier structure — a departmental academic integrity panel and an institutional one — with powers of inquiry and recommendation. And they grade penalties by percentage of similarity: up to ten per cent attracts no penalty; between ten and forty per cent requires a revised submission within a stipulated period; between forty and sixty per cent brings debarment from resubmission for a year; above sixty per cent means cancellation of registration. For faculty, the graded consequences run from withdrawal of the manuscript to denial of annual increments and a bar on supervising research students for two or three years. Quoted work with permission, references, bibliography and standard or common knowledge are excluded from the calculation.

Around this sit other instruments. Electronic deposit of theses in a national repository has been required since the 2009 doctoral regulations, which makes retrospective checking possible for the first time. A curated reference list of journals was introduced in 2018 to steer academics away from predatory outlets, although whitelisting as an approach is contested and its administration has changed over time.

AI-Written Text and the Question of Responsibility

The distinction that matters is not whether a language model was involved. It is whether the reasoning in the submission is the author’s. Using a tool to tighten a sentence, translate a paragraph, or correct grammar in a second language is on the same footing as a copy editor or a colleague who reads a draft. Generating an argument, a literature review or a discussion section and submitting it as one’s own analysis is passing off, whether or not any human text was copied.

Journal policy has converged on one point and moved rapidly on the others. The settled point is that a language model cannot be an author, because it cannot take responsibility for the work, cannot approve a final version, cannot declare a conflict of interest and cannot answer a query about the data. Authorship entails accountability, and accountability requires a person. Beyond that, positions have shifted within short periods: one leading journal moved from barring machine-generated text outright to permitting it with disclosure. Disclosure requirements — a statement of what tool was used and for what — are now common, and the specifics remain unsettled rather than fixed.

Two live failure modes deserve naming. Hallucinated citations are the sharpest: a plausible author, a plausible journal, a plausible year, and no such paper. The most instructive documented instance came from outside academia, when lawyers in a United States federal case filed a brief in 2023 containing fabricated judicial citations produced by a chatbot and were sanctioned for it. The mechanism transfers directly to a literature review, and the duty it creates is the same one examined under professional ethics and the duty to verify. The reviewer’s parallel duty is less discussed. A manuscript sent for review is confidential; pasting it into a commercial model is a disclosure of someone else’s unpublished work to a third party, quite apart from the question of whether a review generated without reading the paper is a review at all. At least one major funding agency has barred the use of generative tools by grant reviewers on precisely this reasoning.

Data, Consent and Correcting the Record

The consent architecture for research on human subjects rests on the Nuremberg Code of 1947, whose first principle is that voluntary consent is essential, and the World Medical Association’s Declaration of Helsinki of 1964 as amended since. India’s operative instruments are the Indian Council of Medical Research’s 2017 national ethical guidelines and the New Drugs and Clinical Trials Rules, 2019, which require registration of ethics committees and provide for compensation in the event of trial-related injury.

Three obligations recur. Prospective ethics-committee review, which cannot be obtained after the fact. Data sharing without exposing subjects, which is where open-science pressure and privacy obligations genuinely conflict; de-identification is weaker than it looks in small populations. And correcting the record, which is the most neglected duty in the entire field. A retraction is not a punishment and a correction is not an admission of dishonesty; both are maintenance of a shared resource. The related duty is to publish negative results, since a literature containing only findings that worked is a biased sample by construction, and every meta-analysis built on it inherits the bias.

The Honest Objections

Similarity percentage is a poor proxy for dishonesty. A thesis can score thirty per cent because its methods section describes a standard assay in the standard words, and eight per cent while containing a stolen central idea expressed freshly. A number that rises with the amount of standard terminology in a field cannot measure intent. Graded penalties tied to that number are administratively convenient and conceptually indefensible.

Detection tools misfire in predictable directions. They flag legal texts, statutory language, taxonomic names, common phrasing and boilerplate. They fail entirely on translated plagiarism, on paraphrase, on ideas taken without words, and on anything not in the comparison corpus — which includes most regional-language scholarship. They also produce false confidence in the opposite direction: a clean report is treated as a certificate of originality.

Punishing individuals while leaving the incentives untouched will not work. If promotion, increment and recruitment continue to count papers, and if a line on a form is worth money regardless of where it appeared, the market will keep supplying lines. Regulations that penalise the researcher and leave the evaluation system unchanged are addressing the last step of a chain they have declined to examine — the same structural point that makes examination fraud recur despite strict penalties for paper leaks.

The AI rules are genuinely in motion. The principle that a model cannot bear responsibility appears settled. Almost everything else — how much editing requires disclosure, where the line between assistance and substitution falls, whether detection tools may be used against students given their documented false-positive rate — is contested and moving. State that it is unsettled rather than reporting a rule that may have changed.

FAQ

What are the three recognised forms of research misconduct? Fabrication, falsification and plagiarism. Honest error and genuine differences of interpretation are deliberately excluded from the definition, because a standard that captured mistakes would make research impossible.

What are questionable research practices? Practices that fall outside the misconduct definition but distort the record: p-hacking, hypothesising after results are known, selective reporting, salami slicing, honorary and ghost authorship, citation manipulation and text recycling. They are more common and more damaging in aggregate than outright fraud.

What penalties do the UGC’s 2018 regulations prescribe? Penalties are graded by similarity percentage: no penalty up to ten per cent, revised submission between ten and forty, debarment from resubmission for a year between forty and sixty, and cancellation of registration above sixty. Faculty face withdrawal of manuscripts, denial of increments and bars on supervising students.

Can an AI system be listed as an author? Journal policies generally say no, on the ground that authorship requires accountability — approving the final version, declaring conflicts, and answering for the data. A system cannot do any of those. Disclosure of use is commonly required instead.

What is a paper mill? A business that produces manuscripts to order and sells authorship slots, often also supplying fake reviewer identities so its own papers pass peer review. Similarity checkers do not detect them, because the text is original.

Why is publishing negative results an ethical duty? Because a literature containing only successful findings is a biased sample. Every review and meta-analysis built on it inherits that bias, and other researchers waste resources repeating work that has already failed unpublished.

Practice Questions

Prelims MCQs

  1. Under the UGC’s 2018 academic integrity regulations, similarity above sixty per cent in a research submission attracts: (a) A warning (b) Resubmission within six months (c) Debarment from resubmission for one year (d) Cancellation of registration — Answer: (d) the four levels run from no penalty up to ten per cent through to cancellation of registration above sixty.
  2. HARKing refers to: (a) Splitting one study into several publications (b) Presenting a hypothesis formed after seeing the results as though it preceded them (c) Adding authors who did not contribute (d) Reusing one’s own published text — Answer: (b) it converts an exploratory finding into an apparently confirmed prediction.
  3. Which is excluded from the similarity calculation under the UGC regulations? (a) The introduction (b) The discussion of results (c) References, bibliography and standard or common knowledge (d) The abstract — Answer: (c) quoted work with permission and standard terminology are also excluded.
  4. The principal reason journal policies refuse to name a language model as an author is that: (a) Its output is not copyrightable (b) It cannot take responsibility for the work or approve the final version (c) Its contribution is always minor (d) It cannot be paid — Answer: (b) authorship standards require accountability, which requires a person.
  5. A paper mill is best described as: (a) A journal that charges a fee without reviewing submissions (b) A business that produces manuscripts to order and sells authorship slots (c) A software tool that detects textual overlap (d) A repository of electronically deposited theses — Answer: (b) predatory journals are a separate phenomenon, though the two often operate together.

Mains Practice Questions

  1. “Similarity percentage is a poor proxy for academic dishonesty.” Examine this criticism of the graded-penalty approach to plagiarism. (150 words)
  2. Questionable research practices do more cumulative damage to the scientific record than fabrication, falsification and plagiarism combined. Critically examine. (250 words)
  3. Discuss the ethics of using generative language tools in research writing, distinguishing assistance with expression from substitution for reasoning. (250 words)
  4. The supervisor’s claim to automatic co-authorship is an abuse of an asymmetry of power for which the affected party has no safe remedy. Discuss, and suggest institutional safeguards. (150 words)
  5. “Regulating the researcher while leaving the evaluation system unchanged addresses the last step of a chain nobody has examined.” Analyse with reference to academic integrity reform in India. (250 words)