The science behind
the service.
Submedit’s editorial protocols are grounded in a formal evidentiary framework developed from the published literature on citation error, reporting inconsistency, distorted claim presentation, and the structural limitations of peer review. This page describes that framework, the architecture that applies it, and the research that informs both.
Two forms of intelligence, one system.
The structural limitation of conventional manuscript review is not a quality problem. It is an architecture problem. An editor reading a complex paper sequentially encounters an irreducible cognitive constraint: as the manuscript grows longer, the attention available for later sections is not equal to the attention applied to earlier ones. An editor scrutinising sentence construction on page 3 has finite resources remaining for verifying a citation on page 28.
This constraint cannot be resolved by engaging more experienced editors, and it is not peculiar to editing. When the same grant applications are scored independently by different qualified reviewers, agreement between them is close to absent8 — evidence that expert judgment applied without supporting structure carries variance that expertise alone does not remove. The problem is architectural, and it requires an architectural answer.
Orchestrated Intelligence™ is that answer. It is not automation with human oversight, and it is not expert review with machine assistance. It is a single system in which two forms of intelligence are each assigned the work the other cannot perform: exhaustive, position-independent analysis on one side, and irreplaceable domain judgment on the other.
The AI cannot judge what a finding means to its field. The expert cannot hold thirty-eight pages at uniform attention. Orchestrated together, each covers precisely what the other cannot.
A team of specialised AI agents processes the complete manuscript before any editor opens the file. Every citation is retrieved and tested against the claim it is attached to. Every numerical value is mapped against each downstream instance. Figures, captions, tables, and supplementary files are reconciled against the body text. The current literature is searched for work published since the draft was written. The analysis does not degrade with position in the document: page 38 receives the same treatment as page 1.
The output is not a summary. It is a systematically derived inventory of every candidate issue, each anchored to a specific location and classified by type and severity. This map is delivered to the human editors before their work begins, which changes the task they face: expert judgment is applied to a pre-structured set of verified targets rather than distributed across a long document in proportion to reading order.
The same framework governs assignment. It analyses the sub-field specificity, citation landscape, methodological context, and target journal of each manuscript, then matches it to available researchers actively publishing at that exact intersection. Determining who is genuinely qualified to edit a given manuscript — not merely who is available — is itself a problem that conventional editing services are not structured to solve.
Each layer is handled by the specialist it requires: a language editor, a sub-field scientist for structure and argument, and a reviewer-level specialist for claim calibration and submission strategy. A final quality-assurance pass then adjudicates every proposed correction against the manuscript itself before delivery, so that no change reaches the author unless it survives independent review.
The sections that follow describe what this architecture is applied to. Each is a distinct class of pre-submission failure, each is documented in the published literature, and each is verified under its own protocol. The purpose is not to duplicate peer review but to precede it: the scientific community invests on the order of 100 million hours in peer review each year,7 and every hour a reviewer spends reconciling a value against a figure is an hour not spent on the scientific evaluation that only they can perform.
One commitment governs the entire system: no client manuscript is used in the training, fine-tuning, or evaluation of any AI model, and no manuscript is exposed to third-party services outside the editorial process. This is binding under our Terms of Service and Privacy Policy.
What makes a citation valid?
A citation is not valid merely because the referenced paper is topically related to the claim, or because it reports findings in the same domain. Conceptual proximity, causal association, and thematic overlap — however reasonable they appear to the author — do not constitute validity in the sense a careful reviewer applies the term.
A citation is valid when the referenced source directly and explicitly supports the exact claim at the point of citation. Conceptual proximity is not sufficient.
Independent meta-analyses across 32,074 references establish that approximately 1 in 6 citations contains a significant error,1,2,3 with 38% of those errors representing systematic misrepresentation of what the source paper demonstrates rather than simple misquotation.5 The rate has not improved since it was first documented in 1984, and the pattern holds across fields and journal tiers.6
The equivocation error
The most frequently identified category of invalidity in our workflow is the equivocation error: evidence established for one construct cited to support a claim about a related but distinct construct. The mechanism is structural rather than negligent. An author reads a paper establishing finding X; later, writing about the related phenomenon Y, they cite that paper as support, reasoning that the two are close enough to be equivalent. A reviewer with deep familiarity with both constructs recognises the inference as unsupported. The citation does not do what it claims to do.
Primary versus secondary support
A citation is primary when the referenced paper presents the finding as its own original result, and secondary when it merely relays a conclusion drawn from a third paper. Secondary citations are not inherently invalid, but they are systematically more fragile: the intermediate paper may have mischaracterised the original, the original may have been superseded, and each additional hop introduces a new point of distortion. Simkin and Roychowdhury estimated, from statistical analysis of copying errors in bibliographies, that roughly 80% of authors do not read the papers they cite but copy citations from other reference lists4 — the mechanistic explanation for why error rates have proved so stable.
Forward and reverse analysis
Every citation–claim pair is assessed in both directions independently. Each direction provides a check on the other’s conclusion, and a citation that passes forward analysis but fails reverse analysis represents the systematic misrepresentation category rather than a simple inaccuracy.
Starting from the manuscript claim, we identify the source cited and ask whether that reference directly supports the claim. The referenced paper is retrieved and read in full — specifically the sections, results, and conclusions relevant to the cited assertion — and the support is assessed for whether it is direct, explicit, and accurately characterised.
Starting from the referenced paper, we ask what that paper actually concludes, then compare its conclusion against the claim it is cited to support. This direction identifies cases where the manuscript's characterisation of the source is inaccurate — including cases where the source draws the opposite conclusion from what the citation implies.
Every discrepancy is returned with a precise annotation specifying the nature of the failure — equivocation, contradiction, secondary relay, or mischaracterisation — the specific passage in the source that is relevant, and the correction required.
Cross-validation across every data surface.
Citation error is the most thoroughly documented category of pre-submission failure, but it is not the most common. Inconsistency between what the text states and what the data show is: an analysis of more than 250,000 reported p-values found that half of published papers contained at least one statistical result inconsistent with its own test statistic and degrees of freedom, and one in eight contained an inconsistency severe enough to affect the stated conclusion.9 Reviewers notice. Inaccurate or inconsistent reported data and defective tables and figures both appear among the ten most frequent reasons reviewers give for rejection.10
The mechanism is mundane. During revision, a value is updated in one location and not propagated to the others where it appears. A p-value revised in the body text remains unchanged in the corresponding table. A percentage stated in the Discussion no longer matches the figure it was derived from. A sample size in the Methods conflicts with the n reported in a supplementary table.
A numerical inconsistency in the Methods causes a reviewer to look more carefully. A second in the Discussion and a third between the text and a figure place the manuscript’s evidentiary reliability in doubt before the scientific evaluation has begun.
Verification proceeds in three directions, because a figure and its caption are separate sources of evidence and either can disagree with the text or with each other.
Does the body text accurately describe what the figure plots or the table reports? This catches directional errors — text reporting an increase where the data show a decrease — and qualifiers such as "significantly" or "markedly" applied to a comparison the data do not support.
Captions and legends are independent data sources. They routinely carry values that appear nowhere else — sample sizes, statistical parameters, chemical formulae, group identifiers — and each must agree exactly with the body text that refers to it.
Does the legend accurately describe what is actually plotted? Panel letters, axis labels, colour assignments, and condition names are verified against the figure as rendered, and every check is applied to supplementary figures and tables with identical rigour.
In parallel, all quantitative claims in the body text are cross-validated against every corresponding data surface: figures, captions, tables, table notes, and supplementary files. The scope covers reported statistics, sample and subject counts, percentage calculations, and any value appearing in more than one location. Each discrepancy is annotated with the conflicting locations and the correction required. Where a discrepancy reflects an update not propagated across all instances, it is flagged for the author to resolve — we do not correct data; we ensure no inconsistency reaches a reviewer unnoticed.
Where an argument loses a reviewer.
Citation and data errors are verifiable against fixed evidence. A manuscript can also fail without containing a single incorrect statement — when the argument it makes is harder to follow than it needs to be, or when its structure does not deliver the finding it was written to deliver. Reviewers report this directly: text that is difficult to follow ranks among the ten most frequent grounds for rejection, alongside overinterpretation of results.10
This layer is assessed by a scientist active in the manuscript’s own sub-field, reading it as a rigorous reviewer would, with the full document in view. The recurring failure modes are these.
An objective stated in the Introduction that the Discussion never revisits, or a conclusion that answers a question the Introduction never posed. Reviewers experience this as an argument that does not close.
A compound, construct, cell line, reagent, instrument, or experimental condition named one way in the Results and differently in the Methods or a figure legend. These accumulate across revision cycles, and each one is a factual error whose correct answer is fixed elsewhere in the manuscript itself.
A specific value asserted in the Abstract or Discussion that cannot be traced to any measurement reported in the Results, Methods, or a figure. The claim may well be correct; what matters is that a reviewer cannot confirm it from the document in front of them.
The central result placed where reader attention is lowest — at the end of a paragraph whose topic sentence points elsewhere, or in a subsection the structure does not signal as important. The finding is present; the manuscript does not make it unmissable.
A sentence whose intended reading is recoverable only on a second pass — an unclear antecedent, a modifier that could attach to two referents, a logical relation between clauses left implicit. Every re-read is attention withdrawn from the science.
Two constraints govern this work. The first is altitude: a manuscript is written at different levels of generality, and a framing statement in an abstract is deliberately general. Importing granular detail into it is itself an error, and a high-level statement is never flagged as unsupported merely because its supporting specifics live where they belong. The second is charity: where a reasonable reading of the author’s wording is supported by the evidence, that reading stands. A revision is proposed only where the manuscript’s own content determines the answer; where the matter turns on judgment, the observation is raised for the author to weigh, not resolved on their behalf.
Calibrating what the evidence will carry.
The distance between what a study shows and what its manuscript claims has a name in the meta-research literature: spin. A systematic review of the phenomenon found spin in more than a quarter of published systematic reviews and meta-analyses, rising as high as 84% in some study categories,12 and the foundational analysis of randomised trials with non-significant primary outcomes found distorted presentation to be the norm rather than the exception.11 The drift is measurable across the literature as a whole: the frequency of promotional language in PubMed abstracts rose sharply over four decades.13
For an author, the risk is asymmetric and specific. A verb that outruns the evidence — demonstrate or prove where suggest, indicate, or are consistent with matches what the data establish — is among the most reliable triggers of a major revision request. Under-claiming carries a quieter cost: a genuine advance that the reviewer never registers as one.
Latitude depends on section
The same sentence can be appropriate in one section and over-reaching in another. Discussion and Conclusions are where an author is entitled to interpret, generalise, and argue for significance; reviewers read them as argument, and reaching somewhat beyond the strict letter of the data is their normal function. Results and Abstract report what was found, and a claim stronger than the data shown is a real risk there. Calibration is applied accordingly, not uniformly.
Calibration is not retreat
A recalibration must never substitute our scientific position for the author’s. Where a claim can be brought into line by qualifier alone — with every piece of the author’s content and every citation left intact — the calibrated sentence is supplied directly, with the reasoning stated. Where bringing it into line would require deleting content, dropping a citation, or narrowing what the sentence asserts, no revision is written: the point is raised, the reviewer’s likely objection is set out, and the author decides how far to go. The author holds the evidence. We identify the exposure.
The same standard governs limitations. A limitation is raised when it challenges the validity, reproducibility, or scope of the paper’s primary contribution — the kind a rigorous reviewer would require addressed before recommending acceptance. Generic model-system caveats that apply to nearly every study in the field are not, unless the manuscript itself claims broad generalisability as part of its novelty. The test is not whether an objection can be raised, but whether the paper would clear review without it.
Minimal intervention, maximal precision.
Language is not a cosmetic layer. A survey of 908 researchers found that scientists who are not native English speakers require substantially more time to write a paper, face markedly higher rates of rejection attributed to writing, and more often avoid presenting their work at international venues for language-related reasons.14 The cost falls on the science, not only on the scientist.
But grammatical correctness is the floor, not the objective. What a scientific manuscript requires of its language is that every sentence describe the science exactly — at the strength the evidence carries, at the scope the study covers, in the terms the field uses with the meaning the field gives them. A sentence can be flawless as English and wrong as science: asserting more than the surrounding results support, naming a construct the Methods define differently, or resolving an ambiguity in a direction the data do not. This is why no sentence is judged in isolation. Each is read within its paragraph, its section, and the argument the manuscript is making — because a correction blind to that context, one that repairs the grammar while leaving the sentence misdescribing the finding, is worse than no correction at all.
The response to that cost is not more editing. It is more accurate editing. The governing principle is minimal intervention: the criterion for changing a sentence is correctness, never stylistic preference. A word that is grammatically correct, clear, and contextually appropriate is not replaced because an alternative exists. Substituting synonyms into competent prose is the most common failure of conventional editing — it produces a document dense with tracked changes and no improvement, and it displaces the author’s voice from their own paper.
What is never altered
Field-specific terminology and nomenclature, numerical values and units, statistical results, citation markers, figure and table references, species and reagent designations, mathematical notation, and any abbreviation the author has explicitly defined: the author’s definition governs for the document. Reference list entries reproduce the published bibliographic record and are left entirely unchanged, including spelling and punctuation variants that would be corrected anywhere else.
Where rewriting is required
Restraint is not passivity. A sentence is rewritten — not patched — when it is genuinely difficult to parse on first reading, when nested clauses bury the point, or when a substantially clearer expression preserves the meaning exactly. Every such change is delivered as a tracked change with an explanatory comment alongside it, stating what was changed and why, so that the author can accept, query, or decline each one with the reasoning in front of them. A document returned as a silent clean copy asks the author to trust the editor. A document returned with its reasoning attached does not require them to.
Every correction produced across all six layers passes a final quality-assurance adjudication before delivery. That pass does not look for new issues; it verifies that nothing incorrect has entered the set — that no proposed change contradicts the manuscript’s own content, reverses the author’s meaning, or conflicts with another correction landing on the same sentence.
See the methodology in action.
Upload your manuscript and receive a review built on the evidentiary framework described above.