A single letter can carry the history of a pandemic. At position 614 of the SARS-CoV-2 spike protein, D stands for aspartic acid and G for glycine. The reference genome from the outbreak’s beginning carried D. Early in 2020, nucleotide 23,403 changed from A to G, replacing that amino acid with glycine. The mutation called D614G crossed continents in months and became almost universal.
That is why a D in a later sequence labeled Delta deserves attention. Delta descended from viruses that already carried G614, then accumulated other changes—including L452R, T478K and P681R—that defined its biological and epidemiological character. A mutation back to the ancestral amino acid is possible. But if G conveyed a fitness advantage, and if D returns in groups at particular times and in particular places, a simple story of random mutation and ordinary community transmission starts to strain.
Hideki Kakeya of the University of Tsukuba’s Institute of Systems and Information Engineering and Yoshihisa Matsumoto of Science Tokyo’s Laboratory for Zero-Carbon Energy examined spike-protein records for 22 lineages and sublineages in NCBI GenBank—a total of 873,143 sequences. Their paper, published in Microbiology Research on February 20, 2026, reports that apparent D614 reversions were strongly concentrated in Delta B.1.617.2 and Omicron BA.2, appeared after each lineage’s main surge and clustered in limited parts of the United States.
The pattern is unusual. Unusual, however, is not a cause. The study did not directly observe transmission chains, cultured viruses, laboratory inventories or patient outcomes. It observed finished sequences in a public database. The mystery therefore lies not only in viral evolution but along a chain that runs from swab to extraction, amplification, sequencing, consensus calling, lineage assignment and metadata submission. At which link did that letter arise?
The finding in one sentence
The authors downloaded GenBank records annotated with Pango lineage names and translated spike proteins in June 2023. Their set included Alpha B.1.1.7, Beta B.1.351, Gamma P.1, Delta B.1.617.2, Lambda C.37, Mu B.1.621, and multiple Omicron branches from BA.1 and BA.2 through BQ.1 and XBB.1.5. The paper calls all 22 “VOCs,” but that is not equivalent to WHO’s formal list of separately designated variants of concern. “Lineages and sublineages” is the more precise description.
They calculated two rates. One divided D614 records by all records in a lineage; the other divided them by the number of observed single-nucleotide variants. Delta was highest under the first measure and BA.2 under the second. Comparing each with the pooled 20 lineages other than Delta and BA.2 produced sequence-level Fisher exact-test P values of 1.89×10−56 for Delta and 1.91×10−39 for BA.2. A supplementary event-level treatment gave 2.34×10−17 and 1.58×10−34, respectively.
A tiny P value is not “the probability that a laboratory event occurred.” It measures how incompatible the observed table is with a statistical null model. Huge datasets can make small differences highly significant, and submissions sharing the same laboratory, month or pipeline are not fully independent. Still, the descriptive conclusion is firm: the D614 records are not evenly scattered among the lineages the authors examined.
- Records carrying D614 were disproportionately assigned to Delta B.1.617.2 and Omicron BA.2.
- The D614 increases followed, rather than led, the main sequence surges for both lineages; the delay was especially long for Delta.
- Delta D614 records concentrated around Michigan and Illinois, while BA.2 D614 records concentrated around New York and New Jersey.
- Co-occurring differences were limited, and many prominent companions were also reversions toward ancestral states.
How D614G swept the world in 2020
The one-letter return matters because D614G supplied the pandemic’s first great story of a selective sweep. Wuhan-Hu-1, the reference sequence collected at the beginning of the outbreak, carried D614. Viruses carrying G were detected in Europe by late January 2020. During the spring, G repeatedly rose in frequency in places where D was already established. By summer it dominated the global population.
A Los Alamos-led group tracked the change across national, regional and city scales in a landmark 2020 Cell paper. People infected with G614 viruses tended to have lower PCR cycle thresholds, consistent with more viral RNA in the upper airway, but the study did not find greater disease severity. Because frequency alone can be distorted by founder effects, the field turned to experimental evidence.
Studies with pseudoviruses, human airway tissues and hamsters found evidence that G614 improved infectivity, replication and transmission. Structural work suggested that replacing the charged aspartic acid with the smaller glycine altered interactions within the spike trimer, increasing the proportion of receptor-binding domains in an accessible “up” state and changing cleavage or S1 shedding. No single mechanism explains the entire advantage; stability, opening, cleavage and spike incorporation all form parts of the picture.
G614 did not create Delta by itself. It became a shared foundation inherited by Alpha, Beta, Gamma, Delta and Omicron, each of which added a distinctive constellation of mutations. A reversion to D614 therefore does not restore a whole 2019 virus. The rest of a Delta or BA.2 genome can remain recognizably modern while one site returns to an ancestral state.
Delta, a descendant of G
B.1.617.2 was first identified in Maharashtra, India, in late 2020 and expanded explosively during the country’s spring 2021 wave. WHO added B.1.617 to the global variant-of-concern list on May 11, 2021, and assigned the Greek label Delta later that month. In England, Delta rose to 98% of sequenced infections by late June. Across the summer and autumn it displaced Alpha in much of the world.
Delta’s spike carried changes in the N-terminal domain as well as L452R, T478K, P681R near the furin-cleavage site and D950N. P681R enhanced spike cleavage and cell fusion in experimental systems; L452R and T478K affected receptor interaction and antibody recognition. Delta’s success came from a constellation of changes operating in an immunological and social landscape—not from one magic mutation.
Omicron displaced Delta near the end of 2021. BA.2, an Omicron sublineage that spread widely in early 2022, also inherited G614. Finding a D614 concentration in two genetically distinct lineages is therefore provocative. But their clusters occurred at different times and in different regions, which makes a single global chain of descent an insufficient explanation on its own.
A reversion is not an ancestral virus resurrected. It is one character in a long manuscript returning to an older spelling.
Does a reverse mutation violate evolution?
No. Evolution has no forward gear and no planned destination. Every replication gives an RNA-dependent RNA polymerase another opportunity to introduce change. Selection, chance, migration and outbreak dynamics determine which variants survive and how often they are sampled. A G-to-A change at the relevant codon can yield D614 again; chemically, that transition is not extraordinary.
The same state arising independently on different branches is called homoplasy. With hundreds of millions of documented infections—and far more viral replication events—rare changes receive innumerable opportunities. Host RNA editing, prolonged infections, animal reservoirs, antiviral pressure and recombination can all complicate a simple branching tree.
What can arise is not necessarily what can spread. If G614 improves transmission, a virus returning to D may be outcompeted and leave few descendants. Kakeya and Matsumoto regarded the combined pattern as anomalous: relative enrichment in two lineages, emergence after the major waves, geographic concentration and limited accompanying diversity.
Recombination supplies another route. When two lineages infect the same cell, the viral polymerase can switch templates and join segments with different histories. A 2022 analysis of roughly 1.6 million genomes conservatively attributed about 2.7% of sampled genomes to detectable recombinant lineages. But explaining D614 through recombination requires whole-genome breakpoints, a plausible donor and a coherent phylogeny. The Tsukuba study analyzed translated spike sequences; it did not reconstruct that full history.
How 873,143 records were filtered
GenBank began in 1982 and now forms part of the International Nucleotide Sequence Database Collaboration with the European Nucleotide Archive and Japan’s DDBJ. The three exchange records daily. During COVID-19, laboratories could submit assembled sequences together with dates, locations, hosts and other metadata, creating an extraordinary open infrastructure for global surveillance.
The authors queried records with Pango annotations and translated spike proteins. For the principal mutation histograms, they removed protein sequences containing insertions or deletions, retaining about 80% to 98% of each lineage’s records. Custom scripts written for Python 3.12 compared amino-acid strings. To identify D614 even when spike lengths differed, they searched for the local motif YQDVN.
The dataset contained 181,878 Alpha records, 34,586 Delta, 107,198 BA.2, 150,965 BA.2.12.1 and 58,409 XBB.1.5, among others. Because lineage counts differed enormously, the two denominators—the share of sequences and the share of mutation events—were intended to show that raw sample size alone did not produce the concentration.
In Delta records carrying D614, almost every major co-occurring difference was itself a reversion, except at residue 95, and co-occurrence was generally sparse. BA.2 carried more co-occurring differences, but prominent ones were again reversions except at residue 408. To the authors, this looked less like one isolated fresh mutation and more like a localized return of an older genetic background.
Time and geography create a second mystery
A mutation that gains a transmission advantage would ordinarily arise during circulation, branch from its parent and leave descendants that connect through time and place. In the paper’s monthly counts, D614 rose after the main registration surge for both Delta and BA.2. The lag for Delta was particularly long.
The maps also diverged from the sampling footprint of each lineage as a whole. Delta D614 records appeared most often around Michigan and Illinois; BA.2 D614 records around New York and New Jersey. The authors argued that independent spontaneous reversions spreading in ordinary community transmission did not readily explain such restricted patterns.
Yet a point on a genomic map is a point in a sampling system before it is a point in a transmission chain. If a state laboratory contributes a large batch, its extraction method, primer scheme, consensus software and lineage-calling practice can look geographically clustered. Neighboring states may send samples to a common center. The place of collection, sequencing and submission need not be the same.
| Observed pattern | Biological possibility | Data-process possibility | Evidence needed |
|---|---|---|---|
| G614 returns to D614 | True point mutation or recombination | Reference bias, mixed infection or consensus error | Raw reads, allele fractions and independent resequencing |
| Increase follows the wave | Late local spread or reintroduction of stored material | Batch submission, pipeline change or confused dates | Separate collection, sequencing and release dates |
| Concentration in a few states | Localized outbreak | Submitting-lab and surveillance-network bias | Facility, county and patient-link analysis |
| Limited accompanying diversity | Older background reintroduced locally | Misclassification, contamination, duplication or assembly bias | Whole genomes, accession audits and controls |
A sequence is not the virus itself
Between a patient’s swab and a GenBank line, information is transformed repeatedly. RNA is extracted and reverse-transcribed. Most surveillance workflows amplify overlapping tiles with PCR. Instruments read many fragments; software aligns them to a reference and chooses a consensus at each position. Low viral load, amplification dropout, mutations under primers, mixed infections and cross-contamination can make that choice difficult.
Two 2026 Nature Methods studies quantified this problem at pandemic scale. They showed that tiled-amplicon sequencing paired with assembly software unaware of its error modes created systematic errors, with waves of artifact following waves of variants. In GenBank consensus genomes, low coverage at some sites produced large numbers of artificial reversions to the reference. Researchers rebuilt about 4.47 million genomes from public reads with an amplicon-aware pipeline and developed models that explicitly handle recurrent error.
Those papers did not declare the D614 cluster—or nucleotide 23,403—an artifact. That site was not among the representative high-error positions they highlighted. Kakeya and Matsumoto also checked coverage around residue 614 and argued that stable local completeness, plus the late rather than early appearance of D614, made a simple missing-data or immature-primer account unlikely.
Still, an analysis of finished spike proteins without raw-read reassembly cannot reveal what fraction of reads supported D, whether D and G coexisted in a sample, or whether controls showed the same pattern. A public consensus sequence is a vital lead. It is not the end of a forensic investigation.
The “older background” hypothesis
The paper concludes that limited diversity, late appearance and geographic clustering are more consistent with localized reintroduction of an older genetic background than with ordinary spontaneous reversion and community spread. It proposes further investigation of whether a laboratory-associated event could be involved. The wording demands care.
The study was motivated partly by a 2025 Zenodo preprint that alleged seven unusually ancestral-like sequences submitted from one U.S. hospital and raised the possibility of laboratory-acquired infections. Five of those sequences lacked D614G, prompting the broader search. That preprint was not an official outbreak investigation and did not establish a facility transmission chain.
Testing a laboratory-associated explanation would require accession-level raw reads, residual RNA or specimens, batch records, primer schemes, instruments, controls, laboratory viral inventories, worker-health records where lawfully and ethically available, patient links and the relationship between collection and submitting sites. Independent teams would need to reproduce the result with criteria specified in advance.
Until such evidence exists, a regional sequence cluster cannot be assigned to an institution or individual. “Laboratory-associated” remains one hypothesis among several—not the verdict delivered by the database.
- The origin of the pandemic or the route of the first human infection in 2019
- That Delta or BA.2 was created in a laboratory
- That an infection occurred at any specific university, hospital or research facility
- That D614 records caused a new outbreak, greater severity, immune escape or vaccine failure
- That every recorded D614 is a true biological change rather than a sequencing, lineage or metadata error
Big data, small fractions and non-independent points
A set of 873,143 sequences has enormous discovery power. It can expose patterns too rare for ordinary studies. But volume does not repair sampling design. GenBank is not a random sample of infections. The probability of a case being sequenced and deposited varies by country, state, month, age, severity, laboratory capacity and submission policy. Heavy contributions by the CDC and other major laboratories are a public-health achievement, but they also shape the archive.
One hundred sequences from the same run, laboratory and week are not always 100 independent observations. A shared technical error can multiply through a batch. Fisher’s exact test compares counts in cells; it does not directly model clustering by facility, batch, state or month. The extreme P values make the descriptive enrichment hard to dismiss, but they do not establish that hundreds of independent biological reversions occurred.
Timing matters too. The data were pulled in June 2023 and the article appeared in 2026. Accession records can be corrected or removed; lineage assignments and assembly methods improve. Pango is intentionally dynamic. A 2026 audit should retrieve the same accessions again, compare version histories and rerun classifications with current tools.
The principal analysis also focused on spike proteins and excluded some sequences with insertions or deletions. Without whole-genome phylogeny, submitter metadata and raw reads, it is difficult to separate many independent reversions from many copies of one ancestor—or from one recurring process error.
Designing an investigation that can be falsified
The next step is an accession-level audit. Freeze the full list of D614 records and exclusion rules; compare the 2023 records with current versions; join collection, release and submission dates; map states, counties, submitters, BioSamples, BioProjects and SRA links; remove duplicates and document updates. The authors’ public Python scripts and mutant list offer a starting point for reproduction.
Second, reassemble every available raw read with a common modern workflow. Report coverage across nucleotide 23,403, base quality, strand balance, allele fraction and proximity to primer ends. Look explicitly for mixed D and G populations. Use an amplicon-aware tool such as Viridian and at least one independent pipeline. A D appearing in only one workflow is a warning; a high-quality D reproduced across pipelines is substantially stronger evidence.
Third, construct a time-resolved whole-genome phylogeny. Do the records form one supported clade or recur independently? Do companion reversions share a pattern? Are there recombinant breakpoints and plausible donors? Are collection dates temporally coherent? Those questions distinguish point reversion, common ancestry, recombination and process artifact.
Fourth, layer in epidemiology and laboratory audit. Examine anonymized patient links, processing batches, controls, reagent lots and specimen handling. A facility-associated event should generate independent corroboration; a process artifact should concentrate around a method or batch. If neither appears, natural local spread or unsampled diversity becomes more plausible.
| Stage | Central question | Strong evidence |
|---|---|---|
| Record audit | Are these unique, current and correctly classified? | Stable accessions, complete version history, coherent metadata |
| Raw-read analysis | Do the underlying reads support D614? | High coverage and quality, agreement across pipelines |
| Whole-genome tree | Independent reversion, common ancestor or recombination? | Strong branch support, temporal continuity, breakpoints |
| Epidemiology and lab audit | Where do patient, facility and batch links converge? | Independent records agreeing with the sequence pattern |
| External replication | Do other teams reach the same result? | Public code, prespecified rules and sensitivity analyses |
The largest evolutionary experiment ever recorded
Before 2020, no single virus had been read so often and so quickly. The Pango nomenclature proposed that year gave genomic epidemiology a shared language for branches such as B.1.1.7 and B.1.617.2. GenBank, DDBJ, ENA, GISAID, Nextstrain and national surveillance networks made invisible movement visible in near real time.
Scale created a new class of problem. Among millions of genomes, an error rate of 0.01% still produces hundreds of records. One primer mismatch, software default or laboratory practice can paint a pattern that resembles evolution. Yet discarding every strange pattern as noise would hide real recombinants, emerging lineages and drug-associated mutational signatures.
The answer is neither to erase anomalies nor instantly turn them into scandals. It is to overlay consensus sequences with reads, metadata, epidemiology and laboratory process, then ask competing hypotheses to make different predictions. Genomic epidemiology is becoming not only the science of reading strings but the science of auditing how those strings were made.
The right question mark between D and G
The value of the Kakeya–Matsumoto paper is not that it supplies a final answer. It makes an easily overlooked pattern visible and reproducible. The apparent D614 records in Delta and BA.2 are not randomly distributed. Their timing, geography and companion mutations have structure, and that structure calls for explanation.
The paper alone cannot choose among all explanations. True reversion, recombination, unsampled transmission, reintroduction of older material, mixed infection, contamination, lineage misassignment, consensus bias and metadata error remain on the table. A laboratory-associated event is a testable possibility, not a judgment.
Understanding G614’s rise in 2020 required convergence across global frequencies, patient data, cell experiments, structural biology and animal transmission. Understanding D614’s apparent return deserves the same standard. A tiny letter in a giant database can be an excellent clue. It cannot narrate an entire event by itself.
Science does not serve the public by rushing to replace a question mark with an exclamation point. It serves by showing which evidence can move the mark. The D at residue 614 is now both a question about viral evolution and a question about how faithfully humanity preserves, audits and trusts the genomic history of a pandemic.
Sources and references
This article is based primarily on the February 20, 2026 paper and the University of Tsukuba release, checked against primary work on D614G, Delta, recombination, GenBank, Pango nomenclature and sequencing error. The authors’ mention of a possible laboratory-associated event is treated as a hypothesis, not proof of a facility infection, pathogenic effect or pandemic origin.
- Kakeya & Matsumoto, Anomalous Emergence of D614 Reverse Mutations in the Delta and Omicron BA.2 Variants, Microbiology Research (2026)
- University of Tsukuba, Observation of Nonrandom Patterns of Spike D614 Reversions (March 3, 2026)
- University of Tsukuba, Japanese research announcement (2026)
- Korber et al., Tracking Changes in SARS-CoV-2 Spike: Evidence that D614G Increases Infectivity, Cell (2020)
- Hou et al., SARS-CoV-2 D614G Variant Exhibits Efficient Replication and Transmission, Science (2020)
- Yurkovetskiy et al., Structural and Functional Analysis of the D614G Spike Variant, Cell (2020)
- Mlcochova et al., SARS-CoV-2 B.1.617.2 Delta Variant Replication and Immune Evasion, Nature (2021)
- Saito et al., Enhanced Fusogenicity and Pathogenicity of the Delta P681R Mutation, Nature (2022)
- WHO, B.1.617 Variant-of-Concern Timeline (2021)
- Rambaut et al., A Dynamic Nomenclature Proposal for SARS-CoV-2 Lineages, Nature Microbiology (2020)
- NCBI, GenBank Overview
- NCBI Virus, Sequence Submission and Source Metadata
- Turakhia et al., Pandemic-Scale Phylogenomics Reveals the SARS-CoV-2 Recombination Landscape, Nature (2022)
- van Dorp et al., Recurrent Mutations, Homoplasy and Sequencing Artefacts, Nature Communications (2020)
- Addressing Pandemic-Wide Systematic Errors in the SARS-CoV-2 Phylogeny, Nature Methods (2026)
- Rate Variation and Recurrent Sequence Errors in Pandemic-Scale Phylogenetics, Nature Methods (2026)
