During the COVID pandemic, scientists sequenced millions of SARS-CoV-2 genomes and built detailed "family trees" (phylogenies) showing how the virus mutated, split into variants, and moved around the world. Those trees guided public-health decisions, variant naming (the Pango system), and claims about how the virus evolved. A 2026 study published in the journal Nature Methods shows that a large share of those genomes contained systematic errors created by the computers themselves.

Most of the sequences were not read as one continuous string of genetic letters. Instead, labs used a common method called tiled amplicon sequencing: they amplified short, overlapping chunks of the viral genome, sequenced those chunks, and then stitched them back together with software. When a chunk failed to sequence properly (a "dropout"), some of the software simply copied letters from the original reference genome, the early Wuhan sequence treated as the ancestor. In other words, where the machine had no real data, it inserted the ancestral letters as if they were present!

That created false "reversions." A real mutation might appear, then later disappear in the record, not because the virus mutated back, but because the computer had filled a gap with the old ancestral letter. These errors arrived in waves. As new variants emerged, they sometimes changed the places where the amplification primers stuck, causing more dropouts. Labs updated their primer schemes, but the lag produced successive waves of systematic mistakes that followed the waves of actual variants.

When researchers went back to the raw sequencing data and rebuilt the genomes with a more careful tool called Viridian, the picture changed. Many of the apparent evolutionary events vanished. Roughly 30,000 samples were reassigned to different Pango lineages. The number of inferred separate introductions of the virus into the United States dropped by about 1,400. The overall phylogenetic tree became cleaner. The biological past had not changed; only the computer-generated record of it had.

For an ordinary person this matters because the detailed evolutionary narrative of SARS-CoV-2 was treated as hard scientific fact. Governments, media, and researchers pointed to specific mutation pathways, transmission chains, and lineage relationships. Some of those details were artifacts of how the data were assembled. The study does not claim the virus never mutated or that sequencing was useless. It shows that a major portion of the published genomic record mixed real biological signal with computer-filled gaps, and that the mixing was systematic rather than random noise.

Modern genomics relies on layers of reconstruction. Finished genomes are rarely pure photographs of nature; they are models built from imperfect data. When those models are then used to tell the story of evolution, transmission, and ancestry, the boundary between observation and inference can blur. The Nature Methods paper makes that boundary visible for the COVID era. It is a reminder that even in high-stakes science, the tools that process the data can quietly write parts of the story.

https://jonfleetwood.substack.com/p/majority-of-covid-genomes-had-systematic