"Whole genome sequencing" is a slightly generous name for what most consumer tests actually do. The DNA is chopped into millions of tiny fragments, each fragment is read, and software reassembles the pieces like a jigsaw puzzle by matching overlapping sequence against a reference genome. This approach — short-read sequencing, dominated by Illumina platforms — is fast, cheap, and extremely accurate for the vast majority of your genome. It is also structurally incapable of resolving certain regions, no matter how many times you resequence them.
Why fragment length actually matters
Imagine trying to reassemble a 1,000-page book by shredding it into 150-character strips and matching overlapping text. If a passage repeats itself — the same sentence appearing in three different chapters — short strips can't tell you which copy came from where. Long, continuous strips can, because they capture enough surrounding context to place the repeat unambiguously.
That's essentially the short-read weakness: the human genome is full of repetitive regions — centromeres, segmental duplications, tandem repeats, and structurally complex areas — where short 150-300bp fragments simply can't be mapped back to a unique location with confidence. Long-read platforms like Oxford Nanopore and PacBio read single DNA molecules continuously, generating reads from several kilobases up to over 100 kilobases, which sails straight through most of these repetitive traps.
| Short-Read (Illumina) | Long-Read (Nanopore / PacBio HiFi) |
|---|---|
| 150-300bp reads | Kilobases to 100kb+ reads |
| Extremely high per-base accuracy | PacBio HiFi now comparable (~99.9%) accuracy; Nanopore historically lower, improving fast |
| Struggles with repetitive/structural regions | Resolves repeats, structural variants, complex regions directly |
| Lower cost per genome | Historically pricier; gap closing fast — some platforms now under $300/genome at scale |
| Cannot natively detect DNA methylation | Can directly detect epigenetic modifications alongside sequence |
What actually gets missed
Structural variants
Large insertions, deletions, duplications, and inversions — some of which are directly disease-causing — are far more reliably detected with long, continuous reads.
Centromeres and complex repeats
Entire chromosomal regions were essentially unreadable with short-read technology alone until long-read methods matured enough to resolve them.
This isn't a hypothetical gap. It's the exact problem the Telomere-to-Telomere (T2T) Consortium was assembled to solve — and long-read sequencing (specifically PacBio HiFi reads) is what finally let researchers complete the last 8% of the human reference genome in 2022, a region that had been unreadable by short-read methods for two full decades.
Should you worry about this for your own results?
For the overwhelming majority of medically actionable findings — carrier status, pharmacogenomics, most single-nucleotide disease variants, ancestry — short-read WGS at good coverage (30x or higher) is genuinely excellent and remains the industry standard for good reason: it's accurate, affordable, and well-validated by decades of clinical use. The gap matters most for structural variants and certain complex, repeat-heavy disease genes, where long-read technology increasingly outperforms.
Coverage and read quality both matter
Dante Labs sequences at high short-read coverage as standard, giving you accurate detection across the vast majority of medically relevant variants in your genome.
Get Your Whole Genome Sequenced → Use code GENOME for 10% off at Dante LabsThe takeaway
"Whole genome sequencing" isn't a single, uniform product — it's a spectrum of technology choices, each with real tradeoffs in cost, speed, and what gets resolved. Understanding whether your test used short-read, long-read, or a hybrid approach is genuinely useful context for interpreting what your results can and can't tell you.