The native tree at risk of
disappearing into its invader
And what it actually takes to detect that
Red mulberry is not only being cut down or crowded out. It is also being absorbed — hybridised, generation by generation, into the introduced white mulberry that grows alongside it. In four studied populations in southern Ontario, more than half of the trees were already hybrids.1 Many of them can pass for red mulberry.
So the practical question is not whether this is happening. It is how you would know — for a particular tree, stand, or batch of seedlings. That turns out to be a precise methods problem, and a tightly constrained one: what any method can detect is fixed by how the marker is inherited and by how many generations deep the hybridisation goes. Get that wrong and you can spend real money on a test that could never answer your question.
This page works through it in order. What the eye can do and where it stops. What limits every molecular test. An audit of the markers that already exist — and why none is yet validated for this problem. Two new markers derived here from 45 published chloroplast genomes and 180 published gene sequences, one maternal and one biparental, both readable on a strip of agarose. Then how to choose a route, and how many trees you actually have to test.
Once you have DNA, telling pure red mulberry from pure white is easy, cheap and definitive — the species differ at many fixed positions, and a single sequencing read settles it. Catching a recent hybrid is nearly as easy: sequence one region both parents contribute to — ITS — and a hybrid shows both parents' bases superimposed at the diagnostic sites, a pattern neither pure species can make. Add one chloroplast read to trace the maternal line. Two reactions, a few dollars each in reagents.
What no cheap test can do is prove a red-looking tree carries no white ancestry at all. A tree several generations of backcrossing deep can look — and sequence — as pure red at any one marker. Ruling that out takes genome-wide data compared against verified pure-parent trees, and, as this page finds, that validated reference set does not yet exist for this species pair.
So: red, white, or recently hybrid — provable and cheap. Certified free of all introgression — not by any test you can buy today.
Substantive claims below carry one of these tags. Computed here means it was calculated from public sequence data and can be reproduced from the accessions given. Unverified means I could not confirm it against a source and you should not rely on it.
A tree can go extinct without anything dying
Red mulberry (Morus rubra) is native to eastern North America, from Ontario and Vermont south to Florida and west to Texas and South Dakota. White mulberry (Morus alba) was brought from Asia for silkworm culture and is now one of the commonest weedy trees on the continent.
The two species hybridise readily, and the hybrids are fertile. Where they grow together, the hybrids carry more white mulberry ancestry than red, and that bias is stronger where white mulberry is more abundant. Each generation of backcrossing dilutes the native genome further, and the endpoint is not a dead tree — it is a stand whose red mulberry ancestry survives only as fragments inside increasingly white-mulberry-like genomes.
The measurement everyone cites comes from Burgess, Morgan, Deverno and Husband, published in Molecular Ecology in 2005.1 They genotyped 184 trees from four populations in southern Ontario where both species grow together, using nuclear markers alongside chloroplast sequence.
| Measure | Value |
|---|---|
| Classified as hybrids (RAPD index) | 53% (98 of 184) |
| Classified as red mulberry | 29% (53) |
| Classified as white mulberry | 18% (33) |
| Range of hybrid frequency across the four sites | 43% – 67% |
| Hybrids with more white than red mulberry markers | 67% |
| Hybrids carrying a white mulberry chloroplast | 68% of 25 |
43 polymorphic RAPD fragments — of which only nine were species-diagnostic, five for white mulberry and four for red — plus chloroplast sequence from an 802 bp window of rbcL, in which the two species differ at just three fixed sites.1 Sampling was stratified: every putative red mulberry was taken, along with a roughly 25% subsample of the putative white and hybrid trees within 25 m of each. The authors note this may overestimate hybrid frequency and underestimate white.1 The paper's abstract gives the hybrid count as 53.7% (n = 99) while its results section and Figure 3 give 53% (n = 98); 53 + 33 + 98 = 184, so the figures here follow the results section. The widely quoted "53.7%" comes from the abstract.
That last line shapes everything that follows. Stated the other way round: 32% of those hybrids carried a red mulberry chloroplast. It is the reason no chloroplast test — including the good one below — can ever be the whole answer.
The chloroplast result rests on 25 hybrids, not 184 — sequencing was done on 42 trees in total. Seventeen of the 25 carried the white mulberry chloroplast type. Burgess and colleagues tested that split against 1:1 and could not reject it (χ² = 2.72, P = 0.099), which is why their own discussion says most hybrids carried the white chloroplast type "although insignificantly".1
Run the interval and the honest range is wide: the fraction of hybrids with a red-type chloroplast — the ones a chloroplast test cannot see — is somewhere between 17% and 52% at 95% confidence. computed here The point estimate is a third. The upper end is half.
"About a third" is the shorthand throughout, as the best available estimate. Read it as a third, and possibly half. Nothing downstream should depend on the difference — and where something would, it is flagged.
Michigan lists red mulberry as state threatened.2 In Ontario, where the species sits at its northern range edge and hybridisation pressure is worst, most known red mulberry sites have white mulberry growing in them.2
A red mulberry that is genetically half white mulberry will still make fruit, still feed birds, and still be counted as a native tree in a plant survey. The loss is invisible without a test, which means every number anyone quotes about how much red mulberry is left depends entirely on how the trees were called. Nobody is systematically screening the trees in any given county, and the trees are not going to be screened by looking at them.
Everything after this is about the second question — given that you need a test, which one, and what will it settle? The answer is more constrained than it first appears — worth understanding before you spend anything.
What you can actually see from the ground
The two pure species are distinguishable by eye. The trouble starts with everything in between — so it matters which characters actually separate the species and which are folklore.
| Character | Red mulberry | White mulberry | Worth? |
|---|---|---|---|
| Hairs on the leaf underside | Erect hairs spread evenly over the whole blade — soft to the touch | Confined to the main veins and the tufts in vein axils | Best character |
| Leaf area | Blade commonly 10–18 cm, sometimes far larger | Blade 2–20 cm — smaller on average, but overlapping | Best measurable |
| Upper leaf surface | Roughened, dull green | Glossy, lustrous | Good |
| Marginal teeth | Small, numerous, pointed | Fewer, larger, blunt | Good, underused |
| Leaf apex | Drawn out to a long point | Acute to blunt | Suggestive |
| Bark texture | Flat thin plates peeling outwards | Firm braided ridges, orange showing in the furrows | Suggestive |
| Petiole length | Overlapping; if anything M. alba is longer | Useless | |
| Style length | Both species effectively lack a style | Useless | |
| Male vs female trees | Both subdioecious; ~10% hermaphrodite, and individuals switch between years | Useless | |
| Fruit colour | Overlapping and non-diagnostic — white mulberry is usually red to black, not white | Useless | |
Compiled from the Flora of North America treatment,8 Nepal, Mayfield & Ferguson 2012,9 and Nepal 2008.10 published
The character to learn is where the hairs sit on the underside — not whether hairs are present. Red mulberry carries erect hairs spread across the whole blade, soft to the touch. White mulberry has them only along the ribs and in the little tufts where veins meet the midrib.
Three things widely believed that are not true
Fruit colour means nothing. Nepal and colleagues put it bluntly: fruit colour is "highly variable within M. alba and non-diagnostic. In fact, in wild populations, fruits of M. alba are usually red to black rather than white."9 The Flora of North America gives white mulberry syncarps as "black, purple, or nearly white."8 More confident misidentification traces to this one piece of folklore than to anything else.
Style length is not diagnostic. This one circulates widely in identification guides. Both species effectively lack a style — Nepal's genus-wide key places M. alba and M. rubra together in the short-or-absent-style half of the genus.10 What older sources call "style length" is a measurement of the stigma arms, and even that is contested.
Whether a tree is male, female or both tells you nothing. Both species are subdioecious, with roughly one tree in ten bearing both sexes,9 and a similar fraction switching sex from one year to the next.10 Treating breeding system as a species character is precisely the error that produced a spurious mulberry species, M. murrayana, later dismantled by exactly that observation.9
Bark colour is not settled. The Michigan abstract calls red mulberry bark "dark-reddish brown",2 while the Flora of North America, the Canadian status report and Nepal all describe it as grey to greyish-tan and put the orange tint on white mulberry, showing in the furrows between firm ridges and on exposed roots.8129 The reproducible part is the texture — flat peeling plates against firm braided ridges — so use that and ignore the colour. unresolved
One more caution that undercuts almost everything in the table: these characters are read from mature leaves on ordinary shoots. Juvenile growth, stump sprouts and vigorous water shoots converge between the species, and as Nepal puts it, "nearly all of the unique characteristics of M. rubra fail in juvenile leaves."9
Morphology sees species. It barely sees ancestry
Burgess and colleagues measured six morphological characters alongside their genetic markers, and found that the pure and hybrid classes differed on all six.1 That sounds like good news for field identification. It is close to the opposite.
The reason is in which classes the characters separate. Of the six, only leaf area and leaf perimeter told all three groups apart. For the other four — number of lobes, sinus depth, and trichome density on both leaf surfaces — white mulberry and the hybrids were statistically indistinguishable from each other, and both differed from red mulberry.1
This is better news than it sounds for one job and much worse for another. Morphology is a decent tool for finding the trees that are not red mulberry — which is what a removal programme needs. It is close to useless for whether a good-looking tree is pure, because the hybrids that most resemble red mulberry are exactly the ones the characters fail on.
Trees that fooled the experts
The clearest example comes from a Kansas population studied by Nepal, where trees were first assigned by an expert on leaf, bud and bark characters and then genotyped.10 The morphological calls did not survive:
| Marker system | Trees called pure by morphology that were genetically admixed |
|---|---|
| Microsatellites | 10 — nine of them called red mulberry, one white |
| RAPD markers | 9 — five called red mulberry, four white |
Of nine trees that morphology had flagged as possible hybrids, only six were confirmed. And the two marker systems agreed with each other on only 44% of the hybrids they found — a reminder that even the genetic answer depends on which markers you use.10
Conservation practice has already absorbed this. Canada's recovery strategy designates critical habitat for trees "confirmed as pure-strain Red Mulberry trees through genetic testing," and lists confirming the genetic purity of morphologically-identified trees as outstanding work.13 The status report records the consequence plainly: "a few of trees previously counted as Red Mulberry were determined to be hybrids and were excluded from subsequent surveys."12
The one quantitative consolation: of everything measurable on a leaf, hair density on the underside is the best single predictor of what the genome actually says. Regressed on its own against hybrid index it explains about 30% of the variation, and in a combined regression of the four characters that individually tracked hybrid index, it was the only one that stayed significant, with the model reaching 40%.1 Thirty percent is a good morphological character. It is not a test.
So morphology can't settle it; a molecular test has to. And which test is the right one is not a matter of preference.
Two things decide what a test can possibly see
Before comparing methods on price or convenience, most of the answer is already fixed by two things: how the marker is inherited, and how many generations of backcrossing have happened. Neither is negotiable, and together they rule out whole categories of test before you spend a dollar.
One: a maternal marker can only ever name one parent
Chloroplasts are inherited maternally in most flowering plants — those in a tree came from the ovule, not the pollen. Burgess and colleagues relied on exactly this when they used mulberry chloroplast DNA to read the maternal lineage of each hybrid.1 It makes chloroplast DNA a superb species marker and a fundamentally limited hybrid marker.
If a red mulberry flower is pollinated by white mulberry, the seedling is a 50/50 nuclear hybrid carrying a pure red mulberry chloroplast. Every chloroplast test ever devised will call that tree red mulberry — correctly, and uselessly. This is not a hypothetical failure mode. It is about a third of the hybrids Burgess found — with a 95% interval running from 17% to 52%, because that result rests on 25 sequenced trees and its 68:32 split could not be distinguished from an even one.1 computed here No amount of money spent on a better chloroplast assay recovers those trees, because the information is not in the molecule.
What a chloroplast test is excellent at is the other direction. A tree that looks like red mulberry but returns a white mulberry chloroplast is definitively not pure, and you have found that out for the price of one PCR — polymerase chain reaction, a routine lab method that makes millions of copies of a target stretch of DNA, enough to work with. Here, that stretch carries the differences between red and white mulberry. Most of the hybrids Burgess sequenced carried a white-type chloroplast, so this catches most of them — though not all, and by a margin the small sample could not pin down.
Because it is maternal and does not recombine, a chloroplast type is a property of a maternal lineage, not of an individual. A red-type chloroplast means the unbroken female line traces back to red mulberry — the immediate mother's own species only if she was herself unadmixed. Every seedling of one mother tree carries her chloroplast, so testing her second seedling tells you exactly nothing you did not learn from the first.
That has a sharp practical edge. If you are looking at a set of related trees — a seed lot, a batch of nursery stock, a stand of root suckers — the chloroplast assay counts mothers, not stems. One test per maternal family is the entire available information, and the highest-value move is not a better assay but keeping track of which seed came from which tree. Where lineage is unknown and material has been mixed, the same assay becomes informative again: it estimates what fraction of the mix has white-type maternal lineages.
Two: detection decays by half with every backcross
Closing the maternal blind spot needs a marker inherited from both parents. But a biparental marker has its own ceiling, and it is arithmetic rather than chemistry.
At a locus where the two species are fixed for different variants, a first-generation hybrid carries one copy of each — heterozygous, unmistakable, detectable with certainty at a single locus. Backcross that hybrid to white mulberry and each offspring has a one-in-two chance of inheriting the red variant at that locus. Backcross again and it is one in four. With n independent fixed-difference markers, the chance a first backcross slips through looking pure is 0.5n:
| Fixed-difference markers | First-generation hybrid | First backcross | Second backcross | Third backcross |
|---|---|---|---|---|
| 1 | 0% | 50% | 75% | 88% |
| 10 | 0% | 0.1% | 5.6% | 26% |
| 20 | 0% | 0.0001% | 0.3% | 6.9% |
| 50 | 0% | negligible | 0.00006% | 0.13% |
The first column is the one people miss. In principle, a single biparental fixed difference detects a first-generation hybrid every time — no probability is involved, because an F1 inherits one copy from each parent and must therefore carry both variants. What one locus cannot do is see deep backcrosses. Ten markers catch first-generation hybrids and first backcrosses. Fifty make it unlikely that anything within three generations slips past. Thousands — which is what sequencing gives you — let you estimate the actual ancestry fraction rather than answering yes or no.
That table describes an idealised single-copy locus: two alleles per individual, fixed between the species, both amplifying equally, both visible in the readout. Real markers fail those assumptions in specific ways, and each failure moves a tree from the left of the table towards the right.
It matters here because the nuclear marker this page ends up recommending is ITS — the internal transcribed spacer of the ribosomal RNA genes — which is not single-copy. It is a tandem array of hundreds to thousands of repeats, and what a PCR returns is a pooled and potentially biased sample of them. An F1 is expected to show both parental repeat classes, but a class can be under-represented through copy-number differences between the parents, primer mismatch, competition during amplification, or partial homogenisation of the array. So the guarantee in the first column is a property of the genetics, not a measured property of the assay — and it has not been measured for this one. Where the page says a marker "detects every F1", read it as detects every F1 whose minority repeat class amplifies above the detection threshold.
Real markers rarely meet those ideal assumptions, which is why the practical guidance asks for many of them. A simulation study by Vähä and Primmer concluded that efficient detection of first-generation hybrids needs 12 to 24 markers, and that "separating backcrosses from purebred parental individuals requires a considerable genotyping effort (at least 48 loci), even when divergence between parental populations is high."6
It would be convenient if most hybrids in the field were F1s, because that is the column every method handles well. The distribution says otherwise. Burgess's hybrids had a mean hybrid index of 0.46 — below the 0.5 an F1 would give — and 67% of them carried more white mulberry genome than red.1 The authors' reading: "some of the hybrids are not F1 crosses; rather they are later generation backcrosses that contain high proportions of M. alba genome."
That makes sense given the history — white mulberry arrived in the early 1600s and mulberry generations are short, under about 15 years, so there has been time for backcrossing. Burgess inferred that at least some of the hybrids were later-generation backcrosses, though the study did not measure what fraction fell in each class. Either way, a one-locus screen is strongest against F1s and weakest against deep backcrosses — so wherever backcrosses are present, it will miss some.
This does not make such a screen worthless — a first-generation cross is exactly what you get from a verified pure mother in a white mulberry pollen cloud, so the F1 case is the common one in a controlled setting. But for a wild tree of unknown pedigree, expect some deep backcrosses in what you are screening, and price the answer accordingly.
That is the measure for everything below: what a method resolves on those two counts, and nothing more.
No test on this page — or anywhere — can prove a tree carries no white mulberry ancestry at all; you cannot rule out a single introgressed gene many generations back. So “pure” here is always shorthand for something narrower: indistinguishable from a red mulberry reference at the resolution of the method you ran.
That makes every “pure”, “clear” and “certify” below conditional on three things — the reference trees you compared against, the markers you used, and the confidence you demanded. A tree that reads as pure on a two-marker gel may not on a genome scan. Ancestry after repeated backcrossing is a gradient, not two bins; the job here is to say how far down that gradient a given test can see, not to draw a line nature does not.
What already exists, and why none is yet a validated panel
The obvious move is to find the published marker panel for this species pair and use it. I went looking and found several partial precedents and no finished one. That needs stating precisely, because "no panel exists" is easy to say and easy to get wrong.
A marker set is ancestry-informative for this problem if it has been validated against three things at once: geographically representative M. rubra reference material, equivalent M. alba reference material, and known hybrids of known generation — F1s, reciprocal F1s, and backcrosses. Without the third, you cannot measure the two numbers that matter: how often the panel calls a real hybrid pure, and how often it calls a pure tree admixed.
Everything below is measured against that bar. Several of these efforts are good work; none of them clears it.
The foundational study used markers that don't transfer between labs
Burgess and colleagues' 2005 paper is still the reference measurement for mulberry hybridisation, and it is where the 53.7% comes from. The nuclear markers were RAPDs — randomly amplified polymorphic DNA — read alongside chloroplast sequence.1 RAPDs are dominant, so a heterozygote is indistinguishable from a homozygote for the present allele, and they are anonymous: the bands are not tied to known sequence. They are also notoriously sensitive to reaction conditions, which is why results generally do not transfer between laboratories. As a published measurement the study stands. The paper even gives its five primers and reaction conditions — but for the reproducibility reason just noted, it is not a protocol you can pick up and rely on.
The numbers underneath set the resolution of the best measurement anyone has. They screened 100 RAPD primers and kept five. Those yielded nine species-diagnostic fragments — five for white mulberry, four for red — plus 34 that were polymorphic but not diagnostic, for the 43 that were scored.1 The hybrid classifications did not rest on the nine alone: they came from a maximum-likelihood index built on all 43 fragments. And because these are dominant, anonymous markers never tested on known crosses, their sensitivity does not follow from the ideal single-locus arithmetic in File 04. The reference sets were cleanly separated — the red set averaged 0.89 on the index (95% interval 0.79–0.93), the white 0.09 (0.05–0.20); the means fall short of 1 and 0 only because most fragments are polymorphic rather than fixed and the index is a probability.1
The microsatellite work exists, and was not designed for this question
There are two relevant efforts. Nepal's 2008 dissertation ran both microsatellites and RAPDs on a Kansas population — the analysis behind File 03, where ten morphologically pure trees came back admixed.10 The more recent one is Schreier and Nepal's survey of 78 trees across six Upper Midwest populations, posted to bioRxiv in July 2026 and not yet peer reviewed.20 preprint They screened 12 markers originally developed in Asian Morus — M. indica, which their paper treats as synonymous with M. alba although Kew currently accepts it as distinct, and M. boninensis. Five amplified cleanly in M. rubra and were used for the analysis.
It is tempting to think that because those markers amplify in both species, they cannot tell them apart. They can. Microsatellites are scored by allele size, not by whether they amplify: to compare two species at a locus, the primers generally have to work in both. Diagnostic power comes from how the allele-size distributions differ, how strongly the parental populations are differentiated, and how many loci you combine. Cross-species transferability is what makes a locus testable in both species, not what disqualifies it.
The real limitation is narrower and the authors state it themselves. The study was designed as an M. rubra population survey, and its own limitations section names "the absence of reference M. alba and confirmed hybrid genotypes," concluding that "additional highly informative nuclear markers are therefore needed to resolve the extent, directionality, and demographic consequences of introgression."20 So whether those five loci are ancestry-informative is not answered in the negative — it is unmeasured, because the design could not measure it.
Where individuals showed extra allelic peaks — the pattern that looks most like introgression — the authors decline to call it: such profiles "do not independently demonstrate allopolyploidy or introgression from M. alba," and what is needed is "species-diagnostic nuclear SNPs or genome-scale data."20
One risk worth flagging without overstating it. Observed heterozygosity came in below expected at every locus in every population, a mean of 0.34 against 0.65, and the authors list null alleles and allele dropout among the possible causes.20 If a primer silently fails on one species' allele, a heterozygous hybrid can read as a homozygous pure tree. But a heterozygote deficit on its own does not establish that — inbreeding, population subdivision (the Wahlund effect) and sampling structure all produce the same signature, and the authors name those too. Since the dataset contained no confirmed hybrids, nothing here shows dropout actually converting hybrids into pure calls. It is a specific failure mode to test for with known crosses, not an observed outcome.
The older result is the more pointed one. In Nepal's data the microsatellite and RAPD systems agreed on only 44% of the hybrids they found.10 Two marker sets, one population, and the ancestry calls largely disagreed — which is a clear demonstration that an unvalidated panel produces answers whose reliability you cannot assess.
Two operational efforts, neither publicly specified
Canada runs the most serious red mulberry identification programme anywhere. Leaf samples go to the University of Guelph Arboretum for genetic testing to call each tree red, white or hybrid, and the recovery strategy designates critical habitat only for trees "confirmed as pure-strain Red Mulberry trees through genetic testing."13 The programme has enough confidence in its calls to exclude trees from surveys on the strength of them.12 But neither the recovery strategy nor the Arboretum's published material names the markers, the loci or the technique. not disclosed in any source I could find
Separately, a US SARE-funded citizen-science project — Red Mulberry Search and Rescue — collected over 100 leaf samples nationally, had DNA extracted at the Ohio University Genomics Facility, and worked with Los Alamos National Laboratory to begin building an M. rubra genome to compare samples against the existing M. alba one.21 It reports classifying samples as rubra, alba or hybrid "with a high degree of accuracy" — while stating the limitation itself: because no complete rubra genome existed, "the results are not necessarily indicative of complete purity of species."21 That is a genome-comparison approach rather than a marker panel, and its classifier is not published in a form anyone can rerun.
Neither of these is a criticism. A conservation programme has no obligation to publish a protocol, and a farmer-grant project has no obligation to release a classifier. It does mean the two most-used methods for this exact question cannot be picked up by a third party.
Which sets the target for what follows. The most useful thing to look for is a fixed difference — a position where every red mulberry carries one variant and every white mulberry another, tied to known sequence so anyone can check the claim, and readable without specialist equipment. The rest of this page derives two, one maternal and one biparental. Neither has been validated against known crosses either, and that is flagged where it matters.
Where the two genomes actually differ
Deriving a fixed difference used to mean a sequencing project. It no longer does — enough Morus sequence is already public that the marker can be found by measurement rather than by bench work, and checked by anyone who wants to repeat it.
In 2025 a group at South Dakota State University published complete chloroplast genomes for 45 mulberry trees collected across eight US states, deposited as GenBank accessions PQ309062–PQ309106.3 That dataset makes it possible to stop guessing and measure which piece of DNA to look at.
I downloaded all 45 and analysed them directly. computed here
Aligning a red mulberry genome (159,423 bp) against a white mulberry one (159,293 bp) gives 421 single-base differences and 696 insertion or deletion events. The largest single indel is 36 bp, and the ten largest are mostly tandem-repeat expansions — a short motif repeated one extra time. Those expand and contract on their own, are prone to assembly error, and make unreliable species markers.
That rules out the laziest possible test. There is no big clean length difference you could see by running a PCR product straight onto a gel. The dependable signal is in substitutions, and substitutions have to be either sequenced or cut with an enzyme.
Which region carries the signal
For each candidate region I counted positions where all 33 unambiguous red mulberries were fixed for one base and all 10 unambiguous white mulberries fixed for another.
| Region | Fixed differences | of which substitutions | Verdict |
|---|---|---|---|
| rpl32–trnL(UAG) | 131 | 14 | The one to use |
| ycf1 | 93 | 24 | Strong but unwieldy |
| ndhF–rpl32 | 68 | 22 | Strong |
| psbE–petL | 51 | 14 | Good |
| trnS–trnG | 36 | 11 | Usable |
| psbA–trnH | 10 | 2 | Too weak |
| trnL–trnF | 8 | 2 | Too weak |
| rbcL (standard barcode) | 5 | 5 | Works, barely |
| matK (standard barcode) | 4 | 4 | Works, barely |
computed here from GenBank PQ309062–PQ309106.
Two results stand out. First, the standard plant barcodes do work — rbcL and matK carry five and four fixed differences respectively. That is unusual for two species in the same genus, and it means a conventional barcoding workflow is not useless here. But with only four or five informative positions, one sequencing error costs you a quarter of your evidence.
That is also, in effect, what Burgess used. Their chloroplast work sequenced an 802 bp window of rbcL and found the two species differing at three fixed sites, with no variation within either species.1 Three sites across 42 trees was enough to call maternal lineage, and it is consistent with the five this analysis finds across the whole gene. The point of what follows is not that rbcL fails — it is that rpl32–trnL carries roughly forty times more signal, which is what makes a restriction digest possible instead of a sequencing run.
Second, rpl32–trnL(UAG) is far ahead at 131 fixed differences. It separated all 43 unambiguous trees perfectly: every red mulberry scored 131 out of 131 red-type positions, every white mulberry 131 out of 131 white-type. No intermediates, no ambiguity.
Two things keep that honest. “Unambiguous” here means the 43 of the 45 whose plastome agrees with their GenBank label; the discordant remainder — a tree labelled one species but carrying the other's plastome — is set aside and discussed in File 07. So “separates all 43 perfectly” describes the accessions that were kept, not a marker no record ever contradicts. And these trees come from eight states rather than the whole range, so the 131 positions are fixed across the genomes sampled here — strong candidate diagnostics, not differences proven fixed across either species everywhere.
The marker that wasn't, and the reference that misleads
A perfect 69 bp marker, which does not exist
Early in the analysis a different region looked ideal. The spacer between rps15 and ycf1 came out at 333 bp in every white mulberry and 405–407 bp in every red mulberry — a 70 bp gap, trivially readable on a gel, no enzyme needed.
It is an artifact. The ycf1 gene is annotated as starting 69 bp further along in the white mulberry records than in the red mulberry ones. The same physical DNA therefore falls inside the gene in one set of records and inside the spacer in the other, and comparing "spacer lengths" compares two different things. Searching all 199 Morus chloroplast genomes in GenBank for the supposedly red-mulberry-specific 69 bp block found it in every one, including all 101 white mulberries. computed here
It is an easy and completely invisible way to invent a marker. Any length difference derived from annotation coordinates rather than from the sequence itself deserves this check.
The NCBI red mulberry reference carries a white-type plastome
Scored against the diagnostic positions it comes out 12 white-type to 2 red-type. Its length, 159,289 bp, sits with white mulberry (159,293 bp) and nowhere near the red mulberry range of 159,396–159,423 bp. GenBank records it as identical to accession OP161259.4 The authors of the 2025 study independently noticed that this accession falls among the Asian species in their phylogeny.3 computed here
Note what that does and does not say about the tree it came from. Everything in File 04 applies here too: a plastome reports maternal lineage, not nuclear ancestry. The source specimen could be a largely M. rubra tree that carries a white mulberry chloroplast through introgression — the gene flow that repeated backcrossing produces — which is not a rare event, it is the majority of the hybrids Burgess sequenced. So the defensible statement is that the record is taxonomically discordant, not that the plant was misidentified.
Either way the practical consequence is the same. Anyone comparing a sample against "the M. rubra reference genome" is comparing it against a white mulberry–type plastome and will get the wrong answer. Use the vouchered PQ309073–PQ309106 series instead.
The same check turned up the mirror image. Accession PQ309072, deposited as M. alba, carries a 131-out-of-131 red mulberry–type plastome — a tree identified as white mulberry whose maternal line was red. computed here That is exactly the bidirectional introgression Burgess described, caught in a modern dataset by accident — the clearest example of why a plastome names a maternal lineage rather than a species.
GenBank holds 249 nucleotide records for M. rubra against 5,097 for M. alba. computed here Some of the red mulberry records carry white mulberry–type sequence and some of the white carry red, which is what you should expect from two species that hybridise freely — the labels record what a collector determined, and an organellar sequence records something narrower. If you are going to compare your tree against a reference, check the reference first.
One PCR, one enzyme, one gel
The 131 fixed differences in rpl32–trnL(UAG) can be read by sequencing. They can also be read for a few dollars with a restriction enzyme, because some of those differences create or destroy an enzyme's recognition site.
I searched the amplicon — the stretch of DNA the PCR copies — for enzymes whose cut count differs consistently between the species, then computed the predicted fragments for all 43 unambiguous trees. One is close to ideal.
HpyCH4III
Across those 43 trees, red mulberry carries one cut site in the amplicon and white mulberry carries three. computed here The resulting patterns are not subtle size shifts needing careful measurement — they are different pictures.
Read it as a shape rather than a measurement. Red mulberry gives one heavy band high on the gel. White mulberry gives a band slightly below it plus an obvious small band near the bottom. A hybrid whose maternal line is white mulberry gives the white mulberry pattern; one whose maternal line is red gives the red pattern. The test reports the maternal chloroplast lineage, and only that.
If HpyCH4III is hard to source, HinfI (cheaper, more widely stocked) separates the two with a busier pattern — red mulberry shows bands near 228 and 168 bp where white shows a single ~425 bp band; SspI and BfaI work too. computed here
The second assay, and this one is biparental
Everything above reads the chloroplast, so it hits the first limit from File 04: it names the maternal chloroplast lineage and nothing else, and is blind to the third of hybrids with a red-type chloroplast. Closing that gap needs a locus inherited from both parents — and there is one that can be run in the same afternoon, on the same DNA extraction.
ITS sits in the nuclear genome, so a tree inherits it from both parents. If the two species carry different ITS variants, a hybrid carries both at once, and both are visible on a gel. That is the codominant signal — both parental variants visible at once — that the arithmetic in File 04 requires: a first-generation hybrid must carry both, so in principle a single locus flags that class — subject to the multicopy caveat in File 04.
I pulled every full-length Morus ITS sequence from GenBank — 90 labelled M. rubra and 90 labelled M. alba — and aligned them. computed here There are 22 near-fixed differences between the species, and one of them creates a restriction site:
The amplicon uses the universal ITS1 and ITS4 primers published by White and colleagues in 1990,16 which are among the most widely used ITS primers across plants and fungi and match Morus directly. verified against the sequences The enzyme is NEB MboI, R0147S, 500 units for $88.00. confirmed
ITS1 TCCGTAGGTGAACCTGCGG
ITS4 TCCTCCGCTTATTGATATGC
Both parental variants have been recovered from a real hybrid
The prediction above would be worth little on its own. A group at the University of Central Missouri did a relevant experiment in 2010, and the result is sitting in GenBank — though what it does and does not show needs stating exactly.
They took a herbarium-vouchered M. alba × M. rubra hybrid — specimen KANU:361918 — cloned its ITS, and sequenced four clones from that one tree.17 I downloaded all four and scored them at my diagnostic site. computed here
| Clone from the one hybrid tree | MboI sites | Reads as | Submitters' own annotation |
|---|---|---|---|
| HQ144170 · clone 1 | 2 | white mulberry type | "Morus alba haplotype" |
| HQ144171 · clone 2 | 2 | white mulberry type | "Morus alba haplotype" |
| HQ144175 · clone 3 | 1 | red mulberry type | "Morus rubra haplotype" |
| HQ144187 · clone 4 | 1 | red mulberry type | "Morus rubra haplotype" |
I scored their pure reference trees too: all eight red mulberry clones carry one MboI site, all three white mulberry clones carry two. computed here The separation is clean.
Those four sequences came from cloning: the ITS product was split into individual molecules, and each was sequenced separately. Nobody has taken bulk PCR product from a hybrid, cut it with MboI, and shown that all three predicted bands are visible on a gel.
The distinction is not pedantry. Cloning recovers a repeat class that is present at any abundance; a digest only shows one that is present at enough abundance to make a visible band. A hybrid whose red-type repeats have been partly outcompeted during amplification could sequence as mixed and still run as a clean white mulberry pattern. So this result establishes the biology — a real hybrid carries both parental ITS variants, at this exact site — and leaves the assay's sensitivity unmeasured.
A second variant at the same locus, from a published study
While checking this, I found a second polymorphism at the same locus. A 2025 paper in Plants genotyped 542 mulberry accessions across the ribosomal region, recovering 158 SNPs and 15 indels, and built a CAPS marker — a PCR-and-enzyme test like the ones here — on a 13 bp insertion in ITS1.18 In my own downloaded set that insertion is present in red mulberry and absent in white. The insertion creates an MstI site; FspI is an isoschizomer recognising the same TGCGCA sequence, and is the more commonly stocked of the two, so it is the one worth ordering.
Be clear about what that paper did and did not do, because it is easy to read it as more supportive than it is. Its CAPS assay was built to separate M. alba and M. notabilis from other Morus species — a taxonomic discrimination — using BstEII and MstI. It was not a red-versus-white hybrid test, it did not use FspI, and it did not measure whether the assay detects mixed parental repeat classes in an F1 or a backcross.18 What it does establish is that the locus carries real, scoreable variation and that a restriction assay on it works at the bench.
Scored against my own downloaded set, the insertion is present in 88 of 88 clean red mulberry sequences and 1 of 90 white mulberry. computed here So the ITS region carries two diagnostic variants, at different positions, each readable with its own enzyme.
It is tempting to treat these as two unlinked nuclear markers — run both, and a hybrid has to fail twice to be missed. It does not work that way. The MboI site and the 13 bp indel sit in the same nuclear ribosomal repeat. They are physically linked, they travel together in the same tandem array, and they share nearly every way this assay can fail: concerted evolution, unequal repeat abundance between the parents, preferential amplification of one repeat class, or a minority class simply falling below the detection threshold.
Anything that hides one variant will tend to hide the other in the same tree. Running both gives you two observations of one locus, which is a useful check against a bad digest or a misread gel — but the errors are correlated, so it does not multiply your confidence the way two independent loci would. Independent confirmation has to come from somewhere else in the genome.
That paper also publishes mulberry-specific ITS primers, which are worth preferring over the universal ones if you are ordering fresh.
ITS sits in hundreds of tandem copies, and those copies can be homogenised over generations by concerted evolution — which would erase the hybrid signal. The worry is that this makes the test fail on older hybrids.
The evidence says homogenisation in Morus is incomplete. The 2025 survey of 542 accessions concluded that the "widespread occurrences of heterogeneous SNPs and InDels" indicate "incomplete concerted evolution of nrDNA" — cloning recovered 26 distinct ITS sequences from 32 clones of a single tree, and 15 to 26 unique sequences per plant in the others.18 A separate study found polymorphic ITS types in 14 of 33 accessions.19 The four cloned molecules show the hybrid above still carried both parental repeat classes, though they do not measure how abundant each was. I found no published case of concerted evolution erasing this red/white distinction — but that absence proves little.
That is more reassuring than it first appears, but it is not a guarantee for a specific tree several generations into backcrossing. Treat a clean result as a provisional negative — no mixed ITS repeat class was detected — not as proof of purity. Until known hybrids are tested, how strong that evidence is stays unknown.
Geographically, Morus celtidifolia, the Texas or mountain mulberry of the southwestern US and Mexico, shares the red-mulberry-type insertion. computed here Within the eastern range of M. rubra the two do not meaningfully overlap, but in the Southwest this test cannot be assumed to separate them.
The same GenBank check turned up the now-familiar problem: two of the 90 sequences deposited as M. rubra carry pure white mulberry ITS at every diagnostic position. computed here They are OR251260 and FJ605516 — and the literature search reached the same two by a different route.
The protocol
- Collect and dry
Young, fully expanded sun leaves. Dry them immediately in silica gel at roughly ten times the tissue mass. This single step matters more than anything else in the workflow — properly dried tissue yields good DNA for years at room temperature, and a leaf left in a warm bag overnight may yield none. Photograph the tree, the bark and both leaf surfaces, and take a GPS point.
- Extract DNA
A silica-column plant kit, or CTAB if you prefer to mix your own. Mulberry leaves are high in polysaccharides and phenolics, so add PVP to a CTAB prep or use a kit with an inhibitor-removal step. A generic animal-tissue kit will disappoint you.
- Amplify rpl32–trnL(UAG)
Published universal primers from Shaw and colleagues.5 I checked both against the actual Morus sequences: computed here
rpL32-F CAGTTCCAAAAAAACGTACTTC
One mismatch to Morus, which reads…CCG…where the primer has…CCA…. It sits at position 8, far from the 3′ end, and will amplify normally.trnL(UAG) CTGCTTCCTAAGAGCAGCGT
Exact match in both species.Expect roughly 1,800 bp. Run 5 µL on a gel to confirm a single clean product before digesting.
- Amplify ITS as well
Same DNA, second tube, primers
ITS1andITS4. Expect roughly 700 bp. Running both loci from one extraction costs one extra tube and doubles what the afternoon tells you. - Digest
Take 10 µL of each product, add buffer and about 5 units of enzyme — HpyCH4III for the chloroplast amplicon, MboI for the ITS amplicon — and hold at 37 °C for an hour. Both enzymes work at the same temperature, so they can share a water bath. Digesting unpurified product straight from the PCR usually works for a yes/no readout, but the leftover PCR components can inhibit the enzyme — NEB recommends cleaning the product up first, and if a digest looks incomplete, dilute or purify it and repeat.
- Run and read
1.5% agarose, alongside a 100 bp ladder, a known white mulberry as a positive control, and a no-template (water) reaction to catch contamination. White mulberry is everywhere; find one in a hedgerow and use it to prove your assay works before you trust it on anything rare. Run an uncut aliquot of each PCR product beside its digest, too — a failed or partial digest can otherwise pass for a genotype. That trap is worst for the ITS assay: uncut ~700 bp product sits almost exactly where a fully cut red mulberry band does, so an under-digested white or hybrid can read as pure red.
Read the two lanes together. The chloroplast lane names the maternal chloroplast lineage. The ITS lane says whether both species are represented in the nuclear genome. Three bands in the ITS lane is the result you are looking for and hoping not to find.
These are candidate screens with sequence-level support, not validated diagnostic tests. Keeping those apart is the whole point.
What is solid. The underlying sequence differences. The chloroplast marker holds across 43 independently sequenced genomes with no exceptions; the ITS marker separates 178 sequences correctly, and both parental variants have been recovered from a genuine vouchered hybrid at the diagnostic site. The locus is real and the polymorphism is real.
What is not. Every fragment size on this page is predicted computationally, and neither digest has been run on a bench by me. unverified as a bench protocol More importantly, neither has been run against known crosses — F1s, reciprocal F1s, backcrosses — which is the only way to measure the two numbers that decide whether a screen is any good: how often it misses a real hybrid, and how often it flags a pure tree. Those numbers are currently unknown for both assays. The published CAPS work at the ITS locus18 shows that a restriction assay there works at the bench, but it was built for species-level taxonomy and never measured hybrid sensitivity.
So treat your first runs as validating the method rather than the trees — that is what the white mulberry control is for — and treat every confidence figure later on this page as conditional on a sensitivity nobody has measured yet.
What the bench actually costs
Both assays can be bought as a service — see the next file — so owning the equipment is a choice rather than a requirement. It is worth making when you expect to run many samples, because the cost per tree collapses to a few dollars and the turnaround drops from weeks to an afternoon. Below what it costs, so the comparison is concrete.
Everything this needs was, twenty years ago, a university facility. It is now four appliances and a shoebox of reagents, and nothing below requires a licence, an institution, or an address that looks like a laboratory.
| Item | Why | Options & price |
|---|---|---|
| Thermocycler | Drives the PCR by cycling temperature — the one non-negotiable instrument. | miniPCR mini8X $695 / mini16X $835; sold to individuals, runs off a laptop.15 A used Bio-Rad or Eppendorf on eBay goes for a fraction. used price unverified |
| Gel rig + viewer | Separates the cut fragments by size so you can read the pattern. | blueGel $309 — tank, power supply and blue-light transilluminator in one.15 Cycler + gel together as the DNA Discovery System, $950–$1,099. |
| 37 °C holder | For the enzyme digest; you also want 65 °C for extraction. | Cozy Cube $199, or a kitchen sous-vide circulator ($0 if you own one) — both hold 37 °C fine. |
| Micropipettes + tips | Measure 1–20 µL accurately; nothing in a kitchen does this. | Three miniPCR H-style $59 each ($177), or Edvotek $95 (lifetime warranty); tips $42 for three 96-racks. Research-grade three-packs run $1,130–$1,380 — not needed here. |
| HpyCH4III enzyme | Cuts the chloroplast amplicon; its site AC^NGT is exactly the diagnostic site. | NEB R0618S, 250 U $83, rCutSmart buffer — 5 U per digest, so 50 trees.14 |
| MboI enzyme | Cuts the ITS amplicon; can show a hybrid outright (site GATC, 37 °C). | NEB R0147S, 500 U $88.14 Sau3AI or DpnII cut the same site. |
| HinfI enzyme optional | Alternative to HpyCH4III — busier pattern, far more units for the money. | NEB R0155S, 5,000 U $77. |
| PCR master mix | 2× mix — add only water, primers and template, removing most first-PCR mistakes. | NEB OneTaq M0482S, $53 / 100 reactions. |
| Four primers | Two pairs — chloroplast and ITS (sequences in the protocol). One order lasts years. | Eurofins $0.42/base, 25 nmol desalted — about $9.24 per 22-mer, ~$37 for all four. price unverified |
| Plant DNA extraction | Mulberry phenolics and polysaccharides inhibit PCR — a generic animal-tissue kit will disappoint. | Column kits: Zymo D6020 $273/50, Qiagen 69104 $326 or 69204 $359. Cheap start: miniPCR X-Tract crude lysate $22/20 — often enough for a multicopy target. Home CTAB + PVP also works. |
| 100 bp ladder | The size reference you read the gel against. | NEB N3231S $71/100 lanes; cheaper if you shop — GoldBio ReadyLadder $49, miniPCR load-ready $66. |
| Gel chemistry | Agarose, buffer, stain. Use GelRed, GelGreen or SYBR with a blue-light viewer — never ethidium bromide under UV. | All-in-one agarose tabs (buffer + stain included) $19/8 gels. Separately: agarose $46, TBE $7.50, GelRed $34, loading dye $30. |
| Silica gel desiccant | Dries the leaves — the cheapest item here and the one that most determines whether anything else works. | Fine 0.5–1.5 mm non-indicating beads: 55 lb drum $79 (~$3.18/kg). Avoid the 3–5 mm flower-drying beads — too little contact area. |
A working bench, everything new: about $1,600.
| Line | Choice | Cost |
|---|---|---|
| Thermocycler + gel rig | miniPCR DNA Discovery System | $950 |
| Pipettes + tips | Three miniPCR H-style, one rack each | $219 |
| Both enzymes | HpyCH4III + MboI | $171 |
| PCR master mix | OneTaq, 100 reactions | $53 |
| Four primers | Eurofins, 25 nmol desalted | $37 |
| DNA extraction | X-Tract buffer, 20 preps | $22 |
| Gels | All-in-one agarose tabs, 8 | $19 |
| Ladder | GoldBio ReadyLadder | $49 |
| Silica gel | 55 lb drum | $79 |
| Total | $1,599 |
Swap in a used thermocycler and gel rig and it lands nearer $800. Swap up to a proper column extraction kit and research-grade pipettes and it passes $2,500. The machines are the whole decision; everything else is noise.
Per-tree running cost after setup is a few dollars for both tests. The expensive part is the first tree; the hundredth is nearly free.
Before buying anything
Two cheaper routes are worth considering first.
Community biology labs already own all of this and will let members run their own projects. Verified monthly rates: BosLab, Somerville MA — $50; SoundBio, Seattle — $55, or $135 to lead your own project; ChiTownBio, Chicago — $75; Counter Culture Labs, Oakland — $100; BUGSS, Baltimore — $100; Genspace, Brooklyn — $110 community, $220 for full project access. confirmed Most ask for a short project proposal and a safety session.
Counter Culture Labs is the notable one for this project: it runs a standing plant biology group, and its fungal group already does ITS extraction and sequencing — which is precisely the second assay on this page.
Or send the reading out. Sequence the amplicon instead of digesting it, and you read all 131 chloroplast positions rather than one enzyme's worth. Psomagen charges $3.50 a reaction; Quintara from $4.00; Eurofins SimpleSeq prepaid kits work out at $6.10; Plasmidsaurus sequences an unpurified amplicon for $15 and takes a credit card with no minimum. confirmed Azenta/GENEWIZ confirms in its own FAQ that individuals can register and pay by card, though it publishes no price.
It is easy to assume sequencing lets you skip the equipment entirely. It does not. Every service in that price band sequences a prepared template — a PCR product or a plasmid. None of them takes raw mulberry genomic DNA and returns the chloroplast region you asked about; the $15 Plasmidsaurus tier wants a linear amplicon, and SimpleSeq and Psomagen's standard Sanger service both assume you supply the product.
So the outsourced route still requires you to extract DNA, amplify the target on a thermocycler, check on a gel that you got one clean product, often clean it up, and supply a sequencing primer. For the 1,800 bp chloroplast amplicon you also need reads from both ends, since conventional Sanger runs out well short of that — so budget two reactions per tree, not one. The ~700 bp ITS amplicon fits in a single read.
What sequencing removes is the restriction digest and the gel readout: no enzymes, no ladder, no transilluminator, and far more information per tree. What it removes entirely is only true if someone else does your PCR — a community lab, a university core, or a collaborator.
So the fair comparison is not "sequencing versus the bench". It is the thermocycler plus a few dollars a read, against the thermocycler plus the gel rig plus enzymes. At three to fifteen dollars a read the sequencing route wins on information per tree and loses on turnaround, and it lowers the equipment bill by roughly the $309 gel rig and the $171 of enzymes rather than by the whole $1,599.
None of this is dangerous work, but two habits matter: keep the DNA stain off your skin and out of the drain, and never run a sample you care about without a known control alongside it. The most common outcome of a first attempt is not a wrong answer — it is a blank gel, which tells you nothing and costs you a sample.
What each method buys you
Now line every method up side by side. Each costs something and resolves something, and the two are not proportional — the cheapest useful step is nearly free, and the last bit of certainty is what costs real money.
| Method | Cost per tree | Detects | Blind to |
|---|---|---|---|
| Leaf hairs, by eye | free | Most pure red mulberry vs everything else | Hybrids, which resemble white mulberry |
| Chloroplast digest rpl32–trnL, HpyCH4III | ~$3 | White-type maternal lineages — most hybrids, on a small sample | Any hybrid with a red-type chloroplast — a third, perhaps half |
| ITS digest MboI, biparental | ~$3 | F1s, and about half of first backcrosses | Deep backcrosses; any repeat class below the detection threshold |
| Sanger, both loci | $7 – $60 plus PCR | Same, plus all 131 chloroplast positions and the exact ITS variants | Same generation limit — more sites, still one locus each |
| Genome skimming | $225 – $1,000 | Ancestry fraction and hybrid index; generation class with enough depth | Very late backcrosses; limited by reference quality, mapping bias and depth |
Per-tree costs assume you can already run a PCR; see File 09 for what a bench costs to build, and note that the Sanger row does not remove that requirement — those services sequence a prepared amplicon, so the thermocycler is needed either way. The chloroplast amplicon needs a read from each end at 1,800 bp, hence two reactions. Vendor prices verified below; the Sanger figures are unverified beyond the two checked by hand.
The useful reading: the two gel assays are not a cheap approximation of sequencing. They answer a different question. The chloroplast assay rejects trees; the ITS assay is aimed at exactly the class of hybrid a first-generation cross produces. What neither can catch is the tree several generations into backcrossing — red-type chloroplast, ITS homogenised, and a quarter of its genome still foreign.
One column is missing from that table, and its absence matters more than anything in it: a sensitivity figure for each row. Nobody has measured how often these assays miss a hybrid that is there, because that requires known crosses and, per File 05, nobody has assembled that reference set for this species pair. The "detects" column describes what each method is designed to reach, not a measured hit rate. Read it that way, and read File 11 before attaching a confidence number to any of it.
For the deeply backcrossed tree, and for anyone who needs a number rather than a flag, you need the last row.
When you need genome-wide ancestry: genome skimming
Send extracted DNA for low-coverage whole-genome sequencing. At one or two times coverage you recover the complete chloroplast genome as a free byproduct — it is present in thousands of copies per cell — and enough nuclear variants to place the tree on a triangle plot of hybrid index against heterozygosity, which distinguishes first-generation hybrids from backcrosses and from pure trees — though placing a tree confidently on that plot, and on the heterozygosity axis in particular, takes appreciably more coverage than recovering the chloroplast does.7 One experiment, potentially both answers, and no wet-lab marker panel to build — though it still needs per-individual genotypes called at many ancestry-informative sites, and a trustworthy set of pure-parent references to define them — which, as File 05 found, does not yet exist in validated form for this species pair.
What it costs, for a real tree, today: Plasmidsaurus publishes $250 for 1 Gb, $500 for 5 Gb, $1,000 for 15 Gb, and lists plants among the organisms the service covers. confirmed Read the tiers rather than the headline, though. The $250 tier is specified for genomes of 20–60 Mb; a 340 Mb mulberry sits in the 300–750 Mb band, which is the $1,000 tier. A gigabase against 340 Mb is about 3× coverage and would probably carry a hybrid index, but it is not the configuration they sell for a genome that size.
The obstacle is not the vendor. It is the DNA. Plasmidsaurus does not accept any intact tissues or organs from animals, plants, insects, fungi, etc.
confirmed A leaf in an envelope is refused. Plant material is taken only as protoplasts with the cell walls stripped, or as high-molecular-weight genomic DNA you extracted yourself — RNase-treated, and never touched by phenol or chloroform.
SeqCenter is the same story approached from the other side. It posts nanopore ligation prices in the open — $150 for 300 Mb, $175 for 600 Mb, $225 for 1.2 Gb, library prep included confirmed — with no quote gate and no institutional account required, which makes it cheaper per gigabase than the tier above. But it wants 60 µL at 40 ng/µL of clean double-stranded DNA, and its own extraction service covers only certain microbes.
So the sequencing is cheap and buyable off a web page. Turning mulberry leaf into a tube of sequenceable DNA is the step that still needs a lab bench or a core facility willing to do the extraction for you, and that step — not the sequencer — is what stands between a private individual and a plant genome.
One complication: there is no red mulberry nuclear genome assembly. NCBI lists zero assemblies and six sequencing runs for the species, against four assemblies for white mulberry. computed here Nuclear reads therefore have to be mapped to the white mulberry reference, which is workable for ancestry estimation but introduces a mild bias and needs someone who knows what they are doing.
That gap is closing on two fronts. Schreier and Nepal state in their preprint that low-coverage genome-skimming data already exist for the species and that development of a high-quality nuclear reference genome is underway using PacBio long-read and Hi-C scaffolding data
, alongside complete chloroplast and mitochondrial genomes from which hundreds of candidate organellar markers are undergoing validation
.20 preprint Separately, the SARE citizen-science project has been working with Los Alamos National Laboratory towards an M. rubra assembly for the same purpose.21 Anyone starting this work now should check whether either has landed before mapping to white mulberry.
Which is a reasonable place to end up. A test that costs a few dollars and rules out most of the problem is worth having, and it is the difference between a shortlist and a guess. Send the survivors for sequencing.
Rejecting is cheap. Certifying never finishes
One question remains, and it is the one that decides what a screening programme costs: how many trees do you have to test? The answer depends entirely on which way you want to be wrong, and the two directions have very different price tags.
Suppose you are looking at a group that should be red mulberry — a stand, a seed lot, a batch of planting stock — and some unknown fraction p of it is admixed. Testing n individuals, the chance of catching at least one hybrid is 1 − (1 − p)n:
| Tested | p = 50% | p = 30% | p = 20% | p = 10% | p = 5% |
|---|---|---|---|---|---|
| 3 | 87.5% | 65.7% | 48.8% | 27.1% | 14.3% |
| 5 | 96.9% | 83.2% | 67.2% | 41.0% | 22.6% |
| 10 | 99.9% | 97.2% | 89.3% | 65.1% | 40.1% |
| 20 | 100% | 99.9% | 98.8% | 87.8% | 64.2% |
| 30 | 100% | 100% | 99.9% | 95.8% | 78.5% |
| 60 | 100% | 100% | 100% | 99.8% | 95.4% |
The left-hand column is not a pessimistic scenario. It is roughly what Burgess measured — 98 of 184 trees, in populations that had been sought out because they held red mulberry.1 Canada's recovery strategy repeats that same 53.7% figure for its core populations,13 which almost certainly makes it the same measurement rather than an independent confirmation — but it does mean the number is what the recovery programme itself plans around. Where that is the true rate, five tests settle it 97 times in 100.
Two cautions on borrowing that number, both from the authors. The sampling took every putative red mulberry but only about a quarter of the surrounding white and hybrid trees, so they warn the figure "is an overestimate of the frequency of hybrids (and underestimate of whites)".1 Pulling the other way, it counts only hybrids that survived to be sampled — "hybridization rates for Morus may be even higher at the time of fertilization."1 So treat 50% as a plausible working figure for a stand where both species grow together, not as a measured constant, and certainly not as transferable to a stand you have not sampled.
Now the other direction. Suppose all n come back clean. What have you proved? Only an upper bound, and it comes down slowly:
| Individuals tested | 5 | 10 | 20 | 30 | 60 | 100 |
|---|---|---|---|---|---|---|
| Upper bound | 45% | 26% | 14% | 9.5% | 4.9% | 3.0% |
Those figures are what you get if every hybrid among the trees you sampled is actually flagged. That is not this assay. Write the assay's sensitivity as s — the probability that a genuinely admixed tree comes back positive — and the detection formula becomes 1 − (1 − p·s)n. Everything in the first table shifts right, and everything in the second bounds only p·s, not p.
We already know s is well below 1, and why. The chloroplast assay misses every hybrid with a red-type chloroplast — a third of them on the best estimate, and up to half at the edge of the confidence interval. The ITS assay misses deep backcrosses by simple arithmetic, and may miss some recent ones through repeat-class bias. Nobody has measured s for either assay, because that needs known crosses, which is exactly the validation File 05 says is missing across this whole field. Until somebody does, these tables bound the frequency of hybrids your test can see, and say nothing rigorous about the biological hybrid fraction.
A second assumption is buried in the exponent. Both formulas treat the n trees as independent draws. Siblings from one mother are not independent — they share her chloroplast entirely, and half her nuclear genome. Neither are root suckers, which may be one clone, or trees clustered in one thicket. Sampling twenty stems from one maternal family is nowhere near twenty independent observations, and the effective sample size can be a small fraction of the stem count. Spread sampling across mothers and across space, or discount n accordingly.
So design the programme to reject, not to certify. Test a handful from anything you are suspicious of and act on the first positive. Save the deep, expensive methods for the small number of trees that survive screening and actually matter — the ones you intend to collect seed from, or protect, or propagate.
And say what you found rather than what you wish you could say. Not this stock is pure red mulberry
, which is not provable by any method on this page, but something a result actually supports: mother tree chloroplast red-type, both ITS variants red-type, twenty offspring screened with no hybrids detected by an assay of unmeasured sensitivity — so the detectable hybrid fraction is under 14% at 95% confidence. That is a much weaker claim. It is also the one the evidence carries.
What to do with a result
Voucher everything. A genetic result attached to a GPS point, a photograph set and a pressed specimen is evidence; the same result attached to a memory is an anecdote. Herbarium sheets can be deposited with a regional herbarium, and most are glad of material from a documented wild population.
Consider sending samples onward. Madhav Nepal's group at South Dakota State University generated the 45-genome dataset this page is built on, works directly on red mulberry hybridisation,3 and has recently posted a population-genetic survey of 78 red mulberries across Kansas, Iowa, Wisconsin and Nebraska as a preprint.20 Their sampling is concentrated in the central US; the eastern and southeastern parts of the range look thin. Contributing tissue from an uncovered population is a real contribution rather than a favour asked.
And if a population comes back clean, that is worth telling your state natural heritage program about, whether or not red mulberry is formally listed where you are.
The words, in one place
Each term is defined in passing where it first appears; gathered here for reference.
| Trichome | A plant hair — here, on the leaf underside; where the hairs sit is the best single field character. |
| Syncarp | The compound mulberry fruit — many tiny fruits fused into one. |
| Allele | One of the alternative versions of the sequence at a locus. |
| Locus | One location in the genome. Two variants close together in the same stretch are not two independent tests. |
| SNP / indel | The two commonest kinds of DNA difference: a single-letter change (SNP), or a short insertion or deletion (indel). |
| Heterozygous / homozygous | Carrying two different versions at a locus, or two copies of the same one. |
| Fixed difference | A position where every sampled member of one species reads one way and every sampled member of the other reads another. Fixed in the sample is not proof it holds everywhere. |
| Introgression | DNA from one species left behind in another after hybrids repeatedly breed back into one parent. |
| Backcross | A hybrid breeding with one of its parent species. Each round roughly halves the other species' share of the genome. |
| F1 | The direct offspring of one red and one white parent — not just any hybrid. |
| Hybrid index | An estimate of how much of a tree's ancestry comes from each parent species, scaled from one to the other. Not a probability that the tree is a hybrid. |
| Plastome / chloroplast type | The chloroplast's DNA, inherited from the seed parent here — so it names the maternal lineage, not the whole tree's ancestry. |
| Biparental | Inherited from both parents, the way nuclear DNA is — unlike the maternal-only chloroplast. |
| ITS (nrDNA array) | A nuclear region present in hundreds to thousands of repeated copies, so one tree can carry several versions at once. |
| Concerted evolution | Processes that slowly make those repeated copies more alike, which can erase a hybrid signal over generations. |
| RAPD | Randomly amplified polymorphic DNA — an early, anonymous fingerprinting marker; dominant and condition-sensitive, so results rarely transfer between labs. |
| Microsatellite | A short repeated DNA motif whose copy number varies between individuals; scored by length, and highly variable. |
| Dominant / codominant | A dominant marker (like RAPD) can't tell a heterozygote from a homozygote; a codominant one (like ITS here) shows both versions at once. |
| Allele dropout | When one version at a locus fails to amplify, making a heterozygote read as homozygous — so a hybrid can read as pure. |
| PCR | Polymerase chain reaction: a lab method that makes millions of copies of a target stretch of DNA. |
| Amplicon | The stretch of DNA a PCR copies. |
| Restriction digest / CAPS | Cutting a PCR product with an enzyme that only cuts where a specific short sequence occurs, giving a band pattern that shows which variant is present. |
| Agarose gel | The jelly-like slab a digest is run on; an electric field pulls the DNA fragments through it and sorts them by size. |
| Sanger sequencing | Reads the exact base sequence of one amplicon; a hybrid shows overlapping peaks at a diagnostic position. |
| Genome skimming | Shallow whole-genome sequencing that reads many fragments a little, rather than assembling a full genome. |
| Sensitivity / specificity | Of real hybrids, the fraction a test catches (sensitivity); of real non-hybrids, the fraction it correctly leaves alone (specificity). |
References
- Burgess, K. S., Morgan, M., Deverno, L. & Husband, B. C. (2005). Asymmetrical introgression between two Morus species (M. alba, M. rubra) that differ in abundance. Molecular Ecology 14: 3471–3483. pubmed.ncbi.nlm.nih.gov/16156816
- Penskar, M. R. (2009). Special Plant Abstract for Morus rubra (red mulberry). Michigan Natural Features Inventory, Lansing, MI. mnfi.anr.msu.edu
- Adhikari, B., Parajuli, S. & Nepal, M. P. (2025). Reporting complete chloroplast genome of endangered red mulberry, useful for understanding hybridization and phylogenetic relationships. Scientific Reports. GenBank PQ309062–PQ309106. pmc.ncbi.nlm.nih.gov/articles/PMC12008415
- Zeng, Q., Chen, M., Wang, S., Xu, X., Li, T., Xiang, Z. & He, N. (2022). Comparative and phylogenetic analyses of the chloroplast genome. Frontiers in Plant Science 13: 1047592. Source of accession OP161259, from which RefSeq NC_070233 is derived. ncbi.nlm.nih.gov/nuccore/NC_070233
- Shaw, J., Lickey, E. B., Schilling, E. E. & Small, R. L. (2007). Comparison of whole chloroplast genome sequences to choose noncoding regions for phylogenetic studies in angiosperms: the tortoise and the hare III. American Journal of Botany 94: 275–288. Source of the rpL32-F and trnL(UAG) primers. doi.org/10.3732/ajb.94.3.275
- Vähä, J.-P. & Primmer, C. R. (2006). Efficiency of model-based Bayesian methods for detecting hybrid individuals under different hybridization scenarios and with different numbers of loci. Molecular Ecology 15: 63–72. pubmed.ncbi.nlm.nih.gov/16367830
- Wiens, B. J. & Colella, J. P. (2025). triangulaR: an R package for identifying AIMs and building triangle plots using SNP data from hybrid zones. Heredity. nature.com/articles/s41437-025-00760-2
- Wunderlin, R. P. (1997). Moraceae: Morus. In Flora of North America North of Mexico, vol. 3. Oxford University Press. efloras.org — genus key, M. rubra, M. alba
- Nepal, M. P., Mayfield, M. H. & Ferguson, C. J. (2012). Identification of eastern North American Morus: taxonomic status of M. murrayana. Phytoneuron 2012-26: 1–6. The clearest published character comparison, and the source of the statement that fruit colour is non-diagnostic. phytoneuron.net (PDF)
- Nepal, M. P. (2008). Systematics and reproductive biology of the genus Morus L. (Moraceae). PhD dissertation, Kansas State University. Source of the Konza Prairie morphology-versus-genotype comparison. krex.k-state.edu
- Burgess, K. S. (2004). The genetic and demographic consequences of hybridization in small plant populations. PhD thesis, University of Guelph. The fuller analysis behind Burgess et al. 2005; listed for further reading. atrium.lib.uoguelph.ca
- COSEWIC (2014). Assessment and Status Report on the Red Mulberry Morus rubra in Canada. Listed under COSEWIC assessments on the species page. Species at Risk Public Registry — Red Mulberry
- Parks Canada Agency (2013). Recovery Strategy for the Red Mulberry (Morus rubra) in Canada. Species at Risk Act Recovery Strategy Series. Listed under Recovery strategies on the same species page. Species at Risk Public Registry — Red Mulberry
- New England Biolabs product catalogue, prices confirmed 2 August 2026: HpyCH4III R0618S, HinfI R0155S, OneTaq 2X Master Mix M0482S, 100 bp DNA Ladder N3231S
- miniPCR bio (Amplyus LLC) store, prices confirmed 2 August 2026. minipcr.com/store
- White, T. J., Bruns, T., Lee, S. & Taylor, J. (1990). Amplification and direct sequencing of fungal ribosomal RNA genes for phylogenetics. In PCR Protocols: A Guide to Methods and Applications, pp. 315–322. Academic Press. The source of the ITS1 and ITS4 primers, now used across plants as well as fungi. doi.org/10.1016/B978-0-12-372180-8.50042-1
- Nikaido, S., Salah, S. M., Ely, J. S., Wolf, H. J. & Raveill, J. A. (2010). Mulberry (Morus: Moraceae) hybridization in eastern North America: morphological and molecular evidence. University of Central Missouri. Unpublished; data deposited as GenBank HQ144170–HQ144187, including four cloned ITS sequences from the vouchered hybrid KANU:361918. ncbi.nlm.nih.gov/nuccore/HQ144170
- Xu, X., Zhang, L., Bi, C., Qin, M., Wang, S., Li, D., He, N. & Zeng, Q. (2025). Extensive nrDNA polymorphism in Morus L. and its application. Plants 14(16): 2570, published 18 August 2025. Open access. 158 SNPs and 15 indels across 542 accessions. Source of the 13 bp ITS1 indel and of the evidence that concerted evolution in Morus is incomplete. Its CAPS assay uses BstEII and MstI and was built to separate M. alba and M. notabilis from other Morus species — a taxonomic discrimination, not a red–white hybrid test. doi.org/10.3390/plants14162570
- Xuan, Y., Wu, Y., Li, P., Liu, R., Luo, Y., Yuan, J., Xiang, Z. & He, N. (2019). Molecular phylogeny of mulberries reconstructed from ITS and two cpDNA sequences. PeerJ 7: e8158. doi.org/10.7717/peerj.8158
- Schreier, S. J. & Nepal, M. P. (2026). Population genetics of native red mulberry at its northwestern boundary suggests postglacial founder effects. bioRxiv preprint, posted 14 July 2026, not certified by peer review. 78 M. rubra individuals across six populations in Kansas, Iowa, Wisconsin and Nebraska. doi.org/10.64898/2026.07.11.737963
- Cornett, J. (2022–2024). Red Mulberry Search and Rescue: preserving genetic diversity for the future of sustainable agroforestry. USDA SARE Farmer/Rancher grant FNC22-1338. Citizen-collected leaf samples, extraction at the Ohio University Genomics Facility, genome comparison against M. alba with Los Alamos National Laboratory. Reports rubra / alba / hybrid calls but states the results are "not necessarily indicative of complete purity of species." projects.sare.org/sare_project/fnc22-1338