Pangenome (mouse)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.11. Desktop beta builds are coming soon.
The mouse strain pangenome aligns eighteen inbred and wild-derived strains from
the Mouse Genomes Project onto the GRCm39 reference with minigraph. GRCm39 is
itself one of the strains, C57BL/6J, which inverts the sign of the best-known
variant in the panel. No locus list has been published for these strains, so we
rank the graph's bubbles, the regions where the strains' paths split and rejoin,
by how many segments each holds, and open the densest. We:
- at Nnt, read a C57BL/6J deletion as sequence the other strains have
- rank the bubbles in the graph and open the densest one that fits in a window, at Dock2
- look the Dock2 bubble up in the hosted index
Every step starts from the mouse graph page on staging.genomes.jbrowse.org, where it stays until a JBrowse 5 host for the graph plugin ships.
The graph view is a beta plugin. We welcome your feedback.
Prerequisites
- htslib (
tabix), to query the hosted index minigraph, to build a graph like this one
Where the data comes from
The reference is UCSC's mm39 and the strains are the Mouse Genomes Project assemblies as UCSC GenArk rehosts them. We host the graph as the index files below.
- mm39 (GRCm39): hgdownload.soe.ucsc.edu/…/mm39.fa.gzhttps://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/mm39.fa.gz
- the eighteen strains, one GenArk folder per GenBank accession: hgdownload.soe.ucsc.edu/…/GCAhttps://hgdownload.soe.ucsc.edu/hubs/GCA/
- segments, the graph's nodes at their reference coordinates: jbrowse.org/…/mouse-mm39-minigraph.segs.bed.gzhttps://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.segs.bed.gz
- links, the edges between segments: jbrowse.org/…/mouse-mm39-minigraph.links.bed.gzhttps://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.links.bed.gz
- bubbles, where the strains' paths split and rejoin: jbrowse.org/…/mouse-mm39-minigraph.bubbles.bed.gzhttps://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz
- the allele inventory, one row per alternative sequence at a bubble: jbrowse.org/…/mouse-mm39-minigraph.alleles.bed.gzhttps://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.alleles.bed.gz
- the whole-chromosome overview, one node per bubble: jbrowse.org/…/mouse-mm39-minigraph.tier10000.segs.bed.gzhttps://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.tier10000.segs.bed.gz and jbrowse.org/…/mouse-mm39-minigraph.tier10000.links.bed.gzhttps://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.tier10000.links.bed.gz
Hosting your own graph describes what each file holds.
Load the graph
We'll load the reference, then the graph and its bubbles. The graph track names
the file prefix build_pangenome_graph.sh writes, and the uris below are our
hosted copy, so swap the prefix for your own build.
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "mm39",
"aliases": ["GRCm39"],
"uri": "https://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/mm39.2bit",
"refNameAliases": {
"uri": "https://jbrowse.org/ucsc/mm39/mm39.chromAlias.txt"
}
}jbrowse add-assembly https://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/mm39.2bit \
--name mm39 \
--alias GRCm39 \
--refNameAliases https://jbrowse.org/ucsc/mm39/mm39.chromAlias.txtGoes in the tracks array of config.json. See Tracks.
{
"type": "GraphTrack",
"trackId": "mouse_minigraph_segments",
"name": "Mouse strain pangenome (rGFA segments)",
"assemblyNames": ["mm39"],
"adapter": {
"type": "RgfaTabixAdapter",
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph",
"coarse": {
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.tier10000",
"aboveBpPerPx": 328
}
},
"displays": [
{ "type": "LinearGraphDisplay" },
{ "type": "LinearBasicDisplay" }
]
}jbrowse add-track-json '{
"type": "GraphTrack",
"trackId": "mouse_minigraph_segments",
"name": "Mouse strain pangenome (rGFA segments)",
"assemblyNames": ["mm39"],
"adapter": {
"type": "RgfaTabixAdapter",
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph",
"coarse": {
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.tier10000",
"aboveBpPerPx": 328
}
},
"displays": [
{ "type": "LinearGraphDisplay" },
{ "type": "LinearBasicDisplay" }
]
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "GraphTrack",
"trackId": "mouse_minigraph_segments",
"name": "Mouse strain pangenome (rGFA segments)",
"assemblyNames": ["mm39"],
"adapter": {
"type": "RgfaTabixAdapter",
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph",
"coarse": {
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.tier10000",
"aboveBpPerPx": 328
}
},
"displays": [
{ "type": "LinearGraphDisplay" },
{ "type": "LinearBasicDisplay" }
]
}The bubbles track reads the same build's bubble index:
Goes in the tracks array of config.json. See Tracks.
{
"type": "FeatureTrack",
"trackId": "mouse_minigraph_bubbles",
"name": "Mouse strain pangenome bubbles",
"assemblyNames": ["mm39"],
"adapter": {
"type": "MinigraphBubbleAdapter",
"uri": "https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz"
}
}jbrowse add-track https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz \
--trackId mouse_minigraph_bubbles \
--name "Mouse strain pangenome bubbles" \
--assemblyNames mm39 \
--adapterType MinigraphBubbleAdapterIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add pangenome graph track, and fill in:
- File type:
Minigraph bubbles (tabix BED) - Path:
https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz - Track name:
Mouse strain pangenome bubbles - Assembly:
mm39
Nnt: a C57BL/6J deletion that appears as an insertion
Click chr13 on the Whole chromosome line of the
portal page, which opens
the whole chromosome with the graph drawn as one node per bubble, a region where
the strains' paths split and rejoin. Type chr13:119,440,000-119,600,000, and
the graph track draws the segments there. Then:
- in the graph track's menu, pick Layout → Force-directed layout and tick Mark bubbles
- turn on the bubbles track in the track selector
C57BL/6J has a multi-exon deletion at Nnt (nicotinamide nucleotide transhydrogenase) that abolishes the protein and makes C57BL/6J mice glucose intolerant. GRCm39 is C57BL/6J, so the backbone of this graph, the reference path, is the strain with the deletion. The graph shows the deletion as sequence that the other strains have and the reference lacks, the opposite sign from the published descriptions.
Ranking the graph's bubbles to find Dock2
The whole-chromosome overview records how many segments each bubble holds, and
more segments mean more variation. Ranking the bubbles by that count finds where
the graph varies most, and matching them to the reference annotation names each
locus.
generatePangenomeLoci.ts
in the genomes.jbrowse.org repo computes the ranking.
The portal page's Loci table shows that ranking. It recovers the vomeronasal
receptor and Speer families and the immunoglobulin heavy chain locus. The rows
above Dock2 are too wide for one window; its row is the densest bubble that
fits in one, inside one intron at chr11:34,516,044-34,560,497. Click its
graph link, then pick Layout → Force-directed layout from the graph
track menu and tick Mark bubbles:
minigraph writes no path lines, so the graph does not record which strains
carry an allele; the allele inventory's firstSeenIn names the first assembly
to contribute it.
Looking up the Dock2 bubble in the hosted bubbles BED
The Dock2 bubble is one row of the hosted bubbles BED, with the same segment count and span the figure labels:
tabix https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz \
'mm39#0#chr11:34516044-34560497'Building a minigraph graph from the strain assemblies
To build a graph like this one, run minigraph once per chromosome over the
sequence for that chromosome from each assembly, reference first so that its
segments take rank 0, the earliest build order:
# -xggs: incremental graph construction, adding each genome in turn
# -c: base-level alignment, which minigraph recommends for graph generation
minigraph -cxggs -t 8 mm39.chr13.fa strain1.chr13.fa strain2.chr13.fa > chr13.gfaPangenome (hosting your own graph) then turns the graph into the files
above with one command, build_pangenome_graph.sh.
build_mouse_pangenome.sh
runs the whole build: it downloads the assemblies, extracts each chromosome
renamed to PanSN (sample#haplotype#contig), runs minigraph per chromosome,
and joins the chromosomes with segment ids renumbered; the alignment takes most
of a day.
The script writes a README.txt beside the data recording the source, the
modifications, the tool versions and the audits that ran. The build stops if the
reference path does not reproduce the reference chromosome lengths, or if
renumbering leaves a duplicate segment id; either failure produces a graph with
wrong coordinates that every later check accepts. Copy these audits into your
own build.
See also
Citations
- Li H, Feng X, Chu C. The design and construction of reference pangenome graphs with minigraph. Genome Biology. 2020;21:265. https://doi.org/10.1186/s13059-020-02168-z
- Wick RR, Schultz MB, Zobel J, Holt KE. Bandage: interactive visualization of de novo genome assemblies. Bioinformatics. 2015;31(20):3350-3352. https://doi.org/10.1093/bioinformatics/btv383
Feedback on this tutorial is welcome: contact us.