Pangenome (mouse), a bubble inside a bubble
A bubble index has one row per bubble, and each row reports where the graph varies and how much. A large bubble covers a whole region. The densest bubble in the mouse strain graph holds hundreds of segments inside one intron of Dock2. No index row describes the inside of a bubble, so we draw the bubble as a force-directed graph and open it to see its contents. Pangenomes beyond human builds this graph and finds the bubble.
The graph view is a beta plugin, and this tutorial covers experimental ideas. Where a step below says the view works a particular way today, the step describes a current limit of the view. We welcome your feedback.
Prerequisites
- the GraphGenomeView plugin, loaded the way the HPRC page loads it
- htslib (
tabix), to query the hosted bubble index as the last section does
Where the data comes from
The data is the mouse strain graph from
Pangenomes beyond human. One minigraph
call per chromosome built it from GRCm39 and eighteen strain assemblies, and we
serve it as rGFA projections.
- the segment and link index: https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.segs.bed.gz and https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.links.bed.gz
- the bubble index: https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz
One bubble, one label
Open the Dock2 window as a graph and pick Force-directed layout. The index lists this window as a single bubble, so the whole cut is that bubble. One label names it in the index's terms: a superbubble, with its segment count and the span of its routes. A bubble that fills the whole drawing gets the label and no halo, because a halo around everything would mark nothing.
The view pins Dock2 under the backbone, and no exon stretch appears anywhere in the cut. The gene track on the graph and the linear view above it both show that the cut is intron.
A superbubble is the index's name for a bubble too big to type. The label gives a segment count and a route range and no kind, since the row's numbers describe the whole region at once.
Open it, and open what is inside
Click the label. The view cuts out the bubble's segments and lays out only those. It then derives bubbles from the popped graph. A backbone node that no edge jumps over is a boundary, and whatever lies between two boundaries is a bubble. Each derived bubble gets a halo and a label, and one of them opens in turn. Part 4 of the HPRC tutorial shows one such level opened under its linear view.
Each level has a button that returns to the level above, so you climb back out in the order you descended. The view types the labels at each level the way the index types a bubble. The type comes from the reference interval a bubble replaces and from the shortest and longest route through it. The layering gives those values for any anchored graph.
The control
Nnt is the window Pangenomes beyond human opens first. It holds one large allele that the other strains carry and the reference lacks. Drawn the same way, Nnt should halo as a plain insertion with nothing to descend into.
Check it against the index
The first figure's halo is one row of the hosted bubble index:
tabix https://jbrowse.org/demos/mouse_pangenome/mouse-mm39-minigraph.bubbles.bed.gz \
'mm39#0#chr11:34516044-34560497'
The label printed this same segment count and route span. No file holds the bubbles inside it. The view derives them from the popped graph's layering each time the level opens, and discards them when the level closes.
See also
References
- Li H, Feng X, Chu C. The design and construction of reference pangenome graphs with minigraph. Genome Biology. 2020;21:265. https://doi.org/10.1186/s13059-020-02168-z
- Wick RR, Schultz MB, Zobel J, Holt KE. Bandage: interactive visualization of de novo genome assemblies. Bioinformatics. 2015;31(20):3350-3352. https://doi.org/10.1093/bioinformatics/btv383
Feedback on this tutorial is welcome: contact us.