Synteny visualization (pairwise minimap2)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.10. Desktop beta builds are coming soon.
Align two assemblies with minimap2 -c --eqx, load the PAF as a synteny track,
and read it whole-genome in a dotplot and base-level in the linear synteny view.
add-track -a takes query,target, the reverse of the minimap2 argument order.
Prerequisites
- a JBrowse 2 instance (see the web quickstart, or the desktop quickstart; the steps below are identical on both, and on Desktop the alignments are local files)
- minimap2
python3, which the build script uses to copy the hub configsnode, for the JBrowse CLI
On Debian/Ubuntu, apt install minimap2 covers the aligner, and node comes
from nodejs.org.
Where the data comes from
Three H. pylori RefSeq assemblies and their genome hubs on genomes.jbrowse.org (26695, CHC155, J99). Each hub serves the sequence minimap2 aligns and a config holding the assembly and its NCBI RefSeq genes.
- 26695 GCF_000307795.1 sequence: https://hgdownload.soe.ucsc.edu/hubs/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1.fa.gz
- 26695 hub config: https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/config.json
- CHC155 GCF_025998455.1 sequence: https://hgdownload.soe.ucsc.edu/hubs/GCF/025/998/455/GCF_025998455.1/GCF_025998455.1.fa.gz
- CHC155 hub config: https://jbrowse.org/hubs/genark/GCF/025/998/455/GCF_025998455.1/config.json
- J99 GCF_000982695.1 sequence: https://hgdownload.soe.ucsc.edu/hubs/GCF/000/982/695/GCF_000982695.1/GCF_000982695.1.fa.gz
- J99 hub config: https://jbrowse.org/hubs/genark/GCF/000/982/695/GCF_000982695.1/config.json
- the alignments, indexed and rehosted beside the merged config the figures open: https://jbrowse.org/demos/hpylori/26695_vs_j99.pif.gz and https://jbrowse.org/demos/hpylori/config.json
Three strains, stacked
Three Helicobacter pylori strains (26695, CHC155, and J99) go from their genome hubs to a stacked three-genome synteny view. The steps work the same on any pair of assemblies.
Aligning the assemblies
minimap2 -c -x asm20 --eqx hpylori_j99.fa.gz hpylori_26695.fa.gz > 26695_vs_j99.paf-x asm20is the assembly preset for the divergence between the genomes.asm5covers up to about 5%; these strains diverge well past that.-cemits the base-level CIGAR the linear synteny view draws from.--eqxsplits CIGAR matches (=) from mismatches (X), so the same track opened in a plain linear genome view draws per-base mismatches. Color by value → Identity reads the PAF's divergence tag or match counts.
JBrowse also loads MUMmer .delta and UCSC
.chain files directly, and
paftools.js has
delta2paf and chain2paf for converting them.
Loading the assemblies and the alignment
Each strain's hub config.json holds a whole JBrowse assembly: the 2bit
sequence, an alias table and the NCBI RefSeq gene track. The build copies each
assembly entry and gene track into its own config as the hub wrote it, relabels
the row, and adds the old short name as an alias so a session can still say
hpylori_26695:
{
"name": "GCF_000307795.1",
"displayName": "H. pylori 26695",
"aliases": ["hpylori_26695"],
"sequence": {
"type": "ReferenceSequenceTrack",
"trackId": "GCF_000307795.1-ReferenceSequenceTrack",
"adapter": {
"type": "TwoBitAdapter",
"uri": "https://hgdownload.soe.ucsc.edu/hubs/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1.2bit",
"chromSizes": "https://hgdownload.soe.ucsc.edu/hubs/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1.chrom.sizes.txt"
}
},
"refNameAliases": {
"adapter": {
"type": "RefNameAliasAdapter",
"refNameColumnHeaderName": "ucsc",
"uri": "https://hgdownload.soe.ucsc.edu/hubs/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1.chromAlias.txt"
}
}
}refNameColumnHeaderName makes the UCSC name canonical, so the views read
NC_018939v1 where the FASTA and the PAF say NC_018939.1, and the alias table
maps one to the other. The alignment then goes on under the hub names:
jbrowse add-track 26695_vs_j99.paf -a GCF_000307795.1,GCF_000982695.1 --load copyThe -a order is query,target, the reverse of the minimap2 argument order:
minimap2 target.fa query.fa becomes add-track -a query,target.
Reading the whole genome in a dotplot
Add → Dotplot view opens the import form in Quick start: pick the track just added and click Launch. Swap? transposes the axes. Manual picks each axis and a synteny file by hand.
To open any of those pieces at base resolution, drag a box across it, then right-click inside the box and choose Linear synteny view.
Stacking the three strains
A band is drawn between adjacent rows only, so a 26695 / CHC155 / J99 stack needs the two adjacent alignments:
minimap2 -c -x asm20 --eqx hpylori_chc155.fa.gz hpylori_26695.fa.gz > 26695_vs_chc155.paf
minimap2 -c -x asm20 --eqx hpylori_j99.fa.gz hpylori_chc155.fa.gz > chc155_vs_j99.pafTake the third assembly from its hub and add both alignments the same way as above, then:
- Add → Linear synteny view, then switch to Manual.
- Pick an assembly per row, with Add row for the third.
- Click the arrow between each adjacent pair to choose that band's synteny track: 26695 against CHC155, then CHC155 against J99.
- Click Launch, and all three strains stack in one view.
A whole strain has about as many genes as its row has pixels, so zoom each row in once with its magnifier, then open each strain's gene track, NCBI RefSeq - RefSeq All (GFF), from its own track selector.
Each panel is a full linear genome view with its own search box, zoom and track selector. See Linear synteny view for ribbon options and URL parameters → linear synteny view for building one from a URL.
Coloring genes by ortholog
In bacteria the gene symbol is effectively the ortholog id, since NCBI reuses
standardized symbols across strains. On each gene track, pick Color by... →
Attribute... from the track menu and enter gene. Every distinct value gets
its own color from one palette, chosen from the value itself, so an ortholog
carries one color down all three panels. Features with no value are grey; most
genes here carry only a locus tag.
The dialog writes the field into the display's color, one line of config on
the hub's gene track:
Goes in the tracks array of config.json. See Tracks.
{
"type": "FeatureTrack",
"trackId": "GCF_000307795.1-ncbiGff",
"name": "NCBI RefSeq - RefSeq All (GFF)",
"assemblyNames": ["GCF_000307795.1"],
"adapter": {
"type": "Gff3TabixAdapter",
"uri": "https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1_ASM30779v1_genomic.gff.gz",
"index": {
"location": {
"uri": "https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1_ASM30779v1_genomic.gff.gz.csi"
},
"indexType": "CSI"
}
},
"displayDefaults": {
"showOnlyGenes": true,
"color": { "field": "gene" }
}
}jbrowse add-track-json '{
"type": "FeatureTrack",
"trackId": "GCF_000307795.1-ncbiGff",
"name": "NCBI RefSeq - RefSeq All (GFF)",
"assemblyNames": ["GCF_000307795.1"],
"adapter": {
"type": "Gff3TabixAdapter",
"uri": "https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1_ASM30779v1_genomic.gff.gz",
"index": {
"location": {
"uri": "https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1_ASM30779v1_genomic.gff.gz.csi"
},
"indexType": "CSI"
}
},
"displayDefaults": {
"showOnlyGenes": true,
"color": { "field": "gene" }
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "FeatureTrack",
"trackId": "GCF_000307795.1-ncbiGff",
"name": "NCBI RefSeq - RefSeq All (GFF)",
"assemblyNames": ["GCF_000307795.1"],
"adapter": {
"type": "Gff3TabixAdapter",
"uri": "https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1_ASM30779v1_genomic.gff.gz",
"index": {
"location": {
"uri": "https://jbrowse.org/hubs/genark/GCF/000/307/795/GCF_000307795.1/GCF_000307795.1_ASM30779v1_genomic.gff.gz.csi"
},
"indexType": "CSI"
}
},
"displayDefaults": {
"showOnlyGenes": true,
"color": { "field": "gene" }
}
}Using PIF for large genomes
A bacterial PAF is small enough to load whole. For a large whole-genome alignment, convert it to PIF (Pairwise Indexed Format) so JBrowse fetches only the alignments in the current viewport:
jbrowse make-pif alignment.paf
jbrowse add-track alignment.pif.gz -a query,target --load copyTroubleshooting
assemblyNames in the wrong order is the common one. JBrowse checks at view
load whether the top row's chromosome names belong to that assembly, and a
warning in the view header names the remedy when they belong to the other row.
A view that scatters its blocks randomly comes from a preset too tight for the
divergence, which leaves only short spurious anchors. Raise it, and check
-c --eqx were passed.
Reproduce it end to end
One script builds everything above,
build_hpylori_synteny.sh:
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_hpylori_synteny.sh
bash build_hpylori_synteny.sh # builds ./hpylori_synteny_build/jbrowse2
npx --yes serve hpylori_synteny_build/jbrowse2 # then open the printed URLThe script downloads each strain's hub sequence and config, aligns the strain
pairs, and writes a config.json with the three hub assemblies and gene tracks,
the pairwise synteny tracks, and a default session stacking all three. It needs
the tools under Prerequisites.
See also
- Synteny track
- Dotplot view
- Linear synteny view
- Synteny on genomes.jbrowse.org
- Comparing one genome's two haplotypes (T2T-HG002)
- Synteny visualization (all-vs-all minimap2)
- Synteny from an ortholog table (grape, peach, cacao)
- Config guide: MAF track
References
- Diesh et al. (2024). Setting Up the JBrowse 2 Genome Browser
- Diesh et al. (2023). JBrowse 2: A Modular Genome Browser with Views of Synteny and Structural Variation
Feedback on this tutorial is welcome: contact us.