Synteny on a circle (human and mouse)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.11. Desktop beta builds are coming soon.
We lay the human and mouse chromosomes around one circle and draw UCSC's hg38-to-mm39 liftOver chain as ribbons between them, so one picture shows where the autosomes have been shuffled and where the X has not. The circle is a circular genome view opened on both assemblies at once. It reads the chain from the indexed copy on jbrowse.org and orders the mouse arc to follow the human one. We add a gene density ring per genome, then read one ribbon and one ring value back out of the indexed alignment file and the bigWig they came from.
Prerequisites
- a JBrowse to open the files in (Web or Desktop)
- to build the density ring for another pair of genomes, htslib (
bgzip,tabix),bedGraphToBigWigandbigWigToBedGraphfrom the UCSC utilities, andnodefor the JBrowse CLI
Where the data comes from
UCSC's hg38-to-mm39 liftOver chain (Kent et al. 2003), the RefSeq curated gene sets of both genomes from genomes.jbrowse.org's copies of the UCSC hubs, and the density bigWig the build script makes of them.
- the chain: hgdownload.soe.ucsc.edu/…/hg38ToMm39.over.chain.gzhttps://hgdownload.soe.ucsc.edu/goldenPath/hg38/liftOver/hg38ToMm39.over.chain.gz
- the same chain as an indexed PAF, every row, which the ribbons draw from and the last section queries: jbrowse.org/…/hg38ToMm39.over.pif.gzhttps://jbrowse.org/ucsc/hg38/liftOver/hg38ToMm39.over.pif.gz
- the hub configs the two assemblies are taken from, those of the hg38 and mm39 pages: jbrowse.org/…/config.jsonhttps://jbrowse.org/ucsc/hg38/config.json and jbrowse.org/…/config.jsonhttps://jbrowse.org/ucsc/mm39/config.json
- RefSeq curated genes, hg38: jbrowse.org/…/ncbiRefSeqCurated.gff.gzhttps://jbrowse.org/ucsc/hg38/ncbiRefSeqCurated.gff.gz
- RefSeq curated genes, mm39: jbrowse.org/…/ncbiRefSeqCurated.gff.gzhttps://jbrowse.org/ucsc/mm39/ncbiRefSeqCurated.gff.gz
- both genomes' gene density in one bigWig, the ring: jbrowse.org/…/hg38ToMm39.genes.gff.density.bwhttps://jbrowse.org/demos/circular_synteny/hg38ToMm39.genes.gff.density.bw
- the finished config, both assemblies and the four tracks: jbrowse.org/…/config.jsonhttps://jbrowse.org/demos/circular_synteny/config.json
Two genomes on one circle
A Circos-style circle shows, for each chromosome of one genome, which chromosomes of the other contain its sequence and how much of each. In human and mouse the autosomes have been cut and rejoined many times since the lineages split, and the X chromosome has not, so a human autosome should fan out across several mouse chromosomes while the two X chromosomes hold one bundle between them. That expectation is the control the figures are read against.
Declaring the human and mouse assemblies
The ribbons name chromosomes, so the circle needs both genomes declared. Each
alias table is the hub's with the build's genome-prefixed names added, so a
track naming 1, chr1 or hg38.chr1 finds chr1:
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "hg38",
"uri": "https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz",
"refNameAliases": {
"uri": "https://jbrowse.org/demos/circular_synteny/hg38.chromAlias.txt"
},
"cytobands": "https://jbrowse.org/genomes/GRCh38/cytoBand.txt"
}jbrowse add-assembly https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz \
--name hg38 \
--refNameAliases https://jbrowse.org/demos/circular_synteny/hg38.chromAlias.txt \
--config '{"cytobands":"https://jbrowse.org/genomes/GRCh38/cytoBand.txt"}'Goes in the assemblies array of config.json. See Assemblies.
{
"name": "mm39",
"uri": "https://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/mm39.2bit",
"refNameAliases": {
"uri": "https://jbrowse.org/demos/circular_synteny/mm39.chromAlias.txt"
}
}jbrowse add-assembly https://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/mm39.2bit \
--name mm39 \
--refNameAliases https://jbrowse.org/demos/circular_synteny/mm39.chromAlias.txtIn JBrowse Desktop, Open new genome on the start screen (or File → Open genome... in a session), then Open from a URL and paste:
https://hgdownload.soe.ucsc.edu/goldenPath/mm39/bigZips/mm39.2bitJBrowse reads the format off the file name. Then fill in:
- Genome name:
mm39 - refName aliases (under More options):
https://jbrowse.org/demos/circular_synteny/mm39.chromAlias.txt
The chain as an indexed alignment
jbrowse.org keeps every liftOver chain UCSC publishes as an indexed PAF, one row per chain, the same files the vertebrates tutorial stacks under a locus. The track names that file and its two assemblies, query first. A chain's query is the genome it lifts to, so for hg38ToMm39 that is mouse:
Goes in the tracks array of config.json. See Tracks.
{
"type": "SyntenyTrack",
"trackId": "hg38ToMm39_liftover",
"name": "hg38 to mm39 liftOver chain",
"assemblyNames": ["mm39", "hg38"],
"adapter": {
"type": "PairwiseIndexedPAFAdapter",
"uri": "https://jbrowse.org/ucsc/hg38/liftOver/hg38ToMm39.over.pif.gz",
"csi": true,
"queryAssembly": "mm39",
"targetAssembly": "hg38"
}
}jbrowse add-track-json '{
"type": "SyntenyTrack",
"trackId": "hg38ToMm39_liftover",
"name": "hg38 to mm39 liftOver chain",
"assemblyNames": ["mm39", "hg38"],
"adapter": {
"type": "PairwiseIndexedPAFAdapter",
"uri": "https://jbrowse.org/ucsc/hg38/liftOver/hg38ToMm39.over.pif.gz",
"csi": true,
"queryAssembly": "mm39",
"targetAssembly": "hg38"
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "SyntenyTrack",
"trackId": "hg38ToMm39_liftover",
"name": "hg38 to mm39 liftOver chain",
"assemblyNames": ["mm39", "hg38"],
"adapter": {
"type": "PairwiseIndexedPAFAdapter",
"uri": "https://jbrowse.org/ucsc/hg38/liftOver/hg38ToMm39.over.pif.gz",
"csi": true,
"queryAssembly": "mm39",
"targetAssembly": "hg38"
}
}That shows it in the linear view. For the synteny view, open Add → Linear synteny view, pick the track under Quick start, and click Launch.
For your own pair, convert the alignment to a PIF. A chain goes through
chain2paf from paftools; a
PAF from minimap2 or wfmash goes straight in. Then add the track with
assemblyNames as query,target:
paftools.js chain2paf pair.over.chain.gz > pair.paf
jbrowse make-pif pair.paf
jbrowse add-track pair.pif.gz -a query,target --load copyA liftOver chain set holds a few hundred chains that cover the genome and tens
of thousands of short ones, most of them repeats and gene copies. Min length
in the view's menu keeps the short ones off the figure, and the sessions below
set it to 100 kb with minAlignmentLength.
Opening both genomes on one circle
The circular view's assembly takes a list, and each assembly lays its
chromosomes out in turn: hg38 takes the first arc of the circle and mm39 the
next, and a ribbon crosses between them. displayedRegionNames is resolved
against each assembly separately, so one list of chromosome names keeps both
genomes' unplaced contigs off the circle. The view reorders the mouse arc when
it opens (next section). The import form's Quick
start opens the same circle from the chain track (Add → Circular view, then
pick the chain track); as a session it is:
Goes at the top level of config.json, replacing any defaultSession there. See Default session.
{
"defaultSession": {
"name": "hg38 and mm39 on one circle",
"views": [
{
"type": "CircularView",
"assembly": ["hg38", "mm39"],
"minAlignmentLength": 100000,
"displayedRegionNames": [
"chr1",
"chr2",
"chr3",
"chr4",
"chr5",
"chr6",
"chr7",
"chr8",
"chr9",
"chr10",
"chr11",
"chr12",
"chr13",
"chr14",
"chr15",
"chr16",
"chr17",
"chr18",
"chr19",
"chrX"
],
"height": 780,
"tracks": ["hg38ToMm39_liftover"]
}
]
}
}jbrowse set-default-session --session - << 'EOF'
{
"name": "hg38 and mm39 on one circle",
"views": [
{
"type": "CircularView",
"assembly": ["hg38", "mm39"],
"minAlignmentLength": 100000,
"displayedRegionNames": [
"chr1",
"chr2",
"chr3",
"chr4",
"chr5",
"chr6",
"chr7",
"chr8",
"chr9",
"chr10",
"chr11",
"chr12",
"chr13",
"chr14",
"chr15",
"chr16",
"chr17",
"chr18",
"chr19",
"chrX"
],
"height": 780,
"tracks": ["hg38ToMm39_liftover"]
}
]
}
EOFhg38 and mm39 share their chromosome names, so the arcs are labelled identically around the circle: the human chromosomes run clockwise from the top and the mouse chromosomes follow, and the view's title bar names the two in that order. Each ribbon takes the ideogram colour of the human chromosome it leaves, so a human chromosome's pieces can be followed to every mouse chromosome that contains one, and a reverse alignment reads as a twist between its two ends. A row narrower than a pixel draws at the share of the pixel it covers, as in the linear synteny view, so the large blocks dominate and the short rows the filter lets through stay faint.
Ordering the second genome
The human and mouse arcs run the same way round the circle, so a mouse chromosome laid out in its native contig order sits opposite the human chromosome it does not align to, and each ribbon crosses the middle to reach its partner.
The circle runs the reorder the linear synteny view and the dotplot run on open: each mouse chromosome takes the human chromosome it shares the most aligned bases with, and the mouse arc follows that order. The circle lays it out mirrored, so the mouse arc's coordinate falls where the human arc's rises. Chords between two arcs cross only where both sides rise together, so most ribbons run between facing stretches of the two arcs. A mouse chromosome that runs antiparallel to its human partner is drawn the other way round again, and the twists left on the figure are the inversions.
Re-order chromosomes in the view's menu runs the same pass on demand, with a
progress bar and a cancel; re-running it on a circle that is already ordered
moves nothing. "autoDiagonalize": false on the view keeps each genome in its
native contig order.
The mouse genome in human chromosomes
Each stretch of a mouse chromosome's ideogram is painted the colour of the human chromosome aligned to it, and grey where nothing is. A mouse chromosome carved from one human chromosome is one colour, and one assembled from several is striped with them, so the mouse arc reads as the mouse karyotype in human pieces.
Hover the mouse chr11 band on the circle opened above. Every ribbon that misses it dims, and the tooltip lists each chromosome aligned to it with the share of it that chromosome covers, which on a mouse chromosome are the colours its band is painted in.
Colouring the ribbons by strand
The chromosome colours show where each piece went; strand shows which way round
it lies. Color by... → Strand in the view's menu paints the reverse
alignments a second colour, and Show legend names the two. A mouse
chromosome that runs antiparallel to its human partner is then one colour along
its whole bundle, and a ribbon of the other colour inside that bundle is an
inversion within it. As a session setting it is "color": { "field": "strand" }
on the view, and "field": "query" is the chromosome colouring the circle opens
with.
The X chromosome as the control
Open the first session above with displayedRegionNames cut to
["chr1", "chr2", "chrX"] to read the control. The two autosomes cross-wire
between the genomes, and the X ribbons run between the two X arcs with nothing
joining either X to an autosome.
The chain rows shorter than the cut are where the X does touch the autosomes: the PIF holds short rows between human chrX and mouse autosomes, each a repeat or a retrocopy a few hundred bases long, and the length filter keeps them off the circle. Lowering Min length in the view's menu brings them back.
Gene density as a ring
Any track that draws in the linear genome view draws on the circle as a ring
inside the ideogram, so the question of whether the conserved blocks are the
gene-rich stretches is one more track. jbrowse make-density counts a GFF3's
top-level features per bin into a bigWig, so a gene is one count however many
transcripts hang under it.
For one genome, make-density takes the GFF3 and its chrom.sizes directly:
jbrowse make-density genes.gff.gz --chrom-sizes genome.chrom.sizes --bin 100000A ring on a two-genome circle comes from one track that names both assemblies,
so both genomes' densities go into one bigWig. The script prefixes each contig
with its genome before counting, and adds the prefixed name to each assembly's
alias table as an alias of the bare name. A chr1 request from the mouse arc
then reaches mm39.chr1, and one from the human arc reaches hg38.chr1.
# one chrom.sizes and one GFF3 for the pair, every contig prefixed by genome
for g in hg38 mm39; do
awk -v g="$g" -F'\t' -v OFS='\t' '{print g "." $1, $2}' "$g.chrom.sizes" >> hg38ToMm39.chrom.sizes
gzip -dc "$g.genes.gff.gz" | awk -v g="$g" -F'\t' -v OFS='\t' '/^#/ {next} {$1 = g "." $1; print}' >> hg38ToMm39.genes.gff
# the hub's alias table, plus the prefixed name as one more alias per row
awk -v g="$g" -F'\t' -v OFS='\t' '/^#/ {print; next} {print $0, g "." $1}' "$g.hub.chromAlias.txt" > "$g.chromAlias.txt"
done
bgzip -f hg38ToMm39.genes.gff
# a whole-genome ring is megabases per pixel, so the bin is 100 kb
jbrowse make-density hg38ToMm39.genes.gff.gz --chrom-sizes hg38ToMm39.chrom.sizes --bin 100000The track is a quantitative track naming both assemblies, and on the circle it is opened as a density strip whose colour is the average over each pixel's bins:
Goes in the tracks array of config.json. See Tracks.
{
"trackId": "hg38ToMm39_gene_density",
"name": "Genes per 100 kb",
"uri": "https://jbrowse.org/demos/circular_synteny/hg38ToMm39.genes.gff.density.bw",
"assemblyNames": ["hg38", "mm39"],
"displayDefaults": {
"mark": "heatmap",
"summaryScoreMode": "avg",
"color": {
"field": "score",
"scale": "threshold",
"range": ["#e01e26", "#d95f02"]
},
"height": 40
}
}jbrowse add-track https://jbrowse.org/demos/circular_synteny/hg38ToMm39.genes.gff.density.bw \
--trackId hg38ToMm39_gene_density \
--name "Genes per 100 kb" \
--assemblyNames hg38,mm39 \
--displayDefaults '{"mark":"heatmap","summaryScoreMode":"avg","color":{"field":"score","scale":"threshold","range":["#e01e26","#d95f02"]},"height":40}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"trackId": "hg38ToMm39_gene_density",
"name": "Genes per 100 kb",
"uri": "https://jbrowse.org/demos/circular_synteny/hg38ToMm39.genes.gff.density.bw",
"assemblyNames": ["hg38", "mm39"],
"displayDefaults": {
"mark": "heatmap",
"summaryScoreMode": "avg",
"color": {
"field": "score",
"scale": "threshold",
"range": ["#e01e26", "#d95f02"]
},
"height": 40
}
}The color is the pair Edit colors/arrangement... in the track menu sets,
one colour below the baseline and one above; a heatmap of counts fades from
white to the one above.
Rings stack inward from the ideogram in the order the view's tracks lists
them, and the ribbons draw inside the innermost ring, so the density entry goes
before the synteny track:
Goes at the top level of config.json, replacing any defaultSession there. See Default session.
{
"defaultSession": {
"name": "hg38 and mm39 with a gene density ring",
"views": [
{
"type": "CircularView",
"assembly": ["hg38", "mm39"],
"minAlignmentLength": 100000,
"displayedRegionNames": [
"chr1",
"chr2",
"chr3",
"chr4",
"chr5",
"chr6",
"chr7",
"chr8",
"chr9",
"chr10",
"chr11",
"chr12",
"chr13",
"chr14",
"chr15",
"chr16",
"chr17",
"chr18",
"chr19",
"chrX"
],
"height": 780,
"tracks": [
{
"trackId": "hg38ToMm39_gene_density",
"type": "LinearWiggleDisplay",
"mark": "heatmap",
"summaryScoreMode": "avg",
"color": {
"field": "score",
"scale": "threshold",
"range": ["#e01e26", "#d95f02"]
},
"height": 40
},
"hg38ToMm39_liftover"
]
}
]
}
}jbrowse set-default-session --session - << 'EOF'
{
"name": "hg38 and mm39 with a gene density ring",
"views": [
{
"type": "CircularView",
"assembly": ["hg38", "mm39"],
"minAlignmentLength": 100000,
"displayedRegionNames": [
"chr1",
"chr2",
"chr3",
"chr4",
"chr5",
"chr6",
"chr7",
"chr8",
"chr9",
"chr10",
"chr11",
"chr12",
"chr13",
"chr14",
"chr15",
"chr16",
"chr17",
"chr18",
"chr19",
"chrX"
],
"height": 780,
"tracks": [
{
"trackId": "hg38ToMm39_gene_density",
"type": "LinearWiggleDisplay",
"mark": "heatmap",
"summaryScoreMode": "avg",
"color": {
"field": "score",
"scale": "threshold",
"range": ["#e01e26", "#d95f02"]
},
"height": 40
},
"hg38ToMm39_liftover"
]
}
]
}
EOFThe ring under a mirrored arc is mirrored with it, so a bin sits under the stretch of ideogram it belongs to whichever way round that chromosome is drawn.
Reading a ribbon back
Cut the ring session's displayedRegionNames to ["chr1", "chr2", "chrX"] and
hover the widest ribbon, the X block that runs reverse between the two genomes,
so it twists. JBrowse fills it in the hover colour, and a tooltip names the
alignment: its span in each genome and which way round the two read.
The two loci in the tooltip come from one row of the PIF. tabix returns that
row at the human coordinate the tooltip starts at:
# the t prefix asks for the row in hg38 coordinates; q would ask in mm39's
tabix https://jbrowse.org/ucsc/hg38/liftOver/hg38ToMm39.over.pif.gz tchrX:10447551-10447552The row's strand column reads -, and its query span on mouse chrX is the other
end of the ribbon. The ring reads back the same way: bigWigToBedGraph prints
the bins under any stretch of either arc, addressed by the prefixed contig name
the bigWig has.
# the first 100 kb of the X block, then the start of human chr19
bigWigToBedGraph -chrom=hg38.chrX -start=10447550 -end=10547550 \
https://jbrowse.org/demos/circular_synteny/hg38ToMm39.genes.gff.density.bw stdout
bigWigToBedGraph -chrom=hg38.chr19 -start=0 -end=300000 \
https://jbrowse.org/demos/circular_synteny/hg38ToMm39.genes.gff.density.bw stdoutA bin's value is the count of genes starting in it, so the pale X ring and the dark chr19 ring are the two counts side by side.
Your own pair of genomes
The script takes any UCSC pair by name. Your own pair needs the same files it writes:
- an alignment as a PIF, sorted and tabix-indexed. jbrowse.org's copy of a UCSC
liftOver chain opens as is, with
"csi": truebeside its URL; a pair with no hosted chain goes through themake-pifcommands above. Whichever way, the synteny track'sassemblyNamesis[query, target], and the reorder reads the same file - both assemblies declared in the config, from a hub entry or the
addassemblyblocks above (jbrowse add-assembly genome.fa) - one bigWig per ring, naming the ring's assembly; for a ring that covers both genomes, one bigWig with prefixed contigs and an alias per assembly as above
Reproduce it end to end
The script works in four steps:
- Fetch the hg38 and mm39 hub configs, chromosome sizes and alias tables, and both RefSeq curated gene sets.
- Prefix every contig with its genome and add the prefixed name to each alias table, so one bigWig can hold both genomes' counts.
- Count genes per 100 kb into that bigWig with
jbrowse make-density, one count per gene however many transcripts it has. - Write the config with the chain track pointing at jbrowse.org's indexed copy, which the ribbons and the reorder both read.
TARGET and QUERY pick another UCSC pair, PIF points the track at a chain
indexed elsewhere, MIN_LENGTH moves the cut, and BASE_URL is the prefix the
finished files will be served under; see Prerequisites.
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_circular_synteny.sh
bash build_circular_synteny.shSee also
- Circular genome view
- Synteny track
- Synteny from liftOver chains (hg38 and eight vertebrates)
- Synteny visualization (pairwise minimap2)
- Gene density and transposon density along a chromosome
Citations
- Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D. Evolution's cauldron: duplication, deletion, and rearrangement in the mouse and human genomes. Proc Natl Acad Sci USA (2003). https://doi.org/10.1073/pnas.1932072100
- Ohno S. Sex Chromosomes and Sex-Linked Genes. Springer (1967), the argument that the mammalian X chromosome's gene content is conserved across species.
- Krzywinski M, Schein J, Birol I, et al. Circos: an information aesthetic for comparative genomics. Genome Research 19:1639-1645 (2009). https://doi.org/10.1101/gr.092759.109
Feedback on this tutorial is welcome: contact us.