Synteny by ancestral linkage group (sponge, comb jelly, jellyfish)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.11. Desktop beta builds are coming soon.
Certain sets of genes have stayed on the same chromosome together since before animals existed, and those sets have names. odp writes, for every pair of genomes it compares, a table of their orthologs with the set each one belongs to and a color for it. JBrowse loads that table as a synteny track, paints every ortholog with its set's color, and sorts one genome's chromosomes by where their orthologs land on the other. In a sponge the colors fall in one block per chromosome along a diagonal, and in a comb jelly each group is spread over several chromosomes. The color mode is a column of the table, so any label a pipeline puts beside an ortholog can drive it.
Prerequisites
- A web browser, to fetch two files from Dryad by hand
python3samtoolscurlandtarnode, for the JBrowse CLI- A running JBrowse instance (the web quickstart or the desktop quickstart)
Where the data comes from
The genomes and the ortholog tables are the Dryad deposit behind Schultz et al. 2023, released CC0.
- Both tarballs,
genomes.tar.gzandsupplementary_information.tar.gz: datadryad.org/…/dryad.dncjsxm47https://datadryad.org/dataset/doi:10.5061/dryad.dncjsxm47 - The Ephydatia assembly, which a script in the deposit fetches from the original host: bitbucket.org/…/Emu_genome_v1.fa.gzhttps://bitbucket.org/EphydatiaGenome/ephydatiagenome/downloads/Emu_genome_v1.fa.gz
The linkage-group label in the ortholog table
Simakov et al. 2022 named a set of gene families the BCnS linkage groups, after
the bilaterians, cnidarians and sponges whose chromosomes have them, and gave
each a letter: A1a, A2, B1, and so on to R. A gene belongs to one of them or to
none. odp ships the groups as a database of protein models, searches every
proteome it is given against them, and writes the group each ortholog landed in
as a column beside it. A synteny track can read that column. Your own genomes
need odp's table with its gene_group and color columns, written when odp
runs with plot_LGs: True, and a .chrom file per genome.
The dotplot uses two genomes and the stack at the end adds four:
RES, the jellyfish Rhopilema esculentum, in the dotplot and the stackEMU, the freshwater sponge Ephydatia muelleri, in the dotplot and the stackHCAandBIN, the comb jellies Hormiphora californensis and Bolinopsis microptera, in the stack onlyBFL, the amphioxus Branchiostoma floridae, in the stack onlyCLAa, a cladorhizid sponge, in the stack only
odp's group database was built from five genomes, three of them on this page: the jellyfish, Ephydatia and amphioxus. The comb jellies and the cladorhizid took no part in it. odp compares genomes two at a time, so the deposit holds one table per pair, each of them every reciprocal best protein hit between the two.
A table's row is the gene pair, the group, where each gene sits, and the color:
rbh EMU_gene RES_gene gene_group EMU_scaf EMU_pos RES_scaf RES_pos ... color
rbh2way_EMU_RES_964 Em0019g38a mRNA.RE04286 A1a EMU19 180964 RES2 13815510 ... #C23D51
rbh2way_EMU_RES_4123 Em0019g57a mRNA.RE14076 None EMU19 290216 RES8 13037026 ... #000000Gene intervals come from odp's .chrom files, since the table's position column
holds one coordinate per gene and would make every feature one base long:
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/rbh_to_blocks.py
# --species sets the column order, anchor first, which the track then follows
# --chrom gives each species' genes their real start and stop
# --attributes names the columns that pass through after the gene ids
python3 rbh_to_blocks.py EMU_RES_xy_reciprocal_best_hits.coloredby_BCnS_LGs.plotted.rbh \
-o RES_EMU.blocks --bed-dir RES_EMU \
--species RES EMU --chrom RES=RES.chrom EMU=EMU.chrom \
--attributes gene_group color break_FETThe helper prints how many of each genome's gene ids its .chrom placed. An
ortholog from outside the BCnS families gets . in the group column, which the
browser draws in grey.
A .blocks row is the gene ids across the two genomes, then the attribute
columns:
mRNA.RE04286 Em0019g38a A1a #C23D51 0.0000
mRNA.RE14076 Em0019g57a . #000000 1.8213Loading the genomes
We'll load the two genomes of the dotplot. The genome FASTA needs a .fai
beside it (samtools faidx), and its sequence names must match the table's
scaffold columns. The stack adds four more assemblies the same way, one block
each.
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "RES",
"displayName": "Rhopilema (jellyfish)",
"uri": "RES.fa"
}jbrowse add-assembly RES.fa \
--name RES \
--displayName "Rhopilema (jellyfish)" \
--load copyIn JBrowse Desktop, Open new genome on the start screen (or File → Open genome... in a session), then Open from a URL and paste, one per line:
RES.fa
RES.fa.faiJBrowse reads the format off the file name. Then fill in:
- Genome name:
RES - Assembly display name (under More options):
Rhopilema (jellyfish)
RES.fa is relative to a config.json. Replace it with its URL or its path on this computer.
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "EMU",
"displayName": "Ephydatia (sponge)",
"uri": "EMU.fa"
}jbrowse add-assembly EMU.fa \
--name EMU \
--displayName "Ephydatia (sponge)" \
--load copyIn JBrowse Desktop, Open new genome on the start screen (or File → Open genome... in a session), then Open from a URL and paste, one per line:
EMU.fa
EMU.fa.faiJBrowse reads the format off the file name. Then fill in:
- Genome name:
EMU - Assembly display name (under More options):
Ephydatia (sponge)
EMU.fa is relative to a config.json. Replace it with its URL or its path on this computer.
Loading the RES-EMU table as a synteny track
attributeColumns makes the last three columns reachable. Each name in it
becomes a color-by mode named after the column, so gene_group becomes a mode;
color is the palette the file puts beside each label, and the menu leaves it
out of the modes. break_FET is odp's test of the row's chromosome pair, which
the stack below reads as opacity.
Goes in the tracks array of config.json. See Tracks.
{
"type": "SyntenyTrack",
"trackId": "RES_EMU",
"name": "RES vs EMU orthologs, colored by BCnS linkage group",
"assemblyNames": ["RES", "EMU"],
"adapter": {
"type": "MCScanBlocksAdapter",
"uri": "RES_EMU.blocks.gz",
"blockAssemblies": ["RES", "EMU"],
"bedLocations": ["RES_EMU.RES.bed.gz", "RES_EMU.EMU.bed.gz"],
"attributeColumns": ["gene_group", "color", "break_FET"]
}
}jbrowse add-track-json '{
"type": "SyntenyTrack",
"trackId": "RES_EMU",
"name": "RES vs EMU orthologs, colored by BCnS linkage group",
"assemblyNames": ["RES", "EMU"],
"adapter": {
"type": "MCScanBlocksAdapter",
"uri": "RES_EMU.blocks.gz",
"blockAssemblies": ["RES", "EMU"],
"bedLocations": ["RES_EMU.RES.bed.gz", "RES_EMU.EMU.bed.gz"],
"attributeColumns": ["gene_group", "color", "break_FET"]
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "SyntenyTrack",
"trackId": "RES_EMU",
"name": "RES vs EMU orthologs, colored by BCnS linkage group",
"assemblyNames": ["RES", "EMU"],
"adapter": {
"type": "MCScanBlocksAdapter",
"uri": "RES_EMU.blocks.gz",
"blockAssemblies": ["RES", "EMU"],
"bedLocations": ["RES_EMU.RES.bed.gz", "RES_EMU.EMU.bed.gz"],
"attributeColumns": ["gene_group", "color", "break_FET"]
}
}RES_EMU.blocks.gz, RES_EMU.RES.bed.gz, RES_EMU.EMU.bed.gz are relative to a config.json. Replace each with its URL or its path on this computer.
That shows it in the linear view. For the synteny view, open Add → Linear synteny view, pick the track under Quick start, and click Launch.
The build script adds one such track per pair: this one for the dotplot, then one for each neighbouring pair of the six-genome stack.
One block per linkage group
Open the jellyfish-against-sponge table as a dotplot and pick gene_group under Color by value on the palette button in the view header; the legend comes up with it. Then Re-order chromosomes on the view menu sorts the vertical genome's chromosomes by where their orthologs land along the horizontal one, turning one block per group into a diagonal. The same view as a session, with the sponge's unplaced scaffolds left off its axis:
Goes at the top level of config.json, replacing any defaultSession there. See Default session.
{
"defaultSession": {
"name": "Ancestral linkage groups",
"views": [
{
"type": "DotplotView",
"displayName": "Rhopilema (jellyfish) vs Ephydatia (sponge)",
"views": [
{ "assembly": "RES" },
{ "assembly": "EMU", "displayedRegionNames": ["EMU*"] }
],
"tracks": ["RES_EMU"],
"color": { "field": "gene_group" },
"autoDiagonalize": true,
"lineWidth": 4,
"height": 860
}
]
}
}jbrowse set-default-session --session - << 'EOF'
{
"name": "Ancestral linkage groups",
"views": [
{
"type": "DotplotView",
"displayName": "Rhopilema (jellyfish) vs Ephydatia (sponge)",
"views": [
{ "assembly": "RES" },
{ "assembly": "EMU", "displayedRegionNames": ["EMU*"] }
],
"tracks": ["RES_EMU"],
"color": { "field": "gene_group" },
"autoDiagonalize": true,
"lineWidth": 4,
"height": 860
}
]
}
EOFInside a block the points fill the square: the order of genes along each chromosome has been shuffled, and which chromosome each gene sits on has not.
Six genomes stacked
Figure 1d of the paper stacks the same tables: one row per genome, the ribbons between neighbours colored by group. The linear synteny view does that with one pair's track per band. The build script loads the pairs in the paper's order, two comb jellies over the jellyfish, amphioxus and two sponges, and the session below sets what the figure needs.
odp tests each pair of chromosomes for more shared orthologs than chance
(Fisher's exact test, corrected for the number of pairs, in the break_FET
column). Like the paper's figure, the stack draws an ortholog on a pair under
0.05 at opacity 0.8 and the rest at 0.15.
Each setting in the session has a menu route except fadeThinAlignmentsMode and
the opacity mapping:
autoDiagonalizesorts each row against its neighbour, working outward fromdiagonalizeAnchorRow. Rows count from 0, so 2 is the jellyfish, where each group sits on one chromosome. Rows → Re-order chromosomes on the view menu runs the same sort and asks for that row.hideUnlabelleddraws only the orthologs in a group: Hide unlabelled rows on the palette menu.drawCurvesbundles the ribbons: Curved lines on the sliders button.fadeThinAlignmentsModeturns off the fade a whole-genome view applies to sub-pixel ribbons, since their colors are what the figure shows. It has no menu item.opacityreadsbreak_FETthrough a threshold at 0.05. Opacity on the sliders button scales both opacities together; the mapping itself has no menu item.
Goes at the top level of config.json, replacing any defaultSession there. See Default session.
{
"defaultSession": {
"name": "Ancestral linkage groups, six genomes",
"views": [
{
"type": "LinearSyntenyView",
"displayName": "Comb jellies, jellyfish, amphioxus, sponges",
"views": [
{ "assembly": "BIN", "displayedRegionNames": ["BIN*"] },
{ "assembly": "HCA", "displayedRegionNames": ["HCA*"] },
{ "assembly": "RES", "displayedRegionNames": ["RES*"] },
{ "assembly": "BFL", "displayedRegionNames": ["BFL*"] },
{ "assembly": "EMU", "displayedRegionNames": ["EMU*"] },
{ "assembly": "CLAa", "displayedRegionNames": ["CLA*_hapA"] }
],
"tracks": [
["BIN_HCA"],
["RES_HCA"],
["RES_BFL"],
["BFL_EMU"],
["EMU_CLAa"]
],
"color": { "field": "gene_group" },
"hideUnlabelled": true,
"autoDiagonalize": true,
"diagonalizeAnchorRow": 2,
"drawCurves": true,
"fadeThinAlignmentsMode": "off",
"opacity": {
"field": "break_FET",
"scale": "threshold",
"domain": [0.05],
"range": [0.8, 0.15]
},
"levelHeights": [130, 130, 130, 130, 130],
"collapseEmptyRows": true
}
]
}
}jbrowse set-default-session --session - << 'EOF'
{
"name": "Ancestral linkage groups, six genomes",
"views": [
{
"type": "LinearSyntenyView",
"displayName": "Comb jellies, jellyfish, amphioxus, sponges",
"views": [
{ "assembly": "BIN", "displayedRegionNames": ["BIN*"] },
{ "assembly": "HCA", "displayedRegionNames": ["HCA*"] },
{ "assembly": "RES", "displayedRegionNames": ["RES*"] },
{ "assembly": "BFL", "displayedRegionNames": ["BFL*"] },
{ "assembly": "EMU", "displayedRegionNames": ["EMU*"] },
{ "assembly": "CLAa", "displayedRegionNames": ["CLA*_hapA"] }
],
"tracks": [
["BIN_HCA"],
["RES_HCA"],
["RES_BFL"],
["BFL_EMU"],
["EMU_CLAa"]
],
"color": { "field": "gene_group" },
"hideUnlabelled": true,
"autoDiagonalize": true,
"diagonalizeAnchorRow": 2,
"drawCurves": true,
"fadeThinAlignmentsMode": "off",
"opacity": {
"field": "break_FET",
"scale": "threshold",
"domain": [0.05],
"range": [0.8, 0.15]
},
"levelHeights": [130, 130, 130, 130, 130],
"collapseEmptyRows": true
}
]
}
EOFIn the second band, between Hormiphora and the jellyfish, far fewer grouped orthologs sit on a significant pair than in the other bands. Each comb jelly chromosome mixes groups that match no jellyfish chromosome, and those orthologs draw faint. The two comb jellies agree with each other in the band above.
Amphioxus and Ephydatia helped build the group database, so the bundles running through them are expected. The cladorhizid took no part, and in the bottom band the groups still reach it as bundles.
Checking the linkage-group counts against the odp tables
The counts behind the pictures come straight from the tables. For one group, tally the chromosomes its orthologs sit on in the genome a table's file name leads with:
# gene_group is column 4; column 5 is the chromosome in the genome the
# file name leads with, EMU here
awk -F'\t' 'NR>1 && $4=="A1a" {print $5}' \
EMU_RES_xy_reciprocal_best_hits.coloredby_BCnS_LGs.plotted.rbh \
| sort | uniq -c | sort -rnChanging the file and the group in that command gives the counts behind the
stack. In Hormiphora (HCA_RES), even the chromosome holding the largest
share of a group holds less than half of it, which is why most of the stack's
second band is faint. In the cladorhizid (CLAa_EMU), which took no part in
building the group database, most of each group still sits on one chromosome.
Reproduce it end to end
Download genomes.tar.gz and supplementary_information.tar.gz from
the Dryad page into
~/Downloads first, since Dryad has no direct download URL. The script needs
the tools under Prerequisites and works in three steps:
- Unpack the genomes and pair tables the page uses from the two tarballs, and fetch the Ephydatia assembly, which the deposit does not redistribute, from its original host.
- Convert each pair's odp table into a
.blockstable and two BEDs, taking each gene's extent from the.chromfiles and keeping the group, color andbreak_FETcolumns. - Write the six assemblies, one track per pair, and the dotplot session.
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_odp_linkage_groups_synteny.sh
DRYAD_DIR=~/Downloads bash build_odp_linkage_groups_synteny.shFor your own genomes, run odp with plot_LGs: True and point the conversion at
the tables it writes under
synteny_analysis/step2-figures/synteny_coloredby_BCnS_LGs/.
See also
- Synteny from an ortholog table (grape, peach, cacao)
- Synteny visualization (OrthoFinder orthogroups)
- Pangenome (HPRC) part 2: haplotypes against each other
- Linear synteny view
External links
- odp, the oxford dot plot toolkit: https://github.com/conchoecia/odp
Citations
- Schultz DT, Haddock SHD, Bredeson JV, Green RE, Simakov O, Rokhsar DS. Ancient gene linkages support ctenophores as sister to other animals. Nature (2023). https://doi.org/10.1038/s41586-023-05936-6
- Simakov O, Bredeson J, Berkoff K, et al. Deeply conserved synteny and the evolution of metazoan chromosomes. Sci Adv (2022). https://doi.org/10.1126/sciadv.abi5884
Feedback on this tutorial is welcome: contact us.