Pangenome (cattle)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.11. Desktop beta builds are coming soon.
The bovine super-pangenome builds twelve assemblies into one graph on the
ARS-UCD1.2 cattle reference, a Hereford that is itself one of the twelve:
taurine and indicine breeds, yak, bison and gaur. Each assembly is a named path
through the graph, and vg deconstruct turns those paths into a VCF, so one
locus reads both as a graph of where sequence is present and absent and as a
callset listing which assemblies have it. We:
- read a whole chromosome with one node per variant region
- at HSPA1A, compare the graph's list of alleles with the callset
- find three published breed and species variants in the callset
Every step starts from the graph's page on staging.genomes.jbrowse.org, where it stays until the graph plugin's JBrowse 5 host ships.
The graph view is a beta plugin. We welcome your feedback.
Prerequisites
python3and htslib (bgzip,tabix), to build the OMIA trackvg, to deconstruct the graph into a callset
Where the data comes from
The graph's index files, the callset and its sample table are hosted beside the graph, and OMIA supplies the curated causal variants:
- the segment and link index: jbrowse.org/…/bovine-arsucd12-minigraph.segs.bed.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.segs.bed.gz and jbrowse.org/…/bovine-arsucd12-minigraph.links.bed.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.links.bed.gz
- the bubble index, where assemblies' paths split and rejoin: jbrowse.org/…/bovine-arsucd12-minigraph.bubbles.bed.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.bubbles.bed.gz
- the allele inventory, one row per alternative sequence at a bubble: jbrowse.org/…/bovine-arsucd12-minigraph.alleles.bed.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.alleles.bed.gz
- the overview used for whole chromosomes, one node per bubble: jbrowse.org/…/bovine-arsucd12-minigraph.tier10000.segs.bed.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.tier10000.segs.bed.gz and jbrowse.org/…/bovine-arsucd12-minigraph.tier10000.links.bed.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.tier10000.links.bed.gz
- the deconstructed callset: jbrowse.org/…/bovine-arsucd12-minigraph.vcf.gzhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.vcf.gz
- the breed and lineage of each assembly in the callset: jbrowse.org/…/bovine-arsucd12-minigraph.samples.tsvhttps://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.samples.tsv
- OMIA's database dump, the source of the curated variant track: omia.org/…/omia.sql.gzhttps://omia.org/static/omia.sql.gz
Hosting your own graph describes each file and how to produce it.
Load the genome and the graph
We'll load the reference the graph is anchored to, then the graph track. Swap
the prefix in uri for your own build of build_pangenome_graph.sh, which
writes the tabix-indexed segments and links and the overview.
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "bosTau9",
"aliases": ["ARS-UCD1.2"],
"uri": "https://hgdownload.soe.ucsc.edu/goldenPath/bosTau9/bigZips/bosTau9.2bit",
"refNameAliases": {
"uri": "https://jbrowse.org/ucsc/bosTau9/bosTau9.chromAlias.txt"
}
}jbrowse add-assembly https://hgdownload.soe.ucsc.edu/goldenPath/bosTau9/bigZips/bosTau9.2bit \
--name bosTau9 \
--alias ARS-UCD1.2 \
--refNameAliases https://jbrowse.org/ucsc/bosTau9/bosTau9.chromAlias.txtGoes in the tracks array of config.json. See Tracks.
{
"type": "GraphTrack",
"trackId": "bovine_minigraph_segments",
"name": "Bovine super-pangenome (rGFA segments)",
"assemblyNames": ["bosTau9"],
"adapter": {
"type": "RgfaTabixAdapter",
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph",
"coarse": {
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.tier10000",
"aboveBpPerPx": 880
}
},
"displays": [
{ "type": "LinearGraphDisplay" },
{ "type": "LinearBasicDisplay" }
]
}jbrowse add-track-json '{
"type": "GraphTrack",
"trackId": "bovine_minigraph_segments",
"name": "Bovine super-pangenome (rGFA segments)",
"assemblyNames": ["bosTau9"],
"adapter": {
"type": "RgfaTabixAdapter",
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph",
"coarse": {
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.tier10000",
"aboveBpPerPx": 880
}
},
"displays": [
{ "type": "LinearGraphDisplay" },
{ "type": "LinearBasicDisplay" }
]
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "GraphTrack",
"trackId": "bovine_minigraph_segments",
"name": "Bovine super-pangenome (rGFA segments)",
"assemblyNames": ["bosTau9"],
"adapter": {
"type": "RgfaTabixAdapter",
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph",
"coarse": {
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.tier10000",
"aboveBpPerPx": 880
}
},
"displays": [
{ "type": "LinearGraphDisplay" },
{ "type": "LinearBasicDisplay" }
]
}Overview of chr23 with one node per variant region
Click chr23 on the Whole chromosome line of the portal page. The whole chromosome opens with the graph drawn as one node per bubble, a stretch where the assemblies' paths split and rejoin.
The curve peaks over BoLA, the cattle major histocompatibility complex (immune genes that vary a lot between breeds). The heat shock gene HSPA1A, the subject of the next section, lies inside it. The portal's Loci table ranks the genome's most variable stretches this way; see Ranking the graph's bubbles.
HSPA1A in the graph and the callset
ARS-UCD1.2 lacks an 11 kb segment beside the heat shock gene HSPA1A, which
contains HSPA1B, its near-identical copy. Leonard et al. (2022) recovered it
in every assembly they built. The graph lists that segment in the allele
inventory. These graphs record no build order, so its firstSeenIn column names
the first assembly in a fixed list, which need not carry the segment.
The callset gives a genotype per assembly. We ran vg deconstruct once per
chromosome over the same graph:
# -p: vg's PackedGraph format, which deconstruct reads
vg convert -g chr1.gfa -p > chr1.vg
# -p: the reference path to decompose against
# -a: nested snarls too, so a bubble inside a bubble gets a record
vg deconstruct -p chr1 -a -t 8 chr1.vg > chr1.vcfThe VCF's CHROM column is the -p path name, which has to equal the assembly's
refName. Its sample names are the three-letter codes on each assembly's path,
and the sample table's first column has to match them; a mismatch leaves the row
unlabelled without an error:
name breed lineage
ANG Angus taurine
BIS Bison bisonThe track config uses the other columns, a breed and a lineage per code:
rows.labelswrites the breed beside each rowrowColortints each row by lineagerows.domainlists the cattle breeds above the wild speciesunit: "haplotype"draws one row per assembly, each being one haplotype, with a second alternate allele in a separate colourshowVariantLanedraws each call once in a lane above the rows, across the reference span it replaces, labelled with its VCF ID (the graph nodes that bound it) and the allele change
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "bovine_pangenome_vcf",
"name": "Bovine super-pangenome variants (12 assemblies)",
"assemblyNames": ["bosTau9"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.vcf.gz",
"samplesTsvLocation": {
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.samples.tsv"
}
},
"displays": [
{
"type": "LinearMultiSampleVariantDisplay",
"unit": "haplotype",
"showVariantLane": true,
"rows": {
"domain": [
"ANG",
"BSW",
"HIG",
"OBV",
"PIE",
"SIM",
"BRA",
"NEL",
"GAU",
"BIS",
"YAK"
],
"labels": {
"ANG": "Angus",
"BSW": "Brown Swiss",
"HIG": "Highland",
"OBV": "Original Braunvieh",
"PIE": "Piedmontese",
"SIM": "Simmental",
"BRA": "Brahman",
"NEL": "Nellore",
"GAU": "Gaur",
"BIS": "Bison",
"YAK": "Yak"
}
},
"rowColor": {
"field": "lineage",
"domain": ["taurine", "indicine", "gaur", "bison", "yak"],
"range": ["#0072B2", "#E69F00", "#009E73", "#CC79A7", "#D55E00"]
}
}
]
}jbrowse add-track-json '{
"type": "VariantTrack",
"trackId": "bovine_pangenome_vcf",
"name": "Bovine super-pangenome variants (12 assemblies)",
"assemblyNames": ["bosTau9"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.vcf.gz",
"samplesTsvLocation": {
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.samples.tsv"
}
},
"displays": [
{
"type": "LinearMultiSampleVariantDisplay",
"unit": "haplotype",
"showVariantLane": true,
"rows": {
"domain": [
"ANG",
"BSW",
"HIG",
"OBV",
"PIE",
"SIM",
"BRA",
"NEL",
"GAU",
"BIS",
"YAK"
],
"labels": {
"ANG": "Angus",
"BSW": "Brown Swiss",
"HIG": "Highland",
"OBV": "Original Braunvieh",
"PIE": "Piedmontese",
"SIM": "Simmental",
"BRA": "Brahman",
"NEL": "Nellore",
"GAU": "Gaur",
"BIS": "Bison",
"YAK": "Yak"
}
},
"rowColor": {
"field": "lineage",
"domain": ["taurine", "indicine", "gaur", "bison", "yak"],
"range": ["#0072B2", "#E69F00", "#009E73", "#CC79A7", "#D55E00"]
}
}
]
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "VariantTrack",
"trackId": "bovine_pangenome_vcf",
"name": "Bovine super-pangenome variants (12 assemblies)",
"assemblyNames": ["bosTau9"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.vcf.gz",
"samplesTsvLocation": {
"uri": "https://jbrowse.org/demos/bovine_pangenome/bovine-arsucd12-minigraph.samples.tsv"
}
},
"displays": [
{
"type": "LinearMultiSampleVariantDisplay",
"unit": "haplotype",
"showVariantLane": true,
"rows": {
"domain": [
"ANG",
"BSW",
"HIG",
"OBV",
"PIE",
"SIM",
"BRA",
"NEL",
"GAU",
"BIS",
"YAK"
],
"labels": {
"ANG": "Angus",
"BSW": "Brown Swiss",
"HIG": "Highland",
"OBV": "Original Braunvieh",
"PIE": "Piedmontese",
"SIM": "Simmental",
"BRA": "Brahman",
"NEL": "Nellore",
"GAU": "Gaur",
"BIS": "Bison",
"YAK": "Yak"
}
},
"rowColor": {
"field": "lineage",
"domain": ["taurine", "indicine", "gaur", "bison", "yak"],
"range": ["#0072B2", "#E69F00", "#009E73", "#CC79A7", "#D55E00"]
}
}
]
}In the chr23 view the portal opened, type chr23:27,508,000-27,536,000, and the
graph track draws the segments around HSPA1A. Turn on the callset and the
allele inventory in the track selector. An insertion has no reference span to
draw along, so in the graph track's menu:
- pick Layout → Force-directed layout
- pick Bubble spread → Compress lengths
The figure shows all three tracks under the RefSeq genes.
The yak row has the reference allele. Leonard et al. built no yak assembly, so their result does not cover it.
Leonard et al. published the bovine graph with a path line per assembly, and
vg deconstruct reads those paths. A graph straight out of minigraph has no
path lines and so no callset to deconstruct; Pangenome (mouse)
starts from such a graph.
Published variants in the callset
Each locus below is a structural variant a paper reported in one of these breeds or species. The rows without it are the control.
The Celtic polled allele
Angus cattle are born without horns. Medugorac et al. (2012) traced the Celtic form of polledness to a 202 bp duplication and insertion on chromosome 1, which OMIA, the Online Mendelian Inheritance in Animals database, curates as OMIA 000483-9913.
build_omia_cattle_variants.sh
writes omia_cattle_variants.gff3.gz from OMIA's nightly database dump. It
keeps the cattle records published on ARS-UCD1.2 or ARS-UCD1.3, which share
every chromosome's coordinates:
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_omia_cattle_variants.sh
bash build_omia_cattle_variants.shAdd the records as a track:
Goes in the tracks array of config.json. See Tracks.
{
"trackId": "omia_cattle_variants",
"name": "OMIA causal variants (cattle)",
"uri": "omia_cattle_variants.gff3.gz",
"assemblyNames": ["bosTau9"],
"displayDefaults": {
"labels": {
"description": "jexl:feature.inheritance"
}
}
}jbrowse add-track omia_cattle_variants.gff3.gz \
--trackId omia_cattle_variants \
--name "OMIA causal variants (cattle)" \
--assemblyNames bosTau9 \
--displayDefaults '{"labels":{"description":"jexl:feature.inheritance"}}' \
--load copyIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"trackId": "omia_cattle_variants",
"name": "OMIA causal variants (cattle)",
"uri": "omia_cattle_variants.gff3.gz",
"assemblyNames": ["bosTau9"],
"displayDefaults": {
"labels": {
"description": "jexl:feature.inheritance"
}
}
}omia_cattle_variants.gff3.gz is relative to a config.json. Replace it with its URL or its path on this computer.
Then open chr1:2,424,000-2,436,000.
OMIA also records the Friesian polled allele, an 80 kb duplication 200 kb further along, which Holstein cattle have. The panel has no Holstein, and the callset holds nothing that size there.
A repeat upstream of KIT in white-headed cattle
Simmental and Hereford cattle have white heads. Milia et al. (2025) tied the
trait to a 14.3 kb segment repeated in tandem upstream of KIT: white-headed
breeds have extra copies, colour-headed breeds a deletion. The Hereford
reference holds one collapsed copy. Open chr6:70,080,000-70,180,000.
The larger box at the left of the Simmental row is an insertion the length of the 14.3 kb segment: Simmental has one more copy than the Hereford reference.
A deletion of TAS2R46 in gaur
TAS2R46 encodes a bitter taste receptor. Leonard et al. (2022) found a 17 kb
deletion in gaur that removes it. Open chr5:98,575,000-98,615,000.
The four cattle rows, Angus, Piedmontese, Brahman and Nellore, hold an allele slightly longer than the reference, and the insertion boxed in each row, between TAS2R46 and the next gene, is most of the difference. Two taurine and two indicine breeds share that allele, and only the gaur's deletes the span.
Building the bovine graph files
Pangenome (hosting your own graph) turns a graph into the files above
with one command, build_pangenome_graph.sh. The published bovine graphs need
three steps beyond it, which
build_bovine_pangenome.sh
runs:
- Join the autosomes. The archive holds one graph per autosome, each numbering its segments from 1, so the script renumbers and concatenates them.
- Recover rGFA tags. The graphs state their coordinates in P lines, and
gfa_paths_to_rgfa.pyconverts them back intoSN/SO/SRtags (each segment's contig, offset and rank). - Deconstruct the callset, with the
vg deconstructcall above.
The whole build takes about half an hour after the download.
The script stops if the reference path does not reproduce bosTau9's chromosome lengths, and a README.txt beside the hosted files records the source and every modification.
Leonard et al. also built pggb and Minigraph-Cactus graphs of the same twelve
assemblies, base-level GFAs with path lines and no rGFA tags.
build_pangenome_graph.sh reads their paths and writes an SM:Z: tag per
segment listing the assemblies that pass through it, which the graph track
shows.
See also
- Pangenome (mouse)
- Pangenome (HPRC) part 1: the graph's alleles and the haplotypes that have them
- Pangenome (hosting your own graph)
Citations
- Leonard AS, Crysnanto D, Fang ZH, Heaton MP, Vander Ley BL, Herrera C, Bollwein H, Bickhart DM, Kuhn KL, Smith TPL, Rosen BD, Pausch H. Structural variant-based pangenome construction has low sensitivity to variability of haplotype-resolved bovine assemblies. Nature Communications. 2022;13:3012. https://doi.org/10.1038/s41467-022-30680-2
- Leonard AS, Crysnanto D, Mapel XM, Bhati M, Pausch H. Graph construction method impacts variation representation and analyses in a bovine super-pangenome. Genome Biology. 2023;24:124. https://doi.org/10.1186/s13059-023-02969-y
- Li H, Feng X, Chu C. The design and construction of reference pangenome graphs with minigraph. Genome Biology. 2020;21:265. https://doi.org/10.1186/s13059-020-02168-z
- Medugorac I, Seichter D, Graf A, Russ I, Blum H, Göpel KH, Rothammer S, Förster M, Krebs S. Bovine polledness: an autosomal dominant trait with allelic heterogeneity. PLoS ONE. 2012;7(6):e39477. https://doi.org/10.1371/journal.pone.0039477
- Milia S, Leonard AS, Mapel XM, Bernal Ulloa SM, Drögemüller C, Pausch H. Taurine pangenome uncovers a segmental duplication upstream of KIT associated with depigmentation in white-headed cattle. Genome Research. 2025;35(4):1041-1052. https://doi.org/10.1101/gr.279064.124
Feedback on this tutorial is welcome: contact us.