LD at a selective sweep (human)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.13. Desktop beta builds are coming soon.
We look at linkage disequilibrium (LD) around the lactase gene, where selection for lactase persistence left one long block of correlated variants. PLINK correlates the phased genotypes and JBrowse draws the triangle from its output. With that view we:
- read the edges of the block against Fst and a genetic map
- compare the swept population's triangle against the pooled release
- cluster a haplotype matrix into the block the triangle draws
- draw each population's allele frequency across the block
Prerequisites
- a JBrowse to paste the tracks into (Web or Desktop)
bcftoolsbuilt with libcurl- htslib (
tabix) curlpython3node, for the JBrowse CLIbedGraphToBigWig, for the Fst track- PLINK 1.9 for the r² tables and PLINK 2.0 for the Fst track and the frequency filter1
Where the data comes from
1000 Genomes 30x high-coverage from NYGC (Byrska-Bishop et al. 2022), called natively on GRCh38.
The build scripts take these files from their URLs, so there is nothing to download by hand.
- phased chromosome 2: ftp.1000genomes.ebi.ac.uk/…/1kGP_high_coverage_Illumina.chr2.filtered.SNV_INDEL_SV_phased_panel.vcf.gzhttps://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/working/20220422_3202_phased_SNV_INDEL_SV/1kGP_high_coverage_Illumina.chr2.filtered.SNV_INDEL_SV_phased_panel.vcf.gz
- the release's own unrelated set, whose SAMPLE_NAME column is
unrelated.samples: ftp.1000genomes.ebi.ac.uk/…/1000G_2504_high_coverage.sequence.indexhttps://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/1000G_2504_high_coverage.sequence.index - populations and superpopulations, narrowed to that unrelated set for
panel.samples(EUR) andrest.samples(everything else): ftp.1000genomes.ebi.ac.uk/…/20130606_g1k_3202_samples_ped_population.txthttps://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/20130606_g1k_3202_samples_ped_population.txt
The gene, ClinVar and recombination tracks come from the hosted UCSC hg38 hub.
Loading the hg38 assembly
The tables, the Fst track and the haplotypes all use GRCh38 coordinates on chr2, so the tracks below go on that assembly.
Goes in the assemblies array of config.json. See Assemblies.
{
"name": "hg38",
"uri": "https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz",
"refNameAliases": {
"uri": "https://s3.amazonaws.com/jbrowse.org/genomes/GRCh38/hg38_aliases.txt"
},
"cytobands": "https://jbrowse.org/genomes/GRCh38/cytoBand.txt"
}jbrowse add-assembly https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz \
--name hg38 \
--refNameAliases https://s3.amazonaws.com/jbrowse.org/genomes/GRCh38/hg38_aliases.txt \
--config '{"cytobands":"https://jbrowse.org/genomes/GRCh38/cytoBand.txt"}'hg38 is one of the genomes JBrowse Desktop hosts, with gene tracks already set up: on the start screen click Show all available genomes and pick it. To load the files of this config instead:
In JBrowse Desktop, Open new genome on the start screen (or File → Open genome... in a session), then Open from a URL and paste, one per line:
https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz
https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz.fai
https://jbrowse.org/genomes/GRCh38/fasta/hg38.prefix.fa.gz.gziJBrowse reads the format off the file name. Then fill in:
- Genome name:
hg38 - refName aliases (under More options):
https://s3.amazonaws.com/jbrowse.org/genomes/GRCh38/hg38_aliases.txt - cytobands (under More options):
https://jbrowse.org/genomes/GRCh38/cytoBand.txt
Drawing the LCT LD triangle from a PLINK table
In the triangle, red means two variants are inherited together and white means independent. The triangle is a matrix turned on its corner, so the vertical axis is the distance between the two variants.
Point an LDTrack at the r² table (squared correlation
per variant pair) PLINK wrote below,
in an hg38 session:
Goes in the tracks array of config.json. See Tracks.
{
"type": "LDTrack",
"trackId": "kgp_lct_ld",
"name": "LCT lactase-persistence LD, 1000G European panel (r²)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "PlinkLDTabixAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_eur.ld.gz"
},
"displayDefaults": {
"variantLayout": "genomic",
"showLegend": true,
"height": 360
}
}jbrowse add-track https://jbrowse.org/demos/popgen/lct_1kg38_chr2_eur.ld.gz \
--trackId kgp_lct_ld \
--name "LCT lactase-persistence LD, 1000G European panel (r²)" \
--assemblyNames hg38 \
--displayDefaults '{"variantLayout":"genomic","showLegend":true,"height":360}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "LDTrack",
"trackId": "kgp_lct_ld",
"name": "LCT lactase-persistence LD, 1000G European panel (r²)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "PlinkLDTabixAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_eur.ld.gz"
},
"displayDefaults": {
"variantLayout": "genomic",
"showLegend": true,
"height": 360
}
}Open this track in JBrowse Web ↗
Open this track in JBrowse Desktop ↗ (JBrowse Desktop 5.0+)
The block is the lactase-persistence sweep
(Bersaglieri et al. 2004).
color.field picks r² or D' (a second LD
measure), and this table has both.
Cutting the LCT region out of the 1000 Genomes VCF
Cut the region twice: over the whole release, and over its European samples (the
panel, where the sweep happened). *.samples files list one ID per line.
# -r fetches the 3.4 Mb region by range request from the 2.5 GB chromosome file
# -e drops symbolic SV records, which are spans; the display correlates
# allele indicators
bcftools view -r chr2:133800000-137200000 -S unrelated.samples \
-e 'ALT[0]~"<"' -Oz -o pooled.vcf.gz \
https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/working/20220422_3202_phased_SNV_INDEL_SV/1kGP_high_coverage_Illumina.chr2.filtered.SNV_INDEL_SV_phased_panel.vcf.gz
tabix -p vcf pooled.vcf.gz
bcftools view -S panel.samples -Oz -o panel.vcf.gz pooled.vcf.gz
tabix -p vcf panel.vcf.gzCorrelating the LCT variants with PLINK
PLINK correlates allele indicators, so we first reduce each slice to biallelic SNVs with one record per position and the IDs dropped:
bcftools view -m2 -M2 -v snps panel.vcf.gz | bcftools norm -d both |
bcftools annotate -x ID -Oz -o panel.snvs.vcf.gzThen pick common variants per cohort and correlate every pair. PLINK's table
holds one row per pair, so the minor allele frequency (MAF) floor keeps it small
enough for a browser to draw. LD across an inversion (mosquitoes) draws a 22 Mb
inversion by thinning the variants to a grid before PLINK correlates them. The
same reduction on pooled.vcf.gz gives pooled.snvs.vcf.gz for the pooled
table.
# 0.35 is a high MAF floor, to keep the variants that tag the block
plink2 --vcf panel.snvs.vcf.gz --double-id --allow-extra-chr --output-chr chrM \
--set-missing-var-ids @:# --maf 0.35 --chr chr2 --write-snplist --out sel
# dprime adds D' beside r2
# --ld-window-r2 0 keeps the uncorrelated pairs, drawn as white cells
# raise both --ld-window and --ld-window-kb, or the defaults clip this block at
# 10 variants or 1 Mb
plink --vcf panel.snvs.vcf.gz --double-id --allow-extra-chr --output-chr chrM \
--set-missing-var-ids @:# --extract sel.snplist \
--r2 dprime --ld-window 999999 --ld-window-kb 4000 --ld-window-r2 0 \
--out lct_1kg38_chr2_eur
# tabix needs tab-separated columns and a commented header; plink pads with
# spaces
awk 'NR==1{$1=$1; print "#" $0; next} {$1=$1; print}' OFS='\t' \
lct_1kg38_chr2_eur.ld | bgzip > lct_1kg38_chr2_eur.ld.gz
tabix -s 1 -b 2 -e 2 -f lct_1kg38_chr2_eur.ld.gzComputing Fst per variant with plink2
Fst, how far apart two groups' allele frequencies sit, comes from plink2 --fst
over panel.samples and rest.samples, as a bigWig for a
quantitative track:
# plink2 takes the two panels as one categorical phenotype, with FID beside IID
{ printf '#FID\tIID\tPOP\n'
awk '{print $1"\t"$1"\tPANEL"}' panel.samples
awk '{print $1"\t"$1"\tREST"}' rest.samples; } > fst_pops.txt
# method=wc is Weir and Cockerham; the plink2 default, Hudson, gives a
# different number
# --output-chr chrM writes the chromosome as chr2, matching the file
plink2 --vcf pooled.vcf.gz --double-id --output-chr chrM --pheno fst_pops.txt \
--fst POP method=wc report-variants vcols=chrom,pos,fst --out fst_site
# 1-based site to bedGraph interval, dropping sites scored nan
awk 'NR>1 && $4!="nan" {printf "%s\t%d\t%d\t%.5f\n",$1,$2-1,$2,$4}' \
fst_site.PANEL.REST.fst.var | sort -k1,1 -k2,2n > fst_site.bedgraph
printf 'chr2\t242193529\n' > hg38.chrom.sizes
bedGraphToBigWig fst_site.bedgraph hg38.chrom.sizes fst.bwThe Fst track lowers the adapter's resolutionMultiplier. Zoomed out, a bigWig
serves summary bins, and a bin's average sinks the few differentiated variants
into the many around them:
Goes in the tracks array of config.json. See Tracks.
{
"type": "QuantitativeTrack",
"trackId": "kgp_lct_fst",
"name": "Fst, European panel vs the other 1000 Genomes samples (Weir & Cockerham)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "BigWigAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_fst_eur_vs_rest.bw",
"resolutionMultiplier": 0.001
}
}jbrowse add-track-json '{
"type": "QuantitativeTrack",
"trackId": "kgp_lct_fst",
"name": "Fst, European panel vs the other 1000 Genomes samples (Weir & Cockerham)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "BigWigAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_fst_eur_vs_rest.bw",
"resolutionMultiplier": 0.001
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "QuantitativeTrack",
"trackId": "kgp_lct_fst",
"name": "Fst, European panel vs the other 1000 Genomes samples (Weir & Cockerham)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "BigWigAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_fst_eur_vs_rest.bw",
"resolutionMultiplier": 0.001
}
}The LCT block against Fst and the deCODE map
Open the region the slice covers with that Fst track, the hub's Recomb Rate -
Recomb. deCODE Avg track, and two copies of the LDTrack above, the second
pointed at https://jbrowse.org/demos/popgen/lct_1kg38_chr2_pooled.ld.gz:
- In the Fst track at the top, the most differentiated sites in the window sit inside the block.
- The block fills the flat span of the deCODE map. The map counts crossovers in sequenced families, so it checks the triangle independently of LD.
- Pooling the swept European samples with populations the sweep never reached lightens the pooled triangle.
Clustering LCT haplotypes into the block the triangle draws
The triangle summarises haplotypes, the variants each chromosome carries along
the block. A track over the six-population VCF draws them below it in
equal-width columns, one row per chromosome. For your own cohort, write a
samples TSV with a name column and a population column and point
samplesTsvLocation at it:
Goes in the tracks array of config.json. See Tracks.
{
"type": "VariantTrack",
"trackId": "kgp_lct_haplotypes",
"name": "1000 Genomes haplotypes across LCT (one row per haplotype)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_6pop.vcf.gz",
"samplesTsvLocation": {
"uri": "https://jbrowse.org/genomes/hg19/1000g.sorted.csv.gz"
}
},
"displays": [
{
"type": "LinearMultiSampleVariantDisplay",
"variantLayout": "columns",
"unit": "haplotype",
"rowColor": "population",
"minorAlleleFrequencyFilter": 0.35,
"forceLoad": true,
"height": 700
}
]
}jbrowse add-track-json '{
"type": "VariantTrack",
"trackId": "kgp_lct_haplotypes",
"name": "1000 Genomes haplotypes across LCT (one row per haplotype)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_6pop.vcf.gz",
"samplesTsvLocation": {
"uri": "https://jbrowse.org/genomes/hg19/1000g.sorted.csv.gz"
}
},
"displays": [
{
"type": "LinearMultiSampleVariantDisplay",
"variantLayout": "columns",
"unit": "haplotype",
"rowColor": "population",
"minorAlleleFrequencyFilter": 0.35,
"forceLoad": true,
"height": 700
}
]
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "VariantTrack",
"trackId": "kgp_lct_haplotypes",
"name": "1000 Genomes haplotypes across LCT (one row per haplotype)",
"assemblyNames": ["hg38"],
"adapter": {
"type": "VcfTabixAdapter",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_6pop.vcf.gz",
"samplesTsvLocation": {
"uri": "https://jbrowse.org/genomes/hg19/1000g.sorted.csv.gz"
}
},
"displays": [
{
"type": "LinearMultiSampleVariantDisplay",
"variantLayout": "columns",
"unit": "haplotype",
"rowColor": "population",
"minorAlleleFrequencyFilter": 0.35,
"forceLoad": true,
"height": 700
}
]
}Open this track in JBrowse Web ↗
Open this track in JBrowse Desktop ↗ (JBrowse Desktop 5.0+)
Cluster the rows by genotype two ways:
- from the track menu, Clustering → Cluster rows by genotype...
- baked into a session with the
runClusteringandclusterRegionmodel properties
- In file order the matrix is a plaid. Clustering puts near-identical chromosomes together, so a swept haplotype forms one slab.
- The ClinVar track marks
rs4988235, which falls below the frequency floor of the matrix. It is the hub's ClinVar track filtered withjexl:feature.phenotypeList=='LACTASE PERSISTENCE'.
Subsampling six populations for the haplotype matrix
Over the whole release each haplotype row falls below a pixel, so the figure
reads 25 samples from each of six populations. The core of the script behind it
is one bcftools call over a list of 150 sample IDs, one per line:
bcftools view -S sub.samples --force-samples -Oz -o lct_1kg38_chr2_6pop.vcf.gz pooled.vcf.gzAllele frequency per population across the LCT block
Twenty-five people from each of the six populations are enough to sort
haplotypes and too few to read a frequency off. So bcftools +fill-tags
computes each population's allele frequency over all of its unrelated samples
and writes it into one INFO field per population. Its -S table holds a sample
ID and a group on each line, tab-separated, so the same command takes any
grouping:
# -S gives fill-tags the groups, and -t AF writes AF_<group> for each
# norm -d both keeps one record per position, so no two bedGraph intervals
# overlap
bcftools view -m2 -M2 -v snps -Ou pooled.vcf.gz |
bcftools norm -d both -Ou |
bcftools +fill-tags -Ou -- -S pops.txt -t AF |
bcftools query \
-f '%CHROM\t%POS0\t%END\t%AF_CEU\t%AF_FIN\t%AF_PJL\t%AF_TSI\t%AF_YRI\t%AF_CHB\n' \
> af.tsv
# one bedGraph per population, columns 4 to 9 of af.tsv, then a bigWig each
printf 'chr2\t242193529\n' > hg38.chrom.sizes
column=4
for pop in CEU FIN PJL TSI YRI CHB; do
cut -f 1-3,$column af.tsv > "af_$pop.bedgraph"
bedGraphToBigWig "af_$pop.bedgraph" hg38.chrom.sizes "lct_1kg38_chr2_af_$pop.bw"
column=$((column + 1))
doneA MultiQuantitativeTrack draws the
six bigWigs as a row each, on one axis from 0 to 1:
Goes in the tracks array of config.json. See Tracks.
{
"type": "MultiQuantitativeTrack",
"trackId": "kgp_lct_population_af",
"name": "Allele frequency per population, 1000 Genomes unrelated samples",
"assemblyNames": ["hg38"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"type": "BigWigAdapter",
"source": "CEU",
"color": "#4e79a7",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_CEU.bw"
},
{
"type": "BigWigAdapter",
"source": "FIN",
"color": "#e15759",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_FIN.bw"
},
{
"type": "BigWigAdapter",
"source": "PJL",
"color": "#76b7b2",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_PJL.bw"
},
{
"type": "BigWigAdapter",
"source": "TSI",
"color": "#59a14f",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_TSI.bw"
},
{
"type": "BigWigAdapter",
"source": "YRI",
"color": "#edc948",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_YRI.bw"
},
{
"type": "BigWigAdapter",
"source": "CHB",
"color": "#f28e2b",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_CHB.bw"
}
]
},
"displayDefaults": {
"height": 300,
"scales": { "y": { "domainMin": 0, "domainMax": 1 } }
}
}jbrowse add-track-json '{
"type": "MultiQuantitativeTrack",
"trackId": "kgp_lct_population_af",
"name": "Allele frequency per population, 1000 Genomes unrelated samples",
"assemblyNames": ["hg38"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"type": "BigWigAdapter",
"source": "CEU",
"color": "#4e79a7",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_CEU.bw"
},
{
"type": "BigWigAdapter",
"source": "FIN",
"color": "#e15759",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_FIN.bw"
},
{
"type": "BigWigAdapter",
"source": "PJL",
"color": "#76b7b2",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_PJL.bw"
},
{
"type": "BigWigAdapter",
"source": "TSI",
"color": "#59a14f",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_TSI.bw"
},
{
"type": "BigWigAdapter",
"source": "YRI",
"color": "#edc948",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_YRI.bw"
},
{
"type": "BigWigAdapter",
"source": "CHB",
"color": "#f28e2b",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_CHB.bw"
}
]
},
"displayDefaults": {
"height": 300,
"scales": { "y": { "domainMin": 0, "domainMax": 1 } }
}
}'In JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "MultiQuantitativeTrack",
"trackId": "kgp_lct_population_af",
"name": "Allele frequency per population, 1000 Genomes unrelated samples",
"assemblyNames": ["hg38"],
"adapter": {
"type": "MultiWiggleAdapter",
"subadapters": [
{
"type": "BigWigAdapter",
"source": "CEU",
"color": "#4e79a7",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_CEU.bw"
},
{
"type": "BigWigAdapter",
"source": "FIN",
"color": "#e15759",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_FIN.bw"
},
{
"type": "BigWigAdapter",
"source": "PJL",
"color": "#76b7b2",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_PJL.bw"
},
{
"type": "BigWigAdapter",
"source": "TSI",
"color": "#59a14f",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_TSI.bw"
},
{
"type": "BigWigAdapter",
"source": "YRI",
"color": "#edc948",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_YRI.bw"
},
{
"type": "BigWigAdapter",
"source": "CHB",
"color": "#f28e2b",
"uri": "https://jbrowse.org/demos/popgen/lct_1kg38_chr2_af_CHB.bw"
}
]
},
"displayDefaults": {
"height": 300,
"scales": { "y": { "domainMin": 0, "domainMax": 1 } }
}
}Open this track in JBrowse Web ↗
Open this track in JBrowse Desktop ↗ (JBrowse Desktop 5.0+)
Open it over chr2:135,844,000-135,858,000, the stretch of MCM6 that holds
rs4988235:
- At
rs4988235the allele is common in the two northern European populations (CEU, FIN) and rarer in PJL (Punjabi) and TSI (Tuscan), the spread Bersaglieri et al. (2004) describe. - The YRI (Yoruba) and CHB (Han Chinese) rows carry common variants elsewhere in the window, at sites where the other four are low.
Reproduce it end to end
build_lct_ld.sh
builds the triangles and the narrow Fst track, and writes a ready-to-serve
config. It:
- keeps the release's unrelated samples, since relatives share long haplotypes whether or not a sweep happened, and splits them into the European panel and everyone else
- cuts a region wide enough to reach past both edges of the block, and prints
mean r² against
rs4988235along it, so the block's ends show in the data - cuts the European panel out of that same slice, so the pooled and panel triangles hold the same variants and differ only in their samples
- correlates each cohort's common variants with PLINK, and scores Fst per variant between the panel and the rest
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_lct_ld.sh
bash build_lct_ld.sh # builds ./lct_ld_build/jbrowse2
npx --yes serve lct_ld_build/jbrowse2 # then open the printed URLThree more scripts build the other files:
build_lct_fst_scan.shwrites the wide Fst track. It scores the same panels with the same estimator over 40 Mb of chr2 with LCT at its middle, one value per variant, and prints wherers4988235ranks across the span.build_lct_haploblock.shbuilds the subsampled haplotype matrix. It takes the same number of unrelated samples from each of CEU, FIN, PJL, TSI, YRI and CHB, which span the lactase-persistence allele from common to absent, so no population outweighs the rest in the clustering. It prints the allele's frequency in each, over the release and over the subsample.build_lct_population_af.shbuilds the per-population frequencies. It reads the pooled slice, takes every unrelated sample of the six populations, and prints the frequencies atrs4988235so the bars can be checked against them.
curl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_lct_fst_scan.sh
bash build_lct_fst_scan.sh # builds ./lct_fst_scan_buildcurl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_lct_haploblock.sh
bash build_lct_haploblock.sh # builds ./lct_haploblock_buildcurl -fO https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_lct_population_af.sh
bash build_lct_population_af.sh # builds ./lct_population_af_buildSee also
- LD across an inversion (mosquitoes)
- Selection scans (Drosophila DGRP)
- Phased trio analysis (1000 Genomes)
- User guide: Variant track
- GWAS / Manhattan track
- Variant track configuration
Citations
- 1000 Genomes Project Consortium (2015). A global reference for human genetic variation
- Bersaglieri et al. (2004). Genetic signatures of strong recent positive selection at the lactase gene
- Byrska-Bishop et al. (2022). High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios
- Halldorsson et al. (2019). Characterizing mutagenic effects of recombination through a sequence-level genetic map
Notes
-
PLINK 1.9 and PLINK 2.0 are separate programs, and the page runs both. plink2 gained
--r2-phasedin its a6 alphas, so on earlier builds the r² step needs PLINK 1.9. ThePlinkLDTabixAdapterreads the column names of either program from the header, so output from either loads. ↩
Feedback on this tutorial is welcome: contact us.