Driving JBrowse with an AI agent (two Drosophila genomes)
The tutorials target the JBrowse v5 beta, and the v4.3.0 release
on the download page lacks some of what they show. To install the
beta of JBrowse Web, run npm install -g @jbrowse/cli@next, then
jbrowse create jbrowse2 --branch v5.0.0-beta.11. Desktop beta builds are coming soon.
An AI agent, given four requests in plain words, aligns two fruit fly species that have no published alignment and inspects the result in JBrowse Desktop. Typed one at a time, the requests ask the agent to:
- align the two genomes with minimap2 and open them side by side, genes and all
- add a dotplot of the pair, restricted to the six chromosome arms
- count alignment blocks by strand, to find where the two genomes run in opposite directions
- navigate the view to what it found
Prerequisites
- JBrowse Desktop, installed and running (see the desktop quickstart)
- an MCP client with a shell of its own, for the alignment step: Claude Code, or Claude Desktop, set up as in Using JBrowse with AI agents
- minimap2
node, for the JBrowse CLI
Where the data comes from
Two GenArk assemblies and their genome hubs on genomes.jbrowse.org (D. simulans, D. mauritiana). Each hub's config holds the 2bit sequence, a chromAlias file, an NCBI RefSeq gene track and a Trix text index.
- D. simulans GCF_016746395.2 sequence: hgdownload.soe.ucsc.edu/…/GCF_016746395.2.fa.gzhttps://hgdownload.soe.ucsc.edu/hubs/GCF/016/746/395/GCF_016746395.2/GCF_016746395.2.fa.gz
- D. simulans hosted config: jbrowse.org/…/config.jsonhttps://jbrowse.org/hubs/genark/GCF/016/746/395/GCF_016746395.2/config.json
- D. mauritiana GCF_004382145.1 sequence: hgdownload.soe.ucsc.edu/…/GCF_004382145.1.fa.gzhttps://hgdownload.soe.ucsc.edu/hubs/GCF/004/382/145/GCF_004382145.1/GCF_004382145.1.fa.gz
- D. mauritiana hosted config: jbrowse.org/…/config.jsonhttps://jbrowse.org/hubs/genark/GCF/004/382/145/GCF_004382145.1/config.json
- the alignment the figures below open, indexed and rehosted: jbrowse.org/…/sim_vs_mau.pif.gzhttps://jbrowse.org/demos/fly_agent_synteny/sim_vs_mau.pif.gz beside the merged config at jbrowse.org/…/config.jsonhttps://jbrowse.org/demos/fly_agent_synteny/config.json
What the agent is driving
Connected to JBrowse Desktop, the agent gets four tools: run_javascript,
open, screenshot and docs. The requests below go through run_javascript,
which runs code against the session through a jb helper library.
Set up the client as in Using JBrowse with AI agents, then run the four requests below.
Ask the agent to align and open the two genomes
The first request, in one sentence:
Align D. simulans GCF_016746395.2 against D. mauritiana GCF_004382145.1 with
minimap2. Run it in the background and poll it, and when it is done index it
and open both genomes side by side, with their gene tracks and the alignment
between them.The aligner it runs:
## The largest chromosome measures 2.4% divergence (de:f), past asm10's 1%.
## --cs writes the difference string the index below carries.
minimap2 -t 8 -cx asm20 --cs mau.fa.gz sim.fa.gz > sim_vs_mau.pafThe whole-genome alignment takes several minutes, longer than the two-minute
timeout on run_javascript, so the request above runs it in the background.
Indexing the PAF lets the browser read one region of it without parsing the whole file:
jbrowse make-pif sim_vs_mau.pafThe config the agent writes merges the two hosted ones, keeping each gene track and adding the alignment as a synteny track.
Check the order of assemblyNames on the synteny track the agent wrote. Your
own pair takes the same track with its uri swapped for your .pif.gz, which
needs its .tbi beside it, and both assembly names loaded:
Goes in the tracks array of config.json. See Tracks.
{
"type": "SyntenyTrack",
"trackId": "sim_vs_mau",
"name": "D. simulans vs D. mauritiana",
"assemblyNames": ["GCF_016746395.2", "GCF_004382145.1"],
"adapter": {
"type": "PairwiseIndexedPAFAdapter",
"uri": "sim_vs_mau.pif.gz",
"assemblyNames": ["GCF_016746395.2", "GCF_004382145.1"]
}
}jbrowse add-track sim_vs_mau.pif.gz \
--trackId sim_vs_mau \
--name "D. simulans vs D. mauritiana" \
--assemblyNames GCF_016746395.2,GCF_004382145.1 \
--load copyIn JBrowse Desktop, or in any running JBrowse Web session, open a view on this track’s assembly, then File → Open track..., choose Add track from pasted JSON, and paste:
{
"type": "SyntenyTrack",
"trackId": "sim_vs_mau",
"name": "D. simulans vs D. mauritiana",
"assemblyNames": ["GCF_016746395.2", "GCF_004382145.1"],
"adapter": {
"type": "PairwiseIndexedPAFAdapter",
"uri": "sim_vs_mau.pif.gz",
"assemblyNames": ["GCF_016746395.2", "GCF_004382145.1"]
}
}sim_vs_mau.pif.gz is relative to a config.json. Replace it with its URL or its path on this computer.
That shows it in the linear view. For the synteny view, open Add → Linear synteny view, pick the track under Quick start, and click Launch.
The query (sim.fa.gz, the second minimap2 argument) comes first and the
target second, the order of the PAF columns. In the other order no chromosome
name resolves and the synteny band draws empty.
Ask for the dotplot
Add a dotplot of the same two assemblies underneath.D. simulans and D. mauritiana each have a few hundred unplaced scaffolds, which interleave the axes if drawn. Naming the arms gives one diagonal:
Restrict both dotplot axes to chr2L, chr2R, chr3L, chr3R, chr4 and chrX.Ask the agent to quantify what restricting the axes drops. Add one more request, to color the dotplot by strand, so a reversed block draws in blue and a forward one in red.
Ask where they disagree
Where do the two genomes run in opposite directions? Answer from the
alignment file, not from the dotplot, and show me the numbers.A reverse-strand block a few hundred kilobases wide is a few pixels at whole-genome zoom. The same information is in the PAF as numbers: aligned bases per arm, split by strand, at MAPQ 30 or better:
awk -F'\t' '
BEGIN {
## the arm each RefSeq accession is, from the chromAlias files
split("NC_052520.2 2L NC_052521.2 2R NC_052522.2 3L NC_052523.2 3R NC_052524.2 4 NC_052525.2 X", a, " ")
split("NC_046667.1 2L NC_046668.1 2R NC_046669.1 3L NC_046670.1 3R NC_046671.1 4 NC_046672.1 X", b, " ")
for (i = 1; i in a; i += 2) q[a[i]] = a[i+1]
for (i = 1; i in b; i += 2) t[b[i]] = b[i+1]
}
## $12 is MAPQ, $5 the strand, $4-$3 the aligned length on the query
$12 >= 30 && ($1 in q) && ($6 in t) && q[$1] == t[$6] {
aligned[q[$1]] += $4 - $3
if ($5 == "-") reverse[q[$1]] += $4 - $3
}
END {
for (k in aligned)
printf "%-3s %6.2f Mb aligned, %5.2f%% reverse\n", k, aligned[k]/1e6, 100*reverse[k]/aligned[k]
}' sim_vs_mau.paf | sort2L 22.15 Mb aligned, 0.28% reverse
2R 20.56 Mb aligned, 5.02% reverse
3L 22.56 Mb aligned, 0.03% reverse
3R 27.03 Mb aligned, 0.01% reverse
4 1.10 Mb aligned, 0.00% reverse
X 21.04 Mb aligned, 4.44% reverseFour arms have essentially no reverse-strand alignment, and they are the control: 2R and X sit an order of magnitude above them.
Grouping the reverse-strand blocks of 5 kb or more, and cutting a group wherever half a megabase passes with none, gives three regions:
2R sim 59,995 - 2,256,808 <-> mau 628,956 - 3,646,198 (2.20 Mb, 75 blocks)
X sim 8,303,553 - 8,752,357 <-> mau 8,530,265 - 8,980,862 (0.45 Mb, 2 blocks)
X sim 21,441,285 - 22,026,996 <-> mau 21,459,277 - 22,872,816 (0.59 Mb, 12 blocks)The 2R region is the largest and the least tidy; the two X regions are smaller and cleaner.
Ask the agent to open the 2R and X regions
Close the gene tracks and take the synteny view to the 2R region.The agent closes both gene tracks, so the view draws the whole-genome alignment
alone, and opens the simulans row at chr2R:1-2,400,000 and the mauritiana row
at chr2R:500,000-3,800,000.
Then ask for the first of the two X regions, chrX:8,100,000-8,950,000 over
chrX:8,330,000-9,180,000, where the reverse blocks are fewer and longer:
What the requests had to tell the agent
Three sentences, and each prevents a failure with no error message:
- Start long work in the background, and let it finish before opening anything. Otherwise a tool call times out and the agent reports a failure that did not happen.
- Restrict the dotplot axes to the arms. Otherwise a few hundred unplaced scaffolds interleave both axes.
- Answer counted from the alignment file. Otherwise the agent describes the dotplot.
Two more come up unprompted: screenshot what it builds, since a wrong track id or empty region still renders as a plausible browser, and say the numbers before it navigates, so what you see is a claim you can check.
The same pipeline as a script
curl -O https://raw.githubusercontent.com/GMOD/jbrowse-components/main/scripts/build_fly_agent_synteny.sh
bash build_fly_agent_synteny.sh fly_agent_synteny_buildThe script downloads both genomes, runs the alignment, indexes it, and writes
config.json to open in Desktop. It needs the tools in
Prerequisites, plus jq and curl.
See also
- Using JBrowse with AI agents
- Recipes for driving JBrowse from an agent
- Driving the live JBrowse session
- Synteny visualization (pairwise minimap2)
Citations
- Chakraborty M, et al. Evolution of genome structure in the Drosophila simulans species complex. Genome Research 31:380-396 (2021). https://doi.org/10.1101/gr.263442.120
- Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34:3094-3100 (2018). https://doi.org/10.1093/bioinformatics/bty191
Feedback on this tutorial is welcome: contact us.