SV visualization
TL;DR: Triage structural variant (SV) candidates in the SV inspector, a combined table and whole-genome circular overview over a variant track's calls, then drill into the alignments (BAM/CRAM) at each breakpoint for the read-level evidence. This page is about what those read patterns mean, and the alignments track guide covers what each control does.
End-to-end walkthroughs: Structural variants (Cancer GIAB),
Complex rearrangements and derivative alleles, Structural variants (1000 Genomes). For how
split (supplementary) alignments, the SA tag, pair orientation, TLEN and
clipping encode SV evidence in the SAM format itself, see
Structural variants and the SAM format.
SV signals in the alignments track
Most of the evidence is in a standard alignments track already. Soft clipping marks a breakpoint edge, where a cluster of reads terminates and their overhanging bases are clipped, and the insertion and clipping indicators flag those clusters above the coverage row before you zoom into the pileup. Coloring by pair orientation or by insert size then lifts the abnormal pairs out of the concordant ones. What each SV type makes of those signals is below.
Pair orientation color scheme
Set the color scheme to Pair orientation from the track menu. JBrowse uses the
same colors as IGV (see the
IGV paired-end alignments guide),
and assumes standard fr (Illumina) read pairs. SOLiD-style orientations are
not supported.
| Color | Name | Value | Description |
|---|---|---|---|
| LR (→ ←, normal proper pair) | #d3d3d3 | Concordant | |
| RL (← →, mates point away from each other) | #0099bb | Abnormal orientation | |
| LL (→ →, both mates forward strand) | #4d9a4d | Abnormal orientation | |
| RR (← ←, both mates reverse strand) | #5555bb | Abnormal orientation | |
| Inter-chromosomal | #af4d19 | Mate maps to a different chromosome; colored distinctly rather than by orientation | |
| Mate unmapped | #000000 | The other end of the pair aligned nowhere, so orientation and insert size say nothing | |
| Split paired-end read (inverted) | #9b30b0 | A paired read's supplementary segment maps opposite-strand to its primary, so the junction is inverted — an inversion or an inverted duplication |
An inversion turns a stretch of sequence around, and every read inside it turns around too. A pair with one end inside and one end outside therefore has one end reversed and one not, so both point the same way: forward-forward across the left junction, reverse-reverse across the right one. Each of those pairs also has an end carried to the far side of the segment, so a green LL bundle and a navy RR bundle span the same stretch.
In an inverted duplication call, green LL, navy RR and magenta split reads are
all the inversion, the last being one read split into alignments that point in
opposite directions. They are a minority of an otherwise concordant grey pileup,
so they cluster at the breakpoints. The duplication carries no orientation
signature at all: where the second copy went is what the call's
INFO.CPX_INTERVALS names, and no read in the pileup states it.
Inversions in long reads
Short paired-end reads infer an inversion, because neither mate spans a breakpoint. A long read spans the whole event and splits into forward and reverse-strand alignments. With View as pairs / link supplementary alignments on, those segments chain onto one row: the inverted middle paints the reverse-strand color between the forward-strand segments either side, and a magenta arc joins the two breakpoints.
Group by... → Split read (SA tag) puts the reads carrying a supplementary
alignment in their own section, where each breaks into three pieces with the
middle one reversed. The section below spans the locus in one piece. The two
sections together are the genotype: a locus where some reads invert and the rest
run through unbroken is one inverted copy and one uninverted, read off the
pileup rather than from the caller's GT.
Insert size color scheme
Set the color scheme to Insert size from the track menu. Reads are colored red (insert larger than expected), pink (smaller than expected), or light grey (normal), and a separate Insert size (gradient) option shades continuously by the magnitude of the deviation. With read arcs or the read cloud enabled, the Insert size option uses threshold-based coloring:
| Color | Name | Value | Description |
|---|---|---|---|
| Mate on a different chromosome | #af4d19 | Suggests an inter-chromosomal event | |
| Insert larger than expected | #ff0000 | Suggests a deletion spanning the pair | |
| Insert smaller than expected | #f582c0 | Suggests an insertion between the pair | |
| Mate unmapped | #000000 | The other end of the pair aligned nowhere, so insert size says nothing |
The expected range is a robust band around the typical insert size.1
Insert size and orientation combines both signals and is often the most informative setting for a general SV scan. A short insert always paints pink, even where the orientation is abnormal, then an abnormal orientation wins (teal RL for a tandem duplication, green LL or dark blue RR for an inversion), and a large insert with normal orientation paints red for the classic deletion.
SV channels
Every scheme above paints one pileup, so an event's evidence arrives mixed into the rows around it. Track menu → Read connections → SV channels (pairs by orientation) takes the same reads apart: each orientation class becomes its own band with its own coverage curve and its own arcs, the concordant pairs drop out of the arcs, and the pileup goes away. The bands are the rows of the orientation table above, and the signatures below are the key to reading them.
Clicking the row again restores an ordinary pileup, with the color scheme untouched in both directions. Structural variants (1000 Genomes) reads a complex 1000 Genomes call band by band.
SV-type signatures
Each SV type leaves a characteristic combination of the signals above. One row alone has artifacts that produce it, so combine several before calling an SV.
| SV type | Read pairs | Coverage | Clipping and arcs |
|---|---|---|---|
| Deletion | red, insert larger than expected | drops between the breakpoints, halves for a heterozygote | clipped reads at both edges, unusually long arcs |
| Insertion | pink, insert smaller than expected | unchanged | clipped reads at one site, a purple insertion indicator, mates unmapped once the insertion outruns the fragment |
| Inversion | green LL and dark blue RR at the junctions | unchanged | clipped reads at both breakpoints, magenta split-read arcs |
| Tandem duplication | teal RL | elevated over the duplicated segment | arcs pointing back upstream across the junction |
| Translocation | rust, mate on another chromosome | unchanged | a cluster of rust reads at one end, arcs drawn as verticals at the view edge |
Zoomed inside an inverted segment the interior reads look concordant, so the junctions are where to look. Turning on Show soft clipping makes an insertion's bases readable on each side of the site, and the clipped bases at an inversion breakpoint often carry the short homology the junction formed on. For a translocation, open the breakpoint split view from the feature details to see both ends at once.
Read arcs
Read arcs draw bezier curves between the ends of paired or split reads. Unusually long arcs relative to their neighbors point to a deletion spanning the pair, and inter-chromosomal connections (drawn as vertical lines at the view edge) flag translocations. With View as pairs on, each mate pair collapses onto a single row joined by its own curve, so the abnormal same-orientation pairs of an inverted duplication read as a coherent bundle.
A read can look concordant (light-grey LR fill) yet still carry a colored connector: the read itself crosses the breakpoint, splitting into a primary and a strand-flipped supplementary alignment, and the arc joining them takes the split-read inversion color rather than the RR-pair blue. That is evidence from one read rather than a pair. Hover any connector for its classification.
Arcs also count. Near-identical curves stack into one line, so a curve per molecule cannot say how many molecules agree, while the arc band draws each junction once and thickens it by the reads behind it. Gene fusion calls and the DNA behind them counts a fusion's support that way.
Read cloud
Read cloud stratifies reads by the log-scaled distance between mates, so how many reads span a breakpoint and which way they point are both countable at a glance. Chains with supplementary alignments are connected by an orange line, and Edit filters in the track menu shows or hides proper pairs and singletons.
Inspecting individual reads
Right-click any read for Linear read vs ref and Dotplot of read vs ref, which lay one read out against every locus it touches. Both are most useful on a long read spanning a breakpoint, where the order the read visits those loci in is the structure of the rearrangement. See one read against the reference.
Reconstructing a derivative allele
A split read is an ordered, oriented list of reference intervals, which is what a derivative allele is. The alignments track menu's Launch → Reconstruct derivative allele... groups the reads in the window by the route their split alignments describe and lists each route with the number of reads that independently take it. Draw as picks which view it opens in: a linear synteny view puts the allele along the bottom and the loci it visits along the top, a breakpoint split view stacks those loci in the order the reads cross them. Either goes into the launching view's place or into a new view below it.
Complex rearrangements and derivative alleles works through both shapes it produces, a two-segment fold-back and a four-segment allele across three chromosomes. Structural variants (Cancer GIAB) runs it over a tumor/normal HiFi pair and checks the route it ranks first against a published benchmark.
What the reconstruction needs
- Long reads, from an aligner that emits SA tags - minimap2 and ngmlr both do. A short-read library ranks each junction on its own and produces no multi-junction route, because no read reaches from one junction to the next. The picker measures the reads on screen, so it prints the median aligned length it found rather than sending you to a wider window.
- An event above about 10 kb, or interchromosomal. Below that an aligner usually writes the event inside one read's CIGAR rather than as a split alignment, which nothing reading SA tags can reach, and the pileup's own deletion marks are where those show. Where the cliff falls is as much the aligner's doing as the browser's.
- Two reads that agree. A route reaches the list once at least two of them cross the same junctions in the same order and orientation. The floor is what keeps a repetitive window from filling the list, not a judgement that a single split read is mismapped.
- The reads actually loaded. A window over the track's byte budget renders
as
force loadwith nothing behind it, and the dialog says so rather than reporting that no route is supported. Narrow the window. - Every locus you want ranked on screen. A read anchored in the window brings its whole SA chain, so the far side of a junction needs no panel of its own, but reads sitting only at that far locus contribute nothing until it is shown.
The same entry on a synteny track reads contigs rather than reads, since a de novo assembly aligned to the reference is the same object at a larger scale. There one contig is enough to list a route, and every locus has to be on screen, because an alignment block names nothing the view has not fetched.
Judging what it lists
A read count ranks the routes, it does not vouch for them. Reads mismapped into a repeat produce a confident-looking route, so the output is a proposal to check against the reads rather than a call.
- The matched normal is the strongest check. A route on its own says a locus has split reads, not that it has an event, and normals do propose routes at windows with nothing somatic in them.
- Look at the top two rows. Where a published junction is recovered it is the first or second row in both validation callsets. A route further down the list is unlikely to be the one you came for.
- Read the segment sizes. A route whose segments are each about one read long is an aligner splitting a short read across the genome rather than an allele, and the segment strip drawn to scale beside each row is what separates the two.
- A row marked "part of a longer route in this list" crosses a run of another row's junctions and stops, so it is consistent with the longer route rather than a rival to it.
- Dozens of routes means a repetitive window, not a complicated allele. The picker says how many it left off the list.
Recall by event size
Recall follows event size, because the reconstruction reads split alignments and an aligner represents a short deletion inside one read's CIGAR instead. Below, COLO829 against its ONT calls (Valle-Inclán et al. 2022) and HG008-T against the C-GIAB draft benchmark (Wagner et al. 2026), scored at every junction rather than at chosen loci:
| Event size | COLO829, ONT | HG008-T, PacBio HiFi |
|---|---|---|
| < 1 kb | 1 / 9 | 4 / 40 |
| 1 - 10 kb | 6 / 10 | 17 / 26 |
| 10 - 100 kb | 16 / 16 | 17 / 18 |
| > 100 kb | 17 / 17 | 54 / 54 |
| interchromosomal | 11 / 11 | 14 / 14 |
Where a junction is recovered, its two ends land within a base or two of the called breakend, and neither matched normal recovers a somatic junction at the same windows.
Reading a split alignment as an ordered list of reference intervals is what long-read SV callers extract before they cluster anything (Sedlazeck et al. 2018), and ordering those intervals into a derivative chromosome is what long-read rearrangement pipelines do with them (Cretu Stancu et al. 2017, Mitsuhashi et al. 2020). This is the visualization half of that lineage, in the manner of Ribbon (Nattestad et al. 2021): it groups those chains and counts them, and where two routes disagree it lists both with their counts rather than choosing.
Breakpoint split view
The breakpoint split view opens synchronized panels side-by-side, each centered on one breakpoint locus. Splines connect supporting reads across the panels, and the variant call is drawn as a colored line with feet indicating directionality. The header bar accepts location searches in either panel.
Hovering a spline shades the reads it joins, so a junction names the alignments that carry it. The shading follows the whole chain, every segment of the read in every panel it visits, and every other spline of the same read thickens alongside it. Untick Show... → Allow clicking alignment squiggles to turn the overlay back into a static picture.
Launching the breakpoint split view
- From the SV inspector - click a feature in the circular overview or the triangle dropdown on any table row. See the SV inspector guide.
- From variant feature details - click a BND or TRA variant in a variant track; the feature details panel has a button to open the split view, automatically loading any open alignment tracks.
- From alignment feature details - click any read with a supplementary alignment, for a split view centered on that read and its supplementary partner.
- From the circular genome view - click a chord's feature details and use the "Open breakpoints in split view" link in its Breakends section.
Multi-hop events
A read with several supplementary alignments visits more than two loci, and the view grows a panel per locus rather than stopping at a pair. Complex rearrangements and derivative alleles follows one such chain across three chromosomes.
Following a chain of breakends
A BND record names one partner, so a launch from one record is two loci however many the rearrangement has. Follow further breakends at each end reaches the rest from the callset: at each end of the chain it looks for another junction leaving from within a kilobase of the same place, and takes it when there is exactly one, adding a panel per hop up to four. It is offered by the two launches that can read the callset, a variant track's right-click menu and a click on a chord drawn from that same file, and only for the stacked shape.
It reads two junctions leaving one locus as one molecule, which the caller does not assert, so it stops where that would be a guess. Two open continuations at a locus stop the walk, since the records cannot say which molecule carries which, and so does a continuation leading back to a locus already on screen. To work from the reads themselves instead, use Reconstruct derivative allele, which ranks whole routes by how many molecules independently take each.
Phasing heterozygous SVs
For a heterozygous SV, confirming that the supporting reads come from a single
haplotype is strong evidence for the call. Where the BAM/CRAM has been
haplotagged (WhatsHap, HiPhase), reads carry an HP tag, and sorting, coloring
or grouping by it from the
track menu clusters each haplotype. Grouping goes furthest, giving each
haplotype its own pileup section, with untagged reads collected in their own so
unphased support stays visible. The
phased trio tutorial covers phased haplotypes
end-to-end.
Working with large SVs
A window large enough to need more data than a single request allows will fail to load reads. For large or inter-chromosomal SVs:
- Survey the region with a BigWig coverage track, or a multi-quantitative track for tumor vs normal. It loads at any scale and makes copy-number changes immediately visible
- Load the call set as a variant track for a compact overview, where clicking a feature navigates to it
- Open the breakpoint split view for the breakpoint loci themselves. Each panel is a local window around one end, so the distance between them does not matter
- Use the SV inspector for whole-genome triage before drilling in
Whole-genome assembly comparison
Where a de novo assembly of the sample is available, aligning it back to the reference with minimap2 and loading the PAF as a synteny track gives a chromosome-scale view of the rearrangements. Complex events appear as off-diagonal blocks in the dotplot view, and dragging over one launches a base-level linear synteny view on the same alignment. The C-GIAB tutorial walks this through with the HG008 phased tumor assembly.
Summary
| Display / setting | How to enable | Best for |
|---|---|---|
| Pileup (default) | Default lower panel | Base-level detail, individual reads |
| Color by pair orientation | Color scheme in track menu | Abnormal orientation patterns (RL/LL/RR) |
| Color by insert size | Color scheme in track menu | Insert size anomalies (pileup) |
| Read arcs | Read connections in track menu | Overview of long-range connections |
| Read cloud | Read connections in track menu | Counting discordant pairs, orientation per read |
| Linear read vs ref | Right-click on any read | Complex alignment of a single long read |
| Reconstruct derivative allele | Launch in the track menu | The route several long reads agree on |
| Breakpoint split view | Feature details or SV inspector | Side-by-side inspection of both breakpoint loci |
| Sort/color by HP tag | Sort/color by tag in track menu | Confirming heterozygous SVs on one haplotype |
| Dotplot view | Launch from the Add menu | Chromosome-scale rearrangements (de novo assembly) |
| Linear synteny view | Add menu or dotplot selection | Base-level alignment between two genomes |
Limitations
- Read-level displays need the reads: the pileup, read arcs and read cloud only render once the view is zoomed in far enough to load them, and a very large SV cannot be spanned in one pileup. Use the routes in working with large SVs
- Repetitive regions: in segmental duplications and repeats, soft clips and orientation anomalies are common artifacts, so a signature there is weak evidence on its own
See also
- User guide: Alignments track
- SV inspector
- Circular genome view
- User guide: Variant track
- Alignments track configuration
- Gallery: structural variant examples
References
- Cretu Stancu et al. (2017). Mapping and phasing of structural variation in patient genomes using nanopore sequencing
- Mitsuhashi et al. (2020). A pipeline for complete characterization of complex germline rearrangements from long DNA reads
- Nattestad et al. (2021). Ribbon: intuitive visualization for complex genomic variation
- Sedlazeck et al. (2018). Accurate detection of complex structural variations using single-molecule sequencing
- Valle-Inclán et al. (2022). A multi-platform reference for somatic structural variation detection
- Wagner et al. (2026). A complete human pancreatic cancer genome
Notes
-
The band is
median ± 3·1.4826·MADrather thanmean ± 3σ, because the long right tail of large inserts inflates the standard deviation and pushes amean − 3σlower bound below zero, where no short insert is ever flagged. ↩