Custom adapters
Extend BaseFeatureDataAdapter, implement getRefNames() and getFeatures()
(an rxjs stream of SimpleFeatures), and register the type in your plugin.
An adapter is a class that fetches and parses your data and returns it in a format JBrowse understands. To display data from a new source with JBrowse's existing gene displays, write a custom adapter. For custom rendering, you'll also need a custom display, which owns the drawing, state, and menus.
Adapter types
| Extend | You supply | It returns |
|---|---|---|
BaseFeatureDataAdapter | getRefNames(), getFeatures() | features overlapping a region — genes, reads, variants. The common case |
BaseRefNameAliasAdapter | getRefNameAliases() | refName aliases, e.g. chr1 for 1 |
BaseSequenceAdapter | getRefNames(), getFeatures(), getRegions() | a region list plus the sequence for a queried region; extends the feature adapter |
BaseTextSearchAdapter | searchIndex() | search-box hits out of a text index |
CytobandAdapter | getData() | cytoband features for the ideogram |
RegionsAdapter | getRegions() | which regions an assembly has, and how long each is |
Two have guides of their own: RefName aliasing and Text search adapters. Supported file types maps every format JBrowse already reads to the adapter that reads it, which is the place to check before writing one.
What a feature adapter implements
Extend BaseFeatureDataAdapter and supply two methods:
getRefNames— the refNames in the file, used for refName renaming.getFeatures— an rxjs observable stream of the features overlapping a region, positions 0-based half-open.
The base class already holds the config, getSubAdapter and the plugin manager
and exposes this.getConf('slotName'), so no constructor is needed unless the
adapter sets up state of its own. Type it on your config schema
(BaseFeatureDataAdapter<MyAdapterConfig>, where MyAdapterConfig comes from
your config schema) so those
getConf reads are typed.
Example feature adapter
Gff3Adapter in full. It parses the whole file up front, so getFeatures is a
lookup:
import {
BaseFeatureDataAdapter,
cachedSetup,
} from '@jbrowse/core/data_adapters/BaseAdapter'
import { fetchAndMaybeUnzip } from '@jbrowse/core/util'
import { openLocation } from '@jbrowse/core/util/io'
import {
groupLinesByRef,
makeFeatureIntervalTreeMap,
} from '@jbrowse/core/util/parseLineByLine'
import { ObservableCreate } from '@jbrowse/core/util/rxjs'
import { parseLinesLazy } from 'gff-nostream'
import { Gff3Feature } from '../Gff3Feature.ts'
import type { Gff3AdapterConfig } from './configSchema.ts'
import type { BaseOptions } from '@jbrowse/core/data_adapters/BaseAdapter'
import type { Feature } from '@jbrowse/core/util/simpleFeature'
import type { NoAssemblyRegion } from '@jbrowse/core/util/types'
export default class Gff3Adapter extends BaseFeatureDataAdapter<Gff3AdapterConfig> {
// the whole file is resident after one load, so the fetch/parse status comes
// from inside the load itself rather than a label wrapped around it
private loadData = cachedSetup({
setup: async (opts: BaseOptions) => {
const buffer = await fetchAndMaybeUnzip(
openLocation(this.getConf('gffLocation'), this.pluginManager),
opts,
)
const { headerLines, linesByRef } = groupLinesByRef(
buffer,
opts.statusCallback,
)
const intervalTreeMap = makeFeatureIntervalTreeMap(
linesByRef,
// lines are already split and comment/FASTA-filtered by
// groupLinesByRef, so feed them straight to parseLinesLazy rather than
// re-joining and re-splitting through parseStringSync.
//
// Lazy because the whole file stays resident for the session: leaving
// column 9 as text rather than an object per attribute is what keeps
// that resident set small (8.5x on GENCODE-shaped input), and the
// render path reads only a handful of attributes anyway — see
// Gff3Feature.
//
// The id the tree entry carries is minted here, beside the feature
// rather than stamped onto it, so the parsed feature is the library's
// shape and Gff3Feature serializes it identically for both adapters.
(lines, refName) =>
parseLinesLazy(lines).map((feature, i) => ({
start: feature.start,
end: feature.end,
feature,
uniqueId: `${this.id}-${refName}-${i}`,
})),
'Parsing GFF data',
)
return { header: headerLines.join('\n'), intervalTreeMap }
},
})
public async getRefNames(opts: BaseOptions = {}) {
const { intervalTreeMap } = await this.loadData(opts)
return Object.keys(intervalTreeMap)
}
public async getHeader(opts: BaseOptions = {}) {
const { header } = await this.loadData(opts)
return header
}
public getFeatures(query: NoAssemblyRegion, opts: BaseOptions = {}) {
// no try/catch: ObservableCreate forwards a rejected callback to
// observer.error itself
return ObservableCreate<Feature>(async observer => {
const { start, end, refName } = query
const { intervalTreeMap } = await this.loadData(opts)
const tree = intervalTreeMap[refName]
if (tree) {
for (const { feature, uniqueId } of tree(opts.statusCallback).search([
start,
end,
])) {
observer.next(new Gff3Feature(feature, uniqueId))
}
}
observer.complete()
}, opts.signal)
}
}cachedSetupmemoizes the parse for every method that awaits it, and clears the memo on rejection so a failed load retries. Range-streaming adapters (BAM, tabix) skip it and read the index per query.- Prefix
uniqueIdwiththis.id: two tracks over one file must not collide.
For an API instead of a file, only the callback body changes: fetch with
opts?.signal, then observer.next(new SimpleFeature(...)) per hit.
To wrap another adapter, resolve it lazily with this.getSubAdapter — it is
async, so it cannot be called from a constructor. For the reference sequence
specifically, don't ask the config for it: declare READS_REFERENCE in the
adapter's adapterCapabilities, and JBrowse builds each instance with the
sequence adapter config of the assembly it is displayed against, one instance
per genome, in sequenceAdapterConfig. getSequenceSubAdapter reads that,
falling back to a configured slot only when one is set. A track then needs no
sequenceAdapter of its own:
public async configure() {
// the assembly's sequence, unless the config names another one
return getSequenceSubAdapter(this, this.getConf('sequenceAdapter'))
}Resolve it in one configure() the other methods await. Use getSubAdapter
directly for a subadapter that is genuinely part of the track's own
configuration. The method is optional on the base class, so it needs the ?.
and a check. dataAdapter comes back as the base union, so cast it.
Larger example:
MCScanAnchorsAdapter.
Registering the adapter type
The class is half of it; a plugin's install() registers the type.
BedGraphAdapter's registration is the whole file:
import AdapterType from '@jbrowse/core/pluggableElementTypes/AdapterType'
import configSchema, { normalizeSnapshot } from './configSchema.ts'
import type PluginManager from '@jbrowse/core/PluginManager'
export default function BedGraphAdapterF(pluginManager: PluginManager) {
pluginManager.addAdapterType(
() =>
new AdapterType({
name: 'BedGraphAdapter',
displayName: 'BedGraph adapter',
normalizeSnapshot,
configSchema,
getAdapterClass: () =>
import('./BedGraphAdapter.ts').then(r => r.default),
}),
)
}getAdapterClass returns a promise, so the adapter's parsing code stays out of
the startup bundle until a track using it opens. (An eager AdapterClass is
still accepted for older plugins; prefer the lazy form.)
The rest are optional:
adapterMetadatais how the adapter presents itself in the "Add track" form:categorygroups it in the dropdown,descriptionis the sentence under it,hiddenFromGUIkeeps it out entirely (right for an adapter only ever nested inside another), andalsoReadsis aRegExpof file names it can read but the extension guess does not hand it.alsoReadsis a form hint only — it does not enterCore-guessAdapterForLocation, so nothing changes about what a file resolves to headlessly or from the CLI.adapterCapabilitiesis a string list other code tests for, e.g.'exportData'(the VCF adapters) or'hasResolution'(bigWig).locationKeynames the config slot holding the primary file location, so import forms can pull the file back out of an existing track's config.normalizeSnapshotexpands a shorthand config —{ type, uri }— to the location slots the schema declares. This defaults to the config schema's ownpreProcessSnapshot, so declaring the shorthand there is enough and the two cannot come apart. Pass one here only to normalize differently before MST builds the config than during it, which nothing in tree needs.
Feature adapter API
getRefNames
Returns the refNames in the file. Used for "refname renaming", optional but useful when files use different conventions (e.g. chr1 vs 1). See reference renaming.
getFeatures
getFeatures(region, options)
Region is a snapshot of this MST model:
export const NoAssemblyRegion = types
.model('NoAssemblyRegion', {
refName: types.string,
start: types.number,
end: types.number,
reversed: types.optional(types.boolean, false),
})
.actions(self => ({
setRefName(newRefName: string): void {
self.refName = newRefName
},
}))
export const Region = types.compose(
'Region',
NoAssemblyRegion,
types.model({
assemblyName: types.string,
}),
)refName/start/end specify the genomic range, half-open and 0-based.
assemblyName is used when your adapter handles multiple assemblies (e.g.
synteny or a multi-assembly REST API).
originalRefName is not on Region. Refname renaming adds it to the object
it passes you — as the sequence adapter's (FASTA) name for the refname you were
queried with, so a CRAM/BAM adapter can fetch the matching reference sequence.
It is added only when a rename actually happened, so it is absent whenever
the track and the assembly already agree. To read it, type the parameter
AugmentedRegion (from @jbrowse/core/util, which is Region plus an optional
originalRefName) and handle the undefined case — the alignments adapters do
exactly this.
The options parameter is BaseOptions (from
@jbrowse/core/data_adapters/BaseAdapter). signal, headers and
statusCallback are the ones a typical adapter forwards; topLevelOnly and
lodMode are requests an adapter may honour or ignore:
export interface BaseOptions {
signal?: AbortSignal
bpPerPx?: number
sessionId?: string
// The single out-of-band status transport. A plain string is an indeterminate
// phase label; a StatusWithProgress object adds a determinate fraction
// (`current`/`total` are units-agnostic — bytes for a download, blocks for an
// unzip, features for a scan). Adapters wrap the raw byte counts from the
// index reader (@gmod/tabix, @gmod/bam, @gmod/cram) into this object form.
statusCallback?: StatusCallback
// What the adapter wants a reader told about this answer that the features
// cannot show — an index SNP no LD record names, say — appended here and
// carried back on the fetch's result to the display's corner notice.
notices?: string[]
headers?: Record<string, string>
// Which side of a pairing to answer getRefNames for; single-assembly
// adapters ignore it.
assemblyName?: string
// Which level-of-detail tier to read, for adapters that expose more than one
// (e.g. PIF's per-row CIGAR fine tier vs its no-CIGAR coarse tier). Absent, the
// fine tier is served; adapters without tiering ignore it entirely.
//
// This is a *resolved* tier, never the user's 'auto' setting: resolving auto
// needs a zoom, and it happens on the main thread in a display getter that
// feeds the fetch cache key (`resolveLodTier` in @jbrowse/synteny-core).
// Resolving it here instead hides a fetch input from that key, which is how
// a zoom across the threshold came to leave a view holding the wrong tier.
lodMode?: LodTier
// The lanes a multi-genome source is asked about, spelled as it spells them.
// A source that declares its lanes (`headerLanes`) reads only these lanes'
// files for its header and refNames as it does for its features; absent,
// every lane.
haplotypes?: string[]
// "I read only top-level features", so an adapter may skip work that exists
// to complete SUBFEATURE lists. A request, not an instruction: only the
// adapter knows whether its format's top-level set is even a function of the
// lines it read — `Gff3TabixAdapter` honours it by dropping the redispatch
// flanks, `GtfTabixAdapter` declines because it synthesizes genes from the
// transcripts it fetched, and both say why at the call.
//
// Set by the pre-fetch density probe, which counts and draws nothing. Not a
// rendering mode: a caller that will lay features out wants the whole tree.
topLevelOnly?: boolean
}Any rpcProps() the display model defines are spread in at the RPC call site,
so a display's user-facing settings reach the adapter under their own names. A
family of adapters that shares such settings declares its own options type
extending BaseOptions, as the comparative adapters do with
ComparativeOptions from @jbrowse/synteny-core:
export interface ComparativeOptions extends BaseOptions {
// The assembly on the *other* side of a synteny band, set by the synteny
// render RPC from the target view. Lets a multi-genome adapter (e.g.
// MultiGenomePAFAdapter) whose config lists all N assemblies isolate the exact
// pair a band draws — `assemblyName` alone can't, since one file backs every
// pair. Pairwise adapters (which already know their pair) ignore it.
targetAssemblyName?: string
// This side of the band, when it is not the region's assembly: a region on
// the anchor of a source that holds every lane inside the anchor's window (a
// pangenome graph, indexed on its reference alone) asks for this lane and
// `targetAssemblyName` aligned to each other, read inside that window. The
// records are on this lane's coordinates. An adapter declares that it
// answers these with the `lanePairsOnAnchor` capability.
queryAssemblyName?: string
// A multi-genome adapter answering a no-target query folds the pairs
// anchored on one query feature into one feature carrying `mates: [...]`,
// each mate with its own pairwise `orientation`. Absent, one `mate`-carrying
// feature per pair.
mateShape?: 'grouped'
// Each alignment record comes back cut to the region that fetched it, on both
// axes, with its alignment string dropped: a liftOver chain spans tens of Mb,
// and a display fitting a lane to whole records saw 80x the window. Honoured
// by `ComparativeAdapterBase.getFeaturesInMultipleRegions`, for records that
// ARE alignments — a gene-pair table's rows are genes and stay whole.
clipToRegion?: boolean
// With `clipToRegion`, each clipped record is further cut at every insertion
// or deletion of this many bp or more, into one record per gap-free run —
// ids and `syntenyId` suffixed per run — so a display that draws a record as
// one straight placement draws the runs the alignment actually has rather
// than a ribbon across a 25 kb indel. Read off the alignment string the clip
// walks anyway; a record with none is one run.
splitAtGapBp?: number
// With `clipToRegion`, each piece keeps its own stretch of the alignment as
// packed ops in `alignmentOps`, for a display that draws the indels between
// two lanes.
keepAlignment?: boolean
}statusCallback is how an adapter reports load progress to the UI (see
RPC and worker system).
Returns an rxjs Observable. Emit features with
observer.next(new SimpleFeature(...)) and finish with observer.complete().
No try/catch needed: ObservableCreate forwards a thrown error (or rejected
async callback) to observer.error().