About

FILER is a functional genomics repository for harmonized, indexed, and searchable human genomic annotation tracks. This page provides an overview of FILER documentation, data access, frequently asked questions, and the data formats and processing resources used by FILER.

Tutorials and documentation

The following resources cover the main ways to use FILER:

  • For examples of using the FILER web interface, see the FILER webserver tutorial.
  • To browse, search, filter, and download FILER tracks, use the FILER Explore page.
  • To install FILER on a local server or cloud instance, download selected or complete track collections, build FILER Giggle indexes, and query a local installation, see Deploying FILER.
  • FILER command-line, deployment, and data-processing code is available from the FILER2 GitHub repository.

Accessing FILER data

FILER supports both interactive track downloads through the website and metadata-driven access for larger or reproducible workflows. The metadata includes download URLs and other information needed to identify and retrieve individual FILER tracks.

Individual tracks and filtered subsets

Individual tracks can be downloaded directly from the FILER Explore page using the Download link associated with a track. The Explore page can also be filtered by genome build, data source, assay, tissue, genomic region, and other metadata fields to identify a custom set of tracks.

For a filtered result set, use the Download metadata option on the Explore page to obtain metadata for the selected tracks. That metadata can be inspected directly or used as an installation manifest for a local FILER deployment. See Deploying FILER for the installation workflow.

Complete FILER metadata

Current track metadata, including processed track download URLs, is available as tab-separated (TSV) templated metadata for the supported genome builds:

Definitions of the track metadata fields are available in the FILER v2 metadata schema (XLSX) and as a JSON metadata schema. For command-line metadata retrieval, filtering, storage estimates, and local installation, see Downloading and understanding FILER metadata.

Frequently Asked Questions (FAQ)

How can I download an individual FILER track?

The simplest method is to locate the track on the FILER Explore page and use its Download link. The track metadata also contains the processed_file_download_url field, which provides the direct URL for the processed FILER track file.

How can I download or deploy a custom subset of FILER tracks?

Use the filters on the FILER Explore page to define the desired track set, then download the metadata for the filtered results. The resulting metadata can be used to retrieve the individual tracks or supplied to the FILER installation tools to build the corresponding local subset. Command-line filtering of the complete FILER metadata is also supported. See Deploying FILER data for examples and installation instructions.

How do I install and query a local FILER instance?

See the Deploying FILER guide. It covers prerequisites, installation of the FILER command-line tools, metadata retrieval, full and subset deployments, Giggle indexing, and genomic interval overlap queries. The corresponding scripts and configuration files are maintained in the FILER2 GitHub repository.

Why does a FILER script fail with declare: -A: invalid option or a Bash version error?

FILER command-line scripts require Bash 4.3 or later. Check the version used to run the scripts with:

bash --version

If an older Bash version is being used, install or select a newer version before running FILER. On macOS, for example, a current Bash can be installed with Homebrew using brew install bash. See the prerequisites section of Deploying FILER for the rest of the command-line requirements.

What is the difference between FILER track metadata and a track file schema?

Track metadata describes the track as a dataset: for example, its data source, assay, genome build, file format, file size, and download location. A track file schema describes the columns contained within the processed track file itself.

The file_format value in the track metadata corresponds to FILER_BED_format in the FILER file-schema table. The current file-schema table is available at filer2.schemas.latest.tsv.

Where can I find explanations of genomic data terms used by FILER?

See the FILER Summary page for descriptions of FILER data categories and collections. For terminology commonly used by functional genomics resources, the ENCODE glossary and ENCODE data terms provide additional background.

Data formats and processing

FILER integrates functional genomics data from multiple source resources and provides processed tracks in a harmonized format for browsing, download, indexing, and genomic interval search. Track-level metadata records the source and characteristics of each processed track, while file schemas describe the structure of the processed files.

Metadata and provenance

The FILER v2 metadata schema (XLSX) describes the FILER track metadata fields. Additional information about the sources and collection of metadata is available in metadata_info.xlsx.

The current complete metadata files and the JSON metadata schema are linked in the Accessing FILER data section above.

Processed file formats

Most processed FILER genomic tracks are distributed in BED-based formats. Reference-panel VCF files and genomic reference sequence FASTA files are exceptions. The current FILER file-schema table describes the columns used by each processed track format: filer2.schemas.latest.tsv.

BED-formatted FILER tracks use 0-based BED coordinates and are coordinate-sorted. FILER uses FILER Giggle for interval indexing and search, and HTSlib/tabix for indexed access to compressed genomic files where applicable.

Data processing and standardization code

Scripts used for FILER data processing and standardization are maintained in the FILER2 data_processing directory. These scripts document the processing used to convert source datasets into FILER-compatible track formats.

Data-processing code is distinct from the scripts used to deploy and query a local FILER installation. For local deployment and querying, see Deploying FILER.