Deploying FILER

FILER can be deployed on a local server or cloud computing instance as either a complete collection for a selected genome build or as a custom subset of FILER tracks. This guide covers installation of the required software, obtaining FILER metadata, downloading and indexing FILER data, and querying a local FILER deployment.

The FILER2 code repository provides the command-line scripts and configuration files used throughout this guide.

Installing prerequisites

FILER deployment and querying require several command-line utilities for downloading, processing, compressing, indexing, and searching genomic data. The main prerequisite is FILER Giggle, the interval-indexing and search engine used by FILER. Additional utilities include bgzip/tabix, samtools, jq, mlr (Miller), mawk, wget, git, and GNU core utilities.

The recommended installation method is Homebrew because it provides the required tools consistently across Linux, macOS, and Linux distributions running under Windows Subsystem for Linux 2 (WSL 2).

Linux, macOS, and Windows WSL 2 using Homebrew

If Homebrew is not already installed, install it using the official installer:

/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"

Follow any shell configuration instructions printed by the Homebrew installer before continuing. Then install FILER Giggle and the remaining command-line prerequisites:

brew install htslib jq miller samtools mawk wget git coreutils pkuksa/tap/filer-giggle

This installs the principal commands used by the FILER scripts, including:

  • giggle — FILER Giggle interval indexing and search;
  • bgzip and tabix — genomic file compression and indexing utilities provided by HTSlib;
  • samtools — utilities for processing genomic data;
  • jq — command-line processing of JSON data;
  • mlr — Miller command-line processing of tabular data;
  • mawk — AWK implementation used for text processing;
  • wget and git — downloading data and retrieving FILER source repositories;
  • coreutils — GNU command-line utilities used by FILER scripts.

Additional macOS setup for GNU core utilities

On macOS, Homebrew installs many GNU core utilities with a g prefix. To make the GNU versions available using their standard command names, add the Homebrew gnubin directory to your PATH:

export PATH="$(brew --prefix coreutils)/libexec/gnubin:$PATH"

To make this setting permanent, add the command to the appropriate shell startup file, such as ~/.zshrc or ~/.bashrc.

Verifying the installation

The following commands can be used to confirm that the principal FILER prerequisites are available on the command line:

command -v giggle
command -v bgzip
command -v tabix
command -v samtools
command -v jq
command -v mlr
command -v mawk
command -v wget
command -v git

Each command should report the path to the corresponding executable. FILER Giggle can also be tested directly with:


giggle --help
giggle, v0.6.3fsbv

Installing prerequisites without Homebrew

The supporting command-line utilities can also be installed using the operating system package manager. FILER Giggle must then be installed separately, as described below.

Ubuntu/Debian

Install the required command-line utilities with:

sudo apt-get update
sudo apt-get install -y tabix jq miller samtools mawk wget git coreutils

The tabix package provides both tabix and bgzip.

Fedora

On current Fedora releases, install the HTSlib command-line tools together with the remaining FILER prerequisites:

sudo dnf install -y htslib-tools jq miller samtools mawk wget git coreutils

On RHEL or CentOS-based systems, package names and availability may depend on the distribution release and enabled repositories (for example, EPEL). Ensure that bgzip, tabix, and the other commands listed above are available before proceeding.

Installing FILER Giggle from source

If Homebrew is not being used, FILER Giggle can be compiled directly from the FILER Giggle source repository.

Build dependencies on Ubuntu/Debian

sudo apt-get install -y gcc make autoconf zlib1g-dev libbz2-dev \
    libcurl4-openssl-dev libssl-dev ruby

Build dependencies on Fedora/RHEL/CentOS

sudo dnf groupinstall "Development Tools"
sudo dnf install -y autoconf zlib-devel bzip2-devel \
    libcurl-devel openssl-devel ruby

Clone and build FILER Giggle

Clone the FILER Giggle repository:

git clone https://github.com/pkuksa/FILER_giggle.git FILER_giggle
cd FILER_giggle

On Linux, compile using:

make

On macOS, first try the standard build. If that fails, clean the build and use the macOS-specific Makefile:

make clean
make -f Makefile.macos

After a successful source build, the FILER Giggle executable is located at:

FILER_giggle/bin/giggle

Verify the compiled executable before proceeding:

./bin/giggle --help

If FILER Giggle was built from source rather than installed through Homebrew, make sure that the FILER configuration file points to the absolute path of this giggle executable.

Installing command-line FILER scripts

The FILER2 repository contains command-line utilities for downloading, installing, indexing, and querying a local FILER deployment. These scripts work with the FILER metadata service to determine which tracks should be downloaded and how they should be organized on disk.

Clone the FILER2 repository into a directory of your choice:

git clone https://github.com/wanglab-upenn/FILER2 FILER2_scripts

The examples below assume that commands are run from the directory containing FILER2_scripts.

No separate compilation step is required for the FILER shell scripts, but the prerequisite command-line tools described above must be installed before downloading or querying FILER data.

FILER uses a configuration file to locate required executables and local data directories. For Homebrew-based installations, the repository provides FILER2_scripts/data/filer.homebrew.ini. If the prerequisite tools were installed manually, create or modify the configuration file so that its paths match your system.

Downloading and understanding FILER metadata

Track metadata

FILER provides track-level metadata describing the datasets and genomic annotation tracks available in each genome build. Each row of the metadata represents a FILER track and contains information such as its data source, assay, file format, file size, and download location.

The metadata also serves as an installation manifest for local FILER deployments. A complete metadata file can be supplied to the FILER installation tools to deploy all FILER tracks for the selected genome build, while a filtered metadata file can be used to install only a selected subset of FILER.

Full FILER metadata, including track download URLs, is available in tab-separated (TSV) templated metadata files for both supported genome builds:

A description of the metadata fields is available in the FILER v2 metadata schema [XLS spreadsheet; June 2026; 14 KB] . The metadata schema is also available in JSON format.

Downloading metadata from the command line

The current metadata can also be downloaded directly from the FILER metadata service. For example:

wget "https://filer2.niagads.org/metadata/hg38/download/tsv" -O filer2.hg38.metadata.template.tsv

wget "https://filer2.niagads.org/metadata/hg19/download/tsv" -O filer2.hg19.metadata.template.tsv

These files can be inspected directly, filtered to select tracks of interest, or passed to install_filer.sh when constructing a local FILER deployment.

Download information in the metadata

The templated metadata contains the information required to retrieve and organize each FILER track. In particular:

  • processed_file_download_url contains the URL of the processed FILER track file.
  • wget_command contains a ready-to-run command for downloading the track and placing it within the appropriate FILER directory hierarchy.
  • file_size contains the size of the processed track file in bytes and can be used to estimate storage requirements before installation.

The FILER directory hierarchy organizes installed tracks according to properties such as data source, assay, data format, and genome build. Normally, users do not need to construct this hierarchy manually; the FILER installation scripts use the information in the metadata to create the appropriate layout.

Track file schemas

The track metadata describes each FILER track as a whole. A separate file schema table describes the columns contained within the individual processed track files.

The latest FILER file-schema table is available at:

https://filer2.niagads.org/metadata/filer2.schemas.latest.tsv

A track schema can be also accessed in JSON format using a track ID, e.g.,
https://filer2.niagads.org/tracks/NGDSHLCVB4QUD7/schema

To determine the schema of a particular FILER track, use the file_format value from the track metadata and match it to the FILER_BED_format column in the file-schema table. The matching row describes the structure of that track format.

The principal schema fields are:

  • FILER_BED_format — FILER file-format identifier, corresponding to file_format in the track metadata.
  • FILER_BED_type — BED type expressed using bedX+Y notation.
  • FILER_BED_schema — semicolon-separated list of fields contained in the processed track file.
  • FILER_BED_total_columns — total number of columns in the processed track file.
  • FILER_BED_autoSQL_schema_files — corresponding AutoSQL schema information for the track format.

Together, the track metadata and file-schema table provide two complementary levels of information: the metadata describes what each track is and where it can be obtained, while the file schema describes how the contents of that track are structured.

Downloading and installing FILER data

A local FILER deployment is created from a FILER metadata file. The metadata determines which tracks will be installed and contains the information needed to download and organize those tracks within the FILER directory hierarchy.

The install_filer.sh script, included with the FILER command-line tools, automates this process. It downloads all tracks listed in the supplied metadata file into the target directory and then creates the FILER Giggle indexes used for genomic interval queries.

Because the metadata file controls which tracks are installed, the same installation procedure can be used to deploy:

  • all FILER tracks for a selected genome build;
  • a single FILER data source;
  • tracks returned by a FILER search; or
  • a custom subset produced by filtering the FILER metadata.

Running the FILER installation script

The general form of the installation command is:

bash FILER2_scripts/install_filer.sh <target_dir> <metadata_file> <config_file>

The three arguments are:

  • target_dir — directory in which the local FILER data hierarchy and indexes will be created.
  • metadata_file — full or filtered FILER metadata describing the tracks to install.
  • config_file — FILER configuration file containing paths to the required command-line tools and other local configuration settings.

The examples below use FILER2_scripts/data/filer.homebrew.ini, which is provided in the FILER2 GitHub repository and is configured for installations in which the prerequisite tools were installed using Homebrew. If the tools were installed manually or in non-standard locations, adjust the configuration file accordingly.

Installing tracks returned by a FILER search

FILER search results can be downloaded directly as templated metadata and supplied to install_filer.sh. This makes it possible to construct a local FILER installation containing only tracks matching a particular biological query.

For example, the following commands retrieve metadata for tracks matching ATAC-seq Brain and install those tracks locally:


wget "https://filer2.niagads.org/search?query=ATAC-seq Brain&outputFormat=tsv" -O filer_metadata.selected.tsv

bash FILER2_scripts/install_filer.sh FILER2_data filer_metadata.selected.tsv FILER2_scripts/data/filer.homebrew.ini

In this example, FILER2_data is the root directory of the new local FILER installation. The tracks described in filer_metadata.selected.tsv will be downloaded beneath this directory and indexed for subsequent FILER queries.

The Homebrew configuration file used in this example is available in the FILER2 repository: FILER2_scripts/data/filer.homebrew.ini .

Installing a custom subset of FILER tracks

A custom installation can also be created by downloading the complete FILER metadata and filtering it before running install_filer.sh. The filtered metadata file then acts as the installation manifest for the local deployment.

For example, the following commands download the current hg38 metadata and create a subset containing entries with the term enhancer. The first metadata row is retained so that the TSV header remains present:

wget "https://filer2.niagads.org/metadata/hg38/download/tsv" -O filer2.hg38.metadata.template.tsv

awk 'NR == 1 || tolower($0) ~ /enhancer/' \
    filer2.hg38.metadata.template.tsv \
    > filer2.hg38.metadata.enhancers.template.tsv

bash FILER2_scripts/install_filer.sh FILER2_data \
    filer2.hg38.metadata.enhancers.template.tsv \
    FILER2_scripts/data/filer.homebrew.ini

More selective filtering can be performed using individual metadata columns when constructing specialized FILER deployments.

Installing an individual FILER data source

Metadata for a single FILER data source can be requested using the dataSource parameter. This is useful when a local installation should contain an entire source collection without downloading unrelated FILER tracks.

For example, the following commands install the hg38 tracks from the MiGA data source:

wget "https://filer2.niagads.org/metadata/hg38/download/tsv?dataSource=MiGA" -O filer2.hg38.MiGA.tsv

bash FILER2_scripts/install_filer.sh FILER2_data filer2.hg38.MiGA.tsv FILER2_scripts/data/filer.homebrew.ini

Deploying all FILER tracks for a genome build

To deploy all FILER tracks for hg38, download the full hg38 metadata file and use it directly as the installation manifest:

wget "https://filer2.niagads.org/metadata/hg38/download/tsv" -O filer2.hg38.metadata.template.tsv

bash FILER2_scripts/install_filer.sh FILER2_data \
    filer2.hg38.metadata.template.tsv \
    FILER2_scripts/data/filer.homebrew.ini

This installs all hg38 tracks represented in the downloaded metadata. A complete FILER deployment requires substantial disk space because storage is needed both for the downloaded track files and for the corresponding Giggle indexes.

The same procedure can be used for hg19 by downloading the hg19 metadata and installing it into an hg19-compatible local deployment.

Estimating storage requirements

Before beginning a large or complete deployment, it is recommended to estimate the total size of the files represented in the metadata. The metadata field file_size contains the size of each track file in bytes.

The following command locates the file_size column by name and reports the approximate total download size in gigabytes:

awk -F $'\t' '
NR == 1 {
    for (i = 1; i <= NF; i++) {
        if ($i == "file_size") {
            file_size_col = i
            break
        }
    }
    next
}
{
    total += $file_size_col
}
END {
    printf "%.2f GB\n", total / 10^9
}' filer2.hg38.metadata.template.tsv

The resulting value depends on the current FILER release and the metadata being installed. Additional storage should be reserved for the Giggle indexes generated during installation.

Installing FILER data collections

FILER tracks can also be browsed by data collection. The FILER data collections table provides another way to identify collections of related tracks that can be selected for local deployment.

Querying FILER data

Once FILER tracks have been downloaded and Giggle-indexed, a local FILER installation can be queried for tracks and genomic records that overlap a set of genomic intervals.

The FILER command-line tools provide data_querying/get_overlaps.sh for performing these searches. The script takes a BED file containing query intervals, searches the appropriate FILER Giggle indexes, combines the results across FILER data collections, and annotates each result with the corresponding FILER track metadata.

Running an example query

The FILER2 repository includes a small test BED file that can be used to verify a local installation. The following example searches the test intervals against all locally installed hg38 FILER Giggle indexes:

bash FILER2_scripts/data_querying/get_overlaps.sh \
    --inBed FILER2_scripts/data/test.20intervals.bed.gz \
    --configFile FILER2_scripts/data/filer.homebrew.ini \
    --outputDir filer_test_overlaps/ \
    --verboseSearch 1 \
    --genomeBuild hg38 \
    --forceOverwrite 1

When --giggleIndexDirList is not specified, the script scans the FILER installation defined by FILERDIR in the configuration file and searches all Giggle indexes corresponding to the requested genome build.

Query parameters

  • --inBed — BED file containing the genomic intervals to query. The current script accepts either an uncompressed BED file or a gzip/bgzip-compressed BED file. The input is coordinate-sorted and bgzip-compressed internally before the Giggle search is performed.
  • --configFile — FILER configuration file. Among other settings, this identifies the local FILER installation, FILER metadata, Giggle executable, and supporting command-line tools.
  • --outputDir — directory in which the query results and working files will be created.
  • --genomeBuild — genome build to search, for example hg38 or hg19. This is required when --giggleIndexDirList is not supplied.
  • --giggleIndexDirList — optional text file containing the absolute paths of specific Giggle index directories to search, one directory per line. This can be used to restrict a query to selected portions of a FILER installation.
  • --verboseSearch — when set to 1, reports individual overlapping records in addition to the track information. Verbose searches generate more detailed output and may be slower.
  • --forceOverwrite — when set to 1, allows reuse of an existing output directory.
  • --tempDir — optional temporary directory used during sorting of the combined results. The default is /tmp. For large searches, this can be changed to a location with more available temporary storage.

Querying your own genomic intervals

To query your own regions, replace the value supplied to --inBed with your BED file. The genomic coordinates must correspond to the genome build being searched.

bash FILER2_scripts/data_querying/get_overlaps.sh \
    --inBed my_regions.bed.gz \
    --configFile FILER2_scripts/data/filer.homebrew.ini \
    --outputDir my_filer_overlaps/ \
    --genomeBuild hg38 \
    --verboseSearch 1

For an hg19 query, use hg19 coordinates and specify --genomeBuild hg19. The corresponding hg19 FILER data and Giggle indexes must be present in the local installation.

Query output

The script produces two combined overlap files in the top-level output directory:

  • filer_overlaps.bed — combined Giggle overlap results from all indexes searched.
  • filer_overlaps.with_meta.bed — the combined overlap results with the corresponding FILER track metadata appended to each result. This is generally the most useful output for downstream analysis because each overlap can be interpreted in the context of its FILER data source, assay, classification, and other track annotations.

The columns at the beginning of each result correspond to the original fields in the input BED file. If the input BED begins with a #-prefixed header, those field names are retained. Otherwise, the query columns are named inputField1, inputField2, and so on.

With the default non-verbose search (--verboseSearch 0), the overlap-specific columns are:

  • trackFile — FILER track file containing the overlap;
  • trackNumIntervals — number of intervals represented for the track in the Giggle result;
  • numOverlaps — number of overlaps with the query interval.

With --verboseSearch 1, individual overlapping genomic records are reported and the overlap-specific columns are:

  • trackFile — FILER track file containing the overlapping record;
  • hitString — the individual overlapping record returned by Giggle. Tabs within this record are represented as @@@ so that the record can be stored within a single output column;
  • numOverlaps — set to 1 for each individual reported hit.

In filer_overlaps.with_meta.bed, the complete FILER metadata columns are appended after these overlap columns.

Results grouped by genomic feature type

In addition to the combined result files, the script uses the FILER metadata field classification to divide the metadata-enriched overlaps into separate files by genomic feature type.

These files are written under:

overlaps_by_feature_type/

with filenames of the form:

filer_overlaps.<classification>.bed

For example, if the query overlaps tracks belonging to several FILER classifications, a separate result file is created for each classification. Each of these files contains the same metadata-enriched columns as filer_overlaps.with_meta.bed, but only for tracks belonging to that feature type.

Additional files in the output directory

The script retains several files used to perform and organize the search:

  • input.bed.gz — coordinate-sorted, bgzip-compressed copy of the input intervals used for the Giggle searches.
  • giggle_index_dirs_for_search.<genomeBuild>.txt — automatically generated list of Giggle index directories when --giggleIndexDirList is not supplied.
  • overlaps/ — per-index Giggle search output. Its directory structure mirrors the corresponding organization of the local FILER installation.
  • overlaps/.../giggle_out.txt — raw/intermediate results produced for each Giggle index searched.

The top-level filer_overlaps.with_meta.bed and the files in overlaps_by_feature_type/ are generally the most useful results for subsequent analysis.

Restricting a search to selected FILER indexes

By default, specifying --genomeBuild searches all matching Giggle indexes found under the local FILER installation. For more targeted searches, a list of index directories can instead be supplied using --giggleIndexDirList.

The list must contain one absolute Giggle index directory per line:

/path/to/FILER/.../giggle_index
/path/to/FILER/.../giggle_index
/path/to/FILER/.../giggle_index

When an explicit index list is supplied, only those indexes are queried and --genomeBuild is not required.