Deploying FILER
FILER can be deployed on a local server or cloud computing instance as either a complete collection for a selected genome build or as a custom subset of FILER tracks. This guide covers installation of the required software, obtaining FILER metadata, downloading and indexing FILER data, and querying a local FILER deployment.
The FILER2 code repository provides the command-line scripts and configuration files used throughout this guide.
Installing prerequisites
FILER deployment and querying require several command-line utilities for
downloading, processing, compressing, indexing, and searching genomic data.
The main prerequisite is FILER Giggle, the interval-indexing
and search engine used by FILER. Additional utilities include
bgzip/tabix, samtools,
jq, mlr (Miller), mawk,
wget, git, and GNU core utilities.
The recommended installation method is Homebrew because it provides the required tools consistently across Linux, macOS, and Linux distributions running under Windows Subsystem for Linux 2 (WSL 2).
Linux, macOS, and Windows WSL 2 using Homebrew
If Homebrew is not already installed, install it using the official installer:
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
Follow any shell configuration instructions printed by the Homebrew installer before continuing. Then install FILER Giggle and the remaining command-line prerequisites:
brew install htslib jq miller samtools mawk wget git coreutils pkuksa/tap/filer-giggle
This installs the principal commands used by the FILER scripts, including:
-
giggle— FILER Giggle interval indexing and search; -
bgzipandtabix— genomic file compression and indexing utilities provided by HTSlib; -
samtools— utilities for processing genomic data; -
jq— command-line processing of JSON data; -
mlr— Miller command-line processing of tabular data; -
mawk— AWK implementation used for text processing; -
wgetandgit— downloading data and retrieving FILER source repositories; -
coreutils— GNU command-line utilities used by FILER scripts.
Additional macOS setup for GNU core utilities
On macOS, Homebrew installs many GNU core utilities with a
g prefix. To make the GNU versions available using their
standard command names, add the Homebrew gnubin directory to
your PATH:
export PATH="$(brew --prefix coreutils)/libexec/gnubin:$PATH"
To make this setting permanent, add the command to the appropriate shell
startup file, such as ~/.zshrc or ~/.bashrc.
Verifying the installation
The following commands can be used to confirm that the principal FILER prerequisites are available on the command line:
command -v giggle
command -v bgzip
command -v tabix
command -v samtools
command -v jq
command -v mlr
command -v mawk
command -v wget
command -v git
Each command should report the path to the corresponding executable. FILER Giggle can also be tested directly with:
giggle --help
giggle, v0.6.3fsbv
Installing prerequisites without Homebrew
The supporting command-line utilities can also be installed using the operating system package manager. FILER Giggle must then be installed separately, as described below.
Ubuntu/Debian
Install the required command-line utilities with:
sudo apt-get update
sudo apt-get install -y tabix jq miller samtools mawk wget git coreutils
The tabix package provides both tabix and
bgzip.
Fedora
On current Fedora releases, install the HTSlib command-line tools together with the remaining FILER prerequisites:
sudo dnf install -y htslib-tools jq miller samtools mawk wget git coreutils
On RHEL or CentOS-based systems, package names and availability may depend
on the distribution release and enabled repositories (for example, EPEL).
Ensure that bgzip, tabix, and the other commands
listed above are available before proceeding.
Installing FILER Giggle from source
If Homebrew is not being used, FILER Giggle can be compiled directly from the FILER Giggle source repository.
Build dependencies on Ubuntu/Debian
sudo apt-get install -y gcc make autoconf zlib1g-dev libbz2-dev \
libcurl4-openssl-dev libssl-dev ruby
Build dependencies on Fedora/RHEL/CentOS
sudo dnf groupinstall "Development Tools"
sudo dnf install -y autoconf zlib-devel bzip2-devel \
libcurl-devel openssl-devel ruby
Clone and build FILER Giggle
Clone the FILER Giggle repository:
git clone https://github.com/pkuksa/FILER_giggle.git FILER_giggle
cd FILER_giggle
On Linux, compile using:
make
On macOS, first try the standard build. If that fails, clean the build and use the macOS-specific Makefile:
make clean
make -f Makefile.macos
After a successful source build, the FILER Giggle executable is located at:
FILER_giggle/bin/giggle
Verify the compiled executable before proceeding:
./bin/giggle --help
If FILER Giggle was built from source rather than installed through
Homebrew, make sure that the FILER configuration file points to the
absolute path of this giggle executable.
Installing command-line FILER scripts
The FILER2 repository contains command-line utilities for downloading, installing, indexing, and querying a local FILER deployment. These scripts work with the FILER metadata service to determine which tracks should be downloaded and how they should be organized on disk.
Clone the FILER2 repository into a directory of your choice:
git clone https://github.com/wanglab-upenn/FILER2 FILER2_scripts
The examples below assume that commands are run from the directory containing
FILER2_scripts.
No separate compilation step is required for the FILER shell scripts, but the prerequisite command-line tools described above must be installed before downloading or querying FILER data.
FILER uses a configuration file to locate required executables and local
data directories. For Homebrew-based installations, the repository provides
FILER2_scripts/data/filer.homebrew.ini. If the prerequisite tools were
installed manually, create or modify the configuration file so that its
paths match your system.
Downloading and understanding FILER metadata
Track metadata
FILER provides track-level metadata describing the datasets and genomic annotation tracks available in each genome build. Each row of the metadata represents a FILER track and contains information such as its data source, assay, file format, file size, and download location.
The metadata also serves as an installation manifest for local FILER deployments. A complete metadata file can be supplied to the FILER installation tools to deploy all FILER tracks for the selected genome build, while a filtered metadata file can be used to install only a selected subset of FILER.
Full FILER metadata, including track download URLs, is available in tab-separated (TSV) templated metadata files for both supported genome builds:
A description of the metadata fields is available in the FILER v2 metadata schema [XLS spreadsheet; June 2026; 14 KB] . The metadata schema is also available in JSON format.
Downloading metadata from the command line
The current metadata can also be downloaded directly from the FILER metadata service. For example:
wget "https://filer2.niagads.org/metadata/hg38/download/tsv" -O filer2.hg38.metadata.template.tsv
wget "https://filer2.niagads.org/metadata/hg19/download/tsv" -O filer2.hg19.metadata.template.tsv
These files can be inspected directly, filtered to select tracks of interest,
or passed to install_filer.sh when constructing a local FILER
deployment.
Download information in the metadata
The templated metadata contains the information required to retrieve and organize each FILER track. In particular:
-
processed_file_download_urlcontains the URL of the processed FILER track file. -
wget_commandcontains a ready-to-run command for downloading the track and placing it within the appropriate FILER directory hierarchy. -
file_sizecontains the size of the processed track file in bytes and can be used to estimate storage requirements before installation.
The FILER directory hierarchy organizes installed tracks according to properties such as data source, assay, data format, and genome build. Normally, users do not need to construct this hierarchy manually; the FILER installation scripts use the information in the metadata to create the appropriate layout.
Track file schemas
The track metadata describes each FILER track as a whole. A separate file schema table describes the columns contained within the individual processed track files.
The latest FILER file-schema table is available at:
https://filer2.niagads.org/metadata/filer2.schemas.latest.tsv
A track schema can be also accessed in JSON format using a track ID, e.g.,
https://filer2.niagads.org/tracks/NGDSHLCVB4QUD7/schema
To determine the schema of a particular FILER track, use the
file_format value from the track metadata and match it to the
FILER_BED_format column in the file-schema table. The matching
row describes the structure of that track format.
The principal schema fields are:
-
FILER_BED_format— FILER file-format identifier, corresponding tofile_formatin the track metadata. -
FILER_BED_type— BED type expressed usingbedX+Ynotation. -
FILER_BED_schema— semicolon-separated list of fields contained in the processed track file. -
FILER_BED_total_columns— total number of columns in the processed track file. -
FILER_BED_autoSQL_schema_files— corresponding AutoSQL schema information for the track format.
Together, the track metadata and file-schema table provide two complementary levels of information: the metadata describes what each track is and where it can be obtained, while the file schema describes how the contents of that track are structured.
Downloading and installing FILER data
A local FILER deployment is created from a FILER metadata file. The metadata determines which tracks will be installed and contains the information needed to download and organize those tracks within the FILER directory hierarchy.
The install_filer.sh script, included with the FILER command-line
tools, automates this process. It downloads all tracks listed in the supplied
metadata file into the target directory and then creates the FILER Giggle
indexes used for genomic interval queries.
Because the metadata file controls which tracks are installed, the same installation procedure can be used to deploy:
- all FILER tracks for a selected genome build;
- a single FILER data source;
- tracks returned by a FILER search; or
- a custom subset produced by filtering the FILER metadata.
Running the FILER installation script
The general form of the installation command is:
bash FILER2_scripts/install_filer.sh <target_dir> <metadata_file> <config_file>
The three arguments are:
-
target_dir— directory in which the local FILER data hierarchy and indexes will be created. -
metadata_file— full or filtered FILER metadata describing the tracks to install. -
config_file— FILER configuration file containing paths to the required command-line tools and other local configuration settings.
The examples below use FILER2_scripts/data/filer.homebrew.ini, which is provided
in the FILER2 GitHub repository and is configured for installations in which
the prerequisite tools were installed using Homebrew. If the tools were
installed manually or in non-standard locations, adjust the configuration
file accordingly.
Installing tracks returned by a FILER search
FILER search results can be downloaded directly as templated metadata and
supplied to install_filer.sh. This makes it possible to construct
a local FILER installation containing only tracks matching a particular
biological query.
For example, the following commands retrieve metadata for tracks matching
ATAC-seq Brain and install those tracks locally:
wget "https://filer2.niagads.org/search?query=ATAC-seq Brain&outputFormat=tsv" -O filer_metadata.selected.tsv
bash FILER2_scripts/install_filer.sh FILER2_data filer_metadata.selected.tsv FILER2_scripts/data/filer.homebrew.ini
In this example, FILER2_data is the root directory of the new
local FILER installation. The tracks described in
filer_metadata.selected.tsv will be downloaded beneath this
directory and indexed for subsequent FILER queries.
The Homebrew configuration file used in this example is available in the FILER2 repository: FILER2_scripts/data/filer.homebrew.ini .
Installing a custom subset of FILER tracks
A custom installation can also be created by downloading the complete FILER
metadata and filtering it before running install_filer.sh.
The filtered metadata file then acts as the installation manifest for the
local deployment.
For example, the following commands download the current hg38 metadata and
create a subset containing entries with the term enhancer.
The first metadata row is retained so that the TSV header remains present:
wget "https://filer2.niagads.org/metadata/hg38/download/tsv" -O filer2.hg38.metadata.template.tsv
awk 'NR == 1 || tolower($0) ~ /enhancer/' \
filer2.hg38.metadata.template.tsv \
> filer2.hg38.metadata.enhancers.template.tsv
bash FILER2_scripts/install_filer.sh FILER2_data \
filer2.hg38.metadata.enhancers.template.tsv \
FILER2_scripts/data/filer.homebrew.ini
More selective filtering can be performed using individual metadata columns when constructing specialized FILER deployments.
Installing an individual FILER data source
Metadata for a single FILER data source can be requested using the
dataSource parameter. This is useful when a local installation
should contain an entire source collection without downloading unrelated
FILER tracks.
For example, the following commands install the hg38 tracks from the
MiGA data source:
wget "https://filer2.niagads.org/metadata/hg38/download/tsv?dataSource=MiGA" -O filer2.hg38.MiGA.tsv
bash FILER2_scripts/install_filer.sh FILER2_data filer2.hg38.MiGA.tsv FILER2_scripts/data/filer.homebrew.ini
Deploying all FILER tracks for a genome build
To deploy all FILER tracks for hg38, download the full hg38 metadata file and use it directly as the installation manifest:
wget "https://filer2.niagads.org/metadata/hg38/download/tsv" -O filer2.hg38.metadata.template.tsv
bash FILER2_scripts/install_filer.sh FILER2_data \
filer2.hg38.metadata.template.tsv \
FILER2_scripts/data/filer.homebrew.ini
This installs all hg38 tracks represented in the downloaded metadata. A complete FILER deployment requires substantial disk space because storage is needed both for the downloaded track files and for the corresponding Giggle indexes.
The same procedure can be used for hg19 by downloading the hg19 metadata and installing it into an hg19-compatible local deployment.
Estimating storage requirements
Before beginning a large or complete deployment, it is recommended to
estimate the total size of the files represented in the metadata.
The metadata field file_size contains the size of each track
file in bytes.
The following command locates the file_size column by name and
reports the approximate total download size in gigabytes:
awk -F $'\t' '
NR == 1 {
for (i = 1; i <= NF; i++) {
if ($i == "file_size") {
file_size_col = i
break
}
}
next
}
{
total += $file_size_col
}
END {
printf "%.2f GB\n", total / 10^9
}' filer2.hg38.metadata.template.tsv
The resulting value depends on the current FILER release and the metadata being installed. Additional storage should be reserved for the Giggle indexes generated during installation.
Installing FILER data collections
FILER tracks can also be browsed by data collection. The FILER data collections table provides another way to identify collections of related tracks that can be selected for local deployment.
Querying FILER data
Once FILER tracks have been downloaded and Giggle-indexed, a local FILER installation can be queried for tracks and genomic records that overlap a set of genomic intervals.
The FILER command-line tools provide
data_querying/get_overlaps.sh for performing these searches.
The script takes a BED file containing query intervals, searches the
appropriate FILER Giggle indexes, combines the results across FILER data
collections, and annotates each result with the corresponding FILER track
metadata.
Running an example query
The FILER2 repository includes a small test BED file that can be used to verify a local installation. The following example searches the test intervals against all locally installed hg38 FILER Giggle indexes:
bash FILER2_scripts/data_querying/get_overlaps.sh \
--inBed FILER2_scripts/data/test.20intervals.bed.gz \
--configFile FILER2_scripts/data/filer.homebrew.ini \
--outputDir filer_test_overlaps/ \
--verboseSearch 1 \
--genomeBuild hg38 \
--forceOverwrite 1
When --giggleIndexDirList is not specified, the script scans
the FILER installation defined by FILERDIR in the configuration
file and searches all Giggle indexes corresponding to the requested genome
build.
Query parameters
-
--inBed— BED file containing the genomic intervals to query. The current script accepts either an uncompressed BED file or a gzip/bgzip-compressed BED file. The input is coordinate-sorted and bgzip-compressed internally before the Giggle search is performed. -
--configFile— FILER configuration file. Among other settings, this identifies the local FILER installation, FILER metadata, Giggle executable, and supporting command-line tools. -
--outputDir— directory in which the query results and working files will be created. -
--genomeBuild— genome build to search, for examplehg38orhg19. This is required when--giggleIndexDirListis not supplied. -
--giggleIndexDirList— optional text file containing the absolute paths of specific Giggle index directories to search, one directory per line. This can be used to restrict a query to selected portions of a FILER installation. -
--verboseSearch— when set to1, reports individual overlapping records in addition to the track information. Verbose searches generate more detailed output and may be slower. -
--forceOverwrite— when set to1, allows reuse of an existing output directory. -
--tempDir— optional temporary directory used during sorting of the combined results. The default is/tmp. For large searches, this can be changed to a location with more available temporary storage.
Querying your own genomic intervals
To query your own regions, replace the value supplied to
--inBed with your BED file. The genomic coordinates must
correspond to the genome build being searched.
bash FILER2_scripts/data_querying/get_overlaps.sh \
--inBed my_regions.bed.gz \
--configFile FILER2_scripts/data/filer.homebrew.ini \
--outputDir my_filer_overlaps/ \
--genomeBuild hg38 \
--verboseSearch 1
For an hg19 query, use hg19 coordinates and specify
--genomeBuild hg19. The corresponding hg19 FILER data and
Giggle indexes must be present in the local installation.
Query output
The script produces two combined overlap files in the top-level output directory:
-
filer_overlaps.bed— combined Giggle overlap results from all indexes searched. -
filer_overlaps.with_meta.bed— the combined overlap results with the corresponding FILER track metadata appended to each result. This is generally the most useful output for downstream analysis because each overlap can be interpreted in the context of its FILER data source, assay, classification, and other track annotations.
The columns at the beginning of each result correspond to the original
fields in the input BED file. If the input BED begins with a
#-prefixed header, those field names are retained. Otherwise,
the query columns are named inputField1,
inputField2, and so on.
With the default non-verbose search
(--verboseSearch 0), the overlap-specific columns are:
trackFile— FILER track file containing the overlap;trackNumIntervals— number of intervals represented for the track in the Giggle result;numOverlaps— number of overlaps with the query interval.
With --verboseSearch 1, individual overlapping genomic records
are reported and the overlap-specific columns are:
trackFile— FILER track file containing the overlapping record;-
hitString— the individual overlapping record returned by Giggle. Tabs within this record are represented as@@@so that the record can be stored within a single output column; numOverlaps— set to1for each individual reported hit.
In filer_overlaps.with_meta.bed, the complete FILER metadata
columns are appended after these overlap columns.
Results grouped by genomic feature type
In addition to the combined result files, the script uses the FILER metadata
field classification to divide the metadata-enriched overlaps
into separate files by genomic feature type.
These files are written under:
overlaps_by_feature_type/
with filenames of the form:
filer_overlaps.<classification>.bed
For example, if the query overlaps tracks belonging to several FILER
classifications, a separate result file is created for each classification.
Each of these files contains the same metadata-enriched columns as
filer_overlaps.with_meta.bed, but only for tracks belonging to
that feature type.
Additional files in the output directory
The script retains several files used to perform and organize the search:
-
input.bed.gz— coordinate-sorted, bgzip-compressed copy of the input intervals used for the Giggle searches. -
giggle_index_dirs_for_search.<genomeBuild>.txt— automatically generated list of Giggle index directories when--giggleIndexDirListis not supplied. -
overlaps/— per-index Giggle search output. Its directory structure mirrors the corresponding organization of the local FILER installation. -
overlaps/.../giggle_out.txt— raw/intermediate results produced for each Giggle index searched.
The top-level filer_overlaps.with_meta.bed and the files in
overlaps_by_feature_type/ are generally the most useful results
for subsequent analysis.
Restricting a search to selected FILER indexes
By default, specifying --genomeBuild searches all matching
Giggle indexes found under the local FILER installation. For more targeted
searches, a list of index directories can instead be supplied using
--giggleIndexDirList.
The list must contain one absolute Giggle index directory per line:
/path/to/FILER/.../giggle_index
/path/to/FILER/.../giggle_index
/path/to/FILER/.../giggle_index
When an explicit index list is supplied, only those indexes are queried and
--genomeBuild is not required.