Functional enrichment analysis determines if specific biological functions or pathways are overrepresented in a set of features (e.g., genes, proteins). These sets often originate from differential expression analysis, while the functions and pathways are derived from databases such as GO, Reactome, or KEGG. In its simplest form, enrichment analysis employs Fisher’s test to evaluate if a given function is enriched in the selection. The null hypothesis asserts that the proportion of features annotated with that function is the same between selected and non-selected features.
Functional enrichment analysis requires downloading large datasets from the aforementioned databases before conducting the actual analysis. While downloading data is time-consuming, Fisher’s test can be performed rapidly. This package aims to separate these two steps, enabling fast enrichment analysis for various feature selections using a given database. It is specifically designed for interactive applications like Shiny. A small Shiny app, included in the package, demonstrates the usage of fenr.
Functional enrichment analysis should not be considered the ultimate answer in understanding biological systems. In many instances, it may not provide clear insights into biology. Specifically, when arbitrary groups of genes are selected, enrichment analysis only reveals the statistical overrepresentation of a functional term within the selection, which may not directly correspond to biological relevance. This package serves as a tool for data exploration; any conclusions drawn about biology require independent validation and further investigation.
fenr can be installed from Bioconductor by using:
The input data for fenr consists of two tibbles or data
frames (see figure below). One tibble contains functional term
descriptions, while the other provides the term-to-feature mapping.
Users can acquire these datasets using the convenient helper functions
like fetch_*, which facilitate data retrieval from GO,
KEGG, BioPlanet, or WikiPathways. Alternatively, users have the option
to provide their own datasets. These tibbles are subsequently
transformed into a fenr_terms object, specifically designed for
efficient data retrieval and access. The resulting object is then
employed by the functional_enrichment function to perform enrichment
analysis for any chosen feature selection.
Overview of fenr data pipeline.
Package fenr and example data are loaded with the following commands:
The initial step involves downloading functional term data.
fenr supports data downloads from Gene Ontology,
Reactome, KEGG, BioPlanet, and
WikiPathways. Custom ontologies can also be used, provided they
are converted into an appropriate format (refer to the
prepare_for_enrichment function for more information). The
command below downloads functional terms and gene mapping from Gene
Ontology (GO):
YEAST-mod is a designation for yeast in Gene Ontology.
The full list of available species designations can be obtain using the
command:
which returns a tibble:
go_species
#> # A tibble: 171 × 7
#> designation species common_name tax_id taxonomic_group group
#> <glue> <chr> <chr> <chr> <chr> <chr>
#> 1 STRCO-uniprot Streptomyces … Streptomyc… 100226 NCBITaxon:2 UniP…
#> 2 MOUSE-mod Mus musculus Mouse 10090 NCBITaxon:7742 MGI
#> 3 RAT-mod Rattus norveg… Rat 10116 NCBITaxon:7742 RGD
#> 4 TRIAD-uniprot Trichoplax ad… Trichoplax… 10228 NCBITaxon:6072 UniP…
#> 5 VACCW-uniprot Vaccinia viru… <NA> 10254 NCBITaxon:10239 UniP…
#> 6 HHV11-uniprot Human herpesv… <NA> 10299 NCBITaxon:10239 UniP…
#> 7 VZVD-uniprot Varicella-zos… <NA> 10338 NCBITaxon:10239 UniP…
#> 8 HCMVA-uniprot Human cytomeg… <NA> 10360 NCBITaxon:10239 UniP…
#> 9 KLENI-uniprot Klebsormidium… Green algae 105231 NCBITaxon:33090 UniP…
#> 10 SYNY3-uniprot Synechocystis… Synechocys… 11117… NCBITaxon:2 UniP…
#> # ℹ 161 more rows
#> # ℹ 1 more variable: annotation_source <chr>go is a list with two tibbles. The first tibble contains
term information:
go$terms
#> # A tibble: 51,975 × 3
#> term_id term_name namespace
#> <chr> <chr> <chr>
#> 1 GO:0000001 mitochondrion inheritance biologic…
#> 2 GO:0000002 obsolete mitochondrial genome maintenance biologic…
#> 3 GO:0000003 obsolete reproduction biologic…
#> 4 GO:0000005 obsolete ribosomal chaperone activity molecula…
#> 5 GO:0000006 high-affinity zinc transmembrane transporter ac… molecula…
#> 6 GO:0000007 low-affinity zinc ion transmembrane transporter… molecula…
#> 7 GO:0000008 obsolete thioredoxin molecula…
#> 8 GO:0000009 alpha-1,6-mannosyltransferase activity molecula…
#> 9 GO:0000010 heptaprenyl diphosphate synthase activity molecula…
#> 10 GO:0000011 vacuole inheritance biologic…
#> # ℹ 51,965 more rowsThe second tibble contains gene-term mapping:
go$mapping
#> # A tibble: 100,465 × 5
#> gene_symbol gene_id db_id term_id evidence
#> <chr> <chr> <chr> <chr> <chr>
#> 1 TFC3 FUN24 S000000001 GO:0001002 IDA
#> 2 TFC3 FUN24 S000000001 GO:0001003 IDA
#> 3 TFC3 FUN24 S000000001 GO:0008301 IDA
#> 4 TFC3 FUN24 S000000001 GO:0000995 IBA
#> 5 TFC3 FUN24 S000000001 GO:0003677 IEA
#> 6 TFC3 FUN24 S000000001 GO:0006383 IDA
#> 7 TFC3 FUN24 S000000001 GO:0006384 IEA
#> 8 TFC3 FUN24 S000000001 GO:0006384 IBA
#> 9 TFC3 FUN24 S000000001 GO:0006384 NAS
#> 10 TFC3 FUN24 S000000001 GO:0042791 IBA
#> # ℹ 100,455 more rowsNote that the mapping can be filtered based on the evidence
code (column evidence) to include only high-quality GO
annotations, before further analysis. Here, we simply use all
annotations.
To make these user-friendly data more suitable for rapid functional enrichment analysis, they need to be converted into a machine-friendly object using the following function:
exmpl_all is an example of gene background - a vector
with gene symbols related to all detections in an imaginary RNA-seq
experiment. Since different datasets use different features (gene id,
gene symbol, protein id), the column name containing features in
go$mapping needs to be specified using
feature_name = "gene_symbol". The resulting object,
go_terms, is a data structure containing all the mappings
in a quickly accessible form. From this point on, go_terms
can be employed to perform multiple functional enrichment analyses on
various gene selections.
In June 2026, Gene Ontology (GO) changed the naming convention for
species-related filenames (see
announcement). Previously used names, such as goa_human
and sgd, became HUMAN-uniprot and
YEAST-mod. This affects how GO species are referenced in
the package. To retain backward compatibility, fenr accepts
legacy names and internally translates them into the new names. The
mapping table included with the package can be accessed with the
following function:
get_go_legacy_mapping()
#> # A tibble: 28 × 2
#> legacy_name designation
#> <chr> <glue>
#> 1 cgd CANAL-mod
#> 2 cgd CANDC-mod
#> 3 cgd DEBHA-uniprot
#> 4 cgd LODEL-uniprot
#> 5 cgd PICGU-uniprot
#> 6 dictybase DICDI-mod
#> 7 ecocyc ECOLI-mod
#> 8 fb DROME-mod
#> 9 japonicusdb SCHJY-mod
#> 10 mgi MOUSE-mod
#> # ℹ 18 more rowsNote that some legacy names (e.g. cgd) map to multiple
new names. In such cases, using cgd will result in an error
message asking the user to choose one new name.
The package includes two pre-defined gene sets.
exmpl_all contains all background gene symbols, while
exmpl_sel comprises the genes of interest. To perform
functional enrichment analysis on the selected genes, you can use the
following single, efficient function call:
The result of functional_enrichment is a tibble with
enrichment results.
enr |>
head(10)
#> # A tibble: 10 × 10
#> term_id term_name N_with n_with_sel n_expect enrichment odds_ratio
#> <chr> <chr> <int> <int> <dbl> <dbl> <dbl>
#> 1 GO:0031931 TORC1 co… 5 5 0.02 333 Inf
#> 2 GO:1905356 regulati… 2 2 0.01 333 Inf
#> 3 GO:0038201 TOR comp… 2 2 0.01 333 Inf
#> 4 GO:0031929 TOR sign… 29 14 0.09 161 927
#> 5 GO:0007584 response… 8 5 0.02 208 725
#> 6 GO:0001558 regulati… 15 7 0.05 155 435
#> 7 GO:0038203 TORC2 si… 6 3 0.02 166 387
#> 8 GO:0038202 TORC1 si… 4 2 0.01 166 366
#> 9 GO:0031932 TORC2 co… 7 3 0.02 143 290
#> 10 GO:0016242 negative… 9 3 0.03 111 193
#> # ℹ 3 more variables: ids <chr>, p_value <dbl>, p_adjust <dbl>The columns are as follows
N_with: The number of features (genes) associated with
this term in the background of all genes.n_with_sel: The number of features associated with this
term in the selection.n_expect: The expected number of features associated
with this term under the null hypothesis (terms are randomly
distributed).enrichment: The ratio of observed to expected.odds_ratio: The effect size, represented by the odds
ratio from the contingency table.ids: The identifiers of features with the term in the
selection.p_value: The raw p-value from the hypergeometric
distribution.p_adjust: The p-value adjusted for multiple tests using
the Benjamini-Hochberg approach.A small Shiny app is included in the package to demonstrate the usage of fenr in an interactive environment. All time-consuming data loading and preparation tasks are performed before the app is launched.
yeast_de is the result of differential expression (using
edgeR) on a subset of 6+6 replicates from Gierlinski
et al. (2015).
The function fetch_terms_for_example uses
fetch_* functions from fenr to download and
process data from GO, Reactome and KEGG. The
step-by-step process of data acquisition can be examined by viewing the
function by typing fetch_terms_for_example in the R
console. The object term_data is a named list of
fenr_terms objects, one for each ontology.
After completing the slow tasks, you can start the Shiny app by running:
To quickly see how fenr works an example can be loaded directly from GitHub:
shiny::runGitHub("bartongroup/fenr-shiny-example")
There’s no shortage of R packages dedicated to functional enrichment analyses, including topGO, GOstats, clusterProfiler, and GOfuncR. Many of these packages provide comprehensive analyses tailored to a single gene selection. While they offer detailed insights, their extensive computational requirements often render them slower.
To illustrate, let’s take the example of the topGO package.
suppressPackageStartupMessages({
library(topGO)
})
#>
#> groupGOTerms: GOBPTerm, GOMFTerm, GOCCTerm environments built.
# Preparing the gene set
all_genes <- setNames(rep(0, length(exmpl_all)), exmpl_all)
all_genes[exmpl_all %in% exmpl_sel] <- 1
# Define the gene selection function
gene_selection <- function(genes) {
genes > 0
}
# Mapping genes to GO, we use the go2gene downloaded above and convert it to a
# list
go2gene <- go$mapping |>
dplyr::select(gene_symbol, term_id) |>
dplyr::distinct() |>
dplyr::group_by(term_id) |>
dplyr::summarise(ids = list(gene_symbol)) |>
tibble::deframe()
# Setting up the GO data for analysis
go_data <- new(
"topGOdata",
ontology = "BP",
allGenes = all_genes,
geneSel = gene_selection,
annot = annFUN.GO2genes,
GO2genes = go2gene
)
#>
#> Building most specific GOs .....
#> ( 3000 GO terms found. )
#>
#> Build GO DAG topology ..........
#> ( 4580 GO terms and 9670 relations. )
#>
#> Annotating nodes ...............
#> ( 6792 genes annotated to the GO terms. )
# Performing the enrichment test
fisher_test <- runTest(go_data, algorithm = "classic", statistic = "fisher")
#>
#> -- Classic Algorithm --
#>
#> the algorithm is scoring 359 nontrivial nodes
#> parameters:
#> test statistic: fisherThough the enrichment test (runTest) is fairly fast,
completing in about 200 ms on an Intel Core i7 processor, the
preparatory step (new call) demands a more substantial time
commitment, around 6 seconds. Considering each additional gene selection
requires a similar amount of time, conducting exploratory interactive
analyses on multiple gene sets becomes notably cumbersome.
fenr separates preparatory steps and gene selection and is designed with interactivity and speed in mind, particularly when users wish to execute rapid, multiple selection tests.
# Setting up data for enrichment
go <- fetch_go(species = "YEAST-mod")
go_terms <- prepare_for_enrichment(go$terms, go$mapping, exmpl_all,
feature_name = "gene_symbol")
# Executing the enrichment test
enr <- functional_enrichment(exmpl_all, exmpl_sel, go_terms)While data preparation with fenr might take a bit longer (around 10 s), the enrichment test is fast, finishing in just about 50 ms. Each new gene selection takes a similar amount of time (with an obvious caveat that a larger gene selection would take more time to calculate).
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 26.04 LTS
#>
#> Matrix products: default
#> BLAS: /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.32.so; LAPACK version 3.12.0
#>
#> locale:
#> [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
#> [3] LC_TIME=en_US.UTF-8 LC_COLLATE=en_US.UTF-8
#> [5] LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8
#> [7] LC_PAPER=en_US.UTF-8 LC_NAME=C
#> [9] LC_ADDRESS=C LC_TELEPHONE=C
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: Etc/UTC
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats4 stats graphics grDevices utils datasets
#> [7] methods base
#>
#> other attached packages:
#> [1] topGO_2.64.0 SparseM_1.84-2 GO.db_3.23.1
#> [4] AnnotationDbi_1.74.0 IRanges_2.46.0 S4Vectors_0.50.1
#> [7] Biobase_2.72.0 graph_1.90.0 BiocGenerics_0.58.1
#> [10] generics_0.1.4 tibble_3.3.1 fenr_1.10.2
#> [13] BiocStyle_2.40.0
#>
#> loaded via a namespace (and not attached):
#> [1] sass_0.4.10 utf8_1.2.6 lattice_0.22-9
#> [4] RSQLite_3.53.3 digest_0.6.39 magrittr_2.0.5
#> [7] grid_4.6.1 evaluate_1.0.5 fastmap_1.2.0
#> [10] blob_1.3.0 jsonlite_2.0.0 DBI_1.3.0
#> [13] BiocManager_1.30.27 httr_1.4.8 purrr_1.2.2
#> [16] Biostrings_2.80.1 jquerylib_0.1.4 cli_3.6.6
#> [19] crayon_1.5.3 rlang_1.3.0 XVector_0.52.0
#> [22] bit64_4.8.2 withr_3.0.3 cachem_1.1.0
#> [25] yaml_2.3.12 otel_0.2.0 tools_4.6.1
#> [28] memoise_2.0.1 dplyr_1.2.1 assertthat_0.2.1
#> [31] buildtools_1.0.0 vctrs_0.7.3 R6_2.6.1
#> [34] png_0.1-9 matrixStats_1.5.0 lifecycle_1.0.5
#> [37] Seqinfo_1.2.0 KEGGREST_1.52.2 bit_4.6.0
#> [40] pkgconfig_2.0.3 pillar_1.11.1 bslib_0.11.0
#> [43] glue_1.8.1 xfun_0.60 tidyselect_1.2.1
#> [46] sys_3.4.3 knitr_1.51 htmltools_0.5.9
#> [49] rmarkdown_2.31 maketools_1.3.2 compiler_4.6.1