40  Package and capability map

This page is a navigation map for the biomedical knowledge-mining stack used in this book. It is organized four ways: by task, by analytical layer, by package, and by typical input/output. The links point to the existing chapters and anchors so that this page can serve as a cross-reference when it is added to the book navigation.

This map separates statistical computation, knowledge-base access, result handling, visualization, and optional language-model interpretation. The layers are complementary: an interpretation report does not replace an enrichment test, and a plot does not recreate the information that was lost before an enrichment result object was constructed.

40.1 1. Index by task

Use this index when the starting point is a biological question or an analysis artifact rather than a package name.

Task Start here Main capability Typical next step
Run a first complete analysis first-workflow.qmd Build a local universe, run GO ORA and GSEA, and retain the result objects for plotting Replace the demonstration inputs with an experiment-specific data contract
Validate the input contract data-contract.qmd Separate gene vectors, ranked lists, universes, term-to-gene mappings, and result classes Choose ORA, GSEA, comparison, or custom enrichment
Choose an analysis route choose-a-route.qmd Match the evidence shape to the null question, function, and result class Connect a biological knowledge source
Prepare identifiers and backgrounds identifier-utilities.qmd Convert identifiers with bitr()/bitr_kegg(), make IDs character-valued, define the universe, and make output readable with setReadable() Run a domain-specific or universal enrichment workflow
Measure similarity between GO terms, genes, or clusters go-semantic-similarity.qmd Build GOSemSim semantic data and calculate term-, gene-, or cluster-level similarity Cluster functionally similar entities or simplify redundant enrichment terms
Measure disease-term similarity do-semantic-similarity.qmd Compare Disease Ontology terms, genes, and gene clusters Continue to disease-enrichment.qmd for disease enrichment
Measure MeSH similarity mesh-semantic-similarity.qmd Prepare a species-specific MeSHDb and compare MeSH terms or genes Continue to mesh-enrichment.qmd
Test a thresholded gene list enrichment-foundations.qmd ORA / one-sided hypergeometric testing against a gene-set collection Inspect enrichResult with enrichplot
Test a complete ranked list enrichment-foundations.qmd GSEA using the direction and magnitude of a ranking Inspect leading-edge genes and GSEA plots
Account for network topology enrichment-network.qmd NSEA, signed NSEA, multi-layer NSEA, and weighted enrichment through enrichit Use nseGO()/nseKEGG() or explanation-ready contribution tables
Integrate multiple omics layers enrichment-multiomics.qmd Early feature-level aggregation, ID harmonization, late pathway-level fusion, and contribution tracing Feed the unified ranking or selected features to enrichment
Analyse GO go-enrichment.qmd groupGO(), enrichGO(), gseGO(), enricher(), and GO-aware topology wrappers Simplify terms, compare groups, or visualize
Analyse KEGG pathways or modules kegg-enrichment.qmd Online KEGG pathways/modules, local GSON, offline databases, and compound analysis Use browseKEGG(), pathview(), or enrichplot
Analyse Reactome reactome-enrichment.qmd Reactome ORA (enrichPathway()) and GSEA (gsePathway()), with viewPathway() Apply general enrichment plots or compare clusters
Analyse Disease Ontology, HPO/MPO, or NCG disease-enrichment.qmd Disease and phenotype ORA/GSEA; NCG is supported by DOSE Use enrichplot or compare profiles
Analyse MeSH mesh-enrichment.qmd MeSH ORA and GSEA with selectable annotation source and category Use general enrichment plots
Analyse a custom or unsupported knowledge base universal-enrichment.qmd Supply TERM2GENE and optional TERM2NAME to enricher()/GSEA(); parse GMT with read.gmt() Store reusable sets as GSON or compare them
Freeze, share, or combine knowledge bases gson.qmd, with KEGG examples in kegg-enrichment.qmd Create, serialize, and combine GSON objects; use gsonList for multiple sources Run one analysis per source and facet the results
Compare conditions, clusters, or experiments comparecluster-analysis.qmd Re-run one enrichment function per gene list with compareCluster() or its formula interface Use comparison-aware dot, network, or map plots from comparecluster-visualization.qmd
Start from genomic coordinates chipseeker-enrichment.qmd Annotate peaks/regions and map them to nearest, host, flanking, or seq2gene()-linked genes Pass the resulting gene sets to clusterProfiler
Reduce redundancy or select explanatory terms go-enrichment.qmd and bayesian-term-selection.qmd simplify() uses semantic similarity; bayes_enrich() uses joint term-gene coverage Keep a compact, interpretable term set
Turn results into figures enrichplot.qmd, summary, networks, GSEA, comparison, specialized Bar, dot, Manhattan, network, map, heatmap, tree, semantic-space, ridge, and GSEA plots Export publication or presentation graphics
Retrieve a PPI context ppi.qmd Fetch STRING interactions for genes or enriched pathways with getPPI() Visualize with ggtangle or provide network evidence to interpretation
Manipulate or import result objects identifier-utilities.qmd, result-objects.qmd, and enrichplot-import.qmd Filter, select, mutate, preserve object classes, or import external result tables Plot, compare, or interpret canonical objects
Request an optional narrative or annotation interpretation.qmd interpret(), interpret_agent(), and hierarchical interpretation using an external LLM Validate every claim against the supplied results before reporting

40.2 2. Index by analytical layer

The packages are easier to combine when their responsibilities are viewed as a pipeline. A package can appear in more than one layer, but each layer has a different contract.

Layer Question answered Principal components Input contract Output contract
1. Experimental feature and ID layer What are the features, and how are they named? Differential-expression lists, ranked statistics, peak coordinates, bitr(), bitr_kegg(), setReadable() Character IDs, or genomic ranges plus annotation context; a ranked list must be a named numeric vector sorted in descending order Harmonized IDs, gene lists, ranked vectors, or genes linked from regions
2. Knowledge and annotation layer What do terms and gene sets mean? OrgDb/GO, KEGG, Reactome, DO/HPO/MPO, MeSH, WikiPathways, MSigDB, custom TERM2GENE/TERM2NAME, and GSON A knowledge source, organism/key type, and optionally a versioned local copy Term-to-gene mappings, term names, ontology graphs, GSON, GOSemSimDATA, or MeSH semantic data
3. Similarity layer How close are terms, genes, or clusters in a knowledge graph? GOSemSim, DOSE, meshes; Resnik, Lin, Rel, Jiang, Wang, and combination methods such as BMA Ontology/MeSH terms or IDs plus prepared semantic data Numeric similarity scores, matrices, and term/group relationships
4. Enrichment engine layer Are selected or ranked features associated with gene sets more than expected? enrichit (ORA, GSEA, NSEA/MNSEA, weighted enrichment, multi-omics aggregation, Bayesian selection) Thresholded genes plus universe, or a ranked vector; network/layer/weight inputs when requested Canonical enrichment result objects, p-values/FDR/NES, leading-edge/core genes, contribution tables, and selected terms
5. Knowledge-base-aware API layer Which biological vocabulary should the engine test? clusterProfiler, DOSE, ReactomePA, meshes, and gson helpers Domain-specific IDs and source parameters, or a universal mapping enrichResult, gseaResult, compareClusterResult, lists of results, or GSON-backed analyses
6. Genomic-coordinate adapter layer Which genes may be regulated by these regions? ChIPseeker and seq2gene() Peaks or genomic regions plus transcript/gene annotation Peak annotation tables and many-to-many region-to-gene mappings suitable for enrichment
7. Comparison and result-shaping layer How do profiles differ, and which terms should be retained? compareCluster(), simplify(), bayes_enrich(), dplyr/data-frame interfaces Gene lists grouped by condition, a result object, or a canonical external table Comparable multi-group objects, reduced term sets, posterior columns, and filtered objects
8. Visualization layer How can the result be inspected or communicated? enrichplot, ggplot2, ggtangle, plus browseKEGG(), pathview(), viewPathway() Canonical enrichment objects with mappings and (for GSEA) ranked lists Plots, networks, maps, and plot-ready tables
9. Optional interpretation layer What biological narrative or label is supported by the result? interpret(), interpret_agent(), interpret_hierarchical() plus aisdk One or more canonical result objects, optional context, fold changes, PPI, and an API-backed model Structured narrative, drivers, mechanisms, confidence, annotation labels, or a report plot; not a replacement for statistics

A useful mental model is:

features/regions -> IDs and background -> knowledge mapping ->
statistical enrichment -> canonical result object -> reduction/comparison ->
visualization -> optional, validated interpretation

40.3 3. Index by package or capability

Package or capability Primary responsibility Representative entry points Does it compute enrichment? Main hand-off
clusterProfiler User-facing orchestration across biological knowledge bases; ID conversion; universal enrichment; comparison; result methods; high-level network-aware wrappers enrichGO(), gseGO(), enrichKEGG(), gseKEGG(), enricher(), GSEA(), compareCluster(), bitr(), nseGO(), nseKEGG() Yes, through domain APIs and delegated engines Produces/consumes canonical enrichment objects; hands them to enrichplot, dplyr, PPI, or interpretation
enrichit Algorithm and result-object engine with minimal dependencies and fast implementations ora(), gsea(), nsea(), mnsea(), weighted enrichment, aggregate_omics(), aggregate_enrichment(), harmonize_ids(), bayes_enrich(), as_enrichResult() Yes; it is the computational backbone for classical, topology-aware, weighted, and multi-omics workflows Returns engine results or canonical objects that clusterProfiler and enrichplot can use
gson Versionable Gene Set Object Notation container and multi-knowledge-base exchange format gson(), gsonList(), write.gson(), read.gson(); ecosystem helpers such as gson_KEGG(), gson_WP(), gson_GO() No; it stores term/gene/name/metadata mappings Supplies local or combined gene sets to enricher(), GSEA(), nseKEGG(), or related APIs
DOSE Disease- and phenotype-oriented semantic similarity and enrichment; includes NCG workflows and demo data doSim(), geneSim(), clusterSim(), enrichDO(), gseDO(), enrichNCG(), gseNCG() Yes, for DO/HPO/MPO and NCG; it also provides similarity Produces disease-aware result objects consumed by enrichplot, comparison, or optional interpretation
ReactomePA Reactome pathway domain adapter enrichPathway(), gsePathway(), viewPathway() Yes, Reactome ORA and GSEA Produces Reactome result objects for enrichplot, comparison, and reporting
meshes MeSH annotation, semantic similarity, and enrichment across many species and MeSH categories meshdata(), meshSim(), geneSim(), enrichMeSH(), gseMeSH() Yes, MeSH ORA and GSEA Produces MeSH semantic data and enrichment objects for enrichplot
GOSemSim GO semantic data and similarity for terms, genes, and gene clusters godata(), goSim(), mgoSim(), geneSim(), mgeneSim(), clusterSim(), mclusterSim() No; its main output is semantic similarity (it supports downstream simplification) Supplies similarity to clustering, redundancy reduction, and interpretation of GO results
enrichplot Visualization and import boundary for enrichment objects barplot(), dotplot(), cnetplot(), emapplot(), heatplot(), treeplot(), ssplot(), gseaplot2(), ridgeplot(), import_*() No; it visualizes or imports results Consumes canonical objects and produces ggplot/network graphics; importers normalize external tables
ChIPseeker Genomic-region annotation and region-to-gene mapping before functional analysis Peak annotation, nearest/host/flanking genes, seq2gene() No; it prepares the gene universe for downstream enrichment Hands gene IDs or mappings to clusterProfiler/enrichit
Optional AI interpretation External, evidence-conditioned narrative, annotation, or phenotype labeling after analysis interpret(), interpret_agent(), interpret_hierarchical() via aisdk No; it does not replace ORA/GSEA/NSEA or alter their p-values Consumes one or more result objects plus context/optional PPI/fold change and returns a structured report

40.3.1 Package boundaries that matter

  • clusterProfiler is the knowledge-aware front door; enrichit is the algorithm engine. The nse*() and mnse*() wrappers preserve ontology and pathway semantics while delegating topology-aware computation.
  • gson is a data contract, not a new statistical test. It is useful when a remote source should be frozen, reused offline, or combined with another knowledge base.
  • GOSemSim, DOSE, and meshes provide semantic similarity in different vocabularies (GO, DO/phenotypes, and MeSH). Similarity can be used before enrichment or after enrichment for term reduction.
  • enrichplot expects the mappings carried by enrichResult, gseaResult, or compareClusterResult; a bare table from another program must first be imported or converted. See the external-result import section.
  • ChIPseeker solves the region-to-gene problem. It does not decide which ontology or enrichment test should be applied to the resulting genes.
  • AI interpretation is an optional, post-analysis consumer. It needs an API-backed model and should be checked against the input terms, genes, ranked statistics, and any external PPI evidence before use.

40.4 5. Index by typical input/output

The table below is a compact routing guide. “Output” describes the object that should be preserved for the next step, rather than only the printed table.

Starting input Typical question Route Expected output Useful cross-link
Character vector of IDs + character universe Are selected genes over-represented? clusterProfiler::enrichGO(), enrichKEGG(), DOSE::enrichDO(), ReactomePA::enrichPathway(), meshes::enrichMeSH(), or enricher() enrichResult with term IDs, descriptions, ratios, p-values, adjusted p-values, and gene mappings ORA algorithm
Named, numeric, decreasing vector Is a gene set shifted toward one end of the ranking? gseGO(), gseKEGG(), gseDO(), gsePathway(), GSEA(), or gseMeSH() gseaResult with ES/NES, significance, ranked list, and leading-edge/core genes GSEA algorithm
Ranked vector + edge list network Does network topology reinforce enrichment? enrichit::nsea() or clusterProfiler::nseGO()/nseKEGG() Network-aware enrichment result with propagated ranking and inferential statistics NSEA section
Named layer-specific vectors + networks + couplings Which pathways are supported across RNA/protein or other network layers? mnsea() or mnseGO()/mnseKEGG() Multi-layer result plus pathway/feature contribution tables and optional subnetwork Multi-layer example
Several p-value or signed-score vectors Can weak evidence be combined before enrichment? aggregate_omics(); then GSEA/NSEA or select_features_for_ora() + ORA Unified feature ranking, selected ORA genes/universe, or an integrated result Multi-omics integration
Multiple enrichment result objects Which layer or assay drives a pathway? aggregate_enrichment() and contribution helpers Pathway-level combined result and source/contribution annotations Late fusion
GO term vectors or gene IDs + OrgDb How similar are functions? GOSemSim::godata() then goSim(), mgoSim(), geneSim(), or clusterSim() Numeric score or similarity matrix; GOSemSimDATA stores semantic data GO semantic similarity
DO term vectors or Entrez IDs How similar are disease associations? DOSE::doSim(), geneSim(), or clusterSim() DO similarity score/matrix or gene/cluster similarity DO semantic similarity
MeSH IDs or Entrez IDs + MeSHDb How similar are biomedical concepts? meshes::meshdata() then meshSim()/geneSim() MeSH semantic data and similarity scores/matrices MeSH semantic similarity
TERM2GENE plus optional TERM2NAME Can a custom annotation be tested? enricher() for ORA or GSEA() for a ranking Canonical enrichment result with custom term IDs/names Universal enrichment
GMT file Can an external gene-set library be used directly? read.gmt() or read.gmt.wp() then enricher()/GSEA() Tidy term-to-gene table and canonical result GMT utilities
GSON object or serialized .gson file Can knowledge-base data be frozen and reused? read.gson() then enricher()/GSEA(); use gsonList() for several sources Versionable GSON or a list of per-source enrichment results GSON workflow
Named list of gene vectors Which functions differ among clusters/conditions? compareCluster(geneCluster = ..., fun = ...) compareClusterResult retaining cluster labels and per-group enrichment Biological theme comparison
One gene-per-row table with grouping columns Can a formula describe a factorial comparison? compareCluster(gene ~ group + condition, data = ..., fun = ...) Formula-based compareClusterResult Formula interface
Peak/region file plus genome annotation Which genes may be regulated by the regions? ChIPseeker annotation and seq2gene() Annotated regions plus many-to-many gene mappings Genomic coordination
enrichResult with redundant GO terms Which terms summarize distinct functions? pairwise_termsim() + simplify(); similarity supplied by GOSemSim Reduced enrichment object preserving the selected representative terms GO term reduction
enrichResult with overlapping candidate terms Which terms jointly explain observed genes? bayes_enrich() then bayes_summary() Enrichment object with posterior, coverage, rank, and active-term columns Bayesian selection
Any canonical enrichment object How should terms and genes be shown? enrichplot (dotplot, barplot, cnetplot, emapplot, treeplot, gseaplot2, etc.) ggplot objects, networks, maps, and plot-ready tables Visualization
External result table Can an output from another tool be plotted here? import_enrichr(), import_gprofiler2(), import_webgestalt(), import_fgsea(), as_enrichResult(), or as_gseaResult() Canonical enrichResult/gseaResult with enough mappings for plotting Importing results
Gene list or enrichment object What interactions connect the genes or terms? clusterProfiler::getPPI() and ggtangle STRING-backed edge/node object and network plot PPI workflow
Enrichment result(s), optional context and fold changes What mechanism, cell type, or phenotype is supported? Optional interpret()/interpret_agent() through aisdk Structured narrative, drivers, evidence, confidence, and possibly an inferred network AI interpretation

40.5 4. Reference routes

Use these pages for the details that do not belong in a task table:

Need Maintained reference
Result classes, identifier contracts, and dispatch Data contract and result-object manipulation
ORA, GSEA, NSEA, and multi-omics methods Enrichment foundations, network-aware enrichment, and multi-omics integration
GSON schema, serialization, metadata, and gsonList GSON chapter
Semantic data and similarity methods GOSemSim, DO similarity, and MeSH similarity
AI validation and claim checking How to read an enrichment result, evidence-guided interpretation, and reproducible reporting
Version, citation, and provenance record Reference appendix

These pages are the maintained reference surface; the capability map routes readers to them without duplicating package-specific API documentation. ## 6. Minimal routing recipes

40.5.1 A thresholded list

IDs -> check/convert IDs -> choose universe -> ORA -> simplify or Bayesian selection
    -> enrichplot -> optional PPI or AI interpretation

Start with ID utilities, choose the relevant enrichment API, and preserve the result object for visualization.

40.5.2 A ranked list

named numeric ranking -> verify direction -> GSEA -> leading edge -> GSEA plots
    -> optional network-aware or AI interpretation

The ranking requirements are summarized in the FAQ, while the statistical choice is covered in GSEA.

40.5.3 A custom or non-model annotation

custom annotation/GMT/GOA/eggNOG -> TERM2GENE (+ TERM2NAME) or GSON
    -> enricher/GSEA -> compareCluster/enrichplot

See universal enrichment, non-model annotation notes, and the GSON section.

40.5.4 Genomic regions

peaks/regions -> ChIPseeker annotation or seq2gene() -> gene IDs
    -> domain-specific or universal enrichment -> visualization

The region-to-gene adapter is ChIPseeker; the downstream package choice belongs to the enrichment question, not to the peak caller.

40.5.5 Optional AI interpretation

validated result object(s) + explicit context + optional fold changes/PPI
    -> interpret()/interpret_agent() -> structured draft -> human evidence check

The model sees only the supplied evidence and selected terms. Check every named gene and pathway against the input object, distinguish association from causal claims, and record model/provider settings when the report is used outside exploration. See AI-assisted interpretation.