38Recipes: leading-edge analysis and non-model organisms
This chapter collects two practical recipes that belong to the main workflow: extracting the genes that drive a GSEA signal, and building custom enrichment inputs for organisms without standard annotation packages.
38.1 Leading Edge Analysis
Leading edge analysis is a powerful feature in GSEA that identifies the core set of genes driving the enrichment signal. It reports three key metrics:
Tags: Percentage of genes contributing to the enrichment score
List: Position in the ranked list where the enrichment score is attained
Signal: Strength of the enrichment signal
DOSE, clusterProfiler, and ReactomePA all support leading edge analysis and can report the core enriched genes that contribute to the enrichment.
38.1.1 Core Enriched Genes Extraction
After performing GSEA, the results object contains a core_enrichment column that lists the core genes responsible for each enriched term:
library(DOSE)
DOSE v4.7.3 Learn more at https://yulab-smu.top/contribution-knowledge-mining/
Please cite:
Guangchuang Yu, Li-Gen Wang, Guang-Rong Yan, Qing-Yu He. DOSE: an
R/Bioconductor package for Disease Ontology Semantic and Enrichment
analysis. Bioinformatics. 2015, 31(4):608-609
The output includes the core enriched genes in Entrez ID format for each significant term.
38.1.2 Enhancing Readability with setReadable
To make the results more interpretable, use setReadable() to convert Entrez IDs to gene symbols:
library(clusterProfiler)
clusterProfiler v4.21.3 Learn more at https://yulab-smu.top/contribution-knowledge-mining/
Please cite:
Guangchuang Yu, Li-Gen Wang, Yanyan Han and Qing-Yu He.
clusterProfiler: an R package for comparing biological themes among
gene clusters. OMICS: A Journal of Integrative Biology. 2012,
16(5):284-287
Attaching package: 'clusterProfiler'
The following object is masked from 'package:stats':
filter
This transformation makes the core enrichment results much more readable and biologically meaningful.
For visualization of leading edge analysis results using cnetplot, please refer to the enrichplot chapter.
38.2 Non-Model Plant Annotation with clusterProfiler
For non-model plants and other organisms lacking standard annotation packages, clusterProfiler can be used with custom annotation data obtained from tools like eggNOG.
38.2.1 Workflow Overview
Annotation with eggNOG: Use the eggNOG web server to annotate protein sequences
Parse eggNOG Results: Extract GO and KEGG annotations using custom scripts
Enrichment Analysis: Use clusterProfiler’s enricher() function with custom annotation data
38.2.2 Key Steps
38.2.2.1 1. eggNOG Annotation
Upload protein sequences to the eggNOG mapper with appropriate parameters for your organism.
# Parse GO ontology filepython parse_go_obofile.py -i go-basic.obo -o go.tb# Parse eggNOG annotations with reference species filteringpython parse_eggNOG.py -i panax_ginseng.annotations \-g go.tb \-O ath,osa \-o output_directory
This generates two key files: - GOannotation.tsv: GO term annotations - KOannotation.tsv: KEGG pathway annotations
38.2.2.3 3. Enrichment Analysis with clusterProfiler
library(clusterProfiler)# Read annotation filesKOannotation <-read.delim("KOannotation.tsv", stringsAsFactors=FALSE)GOannotation <-read.delim("GOannotation.tsv", stringsAsFactors=FALSE)GOinfo <-read.delim("go.tb", stringsAsFactors=FALSE)# Your gene listgene_list <-c("gene1", "gene2", "gene3") # Replace with your actual gene list# GO enrichment (Molecular Function as example)GOannotation_split <-split(GOannotation, GOannotation$level)enricher(gene_list,TERM2GENE = GOannotation_split[['molecular_function']][c(2,1)],TERM2NAME = GOinfo[1:2])# KEGG enrichmentenricher(gene_list,TERM2GENE = KOannotation[c(3,1)],TERM2NAME = KOannotation[c(3,4)])
38.2.3 Advantages
Works for any organism with protein sequences
Uses reliable eggNOG annotation pipeline
Flexible reference species filtering for KEGG
38.2.4 Considerations
Requires intermediate Python scripting
Performance may vary with dataset size
Manual integration of annotation and analysis steps
This approach enables comprehensive functional enrichment analysis for non-model organisms using clusterProfiler’s powerful enrichment capabilities combined with custom annotation data.