From GEO Accession to Results
Type a GSE number, pick your groups, get a volcano plot and enriched pathways. No downloading, unzipping or reformatting spreadsheets.
-
1
GEO organises every study as a Series (
GSE) made of Samples (GSM) measured on a Platform (GPL). For RNA-seq, the count table you need is usually not in the series matrix; it hides in supplementary files or in a counts table NCBI generates itself. -
2
In TransXplorer’s Transcriptome Analysis tab you type a
GSEID and click Fetch. It looks in three places in a fixed order, loads the first usable matrix, and turns the GEO sample annotations into a metadata table. - 3 You choose which annotation column defines your groups. Check that choice: the pre-selected column is a guess. Then it is the normal pipeline: QC, differential expression, volcano plot, pathway enrichment.
What GEO is, in one diagram
Series, Samples, Platforms. Three accession prefixes explain most of it.
The Gene Expression Omnibus (GEO) is NCBI’s public archive for functional genomics data. It was launched in 2000 to hold microarray experiments2 and now also holds a large share of the world’s published bulk RNA-seq studies1,3. Most journals ask authors to deposit expression data there, so the dataset behind a paper you are reading is very often one accession number away.
GEO’s records nest inside each other. A Series (GSE) is one study: its title, summary, experimental design, contributors, linked publication and any files the submitters attached. A Series contains Samples (GSM): one record per biological sample or library, with its own annotations. Each Sample points to a Platform (GPL), which describes the technology: a specific microarray design, or for sequencing, an instrument model and organism. Related Series can be bundled under a SuperSeries, and for sequencing studies the raw reads live next door in the Sequence Read Archive (SRA).
GSM is one airway smooth-muscle cell line from one donor under one treatment. The SuperSeries layer is optional and absent here. Platform and sample details are taken from the public GEO record.
| Prefix | Record | What it holds | TransXplorer accepts it? |
|---|---|---|---|
| GSE | Series | One study, its design, files, and the list of its samples. | Yes. This is the ID you type. |
| GSM | Sample | One library: title, source, characteristics, processing notes. | No (rejected as an invalid format) |
| GPL | Platform | The array design or sequencing instrument and organism. | No |
| GDS | DataSet | Older, GEO-curated subsets of Series, mostly microarray. | No; fetch its parent GSE |
Each sample’s annotations become your experimental groups. If you know that “Untreated” and “Dexamethasone” are written on the GSM records, you already know where TransXplorer will find them, and you can tell when a study was annotated too sparsely to analyse without extra work.
Reading a GEO page in 60 seconds
Six places to look before you click Fetch.
Before importing anything, open the Series page (ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE…) and scan it top to bottom. A minute here tells you whether the study fits your question, how the groups are defined, and whether a count matrix is likely to exist. ⏱ ~60 s
- Title & OrganismIs this your biology, in your species? A Series listing two organisms is a warning sign (see troubleshooting).
- Experiment type“…by high throughput sequencing” means a sequencing study (RNA-seq, but also ChIP-seq and others). “…by array” means microarray.
- Summary & Overall designThis is the experimental design written in words. Here it says four donor cell lines × four treatments, so the samples are paired by donor. Keep that in mind for the batch step.
- PlatformsCount them. One is simple. More than one means mixed instruments, assays or species, and TransXplorer will use only the largest.
- SamplesSample titles often encode the groups (
N61311_Dex). Click one to see its characteristics, thekey: valuelines TransXplorer turns into metadata columns. - Supplementary filesRead the file names.
count/rawis promising;FPKM,RPKM,TPMordiffmeans normalized values or results, not counts. An NCBI-generated counts link, where shown, is the cleanest source.
Where the counts actually live
Series matrix, supplementary files, NCBI-generated counts.
Many people expect to download “the data” from GEO as one tidy table. For microarray studies that mostly works. For RNA-seq it usually doesn’t, and the reason comes from how GEO was originally built.
Every Series has a series matrix file: a text file with all sample annotations at the top and a data table underneath. For microarrays the table holds each sample’s processed intensity per probe, because GEO historically required those values on every Sample record. Sequencing studies are deposited differently: raw reads go to SRA, and processed data are attached as supplementary files in whatever format the submitter chose. So for most RNA-seq Series the series matrix carries the annotations but an empty data table. You can see it in GSE52778: its matrix file lists all 16 sample columns and zero data rows.
The supplementary files are the submitters’ choice, and the choice varies a lot. Some upload a clean raw-count matrix. Many upload only normalized values (FPKM, RPKM, TPM), differential-expression results, one file per sample bundled in a _RAW.tar archive, or Excel sheets with merged header rows. Count-based methods such as DESeq2, edgeR and limma-voom6–8 need the raw integer counts, because they model the count noise themselves. FPKM/TPM tables have already removed the information those models need.
To fill that gap, NCBI now re-processes the raw SRA reads of many human RNA-seq Series through one uniform pipeline and offers the result as NCBI-generated raw counts: one gene-level count table per Series, aligned to the GRCh38 human reference, with GSM accessions as column names3. When this table exists it is usually the best starting point, because it is raw, complete and processed the same way for every study.
“The data are on GEO” does not mean “the counts are on GEO”. A reanalysis that starts from FPKM values and feeds them to a count model will run without complaint and give you misleading p-values. Knowing which of the three sources a matrix came from is the first quality check of any GEO reanalysis.
What happens when you click Fetch
Three sources, a fixed order, and a metadata parser.
TransXplorer’s GEO import uses the Bioconductor package GEOquery4 to talk to GEO, and then applies its own rules to pick a matrix and build a metadata table. Knowing the order of those rules explains almost every result you will see, including the failures.
The rules, in detail
-
Accepted input
^GSE[0-9]+$Spaces are trimmed and lower case is converted to upper case, sogse52778works. An empty box gets “Please enter a GEO accession ID (e.g., GSE52778)”, and anything that is not a Series ID gets “Invalid GEO ID format. Please enter a Series ID like GSE52778.” -
Multi-platform series
largest platformIf the Series spans several platforms, the expression matrix comes from the platform with the most samples. Sample annotations are merged across all platforms (using only the columns they share), and when you confirm groups, any count column without matching metadata is dropped. -
NCBI-generated counts
GRCh38.p13 fileTransXplorer requests the human GRCh38 raw-count file for the accession and checks that what comes back is a real gzip table with at least two numeric sample columns. For non-human series this lookup simply finds nothing and the next source is tried. -
Supplementary files
name-based filterRuns only if the processing notes look like sequencing (words such as RNA-seq, counts, STAR, Salmon, featureCounts). Files whose names suggest DE results or normalized values (edger,deseq,diff,rpkm,fpkm,tpm,cpm,_vs_,logfc…) are skipped. Files named likecount,raw_count,matrixorgene_expressionare preferred, and the largest one is read. A generic.txt/.csv/.tsvfile is only used when there are three or fewer candidates. -
Series matrix fallback
max < 25 ⇒ microarrayIf the matrix table has values and the largest is below 25, the data are treated as log2 microarray intensities and a note recommends limma. Otherwise they are labelled RNA-seq. -
Gene IDs
version suffix removedTrailing version numbers such asENSG00000141510.17→ENSG00000141510are stripped from row names. -
Metadata
GEO2R-style parsingEachcharacteristicsline of the formkey: valuebecomes a column named after the key. Administrative fields (contact, dates, protocols, file links) are dropped, as are columns where every sample has the same value.titleandsource_name_ch1are always kept.
The order is a quality ranking. NCBI-generated raw counts are preferred because they are uniform and raw. Submitter files come next because they vary. The series matrix table is last because, for sequencing studies, it is usually empty or normalized. When a fetch gives you something unexpected, ask which of the three sources it came from.
Walkthrough: GSE52778, from accession to pathways
Airway smooth muscle cells treated with dexamethasone.
GSE52778 comes from Himes et al. (2014)5. Primary airway smooth-muscle cells from four donors were each left untreated or treated with albuterol, dexamethasone, or both, and profiled by RNA-seq. It is a good teaching study: the design is clean, it is paired by donor, and the submitters’ files show exactly the “FPKM only” problem from Section 3. We will compare dexamethasone vs untreated, the comparison the paper focuses on.
- Accession
- GSE52778NCBI GEO · GPL11154
- Samples
- 164 donors × 4 treatments
- Comparison
- Dex vs Untreated4 vs 4, paired by donor
- Organism
- H. sapiensprimary airway smooth muscle
-
Load
Type the accession and click Fetch
Open the Transcriptome Analysis tab. In the left sidebar, under 1. Data Input & Setup, below the file uploader, is a blue box titled Or fetch from GEO. Type
GSE52778and click Fetch.UI replicaOr fetch from GEOGSE52778 ↓ FetchSupports RNA-seq datasets with NCBI-processed counts. Microarray data auto-detected.The button switches to Fetching… and a progress bar walks through the sources from Section 4. For GSE52778 the second source succeeds: NCBI has generated a raw-count table, so TransXplorer never needs the FPKM file. A notification confirms “GSE52778 loaded successfully!” with the matrix size, a data type of RNA-seq counts, and the number of metadata attributes. The sidebar then shows GEO data loaded: GSE52778; its × link clears the data.
-
Assign groups
Choose the grouping column, and don’t trust the default
A dialog titled Select Sample Grouping Column opens. It previews the first five samples’ metadata, offers a Grouping column: dropdown, and below it shows “Preview: N groups detected” with the sample count per group.
TransXplorer pre-selects a column using a simple rule: it prefers columns with 2–6 distinct, short values. For GSE52778 that rule picks
ercc_mix:ch1, a technical flag recording which libraries received ERCC spike-in controls. It has two values, so it looks like a perfect two-group design. It is not your biology. (We checked this by running TransXplorer’s selection rule on the live GSE52778 metadata in October 2026.)UI replica · simplified▦ Select Sample Grouping ColumnGSE52778 loaded: genes × 16 samples
Select the column that defines your experimental groups.Pre-selected (wrong for this study)ercc_mix:ch1▾Preview: 2 groups detected
Group Samples - 9 1 7 Grouping column: (your choice)treatment▾Preview: 4 groups detected
Group Samples Albuterol 4 Albuterol_Dexamethasone 4 Dexamethasone 4 Untreated 4 Cancel✓ Confirm GroupsSwitch the dropdown to
treatment(the parsed column;treatment:ch1holds the same values), check that the preview shows four groups of four, and click Confirm Groups. TransXplorer matches each count column to its metadata row byGSMaccession, assigns the groups, and confirms with “Groups confirmed from treatment: …”.
GSM accession, so the order of samples in the count file does not matter.
-
Annotate
Select the organism
The GEO import does not set the species for you. In the Select Organism section choose Human from Select Organism: *. The asterisk means it is required. The organism controls gene annotation and which pathway databases are offered later.
NCBI-generated tables use numeric NCBI Gene (Entrez) IDs as row names. TransXplorer’s symbol-annotation step recognises numeric Entrez IDs alongside Ensembl IDs, so with the organism set, your result tables can carry gene symbols such as TSC22D3.
-
Design
Account for the donors
Once groups are confirmed, a Batch Effects & Covariates Detection panel appears in the sidebar. Because this data came from GEO, it shows a green box, “GEO metadata available from GSE52778”, with a Use GEO Metadata for Batch Detection button. You don’t need to upload a metadata file.
Here the “batch” is biological: four donors, each contributing one sample to every group. Donor-to-donor differences are often larger than the drug effect, so they belong in the model. Select Manual Selection, click the GEO metadata button, and in the Configure Sample Metadata dialog set Group Column (Required): to
treatmentand Batch Column (Optional): tocell_line, then click ✓ Confirm Metadata Configuration. Alternatively, leave Automatic Detection on and let TransXplorer score the GEO columns for you.When a batch column is in use, the DE model becomes
~ batch + conditionfor all three packages. Because every donor appears in both groups, the donor and treatment effects are not confounded, so TransXplorer keeps the donor term. If a batch column were confounded with the groups, it would warn you and leave it out. For the reasoning behind this, see Understanding Batch Effects. -
Compare
Pick reference and treatment
With four groups, Comparison Method: offers Specific Comparison or All Pairwise Comparisons. Keep Specific Comparison and set Reference Group (Control): to Untreated and Treatment Group (Case): to Dexamethasone. Positive log2 fold changes will then mean “higher after dexamethasone”.
The other defaults are sensible starting points: DEG Analysis Package: edgeR with TMM normalization, an adjusted p-value (FDR) threshold with quick presets (0.01 / 0.05 / 0.1), and Min Log2FC Cutoff: −2 / Max Log2FC Cutoff: 2. With only four donors per group, an |log2FC| cutoff of 1 gives a broader, more typical gene list. Why the three DE packages usually agree is covered in What is Differential Expression?.
-
Run
Run QC, then differential expression
Click Run Exploratory QC first. It normalizes the counts and computes PCA/UMAP and batch diagnostics, and the results open in the Exploratory Analysis results tab. On the PCA, look for samples clustering by donor as well as by treatment; that is the pairing at work. Then click Run DEG Analysis.
If you click Run DEG Analysis before QC, an Exploratory Analysis Required dialog offers Run Exploratory Analysis First. QC only needs to run once per dataset. After that you can run several comparisons, for example Albuterol vs Untreated, without repeating it.
-
Read
Volcano plot and gene tables
Open DEG Results & Visualisation. The All DEGs, Upregulated Genes and Downregulated Genes tabs each have a download button (Download All DEGs and so on). In Volcano Plot, click Generate Volcano Plot to draw an interactive plot you can hover, zoom, and export with Download Plot.
Sanity check against the paper. Himes et al. reported 316 differentially expressed genes for dexamethasone vs untreated (Benjamini–Hochberg adjusted p < 0.05). These included well-known glucocorticoid-responsive genes, DUSP1, KLF15, PER1 and TSC22D3, plus CRISPLD2, which they showed dexamethasone induces at both mRNA and protein level5. Look these up in your table or hover for them on the volcano. Your exact counts and p-values will differ: the authors used their own hg19 alignment and a different DE method, while NCBI’s table is aligned to GRCh38. The direction and rough ranking of these genes is what should agree.
If the classic genes are missing, check three things before blaming the biology: the grouping column (step 2), the reference/treatment direction (step 5), and that the organism is set so symbols resolve (step 3). -
Interpret
Pathway enrichment
Open the Pathway Enrichment results tab. The organism is carried over from your DEG run (“Auto-detected from DEG Analysis”). Under Gene Set to Analyze: keep Both Up & Down-regulated (Recommended) or split by direction. In Select Pathway Databases: the default for human is GO Biological Process 2023; you can add KEGG, Reactome, WikiPathways, MSigDB Hallmark and others. Then click Run Pathway Enrichment. Results appear as an Enrichment Table, an Enrichment Plot and a Gene-Pathway Network.
That is over-representation analysis (ORA) on your significant gene list. For a threshold-free view, the top-level Enrichment Analysis tab has a GSEA mode that accepts a pre-ranked list. Export Download All DEGs and use gene symbol plus log2FC as the ranking. When to prefer which is the subject of GSEA vs ORA.
Troubleshooting: when a GSE doesn’t cooperate
Read the symptom, follow the branch.
GEO holds hundreds of thousands of Series deposited over more than two decades in very different formats, so no importer handles them all. The map below starts from what you actually see after clicking Fetch and leads to a fix. The numbered cards underneath give the detail for each branch.
-
“Invalid GEO ID format”, “not found”, or a network error
Only Series IDs are accepted. A
GSM(one sample),GPL(platform) orGDS(curated DataSet) ID is rejected before anything is downloaded. “GEO dataset … not found” usually means a typo or a Series that is still private. “Network error: Could not connect to GEO” is usually temporary: NCBI’s servers are sometimes busy, so wait a minute and try again. FixCopy the GSE from the paper’s data-availability statement. For a GDS, open it on GEO and use its reference Series. -
“No Count Matrix Found for GSE…”
The dialog says how many samples have metadata and shows the first few rows, but no source produced a usable matrix. Typical causes: NCBI hasn’t generated counts (or the series is not human), and the supplementary files are a per-sample
_RAW.tararchive, an Excel workbook, a normalized-only table (FPKM, RPKM, TPM, CPM), or a file whose columns cover fewer than half the samples. The dialog itself suggests three ways forward: upload your own count matrix with this metadata, process the raw FASTQ files from SRA in the FASTQ Processing tab, or use uniformly processed resources such asrecount3orDEE29,10. FixIf a raw-count file exists in an awkward format, download it, tidy it into genes × samples, and follow Bringing your own data. If only reads exist, see FASTQ to results and the FASTQ Processing tab. -
Microarray series
When the matrix comes from the series-matrix table and its maximum value is below 25, TransXplorer labels it “Microarray (log2-transformed values)” and recommends limma8. Two caveats. First, the platform annotation table is not downloaded, so rows are the array’s probe IDs (for example
1007_s_at), and gene-symbol mapping and enrichment may not work. Second, older arrays deposited on a linear (non-log) scale exceed 25 and are labelled “RNA-seq counts” even though they are not counts. FixFor a serious microarray reanalysis, map probes to genes with the platform’s annotation first, then upload the gene-level table. Set Normalization Method: deliberately; already-normalized data may need No Normalization. -
Normalized-only supplements (FPKM, RPKM, TPM)
Supplementary files whose names contain
fpkm,rpkm,tpmorcpmare skipped, so a series that offers only those ends at “No Count Matrix Found” instead of loading them as counts. A normalized table with a generic file name, or normalized values in the series matrix, can still get through, and the success notification will still say “RNA-seq counts”. Count models (DESeq2, edgeR, limma-voom) assume raw integer counts, and normalized values break their noise model. FixOpen the source file from the GEO page and look at a few values. Raw counts are whole numbers; FPKM/TPM have decimals. If you see decimals, look for NCBI-generated counts or re-quantify from SRA. -
“No matching sample IDs between count data and metadata”
Groups are attached by matching count-matrix column names to
GSMaccessions. NCBI-generated tables always use GSM IDs, but a submitter file might label columnsCtrl_1,KO_2… In that case Confirm Groups cannot link them and stops with this error. If only some columns match, the unmatched ones are dropped. FixDownload the file, rename its columns to the matchingGSMaccessions (the sample titles on the GEO page show the mapping), and upload it with a metadata file whoseSamplecolumn holds the same IDs. -
SuperSeries, multiple platforms and mixed organisms
A SuperSeries bundles several Series, often different assays (RNA-seq plus ChIP-seq, bulk plus single-cell) or species. GEO keeps platforms separate, and a platform is tied to one organism, so a human+mouse study has at least two. When a fetch returns more than one platform, TransXplorer uses the matrix from the platform with the most samples, which may not be the one you want.
FixOn the SuperSeries page, find the list of SubSeries and fetch the specific
GSEfor your assay and species. Then set Select Organism: * to match. -
The pre-selected grouping column is wrong
The default favours columns with 2–6 short distinct values, so spike-in flags, sex, lanes or batch dates can win, as
ercc_mixdid for GSE52778. The preview turns red with “need at least 2 groups” if a column has only one value. FixRead the GEO page’s Overall design first (Section 2), then choose the column that matches it. If no single column defines your groups (for example, genotype × time), build a small metadata file with aSamplecolumn of GSM IDs and a combined group column, upload it under Upload Batch Metadata File:, and pick that column as the Group Column.
In your methods, state the GSE accession, which file the counts came from (NCBI-generated, or a named supplementary file), which metadata column defined the groups, and any covariates you added. Reviewers increasingly re-run public reanalyses, and these four facts are what they need to reproduce yours.
What’s next
Other ways in, and the concepts behind each step.
Other on-ramps into TransXplorer
Go deeper on the concepts
Pick a GSE from a paper you’re reading
Open TransXplorer, go to Transcriptome Analysis, paste the accession into Or fetch from GEO, and click Fetch.
If TransXplorer helps your research, please cite us
The preprint is open access on bioRxiv. A peer-reviewed version is in submission. When reanalysing GEO data, please also cite the original study and GEO itself.
Sources
Every claim about GEO and the example study traces to these.
- NCBI GEO: archive for functional genomics data sets—update. Nucleic Acids Res. 2013;41(D1):D991–D995.
doi:10.1093/nar/gks1193 - Gene Expression Omnibus: NCBI gene expression and hybridization array data repository. Nucleic Acids Res. 2002;30(1):207–210.
doi:10.1093/nar/30.1.207 - NCBI GEO: archive for gene expression and epigenomics data sets: 23-year update. Nucleic Acids Res. 2024;52(D1):D138–D144.
doi:10.1093/nar/gkad965 - GEOquery: a bridge between the Gene Expression Omnibus (GEO) and BioConductor. Bioinformatics. 2007;23(14):1846–1847.
doi:10.1093/bioinformatics/btm254 - RNA-Seq transcriptome profiling identifies CRISPLD2 as a glucocorticoid responsive gene that modulates cytokine function in airway smooth muscle cells. PLoS One. 2014;9(6):e99625. (GEO: GSE52778)
doi:10.1371/journal.pone.0099625 - Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15(12):550.
doi:10.1186/s13059-014-0550-8 - edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010;26(1):139–140.
doi:10.1093/bioinformatics/btp616 - limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43(7):e47.
doi:10.1093/nar/gkv007 - recount3: summaries and queries for large-scale RNA-seq expression and splicing. Genome Biol. 2021;22(1):323.
doi:10.1186/s13059-021-02533-6 - Digital expression explorer 2: a repository of uniformly processed RNA sequencing data. GigaScience. 2019;8(4):giz022.
doi:10.1093/gigascience/giz022 - Massive mining of publicly available RNA-seq data from human and mouse. Nat Commun. 2018;9(1):1366.
doi:10.1038/s41467-018-03751-6
Frequently asked questions
Which GEO accessions can TransXplorer fetch?
GSE followed by digits, for example GSE52778. The box is case-insensitive and trims spaces. Sample (GSM), Platform (GPL) and DataSet (GDS) IDs are rejected with “Invalid GEO ID format”.
Why “No Count Matrix Found” when the GEO page lists files?
Does GEO import work for mouse and other organisms?
Why was the wrong column pre-selected as the grouping column?
ercc_mix:ch1 instead of treatment. Always check the dropdown and the group preview before clicking Confirm Groups.
Can I analyse microarray series from GEO?
How should I cite a reanalysis of GEO data?
GSE accession in your methods, cite GEO (Barrett et al. 2013, doi:10.1093/nar/gks1193), and cite TransXplorer (doi:10.64898/2026.05.15.724657).