Case study · 20 min read Intermediate 33 cancer cohorts

Working with TCGA Cancer Data

Pick a cancer, ask about a gene, and get from tumor-vs-normal expression to Kaplan–Meier survival curves, differential expression, and pathways. Along the way you’ll learn how to read each plot without fooling yourself.

Data pre-loaded, nothing to download Real TCGA-BRCA numbers No code required
TL;DR · The short version
  • 1 TCGA is the largest public collection of tumor RNA-seq linked to patient outcomes: 33 cancer types, each a project such as TCGA-BRCA or TCGA-LUAD. TransXplorer has all 33 pre-loaded under the TCGA Data Analysis menu.
  • 2 Two questions, two pages. Gene Expression Analysis answers “is my gene different in tumors, and does it track with survival?” (box plot and Kaplan–Meier, median split). RNA-seq Analysis answers “what changes genome-wide between tumor and normal?” (DEG plus pathway enrichment).
  • 3 The plots are quick to make and easy to over-read. Many cohorts have few or no normal samples. “Normal” means tissue next to the tumor, not healthy tissue. Survival associations are correlations. Read the caveats in Section 5 before you put a TCGA figure in a paper.
Section 1

What TCGA is, and why it’s worth your time

Eleven thousand tumors, one consistent pipeline, outcomes attached.

Say you’ve found a gene that changes in your cell line or mouse model. The next question from a reviewer, or from you, is usually: does this matter in human patients? The Cancer Genome Atlas is often the fastest way to get a first answer.

TCGA was a US National Cancer Institute and National Human Genome Research Institute program that profiled primary tumors from thousands of patients across 33 cancer types. Each sample was characterised with several technologies: RNA-seq, DNA sequencing, copy number, methylation, and others. Clinical information, including survival follow-up, came with each patient.1 The data now live at the NCI Genomic Data Commons (GDC), which reprocessed everything through a single harmonised pipeline.2

Each cancer type is a project with a short code: TCGA-BRCA is breast invasive carcinoma, TCGA-LUAD is lung adenocarcinoma, TCGA-GBM is glioblastoma. TransXplorer offers all 33 in every TCGA dropdown. The full list is below so you can find your cancer’s code.

Projects
33TCGA-ACC … TCGA-UVM
RNA-seq samples
11,499across all 33 cohorts
Primary tumors
10,152sample type 01
Normal samples
740solid tissue normal, type 11

Counts are tallied from the TCGA expression files TransXplorer uses for its gene-level analyses. The remaining samples are metastatic, recurrent, blood-derived (LAML) or additional primaries.

  • ACCAdrenocortical carcinoma
  • BLCABladder urothelial carcinoma
  • BRCABreast invasive carcinoma
  • CESCCervical squamous cell carcinoma & endocervical adenocarcinoma
  • CHOLCholangiocarcinoma
  • COADColon adenocarcinoma
  • DLBCDiffuse large B-cell lymphoma
  • ESCAEsophageal carcinoma
  • GBMGlioblastoma multiforme
  • HNSCHead & neck squamous cell carcinoma
  • KICHKidney chromophobe
  • KIRCKidney renal clear cell carcinoma
  • KIRPKidney renal papillary cell carcinoma
  • LAMLAcute myeloid leukemia
  • LGGBrain lower grade glioma
  • LIHCLiver hepatocellular carcinoma
  • LUADLung adenocarcinoma
  • LUSCLung squamous cell carcinoma
  • MESOMesothelioma
  • OVOvarian serous cystadenocarcinoma
  • PAADPancreatic adenocarcinoma
  • PCPGPheochromocytoma & paraganglioma
  • PRADProstate adenocarcinoma
  • READRectum adenocarcinoma
  • SARCSarcoma
  • SKCMSkin cutaneous melanoma
  • STADStomach adenocarcinoma
  • TGCTTesticular germ cell tumors
  • THCAThyroid carcinoma
  • THYMThymoma
  • UCECUterine corpus endometrial carcinoma
  • UCSUterine carcinosarcoma
  • UVMUveal melanoma
Why this matters

TCGA is a discovery and validation resource, not a clinical trial. Patients weren’t randomised, treatments varied, and follow-up is uneven between cohorts. That makes it excellent for asking “is this pattern present in human tumors?” and poor at answering “does this gene cause worse outcomes?” Keep that distinction in mind and every plot below will be easier to interpret.

Section 2

Reading a TCGA barcode

Two digits tell you whether a sample is tumor or normal.

Every TCGA sample has a barcode, and you’ll see them as column names in downloaded tables. They look cryptic, but they are a fixed sequence of fields, each separated by a hyphen. You only need to recognise two of them: the participant (which patient) and the sample type (what kind of tissue).

TransXplorer reads these sample types for you. In the RNA-seq analysis, samples are grouped by their sample-type label (“Primary solid Tumor”, “Solid Tissue Normal”, and so on), and the primary tumor group is tested against Solid Tissue Normal as the reference. The single-gene box plot keeps only Primary solid Tumor and Solid Tissue Normal samples. So you rarely have to parse barcodes yourself, but recognising -01 vs -11 lets you sanity-check anything you download.

Why this matters

One patient can contribute more than one sample: a tumor and a matched normal, or two tumor vials (01A, 01B). The barcode is how you notice. TransXplorer’s survival analysis handles this for you: it keeps only primary tumor samples (01, or 03 for primary blood-derived cancers such as LAML) and the first aliquot per 12-character patient ID, so no patient is counted twice. If you build your own survival table from downloaded data, deduplicate on the patient ID first.

Section 3

Where TCGA lives in TransXplorer

One navbar menu, two pages, five outputs.

Open TransXplorer and look for TCGA Data Analysis in the top navigation bar. It’s a dropdown with two pages, and each one answers a different kind of question:

The single-gene tools read a compact, pre-computed table of TPM values per cohort, so a box plot or survival curve usually appears within seconds. The genome-wide analysis loads the full matrix of GDC read counts for the cohort and runs a complete differential expression model. For the large cohorts (BRCA has more than 1,200 RNA-seq samples) it takes noticeably longer, and a progress bar keeps you informed.

Section 4

Worked example: estrogen receptor and HER2 in breast cancer

Five steps through TCGA-BRCA, from one gene to genome-wide pathways.

For a first case study you want genes whose biology is already well established, so you can tell whether the tool is telling you something sensible. Breast cancer gives us two.

ESR1 encodes estrogen receptor alpha. Most breast tumors are ER-positive and express it strongly, while a minority, including most basal-like tumors, express very little.3 ERBB2 encodes HER2. It is amplified and heavily overexpressed in a subset of breast cancers, and that amplification was linked to relapse and survival decades ago.4 Both patterns are “known answers” you can check against what TransXplorer shows you.

Step 1 · ~2 minutes

Is the gene different in tumors? The box plot

TCGA Data Analysis› Gene Expression Analysis› Box Plot Analysis

In the sidebar, set Select Cancer Type to Breast invasive carcinoma (BRCA) and type ESR1 into Enter Gene Symbol. The page opens with COAD and EGFR filled in as an example, so change both. Click Analyze Expression.

TransXplorer keeps only the Primary solid Tumor and Solid Tissue Normal samples and plots expression as log2(TPM + 1). It also runs a two-sample t-test. R’s default is the Welch version, which doesn’t assume equal variances, and that matters here because tumors are far more variable than normals. The Detailed Statistics tab prints the full test output. The sidebar’s Analysis Summary shows how many samples ended up in each group.

What the real numbers say

We ran the box plot calculation on the TCGA-BRCA data TransXplorer uses (1,111 primary tumors, 113 solid tissue normals). Here is what comes out:

GeneNormal medianTumor medianTumor 5th – 95th pctNormal 5th – 95th pctWelch t-test p
ESR15.326.560.61 – 8.972.65 – 6.931.04e-05
ERBB26.386.944.92 – 10.723.16 – 7.392.91e-13
All values are log2(TPM + 1). Computed from the TCGA-BRCA expression file used by the app; if the underlying data release changes, numbers in the app may shift slightly.

Look past the p-values. For ERBB2 the medians differ by only about half a log2 unit, so the typical tumor is barely different from normal. The interesting part is the top of the tumor range: the 95th percentile sits at 10.7, more than three log2 units (roughly 10-fold) above anything common in normal tissue. That tail is the HER2-amplified subset. For ESR1 the important feature is the spread. Tumors run from almost zero to well above normal, which is what you’d expect from a mix of ER-positive and ER-negative cancers.

Why this matters

With more than a thousand tumors, almost any real difference gets a tiny p-value. In a cohort this large, the p-value mostly tells you the sample size was big. To judge whether a difference matters biologically, look at the size of the shift and the shape of the distribution: medians, spread, tails. Sometimes the biomarker signal is in a subgroup, not in the average.

Step 2 · ~5 minutes

Does it track with survival? The Kaplan–Meier curve

TCGA Data Analysis› Gene Expression Analysis› Survival Analysis

Switch to the Survival Analysis sub-tab. It opens on BRCA. Enter your gene in Enter Gene Symbol (the default is EGFR) and click Analyze Survival. Results appear in three tabs: Survival Plot, Expression Distribution, and Detailed Statistics.

How patients are split: the median, fixed in advance

Survival curves compare groups, so a continuous expression value has to be turned into categories. TransXplorer uses the simplest defensible rule. It takes the gene’s TPM values across the cohort’s primary tumors, one sample per patient (normal, metastatic and recurrent samples are left out first), finds the median, and labels patients above it HIGH and the rest (at or below it) LOW. In TCGA-BRCA that gives 1,033 patients with survival data (151 events). The Expression Distribution tab shows exactly this per-patient set: a histogram of TPM on a log10 axis with the median drawn as a dashed red line.

How to read a Kaplan–Meier curve

A Kaplan–Meier curve5 answers one question: of the patients in this group, what fraction are still alive at each point in time? Every curve starts at 1.0 (everyone alive at diagnosis) and can only step down. The clever part is how it handles patients whose follow-up ended while they were still alive. They are censored: they count toward the denominator for as long as they were followed, then drop out quietly. They are never counted as deaths.

Read the plot in this order:

  • Check the caption first. It reads “n = … patients ( … events)”. An event is a death. The statistics run on events, not patients, so 1,000 patients with 40 deaths is a much weaker analysis than it sounds. The sidebar table gives the size of each group.
  • Look at the separation and when it happens. Curves that split early and stay apart tell a different story from curves that only diverge at the far right, where few patients remain and the confidence bands balloon.
  • Median survival is the time at which a curve crosses 0.50. If a curve never gets there, as in many indolent cohorts, the median is “not reached”. That means most patients outlived their follow-up, not that they are immortal.
  • The log-rank p-value tests whether the two curves differ anywhere over the whole follow-up. It says nothing about how big the difference is.
  • The hazard ratio (HR) from the Cox model6 gives the size. It compares the instantaneous death rate in one group with the other.
Read the hazard ratio the right way round

In TransXplorer’s Cox model the HIGH group is the reference, so the reported HR describes LOW relative to HIGH. HR > 1: patients with low expression die at a higher rate, so high expression goes with better survival. HR < 1: low expressers do better, so high expression goes with worse survival. The Detailed Statistics tab prints the full Cox summary; its coefficient is labelled strataLOW, which confirms the direction.

Then check the 95% confidence interval. If it includes 1.0 (for example 0.9–1.7), the data are compatible with no difference at all, however suggestive the curves look.

For ESR1 and ERBB2 in BRCA, run the analysis yourself and apply the checklist: how many events, which direction is the HR, does the confidence interval cross 1? Breast cancer is a good teaching cohort precisely because its overall survival events are relatively few for its size (about 200 of the 1,231 BRCA RNA-seq samples come from patients recorded as deceased), so intervals are wider than the sample size suggests. Remember too that ER-positive and HER2-positive patients receive targeted therapies, which reshape their outcomes. A survival curve in TCGA reflects biology and the treatment era.

Step 3 · the longest step

What changes genome-wide? Tumor-vs-normal DEG

TCGA Data Analysis› RNA-seq Analysis

Single-gene questions are hypothesis-driven. The RNA-seq Analysis page is the hypothesis-free counterpart: it tests every gene in the cohort at once for a difference between tumor and normal. The sidebar has three numbered sections and one button.

  • Project IDTCGA-BRCA
    Any of the 33 cohorts; BRCA is pre-selected. Experimental strategy is fixed to RNA-Seq.
  • Analysis MethodedgeR (QL)
    Choose limma-voom, edgeR (QL) or DESeq2 (the default). All three are standard count-based models.7–9 We use edgeR’s quasi-likelihood test in this walkthrough. Our differential expression guide compares the three.
  • Normalization MethodTMM
    For limma-voom and edgeR: TMM (recommended), RLE or Upper Quartile. DESeq2 always uses its built-in median-of-ratios normalisation.
  • Adj P-Value Threshold0.05
    Cut-off on the multiple-testing-adjusted p-value (false discovery rate).
  • Log2FC Threshold2
    Minimum absolute log2 fold change, so the default of 2 means at least a 4-fold change. A gene counts as up- or downregulated only if it passes both thresholds. With hundreds of samples, the fold-change cut-off does most of the filtering.

Click Run Analysis. TransXplorer loads the cohort’s raw read counts, removes genes with too few counts to test, groups samples by sample type, and fits the model you picked to test the primary tumor group against Solid Tissue Normal as the reference (for BRCA, Primary_solid_Tumor vs Solid_Tissue_Normal), whichever method you choose. A positive log2FC therefore means “higher in primary tumor than in normal”. Results fill four tabs:

  • Dataset Info: summary cards (genes analysed, up, down, total samples), plus Vital Status, Sample Types and Analysis Parameters tables (the last has a Comparison row, e.g. Primary_solid_Tumor vs Solid_Tissue_Normal, above the Reference row), a summary line such as “Project: TCGA-BRCA | Primary_solid_Tumor vs Solid_Tissue_Normal | N significant DEGs identified”, an MA plot, and an interactive PCA of the 1,000 most variable genes.
  • DEG Results: sortable tables for All DEGs, Upregulated and Downregulated. Click a gene row to open its OpenTargets information.
  • Data Visualization: an interactive Volcano Plot with its own fold-change and p-value thresholds, and a Heatmap (Z-scored by row) of the top N genes by p-value, by fold change, or significant DEGs only, with a choice of palette and clustering method.
  • Pathway Analysis: the next step.
Check the Sample Types table before interpreting

Cohorts don’t always contain just two sample types. TCGA-BRCA also has a handful of metastatic samples, and several cohorts (GBM, LGG, OV, LIHC and others) include recurrent tumors. The DE test still compares only the primary tumor group with Solid Tissue Normal, but nine cohorts have no normal samples at all, and the analysis needs at least two groups. Always open Dataset Info → Sample Types to see which groups are present and how many normals there are, and check the Comparison and Reference rows in Analysis Parameters, before you read the DEG list.

Step 4 · ~2 minutes

What biology do the DEGs point to? Pathway enrichment

RNA-seq Analysis› Pathway Analysis

Once the DEG run has finished, the Pathway Analysis tab unlocks. On the left, tick the databases you want: KEGG Pathways, GO Biological Process and Reactome Pathways are pre-selected, and GO Molecular Function, GO Cellular Component, WikiPathways, MSigDB Hallmark and BioCarta are available. On the right, choose the Gene Set to Analyze (both up and down, up only, or down only), a P-value Cutoff (default 0.05) and Min Genes per Pathway (default 3). Then click Run Pathway Analysis.

This is over-representation analysis (ORA). It asks whether your DEG list contains more members of a pathway than chance would predict. For GO, KEGG, Reactome and WikiPathways, TransXplorer uses clusterProfiler10 with Benjamini–Hochberg correction and, importantly, uses every gene that was tested as the background, not the whole genome. MSigDB Hallmark and BioCarta are queried through the Enrichr service. Results come as a Results Table (filter by database, sort by adjusted p-value, combined score or gene count), a Dot Plot, a Bar Plot, and a pathway Network.

Split up- and downregulated genes into separate runs when you can. In a tumor-vs-normal comparison, “both” tends to mix cell-cycle programmes switched on in the tumor with tissue-specific programmes switched off, and the mixed list is harder to interpret. For a ranked, threshold-free alternative, see GSEA vs ORA.

Step 5 · ~1 minute

Take it with you: downloads

Every TCGA output has its own download button:

WhereButtonYou get
DEG Results › All DEGsDownload All ResultsCSV of every tested gene (log2FC, p-values, Ensembl ID, gene symbol, gene type)
DEG Results › Up / DownDownload Upregulated / DownregulatedCSV of genes passing both thresholds
Data Visualization › Volcano PlotDownload PNGVolcano PNG with the top genes labelled (set by “Labels in Download”)
Data Visualization › HeatmapDownloadHeatmap PNG
Pathway Analysis › Results TableDownload ResultsCSV of enriched terms
Pathway Analysis › Dot / Bar / NetworkDownload Plot / DownloadSelf-contained interactive HTML
Box Plot AnalysisDownload PlotBox plot PNG, 300 dpi
Survival AnalysisDownload PlotSurvival plot and expression-distribution PNGs, 300 dpi

The DEG CSV is the most reusable of these. You can open it in a spreadsheet, feed it into your own tools, or compare it with results from your own experiment (see bringing your own data).

Section 5

Caveats: how TCGA results go wrong

Read this before a TCGA panel goes into Figure 1.

TCGA plots are fast to make, look authoritative, and are easy to over-interpret. Most problems come from a few predictable sources.

Caveat 1 · Many cohorts have few or no normal samples

TCGA set out to profile tumors. Normal tissue was collected when it was available, not systematically. Across the 33 RNA-seq cohorts, only 740 of 11,499 samples are solid tissue normal, and they are very unevenly spread:

Caveat 2 · “Normal” is normal adjacent tissue, not healthy tissue

TCGA normals are usually cut from tissue next to the tumor in the same patient. A pan-cancer comparison of normal-adjacent, tumor and truly healthy tissue (from GTEx donors) found that normal-adjacent tissue forms a distinct intermediate state. It shows inflammatory and wound-response signals that healthy tissue lacks.11 In practice, a tumor-vs-normal difference in TCGA is a tumor-vs-adjacent difference. Some changes you care about can be muted, and some are induced in the adjacent tissue itself.

Caveat 3 · Know which survival endpoint you’re using

TransXplorer’s survival analysis uses overall survival. A patient recorded as deceased has an event at days to death; everyone else is censored at days to last follow-up. Overall survival is unambiguous, but it counts deaths from any cause, and in cohorts with good prognosis and short follow-up (thyroid, prostate, testicular) there are very few events. A curated pan-cancer clinical resource, TCGA-CDR, assessed which endpoints are reliable cohort by cohort. It urges caution with overall survival where events are few or follow-up is short, and it recommends progression-free interval as a generally robust alternative.12 Check the events count in the plot caption before you trust a curve.

Caveat 4 · Cutpoint fishing and multiple testing

Splitting at the median is a choice made in advance. Many web tools, and many papers, instead try every possible cutoff and report the one with the smallest p-value. That “optimal cutpoint” approach badly inflates false positives: with enough cutoffs, even a gene unrelated to survival will produce a striking curve at some threshold.13

That last point is the same problem one level up. Testing one gene in 33 cohorts is 33 tests; at p < 0.05 you’d expect one or two “hits” by chance alone. If you screen many genes or cohorts, say how many you tested and correct for it (a Bonferroni or FDR adjustment across your screen), or treat the screen as hypothesis-generating and confirm in an independent dataset.

More traps, quickly

Trap 5

Correlation is not causation

Expression can track tumor subtype, grade, stage or proliferation rather than drive outcome. In breast cancer, many genes “predict survival” simply because they mark the luminal vs basal split.

Fix — ask whether the gene adds anything beyond known subtypes and stage, and test mechanism experimentally.
Trap 6

Tumor purity

A bulk tumor sample is a mix of cancer cells, stroma and immune cells. A gene that looks “down in tumor” may just be expressed by normal epithelium that the tumor has displaced.

Fix — check which cell types express the gene before giving a mechanistic story.
Trap 7

Unpaired comparison of paired data

Some normals come from the same patients as tumors. The box plot’s t-test treats every sample as independent, which is a reasonable first look but not a paired analysis.

Fix — for a publication claim, consider a matched-pairs analysis on the patients with both samples.
Trap 8

Treatment era and confounders

TCGA patients were treated over many years with different therapies. Survival differences can reflect who received which treatment.

Fix — treat univariable KM results as a starting point; adjust for stage and age in a multivariable model before calling a gene prognostic.

Before you publish a TCGA panel: a checklist

  • I checked the Sample Types table and know exactly which groups were compared.
  • The normal group has enough samples to be meaningful, and I described it as normal adjacent tissue.
  • I reported effect sizes (fold change, hazard ratio with 95% CI), not just p-values.
  • I know which direction the hazard ratio runs (LOW relative to HIGH in TransXplorer).
  • I reported the number of events, and the cohort has enough of them for overall survival to be sensible.
  • The cutoff was fixed in advance, and I disclosed how many genes and cohorts I looked at.
  • I treat the result as an association and have a plan to validate it independently.
TCGA is best at telling you a pattern exists in human tumors. It is weakest at telling you why.
Section 6

What’s next

Bring your own data, try another workflow, or go deeper on the methods.

Keep exploring

TCGA is usually the second half of a story: you find something in your own experiment, then ask whether it shows up in patients. The sibling tutorials cover the first half, whether you’re starting from a count table, a public GEO dataset or raw sequencing reads.

Go deeper on the concepts

Try it on your gene now

Open TransXplorer, choose TCGA Data Analysis › Gene Expression Analysis, and you’ll have a tumor-vs-normal box plot and a survival curve within a couple of minutes. No account, no download.

Launch TransXplorer →

References

Primary sources for the data and methods discussed on this page.

  1. Cancer Genome Atlas Research Network, Weinstein JN, et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat Genet. 2013;45:1113–1120. doi:10.1038/ng.2764
  2. Grossman RL, Heath AP, Ferretti V, et al. Toward a shared vision for cancer genomic data. N Engl J Med. 2016;375:1109–1112. doi:10.1056/NEJMp1607591
  3. Cancer Genome Atlas Network. Comprehensive molecular portraits of human breast tumours. Nature. 2012;490:61–70. doi:10.1038/nature11412
  4. Slamon DJ, Clark GM, Wong SG, et al. Human breast cancer: correlation of relapse and survival with amplification of the HER-2/neu oncogene. Science. 1987;235:177–182. doi:10.1126/science.3798106
  5. Kaplan EL, Meier P. Nonparametric estimation from incomplete observations. J Am Stat Assoc. 1958;53:457–481. doi:10.1080/01621459.1958.10501452
  6. Cox DR. Regression models and life-tables. J R Stat Soc Series B. 1972;34:187–220. doi:10.1111/j.2517-6161.1972.tb00899.x
  7. Law CW, Chen Y, Shi W, Smyth GK. voom: precision weights unlock linear model analysis tools for RNA-seq read counts. Genome Biol. 2014;15:R29. doi:10.1186/gb-2014-15-2-r29
  8. Robinson MD, McCarthy DJ, Smyth GK. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010;26:139–140. doi:10.1093/bioinformatics/btp616
  9. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550. doi:10.1186/s13059-014-0550-8
  10. Yu G, Wang LG, Han Y, He QY. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS. 2012;16:284–287. doi:10.1089/omi.2011.0118
  11. Aran D, Camarda R, Odegaard J, et al. Comprehensive analysis of normal adjacent to tumor transcriptomes. Nat Commun. 2017;8:1077. doi:10.1038/s41467-017-01027-z
  12. Liu J, Lichtenberg T, Hoadley KA, et al. An integrated TCGA pan-cancer clinical data resource to drive high-quality survival outcome analytics. Cell. 2018;173:400–416. doi:10.1016/j.cell.2018.02.052
  13. Altman DG, Lausen B, Sauerbrei W, Schumacher M. Dangers of using “optimal” cutpoints in the evaluation of prognostic factors. J Natl Cancer Inst. 1994;86:829–835. doi:10.1093/jnci/86.11.829
  14. Colaprico A, Silva TC, Olsen C, et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016;44:e71. doi:10.1093/nar/gkv1507

TransXplorer’s TCGA module is built on the Bioconductor ecosystem, including TCGAbiolinks14 for working with GDC data.

If TransXplorer helps your research, please cite us

The preprint is open access on bioRxiv. A peer-reviewed version is in submission.

Verma VM, Oler E, Syed H, Han S, Berjanskii M, Mason AL, Wishart DS, Wong GK. TransXplorer: An automated translational discovery platform for RNA-seq data. bioRxiv. 2026. doi:10.64898/2026.05.15.724657
Read the preprint →

Frequently asked questions

Which TCGA cancer types are available?
All 33 TCGA projects, from TCGA-ACC to TCGA-UVM (full list in Section 1). They are pre-loaded on the server, so there’s nothing to download before you start.
How are the HIGH and LOW groups defined?
By the median. Using one primary tumor sample per patient (normal, metastatic and recurrent samples are excluded), patients whose TPM for your gene is above the median are HIGH; the rest are LOW. The cutoff is fixed in advance rather than optimised, which keeps the p-value honest. The Expression Distribution tab shows the histogram with the median marked.
My hazard ratio is 0.7. Is high expression good or bad?
In TransXplorer the HR compares LOW to HIGH (HIGH is the reference). An HR of 0.7 means low expressers have about 30% lower hazard of death, so high expression goes with worse survival. Check that the 95% confidence interval doesn’t include 1 before reading much into it.
Why do I get “Insufficient samples” for the box plot?
The cohort has no Solid Tissue Normal samples. That applies to ACC, DLBC, LAML, LGG, MESO, OV, TGCT, UCS and UVM. Several others have only one to five normals, which is too few for a meaningful comparison.
Is TCGA “normal” the same as healthy tissue?
No. It is mostly normal tissue adjacent to the tumor, which carries its own inflammatory and field-effect signals. Describe it as normal adjacent tissue in figures and text.
Does a significant survival curve mean my gene drives progression?
No. It is an association in observational data. Expression may simply mark a subtype, stage or treatment group. Validate in an independent cohort, adjust for clinical covariates, and test mechanism experimentally.
Can I cite TransXplorer in a paper?
Yes, please do. Preprint: doi:10.64898/2026.05.15.724657. Please also cite TCGA itself and the analysis methods you report.