Tutorial · 8 min read Beginner No code required

Working with Your Own Data

Two files, a few formatting rules, and the handful of mistakes that quietly wreck RNA-seq results. Get these right before you upload and everything downstream gets easier.

CSV, TSV, TXT or XLSX Symbols, Ensembl or Entrez IDs 10-point pre-upload checklist
TL;DR · The short version
  • 1 You need a count matrix (genes in rows, samples in columns, gene IDs in the first column) and, ideally, a sample sheet with one row per sample.
  • 2 The values must be raw read counts. TPM, FPKM, CPM and log-transformed values break the statistics behind DESeq2, edgeR and limma-voom.
  • 3 The sample sheet's Sample column must match the count-matrix column names character for character. Row order doesn't matter; spelling does.
  • 4 Most failed analyses come from six avoidable problems: Excel-mangled gene names, duplicate IDs, normalised values, mismatched names, too few replicates, and leftover annotation columns. Run the checklist before you upload.
Section 1

The two files, and the one thing that links them

A table of numbers, and a table that says what each column is.

Whatever pipeline produced your data (a core facility, a GEO download, your own alignment), bulk RNA-seq analysis starts from the same pair of tables. The first holds the measurements. The second holds the experimental design. Sample names are what join them.

The count matrix has one row per gene and one column per sample. Each cell is the number of sequencing reads assigned to that gene in that sample. The sample sheet (often called metadata or colData) flips the orientation: one row per sample, with columns describing each one, such as which condition it belongs to, which batch it was processed in, and anything else you recorded.

Why this matters

Keeping measurements and design in separate files means you can re-analyse the same counts with a different grouping (say, by sex instead of treatment) without touching the numbers. It also stops design information sneaking into the count table, which is where a lot of upload errors start.

Section 2

Formatting the count matrix

One header row, one ID column, numbers everywhere else.

TransXplorer reads the count matrix with a few simple assumptions. If your file follows them, it loads first time. The same rules appear in the app under File Format Guide & Tips in the upload sidebar.

  • Header rowrow 1 The first row holds column names: a label for the gene column, then one name per sample. No title rows, notes or blank lines above it.
  • Gene ID columngene_id TransXplorer looks for a column named like gene_id, GeneID, Gene, Symbol, gene_symbol, GeneName or ID (any capitalisation). If none is found it uses the first column. Putting IDs first is the safe choice. Rows with an empty ID are dropped.
  • Sample columnsnumeric only Every column other than the gene ID column is treated as a sample. That includes columns you might not think of as samples, such as Length, Chr or a description column, so delete those first.
  • At least two samples≥ 2 columns The upload is rejected with fewer than two numeric sample columns. For a meaningful comparison you need far more than that; see pitfall 5.
  • Gene levelone row per gene Upload gene-level counts. If your quantifier reports transcripts (Salmon, kallisto, RSEM isoforms), summarise to genes first, for example with tximport. The FASTQ Processing tab in TransXplorer does this step for you.

Accepted file types

The Upload Count Matrix (CSV/XLSX/TSV/TXT) button takes four formats. Files up to 500 MB are accepted.

.csv
Comma-separated. The most portable choice and the one we recommend.
.tsv
Tab-separated. What featureCounts and most pipelines write natively.
.txt
Plain text. The separator (tab, comma or semicolon) is detected from the first line.
.xlsx
Excel workbook. Only the first sheet is read. See pitfall 1 before saving from Excel.

Which gene IDs work

Use whatever your pipeline produced, as long as the whole file uses one type. The identifiers are kept as row names for the differential expression step. Later, when modules such as enrichment or network analysis need gene symbols, TransXplorer converts Ensembl and Entrez IDs to symbols using the organism's annotation database.

ID typeExampleNotes
Gene symbol TP53 Easiest to read. Symbols change over time and some are ambiguous, and they are the type Excel damages (pitfall 1).
Ensembl gene ENSG00000141510 Stable and unique, so the safest choice. Mouse IDs (ENSMUSG…) are recognised too.
Ensembl, versioned ENSG00000141510.17 Fine to upload. The .17 version suffix is removed before IDs are mapped to symbols, so you don't need to strip it yourself.
Entrez / NCBI Gene 7157 Purely numeric IDs are recognised as Entrez. Keep them as text if you edit the file in a spreadsheet.
Don't mix ID types

A file where some rows are symbols and others are Ensembl IDs (common after merging tables) will load, but the same gene can appear twice under two names, and downstream mapping gets patchy. Convert everything to one type before upload.

Section 3

Raw counts, TPM or normalised? Use raw counts

The single most common reason an upload gives strange results.

The same gene in the same sample can be reported three different ways depending on where you got the file. Only one of them is what differential expression software expects.

Why DE tools insist on counts

DESeq2 and edgeR model each gene's reads with a negative binomial distribution, whose variance depends on how many reads were counted (Love et al., 2014; Robinson et al., 2010). A gene with 3 reads is far noisier than one with 3,000, and the model uses that difference to decide how much to trust each fold change. limma-voom estimates the same mean–variance relationship from counts. Once you convert to TPM or FPKM, the information about how many reads were actually observed is gone: 3 reads and 3,000 reads can end up as similar-looking TPM values. These tools also do their own normalisation for library size, so feeding them pre-normalised values normalises twice. Wagner et al. (2012) showed that even comparing RPKM values across samples is unreliable, which is one reason TPM was proposed.

Where to find the raw counts

  • featureCounts / HTSeq-countinteger counts Already raw counts. Drop the annotation columns from featureCounts and the __no_feature-style summary rows from HTSeq.
  • STAR --quantMode GeneCountsReadsPerGene.out.tab Pick the column that matches your library strandedness, and remove the four N_ summary rows at the top.
  • Salmon / kallisto / RSEMestimated counts Use NumReads (Salmon), est_counts (kallisto) or expected_count (RSEM), summarised to genes. These have decimals but are on the count scale, which is fine (Soneson et al., 2015). Don't use the TPM column.
  • GEO supplementary filescheck the label Files labelled raw counts or read counts are what you want. Files labelled normalised, FPKM or TPM are not. TransXplorer's Or fetch from GEO box can pull NCBI-processed counts for many RNA-seq series directly; see the GEO tutorial.
Why this matters

Uploading TPM doesn't crash anything, which is what makes it dangerous. You still get a volcano plot and a gene list; they just rest on the wrong statistical assumptions. If a reviewer asks what went into the DE test, "raw counts" is the only answer that holds up. The differential expression concept page goes deeper on the models.

Section 4

The sample sheet: telling TransXplorer what each sample is

One row per sample. One column that has to be spelled exactly right.

The count matrix says how much of each gene there is. The sample sheet says what was done to each sample. Without it, software can only guess which columns to compare.

Keep it small and plain. A minimal sample sheet that works everywhere in TransXplorer looks like this:

samples.csv
Sample,Group,Batch
Ctrl_1,Control,A
Ctrl_2,Control,B
Ctrl_3,Control,C
KO_1,KO,A
KO_2,KO,B
KO_3,KO,C
  • Samplerequired The key column. Name it exactly Sample. In the group dialog TransXplorer can also recognise headers such as sample_id or SampleName, but the batch-metadata upload insists on Sample, so using that name everywhere avoids surprises. Values must match the count-matrix column headers.
  • Groupthe comparison The condition each sample belongs to: Control/KO, Healthy/Disease, time points, and so on. A column named Group, group or Condition is pre-selected as the group column for you.
  • Batchoptional Sequencing run, library-prep day, plate, donor: anything technical that could shift expression. A column named Batch or batch is pre-selected as the batch column.
  • Other covariatesoptional Sex, age, RIN, and similar. These can be chosen under Additional Covariates when you configure the metadata.

Exact match, or no match

Sample names are compared as plain text. TransXplorer doesn't guess that ko_1 probably means KO_1. Each mismatch below would leave a sample unassigned:

No metadata file? Let the names carry the groups

If you skip the sample sheet, TransXplorer tries to read groups from the sample names. It takes the text after the last underscore as the group, so S1_Control becomes Control. It accepts the result only if there are at least two groups and no more groups than half the number of samples. The app's own format hint shows the same pattern: Sample1_Control.

Check the Batch column against the Group column

If every control was sequenced in run A and every treated sample in run B, batch and biology are perfectly confounded. No software can separate them afterwards, and correcting for batch would remove your treatment effect along with it. Spread each condition across batches when you design the experiment. The batch effects concept page explains how TransXplorer detects and corrects batch effects, and when it can't.

Section 5

Uploading in TransXplorer, step by step

All of this happens in the left sidebar of the Transcriptome Analysis tab.

  1. Load the count matrix

    Transcriptome Analysis › 1. Data Input & Setup

    Click Upload Count Matrix (CSV/XLSX/TSV/TXT) and choose your file. TransXplorer checks the file is not empty and under 500 MB, reads it, finds the gene ID column, and confirms there are at least two numeric sample columns. If something is wrong you get a message naming the problem, for example that no numeric expression data was found.

    The same panel offers two shortcuts: Or fetch from GEO (type a GSE accession and click Fetch) and Load Example Counts, which loads the GSE151427 dataset. Loading the example once is a good way to see what a correctly formatted file produces.

  2. Confirm the sample groups

    Dialog › Sample Group Assignment

    A dialog opens as soon as the counts load. You have three ways to define groups:

    • Auto-detected. If the sample names follow the Sample_Group pattern (fig 4), you'll see the proposed groups with a summary. Click Accept These Groups, or Choose Different Method if they're wrong.
    • Manual Assignment. Type group names, click Add Group, then pick a group for each sample from a dropdown.
    • Upload Metadata. Choose your sample sheet (CSV, TSV, TXT or XLSX), then pick which column holds the groups. TransXplorer shows how many samples matched and how many fall in each group.

    Finish with Confirm Groups.

  3. Select the organism

    Sidebar › Select Organism

    The Select Organism: * menu lists Human, Mouse, Rat, Drosophila, Zebrafish, C. elegans, Yeast, Arabidopsis, Chicken, Pig and Cow. Anything else goes under Other, where you can type the organism name plus an optional KEGG organism code and NCBI taxonomy ID. As the app notes, differential expression, QC and WGCNA work for any organism; pathway enrichment and protein networks need the KEGG code or taxonomy ID; drug targets and cell-type deconvolution are human-only.

  4. Decide how to handle batch

    Sidebar › Batch Effects & Covariates Detection

    Once groups are confirmed, a batch panel appears with three options: Automatic Detection (recommended; scores every metadata column with PVCA, silhouette and kBET), Manual Selection (you name the batch column), or Skip.

    For Automatic or Manual, use Upload Batch Metadata File. This file must contain a column headed Sample that matches your count-matrix names. It can be .csv, .tsv, .txt or .xlsx; tab, semicolon or comma separators are detected automatically. A Configure Sample Metadata window then asks for the Group Column (Required), an optional Batch Column, and any Additional Covariates. Your sample sheet from Section 4 can be reused here unchanged.

  5. Set parameters and run

    Sidebar › 2. Analysis Parameters

    Choose the DEG package (edgeR, DESeq2 or limma-voom), the normalisation method and the comparison, then set thresholds. Click Run Exploratory QC first to look at PCA and sample clustering, then Run DEG Analysis. If any group has fewer than two samples, TransXplorer shows a warning that statistical power may be low. For a guided tour of the results screens, see Your First RNA-Seq Analysis.

Section 6

Spot the mistake: six common upload problems

Each one has turned up in real datasets. Most of them load without any error.

#1 gene_idS1S2 TP5315201488 1-Mar233241 2-Sep8795 was MARCH1 · was SEPT2

Excel turned genes into dates

Open a gene list in Excel and symbols such as MARCH1 and SEPT2 silently become 1-Mar and 2-Sep. Ziemann et al. (2016) found this in about a fifth of papers with Excel gene lists, and a 2021 follow-up found the problem had grown. HGNC renamed these genes (MARCHF1, SEPTIN2) partly for this reason, but older annotations still use the old symbols.

In TransXplorerNothing flags it. Once a name has become a date, the original gene can't be recovered from the file.
FixUse Ensembl IDs, or import with the gene column set to Text. Search the file for -Mar, -Sep and -Dec before uploading.
#2 gene_idS1S2 TP5315201488 TP533741 MYC812760 same ID twice → averaged on load

Duplicate gene IDs

Gene symbols aren't unique: some map to several loci, and merged or re-annotated tables can repeat a gene. A count matrix needs exactly one row per ID.

In TransXplorerDuplicate IDs are collapsed into one row by averaging their values on upload, so two different loci become one blended gene.
FixDecide yourself before uploading: switch to Ensembl IDs, or keep the row you trust (often the one with the most reads).
#3 gene_idS1S2 TP5312.4112.38 MYC6.217.44 IL6−3.12−1.37 decimals + negatives = not counts

Normalised values instead of counts

TPM, FPKM, log-CPM, VST: all useful for plotting, none suitable as input to DESeq2, edgeR or limma-voom. This is the most common problem with files downloaded from GEO supplements.

In TransXplorerNo error appears. Values are rounded to whole numbers and negatives set to zero before DE (fig 2), so the results look normal but aren't valid.
FixTrack down the raw count file (Section 3), or fetch the dataset through GEO in TransXplorer.
#4 HEADER SAMPLE KO-1 KO-2 KO_1 KO_2 ✗ ✗ matched: 0 / 2 samples

Sample names don't match

Retyped names, a hyphen that became an underscore, a stray space, KO_1 versus ko_1. Any difference means the metadata row can't be linked to its column (fig 3).

In TransXplorerThe metadata step shows Matched: x / y samples and lists unmatched names. If nothing lines up you'll see No Matching Samples; fewer than two matches stops validation.
FixCopy the header row from the count file and paste it, transposed, into the Sample column.
#5 gene_idControlTreated TP5315202210 MYC8121975 n = 1 n = 1 no replicates → no variance estimate

Too few replicates

DE tests compare the difference between groups with the variation within them. With one sample per group there is no within-group variation to measure. Replicates must be biological (separate animals, donors or cultures), not the same library sequenced twice.

In TransXplorerYou'll see a warning: At least one group has fewer than 2 replicates. Statistical power may be low.
FixPlan for at least three biological replicates per group. Schurch et al. (2016) recommend at least six when you need to catch most changes.
#6 GeneidChrLengthS1S2 TP53chr17251215201488 MYCchr85364812760 IL6chr7119703 Length is numeric → read as a sample

Leftover annotation columns

featureCounts writes Chr, Start, End, Strand and Length before the counts. Other tools add a gene name or description next to the ID. HTSeq-count appends summary rows such as __no_feature.

In TransXplorerEvery column except the gene ID column is treated as a sample, so a numeric Length column turns into a fake sample.
FixKeep only the gene ID column and the count columns, and delete any summary rows before uploading.
Section 7

Pre-upload checklist

Two minutes here saves an afternoon of puzzling over results.

Tick each item as you check your files. Nothing is saved or sent anywhere; the boxes reset when you reload the page. It prints cleanly if you'd like a paper copy at the bench.

Ready to upload?
0 / 10
Work through the list, then head to TransXplorer.
Section 8

What’s next

Other ways in, and what to read once your data is loaded.

Formatting your own counts is one of four ways into TransXplorer. If your data is already public, or you're starting from raw reads, one of the other tutorials may save you the work.

Tutorials
Concepts

Files ready? Upload them now.

Open TransXplorer, go to Transcriptome Analysis and click Upload Count Matrix. Not sure your format is right? Click Load Example Counts first and compare.

Launch TransXplorer →

If TransXplorer helps your research, please cite us

The preprint is open access on bioRxiv. A peer-reviewed version is in submission.

Verma VM, Oler E, Syed H, Han S, Berjanskii M, Mason AL, Wishart DS, Wong GK. TransXplorer: An automated translational discovery platform for RNA-seq data. bioRxiv. 2026. doi:10.64898/2026.05.15.724657
Read the preprint →
References

Sources

Everything cited above, with DOIs.

  1. Ziemann M, Eren Y, El-Osta A. Gene name errors are widespread in the scientific literature. Genome Biology. 2016;17:177. doi:10.1186/s13059-016-1044-7
  2. Abeysooriya M, Soria M, Kasu MS, Ziemann M. Gene name errors: lessons not learned. PLoS Computational Biology. 2021;17(7):e1008984. doi:10.1371/journal.pcbi.1008984
  3. Bruford EA, Braschi B, Denny P, Jones TEM, Seal RL, Tweedie S. Guidelines for human gene nomenclature. Nature Genetics. 2020;52:754–758. doi:10.1038/s41588-020-0669-3
  4. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology. 2014;15:550. doi:10.1186/s13059-014-0550-8
  5. Robinson MD, McCarthy DJ, Smyth GK. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010;26(1):139–140. doi:10.1093/bioinformatics/btp616
  6. Wagner GP, Kin K, Lynch VJ. Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples. Theory in Biosciences. 2012;131:281–285. doi:10.1007/s12064-012-0162-3
  7. Soneson C, Love MI, Robinson MD. Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences. F1000Research. 2015;4:1521. doi:10.12688/f1000research.7563.1
  8. Schurch NJ, Schofield P, Gierliński M, et al. How many biological replicates are needed in an RNA-seq experiment and which differential expression tool should you use? RNA. 2016;22(6):839–851. doi:10.1261/rna.053959.115

Frequently asked questions

Can I upload TPM or FPKM values?
Not for differential expression. DESeq2, edgeR and limma-voom model raw read counts, and TPM/FPKM have already been rescaled so the number of reads actually observed is lost. TransXplorer rounds values to whole numbers before DE, so TPM won't trigger an error, but the results won't be valid. Go back to the raw count file.
My Salmon or RSEM counts have decimals. Is that a problem?
No. Estimated counts from Salmon, kallisto or RSEM are on the read-count scale even though they're fractional, and they're fine to upload. TransXplorer rounds them before the DE step. Use the estimated-count column (NumReads for Salmon), not the TPM column.
Do I need a separate metadata file?
Not always. If sample names end with the group after the last underscore (S1_Control, S2_KO), TransXplorer can detect groups from the names, and you can also assign groups by hand in the Sample Group Assignment dialog. A sample sheet is the most reliable route, and you need one if you want to supply batch information.
Which gene identifiers can I use?
Gene symbols (TP53), Ensembl gene IDs (ENSG00000141510), versioned Ensembl IDs (ENSG00000141510.17) and Entrez IDs (7157). When TransXplorer maps IDs to symbols for downstream tools it strips the version suffix first. Use one ID type for the whole file.
My organism isn't in the list. Can I still use TransXplorer?
Yes. Choose Other in the organism menu. Differential expression, QC and WGCNA work for any organism. Pathway enrichment and protein networks need a KEGG organism code or NCBI taxonomy ID, which you can enter in the same panel. Drug targets and cell-type deconvolution are human-only.
How large can my count file be?
Up to 500 MB. A typical gene-level matrix of 20,000–60,000 genes and a few hundred samples is far smaller than that.