Tutorial · 25 min read Intermediate No code required

Your First FASTQ to Results Workflow

Your sequencing facility sent you a folder of .fastq.gz files. This guide explains what is inside them, what happens to them in TransXplorer, and how to turn them into a count matrix you can analyse.

Up to 15 GB on the server HISAT2 or Salmon Live progress log
TL;DR · The short version
  • 1 A FASTQ file is millions of 4-line records: a read name, the bases, a +, and a quality character for every base. Paired-end runs give you two files per sample, _R1 and _R2.
  • 2 The FASTQ Processing tab runs FastQC → Trimmomatic → FastQC again, then either aligns reads to the genome (HISAT2 + featureCounts) or pseudo-aligns them to transcripts (Salmon + tximport).
  • 3 Both routes end in the same thing: a gene × sample count matrix. Download it and upload it as-is in Transcriptome Analysis to run differential expression.
  • 4 The server takes up to 15 GB of FASTQ per run (typically 5–10 samples). Every organism in the genome menu runs on the server with HISAT2; bigger studies run on your own machine through the Docker option.
  1. Anatomy of a FASTQ file
  2. Single- vs paired-end
  3. The pipeline at a glance
  4. Alignment vs pseudo-alignment
  5. Running it in TransXplorer
  6. Reading FastQC reports
  7. What you get back
  8. From counts to DE
  9. Troubleshooting
Section 1

What’s actually inside a FASTQ file

Four lines per read, repeated tens of millions of times.

A sequencer doesn’t hand you genes. It hands you reads: short strings of A, C, G and T, each one copied from a random fragment of the cDNA in your library. A typical bulk RNA-seq sample has 20–50 million of them, and every one is stored in the same four-line format.

The format, FASTQ, was never formally standardised by one company. It grew out of the Sanger Institute and was later documented carefully by Cock et al. (2010), who also untangled the incompatible early Illumina variants. Today almost every instrument writes the same flavour: plain text, usually gzip-compressed (that’s the .gz), with qualities encoded as Phred+33. TransXplorer reads .fastq, .fq, .fastq.gz and .fq.gz, and you should upload the compressed files as they came — they are four to five times smaller.

Why the quality line matters

Phred scores are the sequencer’s own confidence in each base call. They tend to be high at the start of a read and drift down toward the end as the chemistry tires, and they collapse to # (Q2) when the instrument effectively gives up. That’s why the first thing the pipeline does is look at those scores (FastQC), and the second is to cut off the parts that fall below a threshold (Trimmomatic). Low-quality tails don’t just add noise; mismatches near the end of a read can stop it aligning at all.

Why this matters

Newer instruments (NovaSeq, NextSeq 2000) “bin” their quality scores into just a few values, which is why the example above only uses F, :, , and #. Don’t be alarmed if your quality strings look repetitive. It’s a storage trick, not a fault.

Section 2

Single-end vs paired-end: one file or two?

Same fragment, read from one end or from both.

Before sequencing, your RNA is converted to cDNA, broken into fragments a few hundred bases long, and given short synthetic adapters at both ends. In a single-end run the sequencer reads each fragment from one end only. In a paired-end run it reads the same fragment again from the opposite end, and the two reads go into two separate files: everything from the first end in _R1, everything from the second in _R2, in exactly the same order.

Paired-end data is more informative: two anchored ends make it much easier to place a fragment uniquely and to tell which isoform it came from. Single-end data is cheaper and perfectly adequate for gene-level differential expression. TransXplorer handles both; you just tell it which you have.

Naming your files so TransXplorer can pair them

For paired-end runs, TransXplorer checks the names before it starts. Every file must contain _R1 or _R2 immediately before the extension, there must be as many R1 files as R2 files, and the total must be even. Internally the pipeline sorts the R1 list and the R2 list and pairs them in order, so the part of the name before _R1/_R2 should be identical for both mates. That prefix becomes the sample name in your count matrix.

Will pair correctly

WT_rep1_R1.fastq.gz WT_rep1_R2.fastq.gz KO_rep1_R1.fastq.gz KO_rep1_R2.fastq.gz Sample columns will be called WT_rep1 and KO_rep1. Illumina names such as WT1_S1_L001_R1_001.fastq.gz work as they are; the sample becomes everything before _R1/_R1_001. A group_replicate pattern also makes the later group-assignment step easier.

Will be rejected or mis-paired

WT_rep1_R1.fastq.gz wt_rep1_R2.fastq.gz WT_rep1.R1.fq.gz Mismatched prefixes (WT_ vs wt_) don’t pair: both mates need an identical prefix. Use underscores, not dots, before R1/R2.
Lanes split across files?

Some facilities deliver one sample as several files per lane (_L001, _L002…). Join the lanes of each read before uploading, R1 lanes with R1 lanes and R2 with R2, in the same lane order for both. On macOS or Linux, gzip files can simply be concatenated: cat WT1_L001_R1_001.fastq.gz WT1_L002_R1_001.fastq.gz > WT1_R1.fastq.gz.

Section 3

The pipeline at a glance

Clean the reads, then count them one of two ways.

Whatever you choose later, every run starts with the same three clean-up stages. Then the path forks at the setting called Quantification Method:. Both branches rejoin as a single table of counts. The percentages in the diagram are the values the progress bar in the app shows at each stage, so you can tell exactly where a run is.

Stage by stage

  • 1. Raw-read QCFastQC Scans every uploaded file and writes an HTML report: per-base quality, GC content, duplication, adapter content and more. Nothing is changed; this is the “before” picture. Section 6 shows how to read it.
  • 2. TrimmingTrimmomatic Clips Illumina TruSeq adapter sequence (ILLUMINACLIP), shaves bases below quality 3 from both ends (LEADING:3, TRAILING:3), cuts the read once the average quality in a 4-base window drops below 15 (SLIDINGWINDOW:4:15), and throws away anything shorter than 36 bases (MINLEN:36). For paired-end data, only pairs where both mates survive go forward.
  • 3. Post-trim QCFastQC The same report on the trimmed reads, so you can confirm adapters are gone and the low-quality tails have been removed. This is the “after” picture.
  • 4a. Align + countHISAT2 → featureCounts HISAT2 places each read on the reference genome, allowing it to jump across introns. samtools sorts and indexes the alignments (BAM files), and featureCounts counts how many reads (or, for paired-end, how many fragments) overlap each gene in the annotation.
  • 4b. Pseudo-align + summariseSalmon → tximport Salmon works out which transcripts each read is compatible with and estimates transcript abundance, correcting for GC-content and sequence-specific bias (--gcBias --seqBias) and auto-detecting the library type (-l A). tximport then adds up the transcript estimates of each gene into gene-level counts.
Why this matters

Trimming parameters are a methods-section detail reviewers do ask about. The settings above are the widely used defaults from the Trimmomatic manual (Bolger et al., 2014), applied to every sample identically. You can copy them straight into your methods.

Section 4

Alignment vs pseudo-alignment, explained visually

“Where exactly did this read come from?” vs “Which transcripts could it have come from?”

The fork in the pipeline is really a choice between two questions. Both answer “how much of each gene is there?”, but they get there differently, and the difference shows up in speed, in what you can check, and in how ambiguous reads are handled.

Alignment, in plain words

HISAT2 (Kim et al., 2019) takes every read and finds where it fits on the reference genome, base by base. Because mRNA has had its introns spliced out, a read can start at the end of one exon and finish at the start of the next; HISAT2 is splice-aware, so it can split a read across a gap of thousands of bases. The result is a BAM file of coordinates. featureCounts (Liao et al., 2014) then walks through those coordinates and asks, for each read or fragment, “which gene’s exons does this overlap?” In TransXplorer it counts without a strand restriction, counts paired-end fragments once rather than twice, and, as is featureCounts’ default, leaves out reads that map equally well to several places.

Pseudo-alignment, in plain words

Salmon (Patro et al., 2017) skips the question “where on the genome?” and asks only “which transcripts is this read consistent with?” It chops the read into short words (k-mers), looks them up in an index built from transcript sequences, and records the set of compatible transcripts. A statistical model then divides ambiguous reads among transcripts according to their overall abundance, while also correcting for fragment GC content and sequence-specific biases. Because it never builds full alignments, it is much faster; TransXplorer’s own label describes it as roughly ten times faster. Finally tximport (Soneson et al., 2015) sums transcripts into genes so the output looks exactly like the other route.

How to choose

For a standard gene-level differential expression study in human or mouse, both routes give you a perfectly good count matrix for DESeq2, edgeR or limma-voom. The table below is for when you want to pick deliberately.

If you care about… HISAT2 + featureCounts Salmon + tximport
Speed Slower — full spliced alignment, then sorting and indexing BAMs Faster — labelled “Fast” in the app
Auditing where reads went Yes — per-sample counts of reads assigned to genes vs. reads that hit no gene, multi-mapped, or were ambiguous Less direct — no genome alignment or featureCounts summary
Reads that fit several genes or isoforms Left out (multi-mappers) or reported as ambiguous Shared statistically across compatible transcripts
Bias correction None at the counting step GC and sequence-specific (--gcBias --seqBias)
Organisms on the server Every genome in the menu, including a custom genome + GTF Human (hg38), mouse (mm10), or a custom genome + GTF
Reads outside the annotation (new genes, intronic signal) Visible as unassigned reads in the summary Invisible — only annotated transcripts are in the index
Methods sentence “Reads were aligned with HISAT2 and gene-level counts obtained with featureCounts.” “Transcripts were quantified with Salmon and summarised to genes with tximport.”
Not sure? Start with the default, HISAT2 + featureCounts. Its assignment summary tells you whether your library is healthy, and that is worth more than speed on a first run.
Section 5

Running it in TransXplorer, step by step

One sidebar, top to bottom. Then watch the log.

Open TransXplorer and click the FASTQ Processing tab. A status banner at the top tells you whether the server is free (“System ready - you can start your analysis”). The left sidebar holds every setting, in the order you need them.

  1. Choose an Analysis Mode

    Analysis ModeRun on Server · Run Locally with Docker

    Two options, depending mostly on how much data you have:

    Default

    Run on Server

    • Upload straight from your browser
    • Maximum 15 GB total file size
    • Typically 5–10 samples
    • Both HISAT2 and Salmon available
    Large data

    Run Locally with Docker

    • No upload; your files stay on your machine
    • Unlimited dataset size
    • Pick Windows, macOS or Linux and follow four setup steps
    • HISAT2 + featureCounts, with references for every organism in the menu

    In Docker mode the app walks you through installing Docker, pulling the pipeline image, putting your files in an input_fastq/ folder, and running the script from Download Analysis Script (./run_rnaseq_*.sh; add a number, e.g. ./run_rnaseq_*.sh 16, to use more CPU threads). The rest of this tutorial follows the server route.

  2. Upload your FASTQ files

    Drop FASTQ files here or click to browse.fastq .fq .fastq.gz .fq.gz

    Select all files for the experiment at once (both mates of every pair). Large uploads take a while; until every file has arrived the start button stays grey and reads Files uploading…, with the reminder “Please wait for all files to finish uploading before starting analysis.”

    The Instructions panel just above (click the chevron to expand) repeats the essentials: accepted formats, the SampleName_R1/R2.fastq.gz convention, and the 15 GB server limit.

  3. Tell it what you have

    Sequencing TypeQuantification MethodReference Genome
    • Sequencing Type: Paired-End (default) or Single-End. Paired-End expects matched _R1/_R2 files.
    • Quantification Method: HISAT2 + featureCounts (tagged Standard, the default) or Salmon (tagged Fast). A one-line hint under the choice summarises each.
    • Reference Genome: the organism your reads come from. Default is Human (hg38). Choosing Custom Genome reveals uploads for a Genome FASTA File:, an Annotation GTF/GFF File: and a Genome Name:.

    The genome menu lists eleven options. Which of them can be processed where is summarised in the table below this list.

  4. Start the analysis

    Start Analysis (N files, X GB)

    Once uploads finish, the button turns blue and shows what it’s about to process, for example Start Analysis (8 files, 6.4 GB). Click it. TransXplorer first validates the inputs (file types, R1/R2 pairing, total size, custom genome files) and stops with a clear message if anything is off; see Troubleshooting. If all is well, the button switches to Processing… and the pipeline runs in a separate background process.

  5. Watch the Processing Log

    Processing LogReal-time Processing Log

    The Processing Log tab shows a progress bar with the current stage and a live, colour-coded terminal: [STEP] marks a stage starting, [DONE] (green) a stage or sample finishing, [WARN] something non-fatal, and [ERROR] (red) a failure. Because every sample is logged by name (Trimmed pair: WT_rep1, Aligned: WT_rep1…), you can see exactly how far a long run has got.

    Run time depends on read depth, the number of samples and how busy the server is. Keep the browser tab open while it runs, and download your results before you close it.

Which genomes work where

Reference Genome (menu label) Server · HISAT2 Server · Salmon Docker
Human (hg38)YesYesYes
Mouse (mm10)YesYesYes
Rat (rn6)YesNoYes
Drosophila (dm6) · Zebrafish (danRer11) · C. elegans (wbcel235) · Yeast (r64) · Arabidopsis (araTha) · Chicken (galGal6) · Pig (susScr11)YesNoYes
Custom Genome (your FASTA + GTF)YesYesYes

On the server, each genome is counted against an annotation from the same assembly: GENCODE release 47 for hg38, GENCODE vM25 for mm10 (the last release on GRCm38), Ensembl 104 for rn6 (the last on Rnor_6.0), Ensembl BDGP6.46 for dm6, and Ensembl or index-bundled GTFs for the rest. Chromosome naming differences (chr2L vs 2L) are harmonised automatically, and the log prints the pairing it used, e.g. Reference: dm6 | index: … | annotation: Drosophila_melanogaster.BDGP6.46.111.gtf. If an index or matching annotation is missing, the run stops with a clear error rather than borrowing another organism’s.

For a custom genome with HISAT2, the server builds a HISAT2 index from your FASTA (hisat2-build) and counts against your GTF; large genomes can take a long time to index. With Salmon, it first extracts transcript sequences from your FASTA using your GTF (with gffread) and builds an index from them. In Docker mode, put genome.fa and annotation.gtf in input_fastq/custom_genome/ and the script builds a HISAT2 index for you.

When the server is busy

FASTQ processing is heavy, so the server runs at most two jobs at a time. If both slots are taken when you click Start, you’ll get an Analysis Queued notification with your position in the queue, the number of active jobs, an estimated wait and a job ID, and the banner at the top of the tab shows your position. The app checks for a free slot about every 10 seconds and starts your job automatically when one opens; the log then reads A processing slot is free - starting your queued analysis. Keep the browser tab open while you wait: closing it cancels your queued or running job and frees its slot. Otherwise the banner switches to “Your analysis is currently processing…” with the minutes elapsed.

Why this matters

Choosing the wrong reference genome produces no error at all. Reads from a mouse experiment will still partly align to the human genome; you’ll just get a poor assignment rate and misleading counts. Check the Reference Genome: menu twice before you press Start.

Section 6

Reading your FastQC reports

What’s fine, what’s alarming, and what only looks alarming.

When the run finishes, open the QC Reports tab. A side panel lets you pick a Sample: and a QC Stage: (Pre-Trimming or Post-Trimming); the full FastQC report then opens right inside the page, and Download Report saves it. Flip between the two stages for the same sample. That before/after comparison is the most useful thing on the tab.

FastQC gives each module a pass, warn or fail flag. Treat those as prompts to look, not verdicts. FastQC’s thresholds were designed for random genomic DNA, and RNA-seq libraries break several of its assumptions by design.

Act on it

Per base sequence quality

Medians in the red across much of the read, in raw and trimmed data, mean a poor run. Talk to your sequencing facility before analysing further.

Check post-trim

Adapter content

Rising adapter curves in the raw report are common with short fragments. After trimming they should be flat near zero. If not, the library may use adapters other than TruSeq.

Usually fine

Sequence duplication levels

Often flagged red in RNA-seq. Highly expressed genes are sequenced over and over, so identical reads are expected. This reflects biology, not PCR failure.

Usually fine

Per base sequence content

Wobbly A/C/G/T lines in the first 10–15 bases come from random-hexamer priming during library prep. A known, harmless RNA-seq signature.

Look closer

Per sequence GC content

A shifted or two-humped curve can mean contamination (bacteria, another species) or abundant rRNA. Compare against other samples in the same run.

Look closer

Overrepresented sequences

Adapter dimers, rRNA or mitochondrial transcripts often top this list. A few hits are normal; one sequence making up several percent of reads is not.

Why this matters

The most informative QC is comparative. One sample whose report looks very different from its replicates is far more worrying than every sample sharing the same “failed” duplication module. Outlier samples found here often reappear later as outliers on the PCA plot in the DE workflow.

Section 7

What you get back

Three result tabs and two downloads.

When the run completes you’ll see “Analysis completed successfully!” and the result tabs fill in:

  • Feature CountsGene Expression Count Matrix A searchable table of the count matrix: one row per gene, one column per sample, with the gene ID in the first column. Glance at it to confirm every sample is present and the numbers are whole-number counts.
  • QC SummaryQuality Control Summary The number of genes and samples, the sequencing type and genome you chose, and, for the HISAT2 route, an Alignment Statistics table from featureCounts showing, per sample, how many reads were Assigned to genes and how many fell into the Unassigned categories.
  • Analysis InfoAbout This Analysis A reminder of the tools and outputs, and the next step: take the count matrix to the Transcriptome Analysis section.

The downloads

The Download Results panel at the bottom of the sidebar has two buttons:

  • Count Matrix (CSV)gene_counts_<date>.csv The file you’ll analyse. The first column is Gene_ID (Ensembl gene IDs from the annotation). It is followed by one column per sample, for both HISAT2 and Salmon, so the file uploads straight into Transcriptome Analysis.
  • Full Results (ZIP)qc_bundle_<date>.zip Everything you need for your records and supplementary files, for both quantification methods: the counts, every FastQC report and a one-page HTML summary. HISAT2 runs add the featureCounts alignment summary; Salmon runs add a salmon_tpm.txt table, and their HTML summary includes a Salmon Mapping Statistics table (reads processed, reads mapped and mapping rate per sample).
qc_bundle_2026-10-05.zip ├── fastqc_raw/ FastQC reports before trimming (HTML + zip) ├── fastqc_trimmed/ FastQC reports after trimming ├── gene_counts.csv the count matrix (Gene_ID + one column per sample) ├── alignment_summary.txt HISAT2 runs: featureCounts Assigned / Unassigned per sample ├── salmon_tpm.txt Salmon runs: gene-level TPM └── qc_summary.html genes detected, assignment or mapping rate per sample, report links
Why this matters

qc_summary.html includes an Assignment Rate per sample for HISAT2 runs (the share of reads featureCounts placed on a gene) or a Mapping Rate for Salmon runs. It is the single quickest health check for an RNA-seq library. Replicates should have similar rates; one sample far below the rest deserves a look before it goes into a DE model. Keep the whole ZIP: it is a ready-made supplementary QC record.

Section 8

From counts to differential expression

Same app, next tab: the standard DE workflow takes over.

The count matrix is the hand-off point. From here on the analysis is identical to starting from a published counts file, and everything in the first-analysis walkthrough applies.

  1. Download the CSV

    Count Matrix (CSV)

    Click Count Matrix (CSV). The file holds only Gene_ID and your sample columns, whichever quantification method you used, so there is nothing to tidy.

    Why it matters: the uploader treats the gene column as row names and every other column as a sample. If you add notes or extra columns in a spreadsheet, they will show up as “samples” in the next dialog.

  2. Upload it to Transcriptome Analysis

    Transcriptome Analysis1. Data Input & SetupUpload Count Matrix (CSV/XLSX/TSV/TXT)

    Switch to the Transcriptome Analysis tab and use Upload Count Matrix (CSV/XLSX/TSV/TXT) at the top of the sidebar. The Gene_ID column is recognised automatically, and Ensembl IDs are mapped to gene symbols for the steps that need symbols.

  3. Confirm sample groups

    Sample Group Assignment

    A Sample Group Assignment dialog opens with your sample names (the FASTQ prefixes, e.g. WT_rep1, KO_rep1). TransXplorer tries to detect the groups from the names; check them and confirm.

  4. Add metadata if you have batches

    Batch Effects & Covariates DetectionUpload Batch Metadata File:

    If samples were prepared or sequenced in different batches, upload a metadata file whose Sample column matches the count-matrix column names exactly, plus columns such as sequencing_run or prep_date. Automatic Detection will test whether those batches matter; Understanding Batch Effects explains how.

  5. Configure and run

    Run DEG Analysis

    Pick DESeq2, edgeR or limma-voom, set reference and treatment groups and thresholds, and click Run DEG Analysis. Both FASTQ routes produce raw integer counts, which is exactly what these methods expect; What is Differential Expression? covers the choice.

Don’t mix routes in one analysis

If you process some samples with HISAT2 and others with Salmon (or on different days with different genomes), the method itself becomes a batch effect. Process every sample of an experiment in the same run, or at least with identical settings.

Section 9

Troubleshooting

The messages you might see, and what they usually mean.

Problems show up in two places: as a red notification (“Analysis failed: …”) and as an [ERROR] Pipeline error: … line in the Processing Log. The pipeline stops at the first failing step, so the message tells you both which tool and which sample to look at.

validation › Paired-end files must follow naming pattern: SampleName_R1.fastq.gz and SampleName_R2.fastq.gz (Illumina _R1_001 / _R2_001 names are also accepted)
validation › These samples are missing their R1 or R2 partner: <names>
validation › Paired-end sequencing requires an even number of files

File names or counts don’t pair up

Every paired-end file must have _R1 or _R2 (or Illumina’s _R1_001/_R2_001) just before the extension, and every R1 needs an R2 with the same sample prefix.

  • If the message lists sample names, those samples have only one mate: check you didn’t miss a file when selecting, and that both mates share an identical prefix.
  • Use an underscore, not a dot, before R1/R2 (X.R1.fq.gz is not recognised).
  • If the data is really single-end, switch Sequencing Type: to Single-End.
validation › Total file size exceeds 15GB limit. Please reduce the number of samples or use compressed FASTQ files.

Too much data for the server

  • Make sure you are uploading the .gz files, not decompressed copies.
  • For bigger studies, switch to Run Locally with Docker, which has no size limit.
  • Avoid splitting one experiment into several server runs unless the settings are identical; see the warning in Section 8.
[ERROR] HISAT2 alignment failed for sample 'KO_rep2' (exit code 1)
pattern › <step> failed for sample '<name>' (exit code N)

A tool failed on one sample

The step can be HISAT2 alignment, samtools sort, samtools index or Salmon quantification; a counting failure reads featureCounts quantification failed (exit code N). The exit code is the tool’s own; the sample name is what matters.

  • Truncated or corrupted upload: the most common cause. Re-check the file against the facility’s checksum (md5sum), or test it with gzip -t file.fastq.gz, then upload again.
  • Mates that don’t match: R1 and R2 with different read counts or read names (see below).
  • Exit code 137: usually means the process was stopped for running out of memory. Try again later when the server is quieter, or use Docker.
symptom › alignment fails, or one sample has a strikingly low assignment rate

Mismatched R1/R2 files

Because R1 and R2 are paired by sorted name, a prefix that differs between mates (WT_rep1_R1 with wt_rep1_R2), or a file from the wrong sample, can pair reads that don’t belong together. Mates must also contain the same number of reads in the same order.

  • Use identical prefixes for both mates of every sample.
  • Compare read counts: on macOS/Linux, echo $(( $(gzip -dc X_R1.fastq.gz | wc -l) / 4 )) should give the same number for R1 and R2.
  • Never trim, filter or subsample one mate without the other.
[ERROR] Salmon transcriptome reference not available for genome: rn6 . Use HISAT2 mode or select hg38/mm10.

Salmon picked with an organism it can’t serve

On the server, Salmon covers human (hg38), mouse (mm10) and custom genomes. For any other organism in the menu, switch Quantification Method: to HISAT2 + featureCounts, which runs on the server for every genome.

symptom › run completes, but Assigned is low and Unassigned_* is high in the QC Summary

Low mapping or assignment rate

No error is raised, which is why it’s worth checking the QC Summary and qc_summary.html every time. As a rough rule, a healthy poly-A library assigns most of its reads to genes; replicates far below their siblings, or a whole run well under half, need explaining. The usual suspects:

  • Wrong reference genome: reads from one species aligned to another. Recheck Reference Genome:.
  • Lots of reads outside genes (high Unassigned_NoFeatures): genomic DNA contamination, or total-RNA libraries with many intronic pre-mRNA reads.
  • Ribosomal RNA: incomplete rRNA depletion. Look for rRNA in FastQC’s overrepresented sequences and an odd GC curve.
  • Multi-mapping reads (high Unassigned_MultiMapping): repetitive sequence or gene families; featureCounts leaves these out. Salmon handles them better.
  • Very short inserts or adapter dimers: many reads drop below 36 bases after trimming and are discarded.
  • Degraded RNA: coverage piles up at the 3′ ends of genes; check RIN values with the facility.
Section 10

What’s next

You have counts. Here’s where to take them.

Continue the workflow

Understand the methods downstream

Ready to process your own reads?

Open TransXplorer, click FASTQ Processing, and drop in your _R1/_R2 files. Free, in the browser, nothing to install for server mode.

Launch TransXplorer →

References

The tools in this pipeline are peer-reviewed. Cite them in your methods alongside TransXplorer.

  1. Cock PJA, Fields CJ, Goto N, Heuer ML, Rice PM. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Res. 2010;38(6):1767–1771. doi:10.1093/nar/gkp1137
  2. Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30(15):2114–2120. doi:10.1093/bioinformatics/btu170
  3. Kim D, Paggi JM, Park C, Bennett C, Salzberg SL. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019;37(8):907–915. doi:10.1038/s41587-019-0201-4
  4. Liao Y, Smyth GK, Shi W. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;30(7):923–930. doi:10.1093/bioinformatics/btt656
  5. Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods. 2017;14(4):417–419. doi:10.1038/nmeth.4197
  6. Soneson C, Love MI, Robinson MD. Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences. F1000Research. 2015;4:1521. doi:10.12688/f1000research.7563.2

FastQC is software from the Babraham Institute; cite it by its project page.

If TransXplorer helps your research, please cite us

The preprint is open access on bioRxiv. A peer-reviewed version is in submission.

Verma VM, Oler E, Syed H, Han S, Berjanskii M, Mason AL, Wishart DS, Wong GK. TransXplorer: An automated translational discovery platform for RNA-seq data. bioRxiv. 2026. doi:10.64898/2026.05.15.724657
Read the preprint →

Frequently asked questions

How much FASTQ data can I process on the server?
Up to 15 GB of FASTQ in total per run, which typically covers 5–10 samples. For larger studies, choose Run Locally with Docker: TransXplorer gives you a script that runs a HISAT2 + featureCounts pipeline on your own computer with no size limit. Organism is not a reason to switch: every genome in the menu, including a custom genome, runs with HISAT2 + featureCounts on the server.
Should I choose HISAT2 + featureCounts or Salmon?
Both give a gene-level count matrix suitable for DESeq2, edgeR or limma-voom. HISAT2 + featureCounts reports how many reads landed on genes, which is great for QC. Salmon is faster, corrects GC and sequence bias, and shares ambiguous reads statistically. On the server Salmon covers human (hg38), mouse (mm10) and custom references, while HISAT2 + featureCounts covers every genome in the menu. See How to choose.
How should I name paired-end files?
SampleName_R1.fastq.gz and SampleName_R2.fastq.gz, with an identical prefix for both mates. Files straight off an Illumina instrument, such as SampleName_S1_L001_R1_001.fastq.gz, are accepted as they are; the sample name becomes everything before _R1 or _R1_001.
FastQC failed my duplication and per-base content modules. Is my data bad?
Usually not. FastQC’s thresholds assume random genomic DNA. In RNA-seq, highly expressed genes make identical reads common, and random-hexamer priming skews the first 10–15 bases. Act on low per-base quality, or on adapter content that persists after trimming.
Can I use the count matrix straight away for differential expression?
Yes. Download Count Matrix (CSV), which contains only Gene_ID and one column per sample for both HISAT2 and Salmon, then upload it unchanged with Upload Count Matrix in the Transcriptome Analysis tab. See Section 8.
Can I cite TransXplorer in a paper?
Yes, please do. Preprint: doi:10.64898/2026.05.15.724657. Please also cite the underlying tools listed in the References.