Your First FASTQ to Results Workflow
Your sequencing facility sent you a folder of .fastq.gz files. This guide explains what is inside them, what happens to them in TransXplorer, and how to turn them into a count matrix you can analyse.
-
1
A FASTQ file is millions of 4-line records: a read name, the bases, a
+, and a quality character for every base. Paired-end runs give you two files per sample,_R1and_R2. -
2
The FASTQ Processing tab runs
FastQC→Trimmomatic→FastQCagain, then either aligns reads to the genome (HISAT2+featureCounts) or pseudo-aligns them to transcripts (Salmon+tximport). - 3 Both routes end in the same thing: a gene × sample count matrix. Download it and upload it as-is in Transcriptome Analysis to run differential expression.
- 4 The server takes up to 15 GB of FASTQ per run (typically 5–10 samples). Every organism in the genome menu runs on the server with HISAT2; bigger studies run on your own machine through the Docker option.
What’s actually inside a FASTQ file
Four lines per read, repeated tens of millions of times.
A sequencer doesn’t hand you genes. It hands you reads: short strings of A, C, G and T, each one copied from a random fragment of the cDNA in your library. A typical bulk RNA-seq sample has 20–50 million of them, and every one is stored in the same four-line format.
The format, FASTQ, was never formally standardised by one company. It grew out of the Sanger Institute and was later documented carefully by Cock et al. (2010), who also untangled the incompatible early Illumina variants. Today almost every instrument writes the same flavour: plain text, usually gzip-compressed (that’s the .gz), with qualities encoded as Phred+33. TransXplorer reads .fastq, .fq, .fastq.gz and .fq.gz, and you should upload the compressed files as they came — they are four to five times smaller.
Why the quality line matters
Phred scores are the sequencer’s own confidence in each base call. They tend to be high at the start of a read and drift down toward the end as the chemistry tires, and they collapse to # (Q2) when the instrument effectively gives up. That’s why the first thing the pipeline does is look at those scores (FastQC), and the second is to cut off the parts that fall below a threshold (Trimmomatic). Low-quality tails don’t just add noise; mismatches near the end of a read can stop it aligning at all.
Newer instruments (NovaSeq, NextSeq 2000) “bin” their quality scores into just a few values, which is why the example above only uses F, :, , and #. Don’t be alarmed if your quality strings look repetitive. It’s a storage trick, not a fault.
Single-end vs paired-end: one file or two?
Same fragment, read from one end or from both.
Before sequencing, your RNA is converted to cDNA, broken into fragments a few hundred bases long, and given short synthetic adapters at both ends. In a single-end run the sequencer reads each fragment from one end only. In a paired-end run it reads the same fragment again from the opposite end, and the two reads go into two separate files: everything from the first end in _R1, everything from the second in _R2, in exactly the same order.
_R1 and record number 1 in _R2 are the two ends of the same fragment. Software relies on that order, so the two files of a pair must always travel together and never be re-sorted, filtered, or concatenated separately.
Paired-end data is more informative: two anchored ends make it much easier to place a fragment uniquely and to tell which isoform it came from. Single-end data is cheaper and perfectly adequate for gene-level differential expression. TransXplorer handles both; you just tell it which you have.
Naming your files so TransXplorer can pair them
For paired-end runs, TransXplorer checks the names before it starts. Every file must contain _R1 or _R2 immediately before the extension, there must be as many R1 files as R2 files, and the total must be even. Internally the pipeline sorts the R1 list and the R2 list and pairs them in order, so the part of the name before _R1/_R2 should be identical for both mates. That prefix becomes the sample name in your count matrix.
Will pair correctly
WT_rep1_R1.fastq.gz
WT_rep1_R2.fastq.gz
KO_rep1_R1.fastq.gz
KO_rep1_R2.fastq.gz
Sample columns will be called WT_rep1 and KO_rep1. Illumina names such as WT1_S1_L001_R1_001.fastq.gz work as they are; the sample becomes everything before _R1/_R1_001. A group_replicate pattern also makes the later group-assignment step easier.
Will be rejected or mis-paired
WT_rep1_R1.fastq.gz
wt_rep1_R2.fastq.gz
WT_rep1.R1.fq.gz
Mismatched prefixes (WT_ vs wt_) don’t pair: both mates need an identical prefix. Use underscores, not dots, before R1/R2.
Some facilities deliver one sample as several files per lane (_L001, _L002…). Join the lanes of each read before uploading, R1 lanes with R1 lanes and R2 with R2, in the same lane order for both. On macOS or Linux, gzip files can simply be concatenated: cat WT1_L001_R1_001.fastq.gz WT1_L002_R1_001.fastq.gz > WT1_R1.fastq.gz.
The pipeline at a glance
Clean the reads, then count them one of two ways.
Whatever you choose later, every run starts with the same three clean-up stages. Then the path forks at the setting called Quantification Method:. Both branches rejoin as a single table of counts. The percentages in the diagram are the values the progress bar in the app shows at each stage, so you can tell exactly where a run is.
Using N threads per process), while Trimmomatic uses 4.
Stage by stage
-
1. Raw-read QC
FastQCScans every uploaded file and writes an HTML report: per-base quality, GC content, duplication, adapter content and more. Nothing is changed; this is the “before” picture. Section 6 shows how to read it. -
2. Trimming
TrimmomaticClips Illumina TruSeq adapter sequence (ILLUMINACLIP), shaves bases below quality 3 from both ends (LEADING:3,TRAILING:3), cuts the read once the average quality in a 4-base window drops below 15 (SLIDINGWINDOW:4:15), and throws away anything shorter than 36 bases (MINLEN:36). For paired-end data, only pairs where both mates survive go forward. -
3. Post-trim QC
FastQCThe same report on the trimmed reads, so you can confirm adapters are gone and the low-quality tails have been removed. This is the “after” picture. -
4a. Align + count
HISAT2 → featureCountsHISAT2places each read on the reference genome, allowing it to jump across introns.samtoolssorts and indexes the alignments (BAM files), andfeatureCountscounts how many reads (or, for paired-end, how many fragments) overlap each gene in the annotation. -
4b. Pseudo-align + summarise
Salmon → tximportSalmonworks out which transcripts each read is compatible with and estimates transcript abundance, correcting for GC-content and sequence-specific bias (--gcBias --seqBias) and auto-detecting the library type (-l A).tximportthen adds up the transcript estimates of each gene into gene-level counts.
Trimming parameters are a methods-section detail reviewers do ask about. The settings above are the widely used defaults from the Trimmomatic manual (Bolger et al., 2014), applied to every sample identically. You can copy them straight into your methods.
Alignment vs pseudo-alignment, explained visually
“Where exactly did this read come from?” vs “Which transcripts could it have come from?”
The fork in the pipeline is really a choice between two questions. Both answer “how much of each gene is there?”, but they get there differently, and the difference shows up in speed, in what you can check, and in how ambiguous reads are handled.
N in the CIGAR string is the skipped intron) and then counted against the gene annotation. On the right, the read is never placed on the genome at all; its k-mers simply reveal that it spans the E2–E3 junction, which only T1 and T3 contain.
Alignment, in plain words
HISAT2 (Kim et al., 2019) takes every read and finds where it fits on the reference genome, base by base. Because mRNA has had its introns spliced out, a read can start at the end of one exon and finish at the start of the next; HISAT2 is splice-aware, so it can split a read across a gap of thousands of bases. The result is a BAM file of coordinates. featureCounts (Liao et al., 2014) then walks through those coordinates and asks, for each read or fragment, “which gene’s exons does this overlap?” In TransXplorer it counts without a strand restriction, counts paired-end fragments once rather than twice, and, as is featureCounts’ default, leaves out reads that map equally well to several places.
Pseudo-alignment, in plain words
Salmon (Patro et al., 2017) skips the question “where on the genome?” and asks only “which transcripts is this read consistent with?” It chops the read into short words (k-mers), looks them up in an index built from transcript sequences, and records the set of compatible transcripts. A statistical model then divides ambiguous reads among transcripts according to their overall abundance, while also correcting for fragment GC content and sequence-specific biases. Because it never builds full alignments, it is much faster; TransXplorer’s own label describes it as roughly ten times faster. Finally tximport (Soneson et al., 2015) sums transcripts into genes so the output looks exactly like the other route.
How to choose
For a standard gene-level differential expression study in human or mouse, both routes give you a perfectly good count matrix for DESeq2, edgeR or limma-voom. The table below is for when you want to pick deliberately.
| If you care about… | HISAT2 + featureCounts | Salmon + tximport |
|---|---|---|
| Speed | Slower — full spliced alignment, then sorting and indexing BAMs | Faster — labelled “Fast” in the app |
| Auditing where reads went | Yes — per-sample counts of reads assigned to genes vs. reads that hit no gene, multi-mapped, or were ambiguous | Less direct — no genome alignment or featureCounts summary |
| Reads that fit several genes or isoforms | Left out (multi-mappers) or reported as ambiguous | Shared statistically across compatible transcripts |
| Bias correction | None at the counting step | GC and sequence-specific (--gcBias --seqBias) |
| Organisms on the server | Every genome in the menu, including a custom genome + GTF | Human (hg38), mouse (mm10), or a custom genome + GTF |
| Reads outside the annotation (new genes, intronic signal) | Visible as unassigned reads in the summary | Invisible — only annotated transcripts are in the index |
| Methods sentence | “Reads were aligned with HISAT2 and gene-level counts obtained with featureCounts.” | “Transcripts were quantified with Salmon and summarised to genes with tximport.” |
Running it in TransXplorer, step by step
One sidebar, top to bottom. Then watch the log.
Open TransXplorer and click the FASTQ Processing tab. A status banner at the top tells you whether the server is free (“System ready - you can start your analysis”). The left sidebar holds every setting, in the order you need them.
-
Choose an Analysis Mode
Two options, depending mostly on how much data you have:
DefaultRun on Server
- Upload straight from your browser
- Maximum 15 GB total file size
- Typically 5–10 samples
- Both HISAT2 and Salmon available
Large dataRun Locally with Docker
- No upload; your files stay on your machine
- Unlimited dataset size
- Pick Windows, macOS or Linux and follow four setup steps
- HISAT2 + featureCounts, with references for every organism in the menu
In Docker mode the app walks you through installing Docker, pulling the pipeline image, putting your files in an
input_fastq/folder, and running the script from Download Analysis Script (./run_rnaseq_*.sh; add a number, e.g../run_rnaseq_*.sh 16, to use more CPU threads). The rest of this tutorial follows the server route. -
Upload your FASTQ files
Select all files for the experiment at once (both mates of every pair). Large uploads take a while; until every file has arrived the start button stays grey and reads Files uploading…, with the reminder “Please wait for all files to finish uploading before starting analysis.”
The Instructions panel just above (click the chevron to expand) repeats the essentials: accepted formats, the
SampleName_R1/R2.fastq.gzconvention, and the 15 GB server limit. -
Tell it what you have
- Sequencing Type: Paired-End (default) or Single-End. Paired-End expects matched
_R1/_R2files. - Quantification Method: HISAT2 + featureCounts (tagged Standard, the default) or Salmon (tagged Fast). A one-line hint under the choice summarises each.
- Reference Genome: the organism your reads come from. Default is Human (hg38). Choosing Custom Genome reveals uploads for a Genome FASTA File:, an Annotation GTF/GFF File: and a Genome Name:.
The genome menu lists eleven options. Which of them can be processed where is summarised in the table below this list.
- Sequencing Type: Paired-End (default) or Single-End. Paired-End expects matched
-
Start the analysis
Once uploads finish, the button turns blue and shows what it’s about to process, for example Start Analysis (8 files, 6.4 GB). Click it. TransXplorer first validates the inputs (file types, R1/R2 pairing, total size, custom genome files) and stops with a clear message if anything is off; see Troubleshooting. If all is well, the button switches to Processing… and the pipeline runs in a separate background process.
-
Watch the Processing Log
The Processing Log tab shows a progress bar with the current stage and a live, colour-coded terminal:
[STEP]marks a stage starting,[DONE](green) a stage or sample finishing,[WARN]something non-fatal, and[ERROR](red) a failure. Because every sample is logged by name (Trimmed pair: WT_rep1,Aligned: WT_rep1…), you can see exactly how far a long run has got.Run time depends on read depth, the number of samples and how busy the server is. Keep the browser tab open while it runs, and download your results before you close it.
Which genomes work where
| Reference Genome (menu label) | Server · HISAT2 | Server · Salmon | Docker |
|---|---|---|---|
| Human (hg38) | Yes | Yes | Yes |
| Mouse (mm10) | Yes | Yes | Yes |
| Rat (rn6) | Yes | No | Yes |
| Drosophila (dm6) · Zebrafish (danRer11) · C. elegans (wbcel235) · Yeast (r64) · Arabidopsis (araTha) · Chicken (galGal6) · Pig (susScr11) | Yes | No | Yes |
| Custom Genome (your FASTA + GTF) | Yes | Yes | Yes |
On the server, each genome is counted against an annotation from the same assembly: GENCODE release 47 for hg38, GENCODE vM25 for mm10 (the last release on GRCm38), Ensembl 104 for rn6 (the last on Rnor_6.0), Ensembl BDGP6.46 for dm6, and Ensembl or index-bundled GTFs for the rest. Chromosome naming differences (chr2L vs 2L) are harmonised automatically, and the log prints the pairing it used, e.g. Reference: dm6 | index: … | annotation: Drosophila_melanogaster.BDGP6.46.111.gtf. If an index or matching annotation is missing, the run stops with a clear error rather than borrowing another organism’s.
For a custom genome with HISAT2, the server builds a HISAT2 index from your FASTA (hisat2-build) and counts against your GTF; large genomes can take a long time to index. With Salmon, it first extracts transcript sequences from your FASTA using your GTF (with gffread) and builds an index from them. In Docker mode, put genome.fa and annotation.gtf in input_fastq/custom_genome/ and the script builds a HISAT2 index for you.
When the server is busy
FASTQ processing is heavy, so the server runs at most two jobs at a time. If both slots are taken when you click Start, you’ll get an Analysis Queued notification with your position in the queue, the number of active jobs, an estimated wait and a job ID, and the banner at the top of the tab shows your position. The app checks for a free slot about every 10 seconds and starts your job automatically when one opens; the log then reads A processing slot is free - starting your queued analysis. Keep the browser tab open while you wait: closing it cancels your queued or running job and frees its slot. Otherwise the banner switches to “Your analysis is currently processing…” with the minutes elapsed.
Choosing the wrong reference genome produces no error at all. Reads from a mouse experiment will still partly align to the human genome; you’ll just get a poor assignment rate and misleading counts. Check the Reference Genome: menu twice before you press Start.
Reading your FastQC reports
What’s fine, what’s alarming, and what only looks alarming.
When the run finishes, open the QC Reports tab. A side panel lets you pick a Sample: and a QC Stage: (Pre-Trimming or Post-Trimming); the full FastQC report then opens right inside the page, and Download Report saves it. Flip between the two stages for the same sample. That before/after comparison is the most useful thing on the tab.
FastQC gives each module a pass, warn or fail flag. Treat those as prompts to look, not verdicts. FastQC’s thresholds were designed for random genomic DNA, and RNA-seq libraries break several of its assumptions by design.
Per base sequence quality
Medians in the red across much of the read, in raw and trimmed data, mean a poor run. Talk to your sequencing facility before analysing further.
Adapter content
Rising adapter curves in the raw report are common with short fragments. After trimming they should be flat near zero. If not, the library may use adapters other than TruSeq.
Sequence duplication levels
Often flagged red in RNA-seq. Highly expressed genes are sequenced over and over, so identical reads are expected. This reflects biology, not PCR failure.
Per base sequence content
Wobbly A/C/G/T lines in the first 10–15 bases come from random-hexamer priming during library prep. A known, harmless RNA-seq signature.
Per sequence GC content
A shifted or two-humped curve can mean contamination (bacteria, another species) or abundant rRNA. Compare against other samples in the same run.
Overrepresented sequences
Adapter dimers, rRNA or mitochondrial transcripts often top this list. A few hits are normal; one sequence making up several percent of reads is not.
The most informative QC is comparative. One sample whose report looks very different from its replicates is far more worrying than every sample sharing the same “failed” duplication module. Outlier samples found here often reappear later as outliers on the PCA plot in the DE workflow.
What you get back
Three result tabs and two downloads.
When the run completes you’ll see “Analysis completed successfully!” and the result tabs fill in:
-
Feature Counts
Gene Expression Count MatrixA searchable table of the count matrix: one row per gene, one column per sample, with the gene ID in the first column. Glance at it to confirm every sample is present and the numbers are whole-number counts. -
QC Summary
Quality Control SummaryThe number of genes and samples, the sequencing type and genome you chose, and, for the HISAT2 route, an Alignment Statistics table from featureCounts showing, per sample, how many reads were Assigned to genes and how many fell into the Unassigned categories. -
Analysis Info
About This AnalysisA reminder of the tools and outputs, and the next step: take the count matrix to the Transcriptome Analysis section.
The downloads
The Download Results panel at the bottom of the sidebar has two buttons:
-
Count Matrix (CSV)
gene_counts_<date>.csvThe file you’ll analyse. The first column isGene_ID(Ensembl gene IDs from the annotation). It is followed by one column per sample, for both HISAT2 and Salmon, so the file uploads straight into Transcriptome Analysis. -
Full Results (ZIP)
qc_bundle_<date>.zipEverything you need for your records and supplementary files, for both quantification methods: the counts, every FastQC report and a one-page HTML summary. HISAT2 runs add the featureCounts alignment summary; Salmon runs add asalmon_tpm.txttable, and their HTML summary includes a Salmon Mapping Statistics table (reads processed, reads mapped and mapping rate per sample).
qc_summary.html includes an Assignment Rate per sample for HISAT2 runs (the share of reads featureCounts placed on a gene) or a Mapping Rate for Salmon runs. It is the single quickest health check for an RNA-seq library. Replicates should have similar rates; one sample far below the rest deserves a look before it goes into a DE model. Keep the whole ZIP: it is a ready-made supplementary QC record.
From counts to differential expression
Same app, next tab: the standard DE workflow takes over.
The count matrix is the hand-off point. From here on the analysis is identical to starting from a published counts file, and everything in the first-analysis walkthrough applies.
-
Download the CSV
Click Count Matrix (CSV). The file holds only
Gene_IDand your sample columns, whichever quantification method you used, so there is nothing to tidy.Why it matters: the uploader treats the gene column as row names and every other column as a sample. If you add notes or extra columns in a spreadsheet, they will show up as “samples” in the next dialog.
-
Upload it to Transcriptome Analysis
Switch to the Transcriptome Analysis tab and use Upload Count Matrix (CSV/XLSX/TSV/TXT) at the top of the sidebar. The
Gene_IDcolumn is recognised automatically, and Ensembl IDs are mapped to gene symbols for the steps that need symbols. -
Confirm sample groups
A Sample Group Assignment dialog opens with your sample names (the FASTQ prefixes, e.g.
WT_rep1,KO_rep1). TransXplorer tries to detect the groups from the names; check them and confirm. -
Add metadata if you have batches
If samples were prepared or sequenced in different batches, upload a metadata file whose
Samplecolumn matches the count-matrix column names exactly, plus columns such assequencing_runorprep_date. Automatic Detection will test whether those batches matter; Understanding Batch Effects explains how. -
Configure and run
Pick
DESeq2,edgeRorlimma-voom, set reference and treatment groups and thresholds, and click Run DEG Analysis. Both FASTQ routes produce raw integer counts, which is exactly what these methods expect; What is Differential Expression? covers the choice.
If you process some samples with HISAT2 and others with Salmon (or on different days with different genomes), the method itself becomes a batch effect. Process every sample of an experiment in the same run, or at least with identical settings.
Troubleshooting
The messages you might see, and what they usually mean.
Problems show up in two places: as a red notification (“Analysis failed: …”) and as an [ERROR] Pipeline error: … line in the Processing Log. The pipeline stops at the first failing step, so the message tells you both which tool and which sample to look at.
validation › These samples are missing their R1 or R2 partner: <names>
validation › Paired-end sequencing requires an even number of files
File names or counts don’t pair up
Every paired-end file must have _R1 or _R2 (or Illumina’s _R1_001/_R2_001) just before the extension, and every R1 needs an R2 with the same sample prefix.
- If the message lists sample names, those samples have only one mate: check you didn’t miss a file when selecting, and that both mates share an identical prefix.
- Use an underscore, not a dot, before
R1/R2(X.R1.fq.gzis not recognised). - If the data is really single-end, switch Sequencing Type: to Single-End.
Too much data for the server
- Make sure you are uploading the
.gzfiles, not decompressed copies. - For bigger studies, switch to Run Locally with Docker, which has no size limit.
- Avoid splitting one experiment into several server runs unless the settings are identical; see the warning in Section 8.
pattern › <step> failed for sample '<name>' (exit code N)
A tool failed on one sample
The step can be HISAT2 alignment, samtools sort, samtools index or Salmon quantification; a counting failure reads featureCounts quantification failed (exit code N). The exit code is the tool’s own; the sample name is what matters.
- Truncated or corrupted upload: the most common cause. Re-check the file against the facility’s checksum (
md5sum), or test it withgzip -t file.fastq.gz, then upload again. - Mates that don’t match: R1 and R2 with different read counts or read names (see below).
- Exit code 137: usually means the process was stopped for running out of memory. Try again later when the server is quieter, or use Docker.
Mismatched R1/R2 files
Because R1 and R2 are paired by sorted name, a prefix that differs between mates (WT_rep1_R1 with wt_rep1_R2), or a file from the wrong sample, can pair reads that don’t belong together. Mates must also contain the same number of reads in the same order.
- Use identical prefixes for both mates of every sample.
- Compare read counts: on macOS/Linux,
echo $(( $(gzip -dc X_R1.fastq.gz | wc -l) / 4 ))should give the same number for R1 and R2. - Never trim, filter or subsample one mate without the other.
Salmon picked with an organism it can’t serve
On the server, Salmon covers human (hg38), mouse (mm10) and custom genomes. For any other organism in the menu, switch Quantification Method: to HISAT2 + featureCounts, which runs on the server for every genome.
Low mapping or assignment rate
No error is raised, which is why it’s worth checking the QC Summary and qc_summary.html every time. As a rough rule, a healthy poly-A library assigns most of its reads to genes; replicates far below their siblings, or a whole run well under half, need explaining. The usual suspects:
- Wrong reference genome: reads from one species aligned to another. Recheck Reference Genome:.
- Lots of reads outside genes (high
Unassigned_NoFeatures): genomic DNA contamination, or total-RNA libraries with many intronic pre-mRNA reads. - Ribosomal RNA: incomplete rRNA depletion. Look for rRNA in FastQC’s overrepresented sequences and an odd GC curve.
- Multi-mapping reads (high
Unassigned_MultiMapping): repetitive sequence or gene families; featureCounts leaves these out. Salmon handles them better. - Very short inserts or adapter dimers: many reads drop below 36 bases after trimming and are discarded.
- Degraded RNA: coverage piles up at the 3′ ends of genes; check RIN values with the facility.
What’s next
You have counts. Here’s where to take them.
Continue the workflow
Understand the methods downstream
Ready to process your own reads?
Open TransXplorer, click FASTQ Processing, and drop in your _R1/_R2 files. Free, in the browser, nothing to install for server mode.
References
The tools in this pipeline are peer-reviewed. Cite them in your methods alongside TransXplorer.
- Cock PJA, Fields CJ, Goto N, Heuer ML, Rice PM. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Res. 2010;38(6):1767–1771. doi:10.1093/nar/gkp1137
- Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30(15):2114–2120. doi:10.1093/bioinformatics/btu170
- Kim D, Paggi JM, Park C, Bennett C, Salzberg SL. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019;37(8):907–915. doi:10.1038/s41587-019-0201-4
- Liao Y, Smyth GK, Shi W. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;30(7):923–930. doi:10.1093/bioinformatics/btt656
- Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods. 2017;14(4):417–419. doi:10.1038/nmeth.4197
- Soneson C, Love MI, Robinson MD. Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences. F1000Research. 2015;4:1521. doi:10.12688/f1000research.7563.2
FastQC is software from the Babraham Institute; cite it by its project page.
If TransXplorer helps your research, please cite us
The preprint is open access on bioRxiv. A peer-reviewed version is in submission.
Frequently asked questions
How much FASTQ data can I process on the server?
Should I choose HISAT2 + featureCounts or Salmon?
DESeq2, edgeR or limma-voom. HISAT2 + featureCounts reports how many reads landed on genes, which is great for QC. Salmon is faster, corrects GC and sequence bias, and shares ambiguous reads statistically. On the server Salmon covers human (hg38), mouse (mm10) and custom references, while HISAT2 + featureCounts covers every genome in the menu. See How to choose.
How should I name paired-end files?
SampleName_R1.fastq.gz and SampleName_R2.fastq.gz, with an identical prefix for both mates. Files straight off an Illumina instrument, such as SampleName_S1_L001_R1_001.fastq.gz, are accepted as they are; the sample name becomes everything before _R1 or _R1_001.
FastQC failed my duplication and per-base content modules. Is my data bad?
Can I use the count matrix straight away for differential expression?
Gene_ID and one column per sample for both HISAT2 and Salmon, then upload it unchanged with Upload Count Matrix in the Transcriptome Analysis tab. See Section 8.