Comprehensive RNA-seq Data Analysis
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a set of raw FASTQ files from a personal research project and I want to take them all the way through to biological insight. My primary goal is to identify differentially expressed genes, and I already have both the reference genome and its matching GTF/GFF annotation ready for you.
Here is the pipeline I want to see implemented:
• Initial quality check with FastQC followed by an aggregated MultiQC summary.
• Adapter and low-quality base trimming.
• Alignment to the reference (or transcript-level quantification if you prefer Salmon/kallisto) under a well-documented, reproducible Linux environment.
• Gene-level count matrix generation, then differential expression with DESeq2 in R.
• Exploratory visualisations: PCA, heatmap, and a volcano plot highlighting key DE genes.
• Functional interpretation through GO and KEGG pathway enrichment.
Deliverables must include:
• All processed result files and figure images in publication-ready resolution.
• Tidy tables of counts, normalised expression values, and DESeq2 outputs (padj, log2FC, etc.).
• The exact shell, R, and/or Python scripts or notebooks you ran, with comments.
• A concise, step-by-step report (Markdown, R Markdown, or Jupyter) so I can reproduce every step on my own workstation.
Please use standard RNA-seq tools—the typical stack would be FastQC, Cutadapt/Trim Galore, STAR or HISAT2, Salmon/kallisto, DESeq2, and clusterProfiler—but feel free to suggest sensible alternatives if they improve accuracy or speed. I work comfortably in Linux, so command-line oriented solutions are welcome, and I expect the code to run under a recent Ubuntu or CentOS environment without extensive tweaking.
If this matches your expertise in bioinformatics and you can turn around clear, reproducible results, I’d love to collaborate.
Here is the pipeline I want to see implemented:
• Initial quality check with FastQC followed by an aggregated MultiQC summary.
• Adapter and low-quality base trimming.
• Alignment to the reference (or transcript-level quantification if you prefer Salmon/kallisto) under a well-documented, reproducible Linux environment.
• Gene-level count matrix generation, then differential expression with DESeq2 in R.
• Exploratory visualisations: PCA, heatmap, and a volcano plot highlighting key DE genes.
• Functional interpretation through GO and KEGG pathway enrichment.
Deliverables must include:
• All processed result files and figure images in publication-ready resolution.
• Tidy tables of counts, normalised expression values, and DESeq2 outputs (padj, log2FC, etc.).
• The exact shell, R, and/or Python scripts or notebooks you ran, with comments.
• A concise, step-by-step report (Markdown, R Markdown, or Jupyter) so I can reproduce every step on my own workstation.
Please use standard RNA-seq tools—the typical stack would be FastQC, Cutadapt/Trim Galore, STAR or HISAT2, Salmon/kallisto, DESeq2, and clusterProfiler—but feel free to suggest sensible alternatives if they improve accuracy or speed. I work comfortably in Linux, so command-line oriented solutions are welcome, and I expect the code to run under a recent Ubuntu or CentOS environment without extensive tweaking.
If this matches your expertise in bioinformatics and you can turn around clear, reproducible results, I’d love to collaborate.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.