Biostar Handbook Bioinformatics Tools Collection
Environment: biostar
Installation date: October 2025
Based on: Biostar Handbook official tool list
Overview
This conda environment contains 28 essential bioinformatics tools curated by the Biostar Handbook for computational biology and bioinformatics analysis. All tools are installed from the bioconda and conda-forge channels.
Total installation size: ~540 MB
Installed Tools
Sequence Alignment and Mapping
| Tool | Version | Description |
|---|---|---|
| blast | 2.12.0 | NCBI Basic Local Alignment Search Tool - sequence similarity search |
| bowtie2 | 2.4.5 | Fast and memory-efficient short read aligner |
| bwa | 0.7.17 | Burrows-Wheeler Aligner for mapping sequences to reference genome |
| hisat2 | 2.2.1 | Fast and sensitive alignment program for mapping NGS reads |
| minimap2 | 2.24 | Versatile sequence alignment program for long reads |
| mafft | 7.525 | Multiple alignment program for nucleotide/protein sequences |
Sequence Analysis and Manipulation
| Tool | Version | Description |
|---|---|---|
| samtools | 1.15.1 | Tools for manipulating SAM/BAM alignment files |
| bcftools | 1.15.1 | Tools for variant calling and manipulating VCF/BCF files |
| bedtools | 2.30.0 | Toolset for genome arithmetic and interval operations |
| seqkit | 2.10.1 | Cross-platform toolkit for FASTA/Q file manipulation |
| seqtk | 1.3 | Toolkit for processing sequences in FASTA/Q formats |
| bioawk | 1.0 | AWK extended with biological sequence parsing |
| fastp | 1.0.1 | Ultrafast FASTQ preprocessing tool |
| fastqc | 0.12.1 | Quality control tool for high throughput sequence data |
Variant Analysis
| Tool | Version | Description |
|---|---|---|
| snpeff | 5.0 | Genetic variant annotation and effect prediction |
RNA-seq and Gene Expression
| Tool | Version | Description |
|---|---|---|
| subread | 2.0.1 | Read alignment and quantification (includes featureCounts) |
| trimmomatic | 0.40 | Flexible read trimming tool for Illumina NGS data |
Data Visualization and Conversion
| Tool | Version | Description |
|---|---|---|
| ucsc-bedgraphtobigwig | 377 | UCSC tool to convert bedGraph format to bigWig |
Data Management and Utilities
| Tool | Version | Description |
|---|---|---|
| aria2 | 1.34.0 | Lightweight multi-protocol download utility |
| wget | 1.20.3 | Network downloader for retrieving files via HTTP/HTTPS/FTP |
| ncbi-datasets-cli | 18.9.0 | NCBI Datasets command-line tools for downloading genomic data |
| parallel | 20170422 | GNU parallel for running jobs in parallel |
Data Processing and Text Manipulation
| Tool | Version | Description |
|---|---|---|
| csvkit | 2.1.0 | Suite of utilities for working with CSV files |
| csvtk | 0.31.0 | Cross-platform CSV/TSV toolkit |
| datamash | 1.1.0 | Command-line statistics and text processing |
| jq | 1.5 | Lightweight and flexible command-line JSON processor |
Build Tools
| Tool | Version | Description |
|---|---|---|
| make | 4.4.1 | GNU Make build automation tool |
Perl Libraries
| Tool | Version | Description |
|---|---|---|
| perl-text-csv | 1.33 | Perl module for CSV file manipulation |
Quick Start
Activate the Environment
# Load miniforge3 module first
module load biostar
# Activate biostar environment
# Biostar environment loaded automatically
Deactivate
# Module unloaded automatically
Usage Examples
Quality Control and Preprocessing
# Run FastQC on sequencing data
fastqc sample_R1.fastq.gz sample_R2.fastq.gz
# Trim adapters with Trimmomatic
trimmomatic PE input_R1.fastq.gz input_R2.fastq.gz \
output_R1_paired.fastq.gz output_R1_unpaired.fastq.gz \
output_R2_paired.fastq.gz output_R2_unpaired.fastq.gz \
ILLUMINACLIP:adapters.fa:2:30:10 LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:36
# Fast preprocessing with fastp
fastp -i input_R1.fastq.gz -I input_R2.fastq.gz \
-o output_R1.fastq.gz -O output_R2.fastq.gz
Sequence Alignment
# Index reference genome with BWA
bwa index reference.fasta
# Align paired-end reads
bwa mem reference.fasta reads_R1.fastq.gz reads_R2.fastq.gz > alignment.sam
# Convert SAM to sorted BAM
samtools view -b alignment.sam | samtools sort -o alignment.sorted.bam
samtools index alignment.sorted.bam
# Alignment with Bowtie2
bowtie2-build reference.fasta ref_index
bowtie2 -x ref_index -1 reads_R1.fastq.gz -2 reads_R2.fastq.gz -S alignment.sam
# Long read alignment with minimap2
minimap2 -ax map-ont reference.fasta longreads.fastq.gz > alignment.sam
Variant Calling and Analysis
# Call variants with bcftools
bcftools mpileup -f reference.fasta alignment.sorted.bam | \
bcftools call -mv -Oz -o variants.vcf.gz
# Index VCF file
bcftools index variants.vcf.gz
# Filter variants
bcftools filter -i 'QUAL>20 && DP>10' variants.vcf.gz -o filtered.vcf
# Annotate variants with SnpEff
snpEff -v GRCh38.99 variants.vcf > annotated.vcf
Sequence Manipulation
# FASTA/FASTQ statistics with seqkit
seqkit stats *.fastq.gz
# Extract sequences by ID
seqkit grep -p "gene_name" sequences.fasta
# Convert FASTQ to FASTA
seqkit fq2fa reads.fastq.gz -o reads.fasta
# Reverse complement sequences
seqkit seq -r -p sequences.fasta
# Subsample FASTQ files
seqtk sample -s100 reads.fastq.gz 10000 > subset.fastq
Genomic Interval Operations
# Intersect two BED files
bedtools intersect -a regions1.bed -b regions2.bed > overlap.bed
# Get coverage of features
bedtools coverage -a genes.bed -b alignment.sorted.bam > coverage.txt
# Merge overlapping intervals
bedtools merge -i sorted_intervals.bed > merged.bed
RNA-seq Analysis
# Align RNA-seq reads with HISAT2
hisat2-build reference.fasta genome_index
hisat2 -x genome_index -1 reads_R1.fastq.gz -2 reads_R2.fastq.gz -S alignment.sam
# Count reads per gene with featureCounts (from subread package)
featureCounts -p -t exon -g gene_id -a annotation.gtf -o counts.txt alignment.sorted.bam
BLAST Searches
# Create BLAST database
makeblastdb -in proteins.fasta -dbtype prot -out protein_db
# Run protein BLAST
blastp -query query.fasta -db protein_db -out results.txt
# Run nucleotide BLAST
blastn -query sequences.fasta -db nt -remote -out blast_results.txt
Data Download
# Download with aria2 (multi-threaded)
aria2c -x 8 -s 8 https://example.com/large_file.fastq.gz
# Download NCBI datasets
datasets download genome accession GCF_000001405.40
Text Processing
# CSV manipulation with csvkit
csvstat data.csv
csvcut -c 1,3,5 data.csv > subset.csv
csvgrep -c column_name -m "pattern" data.csv
# CSV manipulation with csvtk
csvtk stats data.csv
csvtk filter -f "score>50" data.csv
csvtk join -f "id" file1.csv file2.csv
# Statistical operations with datamash
datamash mean 1 median 2 < data.txt
# JSON processing with jq
cat metadata.json | jq '.experiments[] | select(.type=="RNA-seq")'
Parallel Processing
# Run BLAST on multiple files in parallel
parallel "blastn -query {} -db nt -out {.}.blast" ::: *.fasta
# Process multiple samples in parallel
parallel -j 4 "fastqc {}" ::: *.fastq.gz
Additional Dependencies
The environment also includes necessary dependencies:
- Python 3.9 with standard scientific libraries
- OpenJDK 11.0.9 (required for FastQC, Trimmomatic, SnpEff)
- Perl 5.22 with essential modules
- Various C/C++ libraries for bioinformatics tools
Getting Help
Tool-specific Help
Most tools have built-in help:
samtools --help
bcftools --help
bedtools --help
seqkit -h
blastn -help
Online Resources
- Biostar Handbook: https://www.biostarhandbook.com
- Bioconda documentation: https://bioconda.github.io
- Individual tool documentation: Check tool-specific websites
Package Management
List installed packages
conda list
Install additional packages
# Install to biostar environment
conda install -n biostar package_name -c bioconda -c conda-forge
Update packages
# Update all packages (use with caution)
conda update --all
# Update specific package
conda update package_name
Notes
- This environment is shared system-wide - do not modify without coordination
- For personal tool installations, create your own conda environment
- All tools are from official bioconda/conda-forge channels
- Environment is based on official Biostar Handbook recommendations (October 2025)
Troubleshooting
Environment activation fails
Make sure to load the miniforge3 module first:
module load biostar
# Biostar environment loaded automatically
Tool not found
Verify the environment is activated:
which samtools # Should point to /sw/miniforge3/.../envs/biostar/bin/samtools
Permission errors
This is a shared environment. For personal modifications, create your own:
conda create -n my_biostar --clone biostar
For questions or issues, contact cluster administrators.
Last updated: October 2025