Skip to content

Miniforge3 Environment - Installation Guide

Version: 20250911 Location: /sw/miniforge3/20250911 Module: miniforge3/20250911


Table of Contents

  1. About This Environment
  2. Biostar Handbook Bioinformatics Environment
  3. Creating Your Own Conda Environment
  4. Creating a Virtual Environment (venv) with Inheritance
  5. Creating a Clean Virtual Environment (venv)
  6. Using Jupyter Notebook and JupyterLab
  7. Remote Development with VSCode
  8. Installed Packages by Thematic Area

About This Environment

This base environment is a general-purpose installation intended as a starting point for scientific computing. It includes commonly used packages for:

  • Scientific computing and data analysis
  • Molecular dynamics and computational chemistry
  • Bioinformatics and structural biology
  • Data visualization and image processing

However, this environment is incomplete by design. It cannot cover all possible use cases and dependencies for every user. We encourage you to:

  • Create your own conda environment in your home directory for full control
  • Create a virtual environment if you only need a few additional packages
  • Customize your workflow according to your specific research needs

Biostar Handbook Bioinformatics Environment

A pre-configured bioinformatics environment is available with 28 essential tools from the Biostar Handbook.

Quick Start

# Load the module
module load miniforge3/20250911

# Load the biostar module
module load biostar

# Tools are now available: samtools, bcftools, blast, bwa, etc.
samtools --version

What's Included

The biostar module provides tools for:

  • Sequence alignment: blast, bowtie2, bwa, hisat2, minimap2, mafft
  • File manipulation: samtools, bcftools, bedtools, seqkit, seqtk, bioawk
  • Quality control: fastqc, fastp, trimmomatic
  • Variant analysis: snpeff
  • RNA-seq: subread (featureCounts), hisat2
  • Data processing: csvkit, csvtk, datamash, jq
  • Downloads: aria2, wget, ncbi-datasets-cli
  • Utilities: parallel, make

Complete Documentation

For detailed information about all 28 tools, usage examples, and workflows:

# View the complete guide
cat /sw/biostar/BIOSTAR_GUIDE.md

# Or with a pager
less /sw/biostar/BIOSTAR_GUIDE.md

The guide includes: - Complete list of tools with descriptions - Usage examples for each tool category - Best practices for bioinformatics workflows - Integration with Slurm job submission

Deactivate

conda deactivate

Creating Your Own Conda Environment

If you need complete control over your Python environment, we recommend creating a local conda environment in your home directory.

Step 1: Load the Module

module load miniforge3/20250911

Step 2: Create a New Environment

Create a new environment with a specific Python version:

# Create environment with Python 3.12
conda create -n myenv python=3.12

# Or with Python 3.11
conda create -n myenv python=3.11

By default, conda environments are stored in ~/.conda/envs/. You can also specify a custom location:

# Create environment in a custom location
conda create -p ~/my-projects/myenv python=3.12

Step 3: Activate the Environment

# Activate by name
conda activate myenv

# Or by path
conda activate ~/my-projects/myenv

Step 4: Install Packages

Use mamba (faster) or conda to install packages:

# Install packages with mamba (recommended - much faster)
mamba install numpy scipy matplotlib pandas

# Or with conda
conda install numpy scipy matplotlib pandas

# Install from specific channels
mamba install -c conda-forge rdkit

# Install using pip (if package not available in conda)
pip install some-package

Step 5: Deactivate the Environment

conda deactivate

Managing Your Environments

# List all environments
conda env list

# Remove an environment
conda env remove -n myenv

# Export environment to file
conda env export > environment.yml

# Create environment from file
conda env create -f environment.yml

Creating a Virtual Environment (venv) with Inheritance

If you need just a few additional packages but want to reuse the packages already installed in the base environment, you can create a virtual environment with --system-site-packages.

Step 1: Load the Module

module load miniforge3/20250911

Step 2: Create venv with System Packages

# Create venv in your home directory
python -m venv ~/my-venv --system-site-packages

This creates a lightweight virtual environment that: - Has access to ALL packages installed in the base environment - Allows you to install additional packages locally - Stores new packages in ~/my-venv/lib/python3.12/site-packages

Step 3: Activate the venv

source ~/my-venv/bin/activate

After activation, your prompt will change to show (my-venv).

Step 4: Install Additional Packages

# Install packages with pip
pip install some-additional-package

# List locally installed packages (excluding inherited ones)
pip list --local

Step 5: Deactivate the venv

deactivate

Verify Package Sources

# Show where a package is installed
python -c "import numpy; print(numpy.__file__)"

# List all packages (including inherited)
pip list

Creating a Clean Virtual Environment (venv)

If you want a completely isolated environment without inheriting any packages from the base environment:

Step 1: Load the Module

module load miniforge3/20250911

Step 2: Create Clean venv

# Create venv WITHOUT system packages
python -m venv ~/my-clean-venv

This creates a fresh environment with: - Only Python standard library - No access to base environment packages - Complete isolation for reproducibility

Step 3: Activate the venv

source ~/my-clean-venv/bin/activate

Step 4: Install Packages

# Upgrade pip first (recommended)
pip install --upgrade pip

# Install packages
pip install numpy scipy matplotlib pandas

# Install from requirements file
pip install -r requirements.txt

Step 5: Deactivate the venv

deactivate

Using Jupyter Notebook and JupyterLab

The base environment includes JupyterLab and ipykernel for interactive computing.

Where to Run JupyterLab

A small notebook on the login node is fine. A large one is not. xlence is also where the scheduler for the whole cluster runs, along with the accounting database and everyone else's interactive sessions: a notebook that takes several gigabytes of memory there slows down every other user. This is not hypothetical — a single notebook using 38 GB of RAM once made the cluster unusable for everybody for hours.

Notebooks belong on a compute node, requested through Slurm and reached through an SSH tunnel. The full step-by-step procedure, written for people who have never used Slurm or an SSH tunnel before, is on its own page:

👉 See: Interactive Jupyter Notebooks

The short version, for those who already know their way around:

# on xlence: request a slice of a compute node
srun -p ngs -c 4 --mem=32G -t 4:00:00 --pty bash -l

# on the compute node
module load miniforge3
PORT=$((10000 + $(id -u) % 10000))
echo "ssh -N -L $PORT:localhost:$PORT -J $USER@xlence.disfeb.unimi.it $USER@$(hostname)"
jupyter lab --no-browser --ip=127.0.0.1 --port=$PORT

# on your own computer, in a second terminal: paste the line printed above

Then open the http://127.0.0.1:<port>/lab?token=... address that JupyterLab prints.

Bind the server to 127.0.0.1, not to 0.0.0.0: with 0.0.0.0 the notebook is reachable from every machine on the cluster's internal network and only the token stands between it and other users.

Using Your Custom Environment in Jupyter

If you created a custom conda environment or venv, you need to register it as a Jupyter kernel:

For Conda Environments

# Activate your environment
conda activate myenv

# Install ipykernel
conda install ipykernel

# Register kernel
python -m ipykernel install --user --name myenv --display-name "Python (myenv)"

For venv Environments

# Activate your venv
source ~/my-venv/bin/activate

# Install ipykernel
pip install ipykernel

# Register kernel
python -m ipykernel install --user --name my-venv --display-name "Python (my-venv)"

List and Remove Kernels

# List available kernels
jupyter kernelspec list

# Remove a kernel
jupyter kernelspec remove myenv

Now when you open JupyterLab, you'll see your custom environment in the kernel selection menu.


Remote Development with VSCode

VSCode's Remote-SSH extension provides excellent support for remote Python development and Jupyter notebooks.

Prerequisites

  1. Install VSCode on your local machine
  2. Install Remote-SSH extension in VSCode
  3. Connect to cluster via Remote-SSH

Setting Up Python Environment in VSCode

Step 1: Connect via Remote-SSH

  1. Open VSCode
  2. Press F1 or Ctrl+Shift+P
  3. Type "Remote-SSH: Connect to Host"
  4. Enter: username@xlence.disfeb.unimi.it

Step 2: Install Python Extension (Remote)

Once connected, install the Python extension in the remote environment: 1. Go to Extensions (Ctrl+Shift+X) 2. Search for "Python" (Microsoft) 3. Click "Install in SSH: xlence"

Step 3: Select Python Interpreter

  1. Open a Python file or create a new one
  2. Press Ctrl+Shift+P and type "Python: Select Interpreter"
  3. Choose one of:
  4. Base environment: /sw/miniforge3/20250911/bin/python
  5. Your conda env: ~/.conda/envs/myenv/bin/python
  6. Your venv: ~/my-venv/bin/python

You can also set the interpreter in workspace settings (.vscode/settings.json):

{
    "python.defaultInterpreterPath": "/sw/miniforge3/20250911/bin/python"
}

Working with Jupyter Notebooks in VSCode

VSCode has excellent built-in support for Jupyter notebooks (.ipynb files).

Step 1: Open or Create a Notebook

Create a new file with .ipynb extension or open an existing notebook.

Step 2: Select Kernel

  1. Click on kernel selection in the top-right corner of the notebook
  2. Choose "Select Another Kernel" → "Python Environments"
  3. Select your desired Python interpreter:
  4. Base forge environment
  5. Your custom conda environment
  6. Your venv

Step 3: Run Cells

  • Click the play button next to each cell
  • Or press Shift+Enter to run cell and advance
  • The kernel will start automatically

Setting Python Environment in Settings

For consistent behavior, add to your workspace or user settings:

{
    "python.defaultInterpreterPath": "/sw/miniforge3/20250911/bin/python",
    "jupyter.notebookFileRoot": "${workspaceFolder}",
    "python.terminal.activateEnvironment": true
}

Activating Module Automatically in Terminal

To automatically load the module when opening a terminal in VSCode, you can:

  1. Add to your ~/.bashrc: bash # Auto-load forge module if [[ -z "$MODULE_LOADED_FORGE" ]]; then module load miniforge3/20250911 export MODULE_LOADED_FORGE=1 fi

  2. Or create a workspace-specific shell script: Create .vscode/tasks.json: json { "version": "2.0.0", "tasks": [ { "label": "Load forge module", "type": "shell", "command": "module load miniforge3/20250911", "problemMatcher": [] } ] }

Troubleshooting VSCode Connection

Problem: Python packages not found Solution: Ensure you selected the correct interpreter and that the module is loaded

Problem: Jupyter kernel not starting Solution: Check that ipykernel is installed in your environment:

pip install ipykernel

Problem: Can't find conda environments Solution: Make sure conda is initialized in your shell:

conda init bash

Installed Packages by Thematic Area

This base environment includes the following packages, organized by scientific domain:

1. Scientific Computing & Data Analysis

Core scientific computing and data manipulation libraries:

  • numpy - Numerical computing with N-dimensional arrays
  • scipy - Scientific computing algorithms (optimization, integration, signal processing)
  • pandas - Data manipulation and analysis with DataFrames
  • scikit-learn - Machine learning algorithms and tools
  • statsmodels - Statistical modeling and econometrics
  • joblib - Parallel computing and caching utilities

2. Data Visualization

Plotting and visualization libraries:

  • matplotlib / matplotlib-base - Comprehensive plotting library
  • seaborn - Statistical data visualization
  • plotnine - Grammar of graphics (ggplot2-style) for Python

3. Image Processing

Computer vision and image manipulation:

  • pillow - Python Imaging Library (image I/O and manipulation)
  • opencv - Computer vision and image processing
  • scikit-image - Image processing algorithms

4. Molecular Dynamics & Simulation

Tools for molecular dynamics trajectory analysis and simulation:

  • mdtraj - Read, write, and analyze MD trajectories
  • mdanalysis - Analysis of molecular dynamics simulations
  • parmed - Parameter/topology file editor for molecular simulations
  • openmm - High-performance molecular simulation toolkit
  • gromacswrapper - Python wrapper for GROMACS

5. Computational Chemistry & Quantum Chemistry

Quantum chemistry, cheminformatics, and molecular modeling:

  • pyscf - Python-based quantum chemistry package
  • rdkit - Cheminformatics and machine learning toolkit
  • cclib - Parse and interpret computational chemistry log files
  • qcengine - Quantum chemistry program executor and IO standardizer
  • geometric - Geometry optimization tool
  • openbabel - Chemical toolbox for file conversion and analysis
  • ase - Atomic Simulation Environment (atomistic simulations)

6. Molecular Visualization

3D molecular structure visualization:

  • pymol-open-source - Molecular visualization system
  • nglview - Interactive molecular viewer for Jupyter notebooks

7. Bioinformatics & Structural Biology

Biological sequence analysis, structural biology, and single-cell genomics:

  • biotite - Computational molecular biology toolkit
  • prody - Protein structural dynamics analysis
  • scanpy - Single-cell analysis in Python
  • anndata - Annotated data matrices for single-cell analysis
  • ete3 - Phylogenomics and tree visualization
  • toytree - Phylogenetic tree plotting

8. Data I/O & File Handling

Reading and writing various scientific and office file formats:

  • netcdf4 - NetCDF file format (used in climate science, MD simulations)
  • xlrd - Read Excel files (.xls)
  • pyxlsb - Read Excel binary format (.xlsb)
  • openpyxl - Read/write Excel 2010+ files (.xlsx)
  • pyexcel - Unified Excel file manipulation
  • python-docx - Create and modify Word documents (.docx)
  • python-pptx - Create and modify PowerPoint presentations (.pptx)
  • odfpy - Read/write OpenDocument Format files (.odf, .ods)

9. Interactive Computing & Jupyter

Interactive development and notebook environments:

  • jupyterlab - Web-based interactive development environment
  • ipykernel - IPython kernel for Jupyter

10. Development Tools

Compilation and package development utilities:

  • cython - C-extensions for Python (compile Python to C)
  • setuptools - Package development and distribution
  • pip - Package installer for Python

11. Parallel Computing

Distributed and parallel computing:

  • mpi4py - Python bindings for MPI (Message Passing Interface)

12. Package Management

Environment and package managers:

  • conda (25.3.1) - Package, dependency, and environment management
  • mamba (2.1.1) - Fast, drop-in replacement for conda
  • python (3.12) - Python interpreter

Complete List of Explicitly Installed Packages

For reference, here is the complete list of packages that were explicitly requested during installation:

dependencies:
  - python=3.12
  - conda==25.3.1
  - mamba==2.1.1
  - pip
  - cython
  - numpy
  - pandas
  - matplotlib
  - seaborn
  - mdtraj
  - mdanalysis
  - ipykernel
  - scipy
  - matplotlib-base
  - setuptools
  - joblib
  - parmed
  - netcdf4
  - scikit-learn
  - mpi4py
  - openmm
  - statsmodels
  - plotnine
  - xlrd
  - pyxlsb
  - openpyxl
  - scikit-image
  - jupyterlab
  - pyscf
  - ase
  - rdkit
  - cclib
  - qcengine
  - geometric
  - openbabel
  - gromacswrapper
  - pymol-open-source
  - nglview
  - pyexcel
  - python-docx
  - python-pptx
  - odfpy
  - pillow
  - opencv
  - ete3
  - toytree
  - biotite
  - prody
  - scanpy
  - anndata

Getting Help

Module Information

# Show module details
module show miniforge3/20250911

# List available modules
module avail

Package Documentation

Most packages have comprehensive documentation online:

  • NumPy/SciPy: https://numpy.org/doc/, https://scipy.org/doc/
  • pandas: https://pandas.pydata.org/docs/
  • matplotlib: https://matplotlib.org/stable/
  • MDAnalysis: https://docs.mdanalysis.org/
  • MDTraj: https://mdtraj.org/
  • RDKit: https://rdkit.org/docs/
  • PySCF: https://pyscf.org/
  • OpenMM: https://openmm.org/
  • Scanpy: https://scanpy.readthedocs.io/

Conda/Mamba Help

# Conda help
conda --help
conda create --help
conda install --help

# Mamba help
mamba --help
mamba install --help

# Search for packages
mamba search package-name

# Show package information
mamba info package-name

Best Practices

  1. Use mamba instead of conda - It's much faster for dependency resolution
  2. Create separate environments for different projects - Avoids dependency conflicts
  3. Export your environment - Use conda env export > environment.yml for reproducibility
  4. Use pip only when necessary - Prefer conda packages to avoid conflicts
  5. Keep environments small - Install only what you need
  6. Document your dependencies - Keep a requirements.txt or environment.yml file

Last Updated: January 2025 Environment Location: /sw/miniforge3/20250911 Module: miniforge3/20250911

For questions or issues, contact the cluster administrators.