Skip to content

InterProScan 5.75-106.0 - User Guide

Cluster: XLence (UNIMI Dipartimento di Scienze Farmacologiche e Biomolecolari) Date: October 12, 2025 Installation: /sw/interproscan/5.75-106.0/


πŸ“¦ Version Information

  • Version: 5.75-106.0
  • Build: 64-bit (requires Java 11+)
  • Release Date: June 2024
  • Size: ~6.7 GB (includes all databases)

πŸ”§ Quick Start

Load the Module

module load interproscan
interproscan.sh --version

Basic Usage

# Analyze protein sequences
interproscan.sh -i my_proteins.fasta -f tsv -o results.tsv

# Multiple output formats
interproscan.sh -i sequences.fasta -f tsv,gff3,json -o output_prefix

# Specific applications only (faster)
interproscan.sh -i proteins.fasta -appl Pfam,SMART -f tsv -o pfam_smart.tsv

πŸ“Š Available Databases & Applications

InterProScan integrates these protein signature databases:

Database Type Description
Pfam HMM Protein families database
PRINTS Fingerprint Protein motif fingerprints
ProDom Domain Protein domain database
SMART HMM Simple Modular Architecture Research Tool
TIGRFAMs HMM Protein families for prokaryotes
PIRSF HMM Protein classification system
SUPERFAMILY HMM Structural assignments for proteins
Gene3D HMM CATH protein domain predictions
HAMAP Profile High-quality Automated Annotation of Microbial Proteomes
ProSitePatterns Regex Protein domains, families and sites
ProSiteProfiles Profile Generalized profiles
Coils Prediction Coiled coil regions
MobiDBLite Prediction Protein disorder and mobility
SFLD HMM Structure-Function Linkage Database

πŸ’‘ Common Use Cases

1. Standard Protein Annotation

# Run all available analyses
interproscan.sh -i input.fasta -f tsv,gff3,html -o annotation

# Output files:
# - annotation.tsv      (Tab-separated values)
# - annotation.gff3     (GFF3 genomic format)
# - annotation.html.tar.gz (Visual HTML report)

2. Domain Architecture Analysis

# Focus on domain databases
interproscan.sh -i proteins.fasta \
  -appl Pfam,SMART,Gene3D,SUPERFAMILY \
  -f tsv -o domain_architecture.tsv

3. Functional Classification

# Get GO terms and pathways
interproscan.sh -i sequences.fasta \
  -f tsv -goterms -pathways \
  -o functional_annotation.tsv

4. High-Throughput Analysis

# Use multiple CPUs for large datasets
interproscan.sh -i large_proteome.fasta \
  -f tsv -cpu 8 \
  -T /scratch/interproscan_tmp \
  -o proteome_results.tsv

🎯 Output Formats

InterProScan supports multiple output formats:

  • TSV (-f tsv): Tab-separated values (default)
  • GFF3 (-f gff3): Genomic feature format
  • JSON (-f json): JSON format
  • XML (-f xml): XML format
  • HTML (-f html): Visual HTML report (compressed)
  • SVG (-f svg): Graphical representation

TSV Output Columns

  1. Protein accession
  2. Sequence MD5 digest
  3. Sequence length
  4. Analysis method
  5. Signature accession
  6. Signature description
  7. Start location
  8. Stop location
  9. E-value
  10. Match status
  11. Date
  12. InterPro accession (if available)
  13. InterPro description (if available)
  14. GO annotations (if -goterms used)
  15. Pathways (if -pathways used)

βš™οΈ Important Options

Performance Tuning

# Number of parallel threads (default: 8)
-cpu 16

# Temporary directory for intermediate files
-T /scratch/tmp_interproscan

# Disable pre-calculation (faster for few sequences)
-dp

Application Selection

# Run specific applications only
-appl Pfam,SMART,Gene3D

# Exclude specific applications
-exclappl Coils,MobiDBLite

# List all available applications
interproscan.sh -appl

Sequence Input

# FASTA file
-i sequences.fasta

# Skip sequences already in output (resume interrupted run)
-resume

# Input type (default: auto-detect)
-t p    # protein
-t n    # nucleic acid

πŸš€ Slurm Job Examples

Single Node Job

#!/bin/bash
#SBATCH --job-name=interproscan
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=32G
#SBATCH --time=24:00:00
#SBATCH --output=interproscan_%j.log

module load interproscan

# Set temp directory to job-specific location
export INTERPROSCAN_TMP=/scratch/$SLURM_JOB_ID

interproscan.sh \
  -i my_proteome.fasta \
  -f tsv,gff3 \
  -goterms -pathways \
  -cpu $SLURM_CPUS_PER_TASK \
  -T $INTERPROSCAN_TMP \
  -o results_$SLURM_JOB_ID.tsv

Array Job for Multiple Files

#!/bin/bash
#SBATCH --job-name=interproscan_array
#SBATCH --array=1-100
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=16G
#SBATCH --time=12:00:00

module load interproscan

# Get input file for this array task
INPUT=$(ls input_files/*.fasta | sed -n ${SLURM_ARRAY_TASK_ID}p)
BASENAME=$(basename $INPUT .fasta)

interproscan.sh \
  -i $INPUT \
  -f tsv \
  -cpu $SLURM_CPUS_PER_TASK \
  -o results/${BASENAME}_interpro.tsv

πŸ” Troubleshooting

Memory Issues

InterProScan can be memory-intensive. Recommended memory per CPU:

  • Small proteins (<500 sequences): 2 GB per CPU
  • Medium datasets (500-5000): 4 GB per CPU
  • Large proteomes (>5000): 8+ GB per CPU
# Reduce parallelism if running out of memory
interproscan.sh -i large.fasta -cpu 4 -o output.tsv

Disk Space Issues

# Set temp directory to location with more space
export INTERPROSCAN_TMP=/scratch/my_interproscan_tmp
mkdir -p $INTERPROSCAN_TMP

interproscan.sh -i proteins.fasta -T $INTERPROSCAN_TMP -o results.tsv

Resume Interrupted Runs

# InterProScan can resume from interrupted runs
interproscan.sh -i sequences.fasta -resume -o results.tsv

Java Issues

# Check Java version (requires Java 11+)
java -version

# If Java issues occur, try setting JAVA_HOME
export JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64

πŸ“ˆ Performance Tips

  1. Choose specific applications instead of running all: bash # Faster: only essential databases interproscan.sh -i input.fasta -appl Pfam,SMART,Gene3D -o output.tsv

  2. Use appropriate CPU count:

  3. Too many CPUs β†’ memory issues
  4. Too few CPUs β†’ slow runtime
  5. Sweet spot: 8-16 CPUs for most datasets

  6. Use fast local storage for temp files: bash -T /local/scratch # faster than network storage

  7. Split large datasets into smaller chunks: bash # Process in batches of 1000 sequences split -l 2000 large.fasta chunk_ # 1000 seqs (2 lines each in FASTA)


πŸ”— Integration with Other Tools

Combine with BLAST Results

# Extract top BLAST hits, then annotate with InterProScan
module load blast interproscan

# 1. BLAST search
blastp -query query.fasta -db nr -out blast.tsv -outfmt 6

# 2. Extract unique subject sequences
cut -f2 blast.tsv | sort -u > subjects.txt
blastdbcmd -db nr -entry_batch subjects.txt -out subjects.fasta

# 3. Annotate with InterProScan
interproscan.sh -i subjects.fasta -f tsv -o annotations.tsv

Parse Results with Python

import pandas as pd

# Read TSV output
df = pd.read_csv('interproscan.tsv', sep='\t', header=None)
df.columns = ['acc', 'md5', 'len', 'db', 'sig_acc', 'sig_desc',
              'start', 'stop', 'evalue', 'status', 'date',
              'interpro_acc', 'interpro_desc', 'go', 'pathways']

# Filter for Pfam domains
pfam = df[df['db'] == 'Pfam']

# Count domain occurrences
domain_counts = pfam['sig_acc'].value_counts()
print(domain_counts.head())

πŸ“š References & Resources

  • Official Documentation: https://interproscan-docs.readthedocs.io/
  • InterPro Website: https://www.ebi.ac.uk/interpro/
  • GitHub: https://github.com/ebi-pf-team/interproscan
  • Publication: Jones P, et al. (2014) Bioinformatics 30(9):1236-1240

πŸ†˜ Support

  • Cluster Support: uliano.guerrini@unimi.it
  • InterProScan Issues: https://github.com/ebi-pf-team/interproscan/issues
  • Module Location: /opt/modulefiles/interproscan/5.75-106.0.lua

Installation Details: - Path: /sw/interproscan/5.75-106.0/ - Module: module load interproscan - Type: Precompiled 64-bit binaries - Dependencies: Java 11+ (system) - Database Version: 106.0 (June 2024)