Skip to content

NAMD 3.0.2 - Installation and Usage Guide

Cluster: XLence (UNIMI Dipartimento di Scienze Farmacologiche e Biomolecolari) Date: October 11, 2025 Location: /sw/namd/


πŸ“¦ Installed Versions

Two NAMD 3.0.2 builds are available, both with CUDA 11.8 GPU support:

1. multicore-CUDA - Single-node optimized

  • Module: namd/3.0
  • Path: /sw/namd/NAMD_3.0.2_Linux-x86_64-multicore-CUDA/
  • Best for: Login nodes, single compute node, testing, development
  • Parallelism: Multi-threaded (shared memory)
  • Performance: ~5-10% faster on single node (no network overhead)
  • Limitations: ❌ No replica-exchange, ❌ No multi-node

2. netlrts-smp-CUDA - Multi-node capable

  • Module: namd-multinode/3.0
  • Path: /sw/namd/NAMD_3.0.2_Linux-x86_64-netlrts-smp-CUDA/
  • Best for: Multi-node cluster jobs, replica-exchange MD
  • Parallelism: MPI-like processes + threading + comm threads
  • Features: Network layer, supports multi-node scaling, replica-exchange
  • Flexibility: βœ… Works on single node (with small overhead)

Both versions work on single nodes! netlrts can run everywhere but multicore is slightly more efficient for pure single-node work.

For detailed comparison see: /opt/admin/claude/docs/namd-versions-comparison.md


πŸš€ Quick Start

Loading the Module

The namd/3.0 module automatically selects the best version based on context:

module load namd/3.0

# Check which version was selected
echo $NAMD_BUILD  # "multicore" or "netlrts"

Selection logic: - Login nodes (interactive) β†’ multicore (best single-node performance) - Slurm job (single node) β†’ netlrts (flexibility for replica-exchange) - Slurm job (multi-node) β†’ netlrts (required for inter-node communication)

Or load specific version:

# Force multicore (single-node only)
module load namd/3.0          # Auto-selects, or force with NAMD_ROOT

# Force netlrts (multi-node capable)
module load namd-multinode/3.0

πŸ’» Usage Examples

1. Single Node (Login Node or Interactive)

# Load module
module load namd/3.0

# Run with 16 CPU cores, 1 GPU
namd3 +p16 +devices 0 simulation.conf

# Run with all available cores
namd3 +auto-provision +devices 0 simulation.conf

# Multi-GPU (e.g., on login nodes with 2 GPUs)
namd3 +p24 +devices 0,1 simulation.conf

Output (multicore):

Charm++> Running in Multicore mode: 16 threads (PEs)
Info: NAMD 3.0.2 for Linux-x86_64-multicore-CUDA
Pe 0 binding to CUDA device 0 on xlence: 'NVIDIA GeForce RTX 2060 SUPER'

2. Single Compute Node (via Slurm)

#!/bin/bash
#SBATCH --job-name=namd_single
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=10
#SBATCH --gres=gpu:2
#SBATCH --time=24:00:00

module load namd/3.0

# Run with 10 cores, 2 GPUs
namd3 +p10 +devices 0,1 simulation.conf

Output (netlrts in Slurm):

Charm++> Running in SMP mode: 1 processes, 10 worker threads + 1 comm thread
Info: NAMD 3.0.2 for Linux-x86_64-netlrts-smp-CUDA

3. Multi-Node Cluster Job (via Slurm)

Requires netlrts version:

#!/bin/bash
#SBATCH --job-name=namd_multi
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=10
#SBATCH --gres=gpu:2
#SBATCH --time=48:00:00

module load namd-multinode/3.0

# Multi-node run: 4 nodes Γ— 10 cores Γ— 2 GPUs = 40 PEs total
srun namd3 +ppn 10 +devices 0,1 simulation.conf

Key options: - +ppn N: Number of PEs (worker threads) per node - +devices 0,1: Use GPUs 0 and 1 on each node - srun: Slurm launcher (alternative to charmrun)


4. Replica-Exchange Molecular Dynamics (REMD)

IMPORTANT: Replica-exchange requires netlrts version. Multicore doesn't support it.

Single-Node REMD

#!/bin/bash
#SBATCH --job-name=namd_remd
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=16
#SBATCH --gres=gpu:1
#SBATCH --time=48:00:00

module load namd-multinode/3.0

# 8 replicas, 2 cores per replica = 16 total cores
charmrun namd3 ++local +p16 +replicas 8 +stdout rep-%d.log simulation.conf

Output files: rep-0.log, rep-1.log, ..., rep-7.log

Multi-Node REMD

#!/bin/bash
#SBATCH --job-name=namd_remd_multi
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=10
#SBATCH --gres=gpu:2
#SBATCH --time=48:00:00

module load namd-multinode/3.0

# 16 replicas across 4 nodes, 4 replicas per node
srun -n 4 namd3 +ppn 10 +replicas 16 +stdout rep-%d.log simulation.conf

πŸ”§ Important Options

CPU Control

Option Description
+p N Use N processor elements (threads/cores)
+auto-provision Use all available CPU cores
++local Force single-node mode (with charmrun)
+ppn N Cores per node (for multi-node jobs)

GPU Control

Option Description
+devices 0 Use GPU 0 only
+devices 0,1 Use GPUs 0 and 1
+devices all Use all available GPUs (default)
+ignoresharing Allow multiple NAMD processes per GPU

Replica Exchange

Option Description
+replicas N Number of replicas (must evenly divide +p)
+stdout file-%d.log Separate output file per replica

πŸ“ Directory Structure

/sw/namd/
β”œβ”€β”€ NAMD_3.0.2_Linux-x86_64-multicore-CUDA/
β”‚   β”œβ”€β”€ namd3                    # Main executable
β”‚   β”œβ”€β”€ charmrun                 # Minimal launcher stub
β”‚   β”œβ”€β”€ psfgen                   # Structure file generator
β”‚   β”œβ”€β”€ lib/                     # Tcl libraries, scripts
β”‚   β”‚   β”œβ”€β”€ replica/             # Replica exchange examples
β”‚   β”‚   └── namdcph/             # Constant pH scripts
β”‚   β”œβ”€β”€ announce.txt             # Release notes
β”‚   β”œβ”€β”€ notes.txt                # Detailed documentation
β”‚   └── license.txt              # License
β”‚
β”œβ”€β”€ NAMD_3.0.2_Linux-x86_64-netlrts-smp-CUDA/
β”‚   β”œβ”€β”€ namd3                    # Main executable
β”‚   β”œβ”€β”€ charmrun                 # Full parallel launcher
β”‚   β”œβ”€β”€ sortreplicas             # Replica exchange analysis tool
β”‚   └── (same structure as multicore)
β”‚
└── GUIDE.md                     # This file

πŸ§ͺ Testing Your Installation

Basic Functionality Test

module load namd/3.0

# Test NAMD binary
namd3 +p1 2>&1 | grep "Info: NAMD"
# Expected: Info: NAMD 3.0.2 for Linux-x86_64-[multicore|netlrts]-CUDA

# Test GPU detection
namd3 +p1 2>&1 | grep "binding to CUDA device"
# Expected: Pe 0 binding to CUDA device 0 on <hostname>: 'NVIDIA GeForce RTX...'

Small Test Simulation

# Copy example input files
cp -r $NAMDDIR/lib/replica/alanin/ ~/test_namd/
cd ~/test_namd/alanin/

# Run short test
namd3 +p4 +devices 0 alanin.namd

Expected runtime: ~10 seconds for short test.


🌐 GPU Support

CUDA Version: 11.8 (built-in, no module needed)

Supported GPUs on XLence cluster: - Login nodes: 2Γ— NVIDIA GeForce RTX 2060 SUPER (8 GB, Compute 7.5) - Compute nodes (node1-5): 2Γ— NVIDIA GeForce RTX 2080 Ti (11 GB, Compute 7.5) - NGS node: 1Γ— NVIDIA GeForce RTX 2080 Ti (11 GB, Compute 7.5)

Total: 13 GPUs available cluster-wide.


βš™οΈ Advanced: Choosing Version Manually

If you need to override automatic selection:

# Force multicore version
export NAMD_ROOT=/sw/namd/NAMD_3.0.2_Linux-x86_64-multicore-CUDA
export PATH=$NAMD_ROOT:$PATH

# Force netlrts version
export NAMD_ROOT=/sw/namd/NAMD_3.0.2_Linux-x86_64-netlrts-smp-CUDA
export PATH=$NAMD_ROOT:$PATH

Or run directly without module:

/sw/namd/NAMD_3.0.2_Linux-x86_64-multicore-CUDA/namd3 +p16 simulation.conf

πŸ“Š Performance Tips

  1. Single node: Use multicore version for ~5-10% better performance
  2. GPU acceleration: Always use +devices to enable GPU (enabled by default)
  3. CPU count: Match +p to physical cores (avoid hyperthreading for MD)
  4. Multi-GPU: Use 1 GPU per 4-6 CPU cores for best balance
  5. Replica exchange: Use +replicas that evenly divides +p
  6. Large systems: Consider GPU-resident mode (see NAMD 3.0 user guide)

Performance Comparison

Scenario multicore netlrts Best Choice
Single node, no replica-exchange βœ… ~5-10% faster ⚠️ Baseline multicore
Single node, replica-exchange ❌ Not supported βœ… Only option netlrts
Multi-node ❌ Not supported βœ… Only option netlrts

πŸ› Troubleshooting

"No simulation config file specified"

Cause: Missing input file. Solution: Provide .namd or .conf file as argument: namd3 +p16 simulation.conf

"Cannot find CUDA device"

Cause: GPU not available or wrong device ID. Solution: Check available GPUs with nvidia-smi, use correct device ID with +devices N

"Charm++ error: CmiAbort called"

Cause: Memory error, simulation crash, or resource exhaustion. Solution: Check log file for error messages, reduce system size, or request more memory.

Slow single-node performance

Cause: Using netlrts version interactively on login node. Solution: Module should auto-select multicore for login nodes. Verify with echo $NAMD_BUILD.

Replica exchange not working

Cause: Using multicore version (doesn't support replica exchange). Solution: Use module load namd-multinode/3.0 to force netlrts version.

Multi-node job hangs or fails

Cause: Network configuration or nodelist issues. Solution: Use srun with Slurm instead of manual charmrun nodelist.

"Number of replicas must divide +p"

Cause: Replica count doesn't evenly divide processor count. Solution: Adjust +p or +replicas so one divides the other evenly.


πŸ“š Documentation and Support

Local Documentation: - Version comparison: /opt/admin/claude/docs/namd-versions-comparison.md - Release notes: $NAMDDIR/announce.txt - Detailed notes: $NAMDDIR/notes.txt

Online Resources: - NAMD Homepage: http://www.ks.uiuc.edu/Research/namd/ - NAMD 3.0 Documentation: http://www.ks.uiuc.edu/Research/namd/3.0/ - User Guide (PDF): http://www.ks.uiuc.edu/Research/namd/3.0/ug/ - Tutorials: http://www.ks.uiuc.edu/Training/Tutorials/ - Mailing List: NAMD-L (namd-l@ks.uiuc.edu)

Charm++ Documentation: - Manual: https://charm.readthedocs.io/ - Running programs: https://charm.readthedocs.io/en/latest/charm++/manual.html#running-charm-programs


πŸ“‹ Version-Specific Details

Mono-Node Version (namd/3.0)

Charm++ Runtime: multicore (shared memory threading)

Architecture:

Charm++> Running in Multicore mode: N threads (PEs)
Charm++> Running on 1 hosts (X sockets x Y cores x Z PUs = Total-way SMP)

Characteristics: - Multi-threaded using shared memory (OpenMP-like) - Single-node ONLY - cannot communicate across multiple nodes - No network layer - lower overhead - ~5-10% faster than netlrts on single node - ❌ Does NOT support replica-exchange MD

When to Use: - Running on a single node (login or compute) - System fits on 1-2 GPUs - No replica exchange needed - Want maximum single-node performance - Testing and development

Switching to multi-node version:

module unload namd
module load namd-multinode/3.0

Multi-Node Version (namd-multinode/3.0)

Charm++ Runtime: netlrts-smp (Network + SMP + comm threads)

Architecture:

Charm++> Running in SMP mode: N processes, M worker threads + K comm threads per process
Charm++> The comm. thread both sends and receives messages
Charm++> scheduler running in netpoll mode

Characteristics: - Multi-node capable via TCP/IP network layer - Hybrid parallelism: MPI-like processes + threading within each process - Dedicated communication threads for network I/O - βœ… Required for replica-exchange MD (REMD) - βœ… Can also run on single node (with ~5-10% overhead)

When to Use: - Running across multiple compute nodes - Need replica-exchange MD (REMD) - REQUIRED - Need >2 GPUs distributed across nodes - Large production simulations requiring cluster-wide resources - Multi-copy algorithms (FEP with multiple windows, umbrella sampling)

Replica Exchange Tools: - sortreplicas: Analysis tool for replica exchange trajectories - Enhanced lib/replica/ examples

Switching to single-node version:

module unload namd-multinode
module load namd/3.0

πŸ“ Citation

If you use NAMD in your research, please cite:

Phillips, J.C., Hardy, D.J., Maia, J.D.C. et al. Scalable molecular dynamics on CPU and GPU architectures with NAMD. J. Chem. Phys. 153, 044130 (2020). doi:10.1063/5.0014475


πŸ”„ Version History

Date Version Changes
2025-10-11 3.0.2 Initial installation, both multicore and netlrts builds

For questions or issues contact: Cluster administrators Installation date: October 11, 2025 Installed by: Claude Code + Uliano Guerrini