NAMD 3.0.2 - Installation and Usage Guide
Cluster: XLence (UNIMI Dipartimento di Scienze Farmacologiche e Biomolecolari)
Date: October 11, 2025
Location: /sw/namd/
π¦ Installed Versions
Two NAMD 3.0.2 builds are available, both with CUDA 11.8 GPU support:
1. multicore-CUDA - Single-node optimized
- Module:
namd/3.0 - Path:
/sw/namd/NAMD_3.0.2_Linux-x86_64-multicore-CUDA/ - Best for: Login nodes, single compute node, testing, development
- Parallelism: Multi-threaded (shared memory)
- Performance: ~5-10% faster on single node (no network overhead)
- Limitations: β No replica-exchange, β No multi-node
2. netlrts-smp-CUDA - Multi-node capable
- Module:
namd-multinode/3.0 - Path:
/sw/namd/NAMD_3.0.2_Linux-x86_64-netlrts-smp-CUDA/ - Best for: Multi-node cluster jobs, replica-exchange MD
- Parallelism: MPI-like processes + threading + comm threads
- Features: Network layer, supports multi-node scaling, replica-exchange
- Flexibility: β Works on single node (with small overhead)
Both versions work on single nodes! netlrts can run everywhere but multicore is slightly more efficient for pure single-node work.
For detailed comparison see: /opt/admin/claude/docs/namd-versions-comparison.md
π Quick Start
Loading the Module
The namd/3.0 module automatically selects the best version based on context:
module load namd/3.0
# Check which version was selected
echo $NAMD_BUILD # "multicore" or "netlrts"
Selection logic: - Login nodes (interactive) β multicore (best single-node performance) - Slurm job (single node) β netlrts (flexibility for replica-exchange) - Slurm job (multi-node) β netlrts (required for inter-node communication)
Or load specific version:
# Force multicore (single-node only)
module load namd/3.0 # Auto-selects, or force with NAMD_ROOT
# Force netlrts (multi-node capable)
module load namd-multinode/3.0
π» Usage Examples
1. Single Node (Login Node or Interactive)
# Load module
module load namd/3.0
# Run with 16 CPU cores, 1 GPU
namd3 +p16 +devices 0 simulation.conf
# Run with all available cores
namd3 +auto-provision +devices 0 simulation.conf
# Multi-GPU (e.g., on login nodes with 2 GPUs)
namd3 +p24 +devices 0,1 simulation.conf
Output (multicore):
Charm++> Running in Multicore mode: 16 threads (PEs)
Info: NAMD 3.0.2 for Linux-x86_64-multicore-CUDA
Pe 0 binding to CUDA device 0 on xlence: 'NVIDIA GeForce RTX 2060 SUPER'
2. Single Compute Node (via Slurm)
#!/bin/bash
#SBATCH --job-name=namd_single
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=10
#SBATCH --gres=gpu:2
#SBATCH --time=24:00:00
module load namd/3.0
# Run with 10 cores, 2 GPUs
namd3 +p10 +devices 0,1 simulation.conf
Output (netlrts in Slurm):
Charm++> Running in SMP mode: 1 processes, 10 worker threads + 1 comm thread
Info: NAMD 3.0.2 for Linux-x86_64-netlrts-smp-CUDA
3. Multi-Node Cluster Job (via Slurm)
Requires netlrts version:
#!/bin/bash
#SBATCH --job-name=namd_multi
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=10
#SBATCH --gres=gpu:2
#SBATCH --time=48:00:00
module load namd-multinode/3.0
# Multi-node run: 4 nodes Γ 10 cores Γ 2 GPUs = 40 PEs total
srun namd3 +ppn 10 +devices 0,1 simulation.conf
Key options:
- +ppn N: Number of PEs (worker threads) per node
- +devices 0,1: Use GPUs 0 and 1 on each node
- srun: Slurm launcher (alternative to charmrun)
4. Replica-Exchange Molecular Dynamics (REMD)
IMPORTANT: Replica-exchange requires netlrts version. Multicore doesn't support it.
Single-Node REMD
#!/bin/bash
#SBATCH --job-name=namd_remd
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=16
#SBATCH --gres=gpu:1
#SBATCH --time=48:00:00
module load namd-multinode/3.0
# 8 replicas, 2 cores per replica = 16 total cores
charmrun namd3 ++local +p16 +replicas 8 +stdout rep-%d.log simulation.conf
Output files: rep-0.log, rep-1.log, ..., rep-7.log
Multi-Node REMD
#!/bin/bash
#SBATCH --job-name=namd_remd_multi
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=10
#SBATCH --gres=gpu:2
#SBATCH --time=48:00:00
module load namd-multinode/3.0
# 16 replicas across 4 nodes, 4 replicas per node
srun -n 4 namd3 +ppn 10 +replicas 16 +stdout rep-%d.log simulation.conf
π§ Important Options
CPU Control
| Option | Description |
|---|---|
+p N |
Use N processor elements (threads/cores) |
+auto-provision |
Use all available CPU cores |
++local |
Force single-node mode (with charmrun) |
+ppn N |
Cores per node (for multi-node jobs) |
GPU Control
| Option | Description |
|---|---|
+devices 0 |
Use GPU 0 only |
+devices 0,1 |
Use GPUs 0 and 1 |
+devices all |
Use all available GPUs (default) |
+ignoresharing |
Allow multiple NAMD processes per GPU |
Replica Exchange
| Option | Description |
|---|---|
+replicas N |
Number of replicas (must evenly divide +p) |
+stdout file-%d.log |
Separate output file per replica |
π Directory Structure
/sw/namd/
βββ NAMD_3.0.2_Linux-x86_64-multicore-CUDA/
β βββ namd3 # Main executable
β βββ charmrun # Minimal launcher stub
β βββ psfgen # Structure file generator
β βββ lib/ # Tcl libraries, scripts
β β βββ replica/ # Replica exchange examples
β β βββ namdcph/ # Constant pH scripts
β βββ announce.txt # Release notes
β βββ notes.txt # Detailed documentation
β βββ license.txt # License
β
βββ NAMD_3.0.2_Linux-x86_64-netlrts-smp-CUDA/
β βββ namd3 # Main executable
β βββ charmrun # Full parallel launcher
β βββ sortreplicas # Replica exchange analysis tool
β βββ (same structure as multicore)
β
βββ GUIDE.md # This file
π§ͺ Testing Your Installation
Basic Functionality Test
module load namd/3.0
# Test NAMD binary
namd3 +p1 2>&1 | grep "Info: NAMD"
# Expected: Info: NAMD 3.0.2 for Linux-x86_64-[multicore|netlrts]-CUDA
# Test GPU detection
namd3 +p1 2>&1 | grep "binding to CUDA device"
# Expected: Pe 0 binding to CUDA device 0 on <hostname>: 'NVIDIA GeForce RTX...'
Small Test Simulation
# Copy example input files
cp -r $NAMDDIR/lib/replica/alanin/ ~/test_namd/
cd ~/test_namd/alanin/
# Run short test
namd3 +p4 +devices 0 alanin.namd
Expected runtime: ~10 seconds for short test.
π GPU Support
CUDA Version: 11.8 (built-in, no module needed)
Supported GPUs on XLence cluster: - Login nodes: 2Γ NVIDIA GeForce RTX 2060 SUPER (8 GB, Compute 7.5) - Compute nodes (node1-5): 2Γ NVIDIA GeForce RTX 2080 Ti (11 GB, Compute 7.5) - NGS node: 1Γ NVIDIA GeForce RTX 2080 Ti (11 GB, Compute 7.5)
Total: 13 GPUs available cluster-wide.
βοΈ Advanced: Choosing Version Manually
If you need to override automatic selection:
# Force multicore version
export NAMD_ROOT=/sw/namd/NAMD_3.0.2_Linux-x86_64-multicore-CUDA
export PATH=$NAMD_ROOT:$PATH
# Force netlrts version
export NAMD_ROOT=/sw/namd/NAMD_3.0.2_Linux-x86_64-netlrts-smp-CUDA
export PATH=$NAMD_ROOT:$PATH
Or run directly without module:
/sw/namd/NAMD_3.0.2_Linux-x86_64-multicore-CUDA/namd3 +p16 simulation.conf
π Performance Tips
- Single node: Use multicore version for ~5-10% better performance
- GPU acceleration: Always use
+devicesto enable GPU (enabled by default) - CPU count: Match
+pto physical cores (avoid hyperthreading for MD) - Multi-GPU: Use 1 GPU per 4-6 CPU cores for best balance
- Replica exchange: Use
+replicasthat evenly divides+p - Large systems: Consider GPU-resident mode (see NAMD 3.0 user guide)
Performance Comparison
| Scenario | multicore | netlrts | Best Choice |
|---|---|---|---|
| Single node, no replica-exchange | β ~5-10% faster | β οΈ Baseline | multicore |
| Single node, replica-exchange | β Not supported | β Only option | netlrts |
| Multi-node | β Not supported | β Only option | netlrts |
π Troubleshooting
"No simulation config file specified"
Cause: Missing input file.
Solution: Provide .namd or .conf file as argument: namd3 +p16 simulation.conf
"Cannot find CUDA device"
Cause: GPU not available or wrong device ID.
Solution: Check available GPUs with nvidia-smi, use correct device ID with +devices N
"Charm++ error: CmiAbort called"
Cause: Memory error, simulation crash, or resource exhaustion. Solution: Check log file for error messages, reduce system size, or request more memory.
Slow single-node performance
Cause: Using netlrts version interactively on login node.
Solution: Module should auto-select multicore for login nodes. Verify with echo $NAMD_BUILD.
Replica exchange not working
Cause: Using multicore version (doesn't support replica exchange).
Solution: Use module load namd-multinode/3.0 to force netlrts version.
Multi-node job hangs or fails
Cause: Network configuration or nodelist issues.
Solution: Use srun with Slurm instead of manual charmrun nodelist.
"Number of replicas must divide +p"
Cause: Replica count doesn't evenly divide processor count.
Solution: Adjust +p or +replicas so one divides the other evenly.
π Documentation and Support
Local Documentation:
- Version comparison: /opt/admin/claude/docs/namd-versions-comparison.md
- Release notes: $NAMDDIR/announce.txt
- Detailed notes: $NAMDDIR/notes.txt
Online Resources: - NAMD Homepage: http://www.ks.uiuc.edu/Research/namd/ - NAMD 3.0 Documentation: http://www.ks.uiuc.edu/Research/namd/3.0/ - User Guide (PDF): http://www.ks.uiuc.edu/Research/namd/3.0/ug/ - Tutorials: http://www.ks.uiuc.edu/Training/Tutorials/ - Mailing List: NAMD-L (namd-l@ks.uiuc.edu)
Charm++ Documentation: - Manual: https://charm.readthedocs.io/ - Running programs: https://charm.readthedocs.io/en/latest/charm++/manual.html#running-charm-programs
π Version-Specific Details
Mono-Node Version (namd/3.0)
Charm++ Runtime: multicore (shared memory threading)
Architecture:
Charm++> Running in Multicore mode: N threads (PEs)
Charm++> Running on 1 hosts (X sockets x Y cores x Z PUs = Total-way SMP)
Characteristics: - Multi-threaded using shared memory (OpenMP-like) - Single-node ONLY - cannot communicate across multiple nodes - No network layer - lower overhead - ~5-10% faster than netlrts on single node - β Does NOT support replica-exchange MD
When to Use: - Running on a single node (login or compute) - System fits on 1-2 GPUs - No replica exchange needed - Want maximum single-node performance - Testing and development
Switching to multi-node version:
module unload namd
module load namd-multinode/3.0
Multi-Node Version (namd-multinode/3.0)
Charm++ Runtime: netlrts-smp (Network + SMP + comm threads)
Architecture:
Charm++> Running in SMP mode: N processes, M worker threads + K comm threads per process
Charm++> The comm. thread both sends and receives messages
Charm++> scheduler running in netpoll mode
Characteristics: - Multi-node capable via TCP/IP network layer - Hybrid parallelism: MPI-like processes + threading within each process - Dedicated communication threads for network I/O - β Required for replica-exchange MD (REMD) - β Can also run on single node (with ~5-10% overhead)
When to Use: - Running across multiple compute nodes - Need replica-exchange MD (REMD) - REQUIRED - Need >2 GPUs distributed across nodes - Large production simulations requiring cluster-wide resources - Multi-copy algorithms (FEP with multiple windows, umbrella sampling)
Replica Exchange Tools:
- sortreplicas: Analysis tool for replica exchange trajectories
- Enhanced lib/replica/ examples
Switching to single-node version:
module unload namd-multinode
module load namd/3.0
π Citation
If you use NAMD in your research, please cite:
Phillips, J.C., Hardy, D.J., Maia, J.D.C. et al. Scalable molecular dynamics on CPU and GPU architectures with NAMD. J. Chem. Phys. 153, 044130 (2020). doi:10.1063/5.0014475
π Version History
| Date | Version | Changes |
|---|---|---|
| 2025-10-11 | 3.0.2 | Initial installation, both multicore and netlrts builds |
For questions or issues contact: Cluster administrators Installation date: October 11, 2025 Installed by: Claude Code + Uliano Guerrini