Interactive Jupyter Notebooks
This page explains how to run a Jupyter notebook on a compute node while viewing it in the browser on your own computer, as if it were running locally.
No prior knowledge of Slurm or SSH tunnels is required: everything is explained below. If something does not work, jump to "Troubleshooting" at the end of the page.
Large Notebooks
For a quick look at some data, editing a notebook, or plotting a small file, running
JupyterLab on xlence is perfectly fine. That is what a login node is for.
It stops being fine once the notebook starts doing real work. xlence is also where the
scheduler for the whole cluster runs, along with the accounting database and everyone
else's interactive sessions: a notebook that takes several gigabytes of memory there slows
down every other user. This is not hypothetical — a single notebook using 38 GB of RAM once
made the cluster unusable for everybody for hours.
A rule of thumb: if your notebook loads a large dataset, trains anything, or runs for more than a few minutes, move it to a compute node. That is what the rest of this page is about. Using the compute nodes is not red tape — it is what they are there for, and they are much bigger than the login node.
Before You Start
You need:
- your cluster username (in the form
name.surname, the same as your university account); - to be connected to the university network, or to the university VPN if you are working from home. From any other network the cluster will not answer: this is not a fault, it is deliberate;
- a terminal:
- Windows: open Terminal or PowerShell from the Start menu. The
sshcommand is already there, nothing to install; - macOS: open Terminal (in Applications -> Utilities);
- Linux: your usual terminal.
- Windows: open Terminal or PowerShell from the Start menu. The
Read This First: The Two Terminals
This is the part that confuses people most, so let us be explicit.
You need two terminal windows open at the same time, and they do two different jobs:
| What it is for | Which machine it is on | |
|---|---|---|
| TERMINAL 1 | running the notebook | it changes three times: starts on your computer, then xlence, then ngs |
| TERMINAL 2 | holding the tunnel open | stays on your computer the whole time |
Terminal 1 moves from machine to machine as you run the commands. For this reason every command block on this page has a label above it telling you which terminal you are in and which machine you are on. Always check that label before pasting.
How to Tell Where You Are
Look at the prompt, the text to the left of your cursor:
| What you see | Where you are |
|---|---|
| your computer's usual prompt | on your own computer |
name.surname@xlence:~$ |
on xlence, the login node |
name.surname@ngs:~$ |
on ngs, a compute node |
If in doubt, at any moment you can type:
hostname
and it will tell you the name of the machine you are currently on.
What We Are Going to Do
- Ask the cluster scheduler (Slurm) for a slice of a compute node.
- Start JupyterLab on it.
- Build a tunnel that carries the web page to your browser.
Step 1 — Log In to the Cluster
TERMINAL 1 — you are on your own computer
Open your first terminal window and type:
ssh name.surname@xlence.disfeb.unimi.it
replacing name.surname with your username.
The very first time you will be asked something like "The authenticity of host ... can't be
established. Are you sure you want to continue connecting?". Answer yes and press Enter.
This is normal and happens only once.
Then type your password. Nothing appears as you type, not even asterisks — that is normal, just type it and press Enter.
How to check it worked: the prompt now ends with @xlence:~$.
Step 2 — Request a Slice of a Compute Node
TERMINAL 1 — you are now on
xlence
srun -p ngs -c 4 --mem=32G -t 4:00:00 --pty bash -l
What each part means:
| part | meaning |
|---|---|
srun |
"give me some resources and run something on them" |
-p ngs |
on the ngs partition, the node with lots of memory and no GPU |
-c 4 |
4 CPU cores |
--mem=32G |
32 GB of memory |
-t 4:00:00 |
for 4 hours (format hours:minutes:seconds) |
--pty bash -l |
"give me an interactive shell", i.e. a prompt to type into |
Choose these numbers honestly. The ngs node has 18 cores and 251 GB of memory in
total, and it is shared. Ask for what you actually need: ask for everything and you block
your colleagues; ask for too little and your work runs slowly or gets killed.
The time limit matters. When -t expires your notebook is shut down without warning, so
leave some margin. But do not inflate it either: until you release the resources, nobody
else can use them.
How to check it worked: after a few seconds the prompt changes and ends with @ngs:~$.
The same terminal window has taken you to a different machine.
If instead it sits there printing something like
srun: job ... queued and waiting for resources, the node is busy and you are in the queue. Wait, or press Ctrl+C and try again asking for less.
Step 3 — Start JupyterLab
TERMINAL 1 — you are now on
ngs, the compute node
Run these commands one at a time, in this same window.
3a. Make Python and Jupyter available:
module load miniforge3
(If you have your own conda environment, add conda activate myenvironment now.)
3b. Work out your port number and print the command you will need in a moment:
PORT=$((10000 + $(id -u) % 10000))
echo "ssh -N -L $PORT:localhost:$PORT -J $USER@xlence.disfeb.unimi.it $USER@$(hostname)"
That second line prints a command without running it. It will look something like:
ssh -N -L 12345:localhost:12345 -J name.surname@xlence.disfeb.unimi.it name.surname@ngs
Select it and copy it now. You will need it in Step 4, in the other terminal.
That number (
12345in the example) is a port: a number identifying your notebook. It is derived from your user ID, so it is always the same for you and never collides with your colleagues'.
3c. Start the notebook server:
jupyter lab --no-browser --ip=127.0.0.1 --port=$PORT
It will print a number of lines. At the end you will find an address like this:
To access the server, open this file in a browser:
...
Or copy and paste one of these URLs:
http://127.0.0.1:12345/lab?token=a1b2c3d4e5f6a1b2c3d4e5f6...
Copy that address in full, token included. You will need it in Step 5.
At this point you have copied two things: the command from 3b and the address from 3c. Keep both somewhere handy, for example in a text editor.
This window is now busy running Jupyter and will not give you a prompt back. That is correct. Do not close it and do not press Ctrl+C: as long as it stays open, your notebook is alive.
Step 4 — Open the Tunnel
TERMINAL 2 — open a NEW window; you are back on your own computer
Leave Terminal 1 exactly where it is, with Jupyter running.
Open a second terminal window on your own computer and paste the command you copied in Step 3b:
ssh -N -L 12345:localhost:12345 -J name.surname@xlence.disfeb.unimi.it name.surname@ngs
It will ask for your password (once or twice — that is normal, it is making two hops, first
to xlence and then to the compute node).
Then it will appear to hang: no message, no prompt coming back. That is correct. That terminal is not stuck, it is holding the tunnel open. Leave it alone and do not close it.
What is a tunnel, in plain words? The notebook runs on a machine your browser cannot reach. The tunnel says: "anything arriving on port 12345 of my computer, carry it to port 12345 of that machine over there". From then on your browser can pretend the notebook is local.
You now have two windows open, and both must stay open:
| What you see in it | |
|---|---|
| TERMINAL 1 | Jupyter's log lines, no prompt |
| TERMINAL 2 | nothing at all, no prompt |
They both look frozen. They are both working.
Step 5 — Open Your Browser
In your BROWSER — no terminal involved
Paste into the address bar the full address Jupyter printed in Step 3c, the one with
token=...:
http://127.0.0.1:12345/lab?token=a1b2c3d4e5f6...
JupyterLab opens. It is running on the cluster even though it looks local: the files you see are the ones in your cluster home directory, and the computation uses the cores you requested.
The token is your key: without it the page will not open. Do not share it.
When You Are Done — Shut Down Properly, In This Order
1. In the browser, save your notebooks and close the tab.
TERMINAL 2 — on your own computer
2. Press Ctrl+C. The tunnel closes and your prompt comes back. You can close this window now.
TERMINAL 1 — on
ngs
3. Press Ctrl+C. Jupyter asks for confirmation:
Shutdown this Jupyter server (y/[n])? — answer y and press Enter. You get the
@ngs:~$ prompt back.
4. Type:
exit
This releases the compute node resources and makes them available to others. The prompt
goes back to @xlence:~$.
TERMINAL 1 — you are back on
xlence
5. Type:
exit
again, to leave the cluster. Your own computer's prompt returns.
Step 4 is the one that really matters. If you just close the window without exiting, the resources stay reserved until the time limit you asked for expires, and nobody else can use them.
If In Doubt, Check
TERMINAL 1 — on
xlence(or reconnect with ssh)
squeue -u $USER
If no rows appear, you have released everything. If a row appears and you want to end it:
scancel <job number>
where <job number> is the number in the first column.
Troubleshooting
| What you see | In which terminal | What it means | What to do |
|---|---|---|---|
ssh: connect to host ... Connection timed out |
1 or 2 | you are not on the university network | connect to the university VPN |
Could not resolve hostname xlence.disfeb.unimi.it |
1 or 2 | your computer cannot resolve the name | use the numeric address instead: 159.149.32.133 (also inside the tunnel command) |
Permission denied, please try again |
1 or 2 | wrong password or username | try again; if it persists, contact the administrator |
Access denied by pam_slurm_adopt: you have no active jobs on this node |
2 | you are opening the tunnel but your job on the node is not running | Terminal 1 must still be open with Jupyter running. If it died, start again from Step 2 |
srun: job ... queued and waiting for resources |
1 | the node is busy | wait, or press Ctrl+C and retry asking for fewer cores/less memory |
Address already in use when starting Jupyter |
1 | you already have a notebook running somewhere | close the other one, or use --port=$((PORT+1)) — and remember to change the number in the tunnel command too |
bind: Address already in use when opening the tunnel |
2 | that port number is already taken on your own computer | close the old tunnel, or change only the first number: -L 12346:localhost:12345, and then use http://127.0.0.1:12346/lab?token=... in the browser |
| Browser says the site cannot be reached | — | the tunnel is not running | look at Terminal 2: if your prompt came back, the tunnel dropped. Run the command again |
| I lost the token | — | scroll up in Terminal 1 and copy it again from Jupyter's startup lines | |
| The notebook dies on its own after a while | 1 | the -t time limit expired, or you exceeded the memory you asked for |
start again from Step 2 with more time or more --mem |
| Everything closed when I shut my laptop | — | the SSH connection dropped and the notebook with it | see "Notebooks That Need to Run for a Long Time" below |
Useful Things to Know
Installing Your Own Python Packages
TERMINAL 1 — on
ngs(after Step 2), or onxlence
module load miniforge3 gives you a shared base environment which you cannot modify.
To have your own packages, create a personal environment once:
module load miniforge3
conda create -n myenvironment python=3.12 jupyterlab pandas matplotlib
conda activate myenvironment
From then on, in Step 3a, add conda activate myenvironment after module load miniforge3.
The environment lives in your home directory and takes up space: do not create one per project without a reason.
See also Miniforge3 for registering an environment as a Jupyter kernel, so you can switch kernels from inside JupyterLab.
If You Need a GPU
The ngs node has none. GPUs are on node1-node5 (two each). In Step 2, in Terminal 1,
use this instead:
srun -p normal -c 4 --mem=32G --gres=gpu:1 -t 4:00:00 --pty bash -l
Always ask for --gres=gpu:1 if you use the GPU, even if it seems to work without it:
without that request you are using a card the scheduler has assigned to somebody else, and
their work suffers for it.
Note that node1-node5 have 10 cores and 125 GB each, less than ngs. Nothing else in
the procedure changes: the command printed in Step 3b automatically contains the right node
name.
See also GPU Jobs.
Notebooks That Need to Run for a Long Time
With the method on this page, the notebook dies if you close your laptop or lose your
connection. For long runs there is a different approach: the notebook is submitted as a
batch job with sbatch and you connect to it when you need it. Ask the administrator for
the procedure — it is written down and ready.
Being a Good Neighbour
- Never run heavy work on
xlence. It is everyone's way in. - Always ask for
--memand-t. Without--memthe scheduler assigns you a default allocation that may be far larger than you need, taking it away from others. - Release your resources when you are done, with
exit. - Check now and then with
squeue -u $USERthat you have not left anything running.
Quick Reference
TERMINAL 1 (your computer)
ssh name.surname@xlence.disfeb.unimi.it
TERMINAL 1 (now on xlence)
srun -p ngs -c 4 --mem=32G -t 4:00:00 --pty bash -l
TERMINAL 1 (now on ngs)
module load miniforge3
PORT=$((10000 + $(id -u) % 10000))
echo "ssh -N -L $PORT:localhost:$PORT -J $USER@xlence.disfeb.unimi.it $USER@$(hostname)"
-> COPY the printed line
jupyter lab --no-browser --ip=127.0.0.1 --port=$PORT
-> COPY the http://127.0.0.1:.../lab?token=... address
-> leave this window open
TERMINAL 2 (new window, your computer)
paste the copied line here
-> leave this window open
BROWSER
paste the address with the token
TO SHUT DOWN, in this order
TERMINAL 2 : Ctrl+C
TERMINAL 1 : Ctrl+C, then y, then exit (releases the node), then exit (leaves the cluster)