bash: Module 2 Scripting and automation on HPC systems
Goal
Cluster exercise: turn the red-spruce download into a re-runnable, self-checking workflow
In the introductory module you typed commands one at a time. That is how everyone learns, but it does not scale: you cannot remember exactly what you typed three weeks ago, a collaborator cannot rerun it, and a typo in step 4 silently ruins step 9.
In this module you will put those same steps into Bash scripts that check their own work, stop on errors, record what they did, and can be run again safely. Along the way you will process many files at once, verify data with checksums, move files on and off the cluster, set permissions for a shared lab directory, and submit your first Slurm job array.
You will keep working in the same cluster-exercise/ project and the same red-spruce fitness-trait dataset.
Before you begin
This module assumes you have finished Getting started on HPC Linux systems and are comfortable with cd, ls, mkdir -p, curl, head, less, grep, pipes, and > / >>.
Redirecting errors (2>, 2>&1), tee, /dev/null, here-documents, quoting, and globs are covered in the Bash reference page. This module uses them without re-explaining them.
Log in, then move to the project root:
cd ~/projects/cluster-exercise
pwdIf you no longer have the project, recreate it in two lines:
mkdir -p ~/projects/cluster-exercise/{data/raw,results,logs,src}
cd ~/projects/cluster-exerciseYou do not need to download the data again by hand. The first script you write will do it.
Learning objectives
By the end of this module, you should be able to:
- Read and use exit status (
$?) and chain commands with&&and||. - Write conditions with
[[ ... ]]and(( ... )), and branch withifandcase. - Manipulate filenames and paths with parameter expansion, without calling extra programs.
- Write
forandwhile readloops, including loops that read tab-delimited rows. - Write a script with a shebang,
set -euo pipefail, functions, arguments, and a usage message. - Make scripts idempotent: safe to run more than once.
- Verify data integrity with
sha256sum. - Operate on many files with
find,-exec, andxargs. - Compress and archive results with
gzipandtar. - Transfer data with
scpandrsync. - Manage your environment with
PATH,module, and~/.bashrc. - Set permissions for a shared lab directory.
- Keep work alive across disconnections with
tmux. - Submit a Slurm job array driven by a list file.
A surprise in the data: when line counts lie
Start with a puzzle. Count the lines in the raw file two different ways:
wc -l data/raw/red_spruce_fitness_traits.txt
awk 'END {print NR}' data/raw/red_spruce_fitness_traits.txt326 data/raw/red_spruce_fitness_traits.txt
327
The two counts differ by one. The reason is that wc -l does not count lines; it counts newline characters. Look at the last few bytes of the file:
tail -c 20 data/raw/red_spruce_fitness_traits.txt | od -c0000000 0 7 5 6 5 5 3 6 6 \t 1 . 0 0 1 7
0000020 1 3 0 9
0000024
A normal text file ends with \n. This one ends with 9: the final row has no trailing newline, so wc -l misses it. The file actually holds 1 header + 326 data rows = 327 lines, so “wc -l minus 1” gives 325, not 326.
Files exported from spreadsheets, R, or Python often look like this. It matters because while read loops also skip the last line of such files, which you will fix later in this module.
awk 'END {print NR}' file.txt # counts lines, with or without a final newline
grep -c '' file.txt # same idea: counts every line, including the lastUse wc -l for quick checks, but check these alternatives whenever an exact count matters.
Exit status and chaining commands
Every command returns an exit status when it finishes: 0 means success, and anything from 1 to 255 means some kind of failure. Bash stores the last exit status in $?.
ls data/raw/red_spruce_fitness_traits.txt > /dev/null
echo $?
ls missing.txt 2> /dev/null
echo $?0
2
Scripts make decisions based on exit status, so learn these operators:
| Operator | Meaning | Example |
|---|---|---|
a ; b |
Run a, then b, regardless of the result |
date ; hostname |
a && b |
Run b only if a succeeded |
mkdir -p results && cd results |
a \|\| b |
Run b only if a failed |
cd results \|\| exit 1 |
! a |
Invert the exit status of a |
! grep -q NA file.txt |
grep -q (“quiet”) prints nothing and only sets the exit status. That makes it ideal for asking yes-or-no questions:
grep -q "VT" data/raw/red_spruce_fitness_traits.txt && echo "found VT"
grep -q "ZZ" data/raw/red_spruce_fitness_traits.txt || echo "no ZZ rows"found VT
no ZZ rows
a && b || c is not a true if/else
a && b || c also runs c if a succeeds but b fails. When you need real branching, use if, covered next.
Tests and conditionals
File, string, and number tests
[[ ... ]] is Bash’s test command. It returns exit status 0 (true) or 1 (false).
| Test | True when | Example |
|---|---|---|
-e path |
Path exists (file or directory) | [[ -e data/raw ]] |
-f path |
Path is a regular file | [[ -f src/get_data.sh ]] |
-d path |
Path is a directory | [[ -d results ]] |
-s path |
File exists and is not empty | [[ -s data/raw/red_spruce_fitness_traits.txt ]] |
-r / -w / -x |
Readable / writable / executable | [[ -x src/get_data.sh ]] |
a == b |
Strings are equal (right side may be a glob) | [[ $file == *.txt ]] |
a != b |
Strings differ | [[ $loc != VT ]] |
-z s / -n s |
String is empty / non-empty | [[ -z $1 ]] |
s =~ regex |
String matches an extended regex | [[ $col =~ ^[0-9]+$ ]] |
a && b, a \|\| b, ! a |
Combine tests | [[ -d logs && -d src ]] |
For numbers, use arithmetic evaluation (( ... )), which reads like ordinary math:
n=326
(( n > 300 )) && echo "large enough"
(( n % 2 == 0 )) && echo "even"[[ ]] and not [ ]?
[ ] (also called test) is the older POSIX form. It splits unquoted variables into words, so an empty or space-containing variable breaks it. [[ ]] is Bash-specific, does not split variables, and supports &&, ||, and =~. In Bash scripts, prefer [[ ]] for text and files, and (( )) for numbers.
if, elif, else
f=data/raw/red_spruce_fitness_traits.txt
if [[ ! -e $f ]]; then
echo "Missing: download the data first"
elif [[ ! -s $f ]]; then
echo "File exists but is empty: the download probably failed"
else
echo "Ready: $f"
fiThe structure is always if CONDITION; then ... fi. The condition can be any command at all, not just [[ ]]. Bash runs it and checks its exit status:
if grep -q '<!DOCTYPE html' "$f"; then
echo "This is a web page, not data"
ficase: matching one value against many patterns
case is cleaner than a long elif chain when you compare a single value against several patterns:
for loc in VT NC NB; do
case $loc in
VT|NH|ME|MA) echo "$loc: New England" ;;
NB) echo "$loc: Canada" ;;
*) echo "$loc: other" ;;
esac
doneVT: New England
NC: other
NB: Canada
| separates alternatives, *) is the catch-all, and every branch ends with ;;.
Parameter expansion: editing strings without extra programs
You already know $var and ${var}. Bash can also trim, replace, and supply defaults inside the braces. This is the standard way to build output filenames from input filenames.
f=data/raw/red_spruce_fitness_traits.txt
name=${f##*/}
echo "${f##*/}" # strip everything up to the last /
echo "${f%/*}" # strip the last / and everything after it
echo "${f%.txt}" # strip a suffix
echo "${name%.txt}.tsv" # swap the extension
echo "${name/fitness/height}"
echo "${#name}" # length in charactersred_spruce_fitness_traits.txt
data/raw
data/raw/red_spruce_fitness_traits
red_spruce_fitness_traits.tsv
red_spruce_height_traits.txt
29
| Expansion | Result | Memory aid |
|---|---|---|
${var#pat} / ${var##pat} |
Remove the shortest / longest match from the start | # is to the left of $ on a keyboard |
${var%pat} / ${var%%pat} |
Remove the shortest / longest match from the end | % is to the right |
${var/old/new} |
Replace the first match (// replaces all) |
like sed s/old/new/ |
${#var} |
Length | “number of” |
${var:-default} |
Use default if var is unset or empty |
optional arguments |
${var:?message} |
Stop with message if var is unset or empty |
required settings |
${var,,} / ${var^^} |
Lowercase / uppercase | Bash 4+ |
$(( expr )) |
Integer arithmetic | $(( 326 / 4 )) gives 81 |
$(( 326 / 63 )) gives 5. Bash has no decimals. For decimal math, use awk:
awk 'BEGIN {printf "%.3f\n", 326 / 63}'5.175
Loops
for loops
A for loop runs its body once for each word in a list:
for i in {1..3}; do
echo "replicate_$i"
donereplicate_1
replicate_2
replicate_3
The list can be a glob, which is the safe way to loop over files:
for f in results/*.txt; do
[[ -e $f ]] || continue # skip if the glob matched nothing
echo "$f has $(wc -l < "$f") lines"
donecontinue skips to the next item, and break leaves the loop entirely.
Zero-padded sample IDs are common in sequencing work:
for i in $(seq 1 12); do
printf 'S%02d\n' "$i"
done | head -n 3S01
S02
S03
$(ls)
for f in $(ls) splits filenames on spaces and breaks on unusual names. Use a glob (for f in *.txt) or find (shown later) instead.
while read: processing a file line by line
To read a file one line at a time, use:
while IFS= read -r line; do
echo "$line"
done < input.txt| Piece | Why it is there |
|---|---|
IFS= |
Keeps leading and trailing whitespace |
-r |
Treats backslashes literally |
< input.txt |
Feeds the file to the whole loop |
Now count the lines in the red-spruce file this way:
n=0
while IFS= read -r line; do
n=$((n + 1))
done < data/raw/red_spruce_fitness_traits.txt
echo "$n"326
The loop is one line short, for the same reason wc -l was: read returns failure on a final line that has no newline, and the loop ends before processing it. The standard fix is to also continue when read has filled the variable:
n=0
while IFS= read -r line || [[ -n $line ]]; do
n=$((n + 1))
done < data/raw/red_spruce_fitness_traits.txt
echo "$n"327
Splitting tab-delimited rows into fields
Set IFS to a tab and give read several variable names. Each field goes into one variable, and the last variable receives everything that remains:
tail -n +2 data/raw/red_spruce_fitness_traits.txt |
while IFS=$'\t' read -r family pop tree loc rest || [[ -n $family ]]; do
echo "$family is from $loc"
done | head -n 3AB_05 is from TN
AB_08 is from TN
AB_12 is from TN
tail -n +2 means “start from line 2”, which skips the header. $'\t' is Bash’s notation for a literal tab character.
Counting with an associative array
Bash 4+ supports associative arrays, which store values under text keys. This loop counts trees per Location:
declare -A count
{
read -r header # consume the header line
while IFS=$'\t' read -r family pop tree loc rest || [[ -n $family ]]; do
count[$loc]=$(( ${count[$loc]:-0} + 1 ))
done
} < data/raw/red_spruce_fitness_traits.txt
for k in "${!count[@]}"; do
printf '%s\t%s\n' "$k" "${count[$k]}"
done | sort | head -n 4MA 19
MD 10
ME 16
NB 6
"${!count[@]}" expands to all keys, and "${count[$k]}" to one value. This works, but Bash loops are slow on large files. The next module does the same task with awk in one line, thousands of times faster on a VCF or a large phenotype table. Use Bash loops to orchestrate files and programs, and awk to compute over rows.
Writing scripts
Anatomy of a robust script
A script is a text file of commands. A good one has five parts:
#!/usr/bin/env bash # 1. shebang: run this file with bash
# name.sh — one-line purpose # 2. header: what, and how to run it
# Usage: bash src/name.sh ARG
set -euo pipefail # 3. strict mode: stop on errors
in="data/raw/..." # 4. settings at the top, in variables
out="results/..."
# 5. the work, ideally in small named stepsStrict mode is three settings:
| Setting | Effect | Without it |
|---|---|---|
set -e |
Exit as soon as a command fails | The script carries on and builds results on a failed step |
set -u |
Treat an unset variable as an error | A typo like $ouput silently becomes an empty string |
set -o pipefail |
A pipeline fails if any command in it fails | missing_cmd \| sort “succeeds” because sort succeeded |
set -u and rm
Without set -u, rm -rf "$outdir/" with an unset outdir becomes rm -rf "/". Strict mode turns that disaster into an error message. Always use it in scripts that delete or overwrite files.
Script 1: download and verify the data
Create src/get_data.sh with a text editor (nano src/get_data.sh, vim src/get_data.sh, or your editor’s remote mode):
src/get_data.sh
#!/usr/bin/env bash
# get_data.sh — download the red-spruce fitness-trait table and verify it.
# Usage: bash src/get_data.sh (run from the project root)
set -euo pipefail
URL="https://raw.githubusercontent.com/anoobvinu07/Genomic_assisted_selection/master/data/FitnessTraits_GeneticParameters_RedSpruce.txt"
OUT="data/raw/red_spruce_fitness_traits.txt"
LOG="logs/get_data_$(date +%Y%m%d_%H%M%S).log"
log() { printf '[%s] %s\n' "$(date '+%F %T')" "$*" | tee -a "$LOG"; }
die() { log "ERROR: $*"; exit 1; }
# 1. Refuse to run from the wrong directory
[[ -d data/raw && -d logs ]] || { echo "Run this from the project root (cluster-exercise/)." >&2; exit 1; }
# 2. Download only if the file is missing or empty (safe to re-run)
if [[ -s "$OUT" ]]; then
log "Already present: $OUT (skipping download)"
else
log "Downloading $URL"
curl -fsSL -o "$OUT" "$URL" || die "download failed"
fi
# 3. Sanity checks
[[ -s "$OUT" ]] || die "$OUT is empty"
if head -n 1 "$OUT" | grep -qi '<!doctype html'; then
die "$OUT is a web page, not data. Use the raw.githubusercontent.com URL."
fi
n_cols=$(head -n 1 "$OUT" | awk -F'\t' '{print NF}')
n_bad=$(awk -F'\t' -v n="$n_cols" 'NF != n' "$OUT" | wc -l)
(( n_bad == 0 )) || die "$n_bad rows do not have $n_cols columns"
n_rows=$(awk 'END {print NR - 1}' "$OUT")
log "OK: $n_rows data rows x $n_cols columns"
# 4. Record a checksum so you can prove later that the raw file never changed
sha256sum "$OUT" > "$OUT.sha256"
log "Checksum written to $OUT.sha256"What is new here:
| Line | What it does |
|---|---|
log() { ...; } |
Defines a function. $* is all the arguments joined into one string. |
tee -a "$LOG" |
Shows each message on screen and appends it to a time-stamped log file |
>&2 |
Sends the error message to standard error, where error messages belong |
curl -fsSL |
-f fails on HTTP errors such as 404 instead of saving the error page; -s hides the progress bar; -S still shows errors; -L follows redirects |
if [[ -s ... ]] |
Makes the script idempotent: the second run skips the download |
awk ... 'NF != n' |
Prints any row whose column count differs from the header’s (a preview of the next module) |
Run it twice:
bash src/get_data.sh
bash src/get_data.sh[2026-10-08 18:06:45] Downloading https://raw.githubusercontent.com/anoobvinu07/Genomic_assisted_selection/master/data/FitnessTraits_GeneticParameters_RedSpruce.txt
[2026-10-08 18:06:45] OK: 326 data rows x 19 columns
[2026-10-08 18:06:45] Checksum written to data/raw/red_spruce_fitness_traits.txt.sha256
[2026-10-08 18:06:45] Already present: data/raw/red_spruce_fitness_traits.txt (skipping download)
[2026-10-08 18:06:45] OK: 326 data rows x 19 columns
[2026-10-08 18:06:45] Checksum written to data/raw/red_spruce_fitness_traits.txt.sha256
Now test the guard by running it from the wrong place:
cd /tmp
bash ~/projects/cluster-exercise/src/get_data.sh
echo "exit status: $?"
cd -Run this from the project root (cluster-exercise/).
exit status: 1
Testing the failure paths is as important as testing the success path. A script that fails loudly is far better than one that “succeeds” on the wrong input.
Making a script executable
So far you have run scripts with bash script.sh. To run one directly, give it execute permission:
chmod +x src/get_data.sh
./src/get_data.shThe shebang line tells Linux which interpreter to use. #!/usr/bin/env bash finds bash on your PATH, which is more portable than hard-coding /bin/bash. The ./ is required because the current directory is deliberately not on your PATH.
Script 2: arguments and a usage message
Hard-coded filenames make a script single-use. Positional arguments make it reusable:
| Variable | Meaning |
|---|---|
$0 |
The script’s own name |
$1, $2, … |
First, second, … argument |
$# |
Number of arguments |
"$@" |
All arguments, each kept as a separate word (use this when passing them on) |
src/count_column.sh
#!/usr/bin/env bash
# count_column.sh — count how many rows share each value in one column.
# Usage: bash src/count_column.sh FILE COLUMN_NUMBER
set -euo pipefail
usage() {
echo "Usage: $(basename "$0") FILE COLUMN_NUMBER" >&2
echo "Example: $(basename "$0") data/raw/red_spruce_fitness_traits.txt 4" >&2
exit 1
}
(( $# == 2 )) || usage # exactly two arguments
file=$1
col=$2
[[ -f $file ]] || { echo "No such file: $file" >&2; exit 1; }
[[ $col =~ ^[0-9]+$ ]] || { echo "COLUMN_NUMBER must be a positive integer, got: $col" >&2; exit 1; }
header=$(head -n 1 "$file" | cut -f "$col")
echo "# counts of '$header' (column $col) in $file"
tail -n +2 "$file" | cut -f "$col" | sort | uniq -c | sort -rnchmod +x src/count_column.sh
./src/count_column.sh data/raw/red_spruce_fitness_traits.txt 5# counts of 'Region' (column 5) in data/raw/red_spruce_fitness_traits.txt
178 C
101 E
47 M
Calling it incorrectly prints help instead of a cryptic error:
./src/count_column.sh
./src/count_column.sh data/raw/red_spruce_fitness_traits.txt abcUsage: count_column.sh FILE COLUMN_NUMBER
Example: count_column.sh data/raw/red_spruce_fitness_traits.txt 4
COLUMN_NUMBER must be a positive integer, got: abc
Functions and local
Functions group steps under a name. Arguments work exactly as in scripts ($1, $#, "$@"). Use local so a function’s variables do not overwrite variables elsewhere in the script:
n_data_rows() {
local file=$1
awk 'END {print NR - 1}' "$file"
}
echo "Rows: $(n_data_rows data/raw/red_spruce_fitness_traits.txt)"Rows: 326
Define functions before you call them, usually near the top of the script.
Data integrity with checksums
A checksum is a short fingerprint of a file’s exact bytes. If even one byte changes, the checksum changes. Your get_data.sh already saved one:
cat data/raw/red_spruce_fitness_traits.txt.sha256ef503c4d9354e6f44905c3508ed3d26ea34b38b391224a77689ff70a871b98c5 data/raw/red_spruce_fitness_traits.txt
At any later time, verify the file still matches:
sha256sum -c data/raw/red_spruce_fitness_traits.txt.sha256data/raw/red_spruce_fitness_traits.txt: OK
If the file had been edited or truncated, you would see FAILED, and sha256sum -c would exit with a non-zero status, which a script can act on.
Use checksums:
- After every large transfer, especially FASTQ, BAM, and VCF files.
- When a sequencing facility provides
md5files. Check them withmd5sum -c file.md5. - Before publishing, to document exactly which data your results came from.
Working with many files
find: select files by name, type, size, or age
find results -name '*.tsv' # by name (quote the pattern)
find . -type d -name logs # directories only
find data -type f -size +100M # larger than 100 MB
find logs -type f -mtime +30 # modified more than 30 days ago
find . -maxdepth 2 -type f -name '*.sh' # limit how deep to searchActing on what find finds
-exec ... {} + runs a command on the files found. {} stands for the filenames, and + passes many of them per call:
find logs -type f -name '*.log' -exec ls -lh {} +xargs does the same for any list arriving on standard input. Pair find -print0 with xargs -0 so that filenames containing spaces survive:
find results -name '*.tsv' -print0 | xargs -0 wc -lxargs -P N runs up to N commands in parallel, for example compressing many files at once:
find data/fastq -name '*.fastq' -print0 | xargs -0 -n 1 -P 4 gzipxargs -P, GNU parallel, and similar tools can saturate a shared login node. Use them inside a Slurm job that has requested the same number of CPUs, for example #SBATCH --cpus-per-task=4 with -P 4.
Script 3: split the dataset into one file per Location
This script combines a loop, field splitting, the trailing-newline fix, and an idempotent “start clean” step:
src/split_by_location.sh
#!/usr/bin/env bash
# split_by_location.sh — write one TSV per Location (column 4), each with the header.
# Usage: bash src/split_by_location.sh
set -euo pipefail
in="data/raw/red_spruce_fitness_traits.txt"
outdir="results/by_location"
mkdir -p "$outdir"
rm -f "$outdir"/*.tsv # start clean: this script appends below
header=$(head -n 1 "$in")
# Read the data rows (skip header with tail). "|| [[ -n $family ]]" keeps the
# final line even though the file has no trailing newline.
tail -n +2 "$in" | while IFS=$'\t' read -r family pop tree loc rest || [[ -n $family ]]; do
out="$outdir/${loc}.tsv"
[[ -f $out ]] || printf '%s\n' "$header" > "$out" # write header once per file
printf '%s\t%s\t%s\t%s\t%s\n' "$family" "$pop" "$tree" "$loc" "$rest" >> "$out"
done
# Report
for f in "$outdir"/*.tsv; do
printf '%-6s %4d rows\n' "$(basename "$f" .tsv)" "$(( $(wc -l < "$f") - 1 ))"
donebash src/split_by_location.shMA 19 rows
MD 10 rows
ME 16 rows
NB 6 rows
NC 22 rows
NH 23 rows
NY 53 rows
PA 47 rows
TN 14 rows
VA 5 rows
VT 61 rows
WV 50 rows
The per-Location counts add up to 326, and this also answers the introductory module’s Vermont challenge exactly: 61 rows have VT in the Location column.
Check that nothing was lost. This command totals every output file, headers included:
find results/by_location -name '*.tsv' -print0 | xargs -0 wc -l | tail -n 1 338 total
That is 326 data rows plus 12 headers. The rm -f line is what makes the script safe to rerun: without it, a second run would append every row again and double the files.
Compression and archives
| Task | Command |
|---|---|
| Compress, keeping the original | gzip -k file.tsv |
| Decompress | gunzip file.tsv.gz |
| Read without decompressing | zcat file.tsv.gz \| head, or zless file.tsv.gz |
| Create an archive of a directory | tar -czf archive.tar.gz dir/ |
| List an archive’s contents | tar -tzf archive.tar.gz |
| Extract an archive | tar -xzf archive.tar.gz |
| Extract into a specific directory | tar -xzf archive.tar.gz -C target/ |
Text data compresses well:
gzip -k results/by_location/VT.tsv
ls -lh results/by_location/VT.tsv* | awk '{print $5, $9}' # size and name only9.3K results/by_location/VT.tsv
3.5K results/by_location/VT.tsv.gz
Bundle the split files into a single archive to send to a collaborator. -C results means “change into results/ first”, so the archive does not contain the results/ prefix:
tar -czf results/by_location.tar.gz -C results by_location
tar -tzf results/by_location.tar.gz | head -n 4by_location/
by_location/VA.tsv
by_location/NB.tsv
by_location/PA.tsv
The tar flags read as words: create / table of contents / extract, gzip, file name follows. Always list (-t) an archive before extracting one you did not create, so it cannot scatter files over your working directory.
Moving data on and off the cluster
Run these from your own computer, not from the cluster, unless your site documents otherwise. Replace user@cluster.example.edu with your login.
# copy one file up to the cluster
scp results.tsv user@cluster.example.edu:~/projects/cluster-exercise/results/
# copy one file down from the cluster
scp user@cluster.example.edu:~/projects/cluster-exercise/results/fitness_by_location.tsv .For directories and large data, use rsync. It copies only what has changed and can resume an interrupted transfer:
# preview first: -n (dry run) shows what would be copied
rsync -avhn --progress results/ user@cluster.example.edu:~/projects/cluster-exercise/results/
# then run it for real
rsync -avh --progress results/ user@cluster.example.edu:~/projects/cluster-exercise/results/| Flag | Meaning |
|---|---|
-a |
Archive mode: recurse into directories, keep timestamps and permissions |
-v |
Verbose |
-h |
Human-readable sizes |
-n |
Dry run: show the plan, change nothing |
--progress |
Per-file progress |
--delete |
Also delete destination files that are missing from the source. Powerful; always use with -n first. |
rsync -a results/ dest/ copies the contents of results into dest/. rsync -a results dest/ copies the directory itself, creating dest/results/. A dry run (-n) shows the difference before anything is written.
For very large transfers (hundreds of gigabytes or more), use your institution’s data-transfer nodes or Globus, if available, rather than a login node.
Your shell environment
PATH and finding programs
When you type a command, Bash searches the directories listed in PATH, in order:
echo "$PATH" | tr ':' '\n' # one directory per line
command -v Rscript # where does this command come from?
type -t cd # builtin, file, alias, or function?If command -v tree prints nothing, the program is not on your PATH: either it is not installed, or it sits in a module you have not loaded.
Environment modules
Most clusters provide software through environment modules (often Lmod). Loading a module edits your PATH and related variables:
module avail # what is installed
module spider samtools # search all versions (Lmod)
module load r # load a module (names vary by cluster)
module list # what is currently loaded
module purge # unload everythingIn scripts and job files, load explicit versions, such as module load r/4.4.1, so results are reproducible after the default version changes.
~/.bashrc: aliases and personal functions
~/.bashrc runs every time you open an interactive shell. It is the place for shortcuts:
~/.bashrc (excerpt)
# Safer defaults
alias rm='rm -i' # ask before deleting
alias ll='ls -lhF'
# Jump to the project
alias cdp='cd ~/projects/cluster-exercise'
# Show my jobs in the Slurm queue
alias myq='squeue -u $USER'
# Count records correctly, even without a final newline
nrec() { awk 'END {print NR}' "$@"; }Apply your changes to the current session with:
source ~/.bashrc~/.bashrc small
Do not module load analysis software in ~/.bashrc. It slows down every login, and it hides dependencies from your job scripts. Load modules inside each script that needs them.
Long-running sessions with tmux
If your SSH connection drops, anything running in that terminal is killed. tmux keeps a session alive on the server, so you can disconnect and reattach later.
| Action | Command or keys |
|---|---|
| Start a named session | tmux new -s spruce |
| Detach, leaving it running | Ctrl+b, then d |
| List sessions | tmux ls |
| Reattach | tmux attach -t spruce |
| New window inside tmux | Ctrl+b, then c |
| Next / previous window | Ctrl+b, then n / p |
| Split pane vertically / horizontally | Ctrl+b, then % / “ |
| Close a session | exit inside it, or tmux kill-session -t spruce |
Many clusters have several login nodes. A tmux session lives on one of them, so note its hostname (hostname) and SSH to that same node to reattach. Use tmux for interactive work, editors, and monitoring. Heavy computation still belongs in Slurm jobs.
From script to scheduler: a first job array
A job array runs the same script many times, with a different SLURM_ARRAY_TASK_ID (1, 2, 3, …) in each task. The standard pattern is a list file with one item per line, where each task picks the line matching its ID.
Create the list of Locations:
cut -f 4 data/raw/red_spruce_fitness_traits.txt | tail -n +2 | sort -u > results/locations.txt
wc -l < results/locations.txt
sed -n '3p' results/locations.txt12
ME
sed -n '3p' prints only line 3. The next module covers sed in depth.
Now write the job script:
src/per_location.slurm
#!/bin/bash
#SBATCH --job-name=per_location
#SBATCH --array=1-12 # one task per line of results/locations.txt
#SBATCH --time=00:05:00
#SBATCH --mem=1G
#SBATCH --cpus-per-task=1
#SBATCH --output=logs/%x_%A_%a.out # %x job name, %A array job ID, %a task ID
#SBATCH --account=your_account # site-specific: ask your cluster's support team
##SBATCH --partition=standard # site-specific; remove one # to enable
set -euo pipefail
cd "$SLURM_SUBMIT_DIR" # the directory you ran sbatch from
loc=$(sed -n "${SLURM_ARRAY_TASK_ID}p" results/locations.txt)
in="results/by_location/${loc}.tsv"
echo "Task ${SLURM_ARRAY_TASK_ID} on $(hostname): ${loc}"
[[ -s $in ]] || { echo "Missing $in; run src/split_by_location.sh first" >&2; exit 1; }
echo "Rows: $(( $(wc -l < "$in") - 1 ))"Submit and monitor it from the project root:
sbatch src/per_location.slurm
squeue -u "$USER"
ls logs/
cat logs/per_location_*_3.outYou can test the selection logic on the login node, without Slurm, by setting the variable yourself:
for SLURM_ARRAY_TASK_ID in 1 12; do
loc=$(sed -n "${SLURM_ARRAY_TASK_ID}p" results/locations.txt)
echo "task $SLURM_ARRAY_TASK_ID -> $loc: $(( $(wc -l < results/by_location/${loc}.tsv) - 1 )) rows"
donetask 1 -> MA: 19 rows
task 12 -> WV: 50 rows
The same pattern scales from 12 Locations to 1,000 samples: put one sample ID or FASTQ path per line in a list file and set --array=1-1000. Add %20 (as in --array=1-1000%20) to cap how many tasks run at once.
Checkpoint exercises
Work from inside cluster-exercise/. Try each exercise before opening its solution.
Print
readyonly ifresults/locations.txtexists and is not empty.NoteSolution[[ -s results/locations.txt ]] && echo readyGiven
f=results/by_location/VT.tsv, use parameter expansion alone to printVT.NoteSolutionf=results/by_location/VT.tsv name=${f##*/} # VT.tsv echo "${name%.tsv}" # VTWrite a
forloop that prints each file inresults/by_location/together with its number of data rows (lines minus the header).NoteSolutionfor f in results/by_location/*.tsv; do echo "$f $(( $(wc -l < "$f") - 1 ))" doneThese files were written by
printf '...\n', so they do end in a newline andwc -lis accurate here.Edit one byte of a copy of the raw file, then show that its checksum no longer matches the original’s.
NoteSolutioncp data/raw/red_spruce_fitness_traits.txt /tmp/copy.txt sed -i 's/AB_05/AB_06/' /tmp/copy.txt sha256sum data/raw/red_spruce_fitness_traits.txt /tmp/copy.txtThe two fingerprints are completely different, even though only one character changed.
Use
findto list every.shfile in the project with its size.NoteSolutionfind . -type f -name '*.sh' -exec ls -lh {} +Create
results/by_location.tar.gzand list its contents without extracting it.NoteSolutiontar -czf results/by_location.tar.gz -C results by_location tar -tzf results/by_location.tar.gzRun
./src/count_column.shwith no arguments. What is printed, and what is the exit status?NoteSolution./src/count_column.sh; echo "exit: $?"It prints the two usage lines on standard error and exits with status
1.
Challenge: a raw-data guard
Write src/check_raw.sh that:
- Exits with an error if
data/raw/red_spruce_fitness_traits.txt.sha256does not exist. - Runs
sha256sum -cquietly, and printsraw data unchangedon success, orRAW DATA CHANGEDon standard error with exit status 1 on failure. - Is called as the first step of every other script you write.
src/check_raw.sh
#!/usr/bin/env bash
# check_raw.sh — stop if the raw data no longer matches its recorded checksum.
set -euo pipefail
sum="data/raw/red_spruce_fitness_traits.txt.sha256"
[[ -f $sum ]] || { echo "No checksum file: $sum (run src/get_data.sh)" >&2; exit 1; }
if sha256sum --quiet -c "$sum"; then
echo "raw data unchanged"
else
echo "RAW DATA CHANGED: do not trust downstream results" >&2
exit 1
fiCall it from other scripts with bash src/check_raw.sh. Because those scripts use set -e, a failed check stops them immediately.
Wrap up
Your project now looks like this:
cluster-exercise/
├── data/
│ └── raw/
│ ├── red_spruce_fitness_traits.txt
│ └── red_spruce_fitness_traits.txt.sha256
├── logs/
│ └── get_data_YYYYMMDD_HHMMSS.log
├── results/
│ ├── by_location/ # 12 TSV files, one per Location
│ ├── by_location.tar.gz
│ └── locations.txt
└── src/
├── check_raw.sh
├── count_column.sh
├── get_data.sh
├── per_location.slurm
└── split_by_location.sh
Before moving on, make sure you can answer these questions:
- Why does
wc -lreport 326 lines for a file that contains 327? - What is the difference between
a && banda || b? - What do each of
-e,-u, and-o pipefailprotect you from? - What does
${f%.txt}return whenf=data/raw/x.txt? - Why does
while readneed|| [[ -n $line ]]for this dataset? - What makes a script idempotent, and which line makes
split_by_location.shidempotent? - When would you use
rsync -n? - What does
chmod 750allow, and for whom? - How does a Slurm array task know which Location to process?
In the next module, text processing with grep, sed, and awk, you will replace the slow Bash loop from this page with one-line awk programs, and compute per-Location summaries directly from the raw file.