Bash Scripting for Beginners: A Complete Hands-On Guide

Think about what you type every time you log into a Linux server. You check disk space, then memory, then whether the services you care about are up. It takes a few minutes, and you do it again the next day, and the day after that.
A Bash script is those commands saved in a file, with logic around them so the file can make decisions for you. That's the whole idea, and it's why Bash is still one of the most useful skills in IT. Every Linux server has it, every CI/CD pipeline runs it, and most cloud automation you'll read at work is either Bash or calls Bash somewhere.
This guide takes you from "I can type ls" to writing scripts you'd trust to run unattended. It's long on purpose, and here's how it's laid out:
- The shell itself: how commands, output, errors and pipes work.
- The language: variables, quoting, conditions, loops, functions, arrays and arguments.
- Real-world skills: text processing, error handling and debugging.
- A project: a server health-check script that uses everything above, scheduled with cron.
How to use this guide:
- Run every example yourself. The output shown is what Ubuntu 24.04 with Bash 5.2 actually printed.
- Type the examples instead of pasting them. The typos you make, and the errors Bash gives you for them, are how you learn what its messages mean.
Where to Run Bash
You need a Linux shell, and there's a good chance you already have one.
- Azure Cloud Shell runs Bash in your browser. Open shell.azure.com and pick Bash.
- WSL on Windows gives you a real Ubuntu install. Run
wsl --installin an admin PowerShell window and restart. - Any Linux VM, whether it's in Azure, AWS or a box under your desk.
Check which shell and version you have:
echo $SHELL
bash --version | head -1
/bin/bash
GNU bash, version 5.2.21(1)-release (x86_64-pc-linux-gnu)
Gotcha: The Terminal on a Mac defaults to zsh, and the Bash that ships with macOS is version 3.2 from 2007. Associative arrays,
mapfile,${var,,}and several other features in this guide don't exist in Bash 3.2, and commands likefreeanddf --outputaren't on macOS at all. Use Cloud Shell, WSL or a Linux VM while you're learning.
Part 1: How the Shell Works
Commands, options and arguments
Every command you type has the same shape: the command, then options that change how it behaves, then arguments that tell it what to work on.
ls -lh /var/log
ls is the command, -lh is two options (-l for long format and -h for human-readable sizes), and /var/log is the argument. When you don't know what a command does, use one of these:
man ls # the full manual page (press q to quit)
ls --help # a shorter summary most commands support
type cd # tells you whether something is a program, a builtin or an alias
type is more useful than it looks. Some commands, like cd, echo and read, are builtins, which means they're part of Bash itself rather than separate programs. That's why man cd often finds nothing; use help cd for builtins instead.
Output, errors and redirection
Every command has three standard streams:
| Stream | Number | What it is |
|---|---|---|
| stdin | 0 | Input, usually your keyboard |
| stdout | 1 | Normal output |
| stderr | 2 | Error messages |
Both stdout and stderr show up on your screen by default, so they look the same. The difference shows up when you redirect them:
echo "first line" > notes.txt # > creates or overwrites a file
echo "second line" >> notes.txt # >> appends to the end
cat notes.txt
ls /etc/hostname /nope > out.txt 2> err.txt
echo "--- out.txt"; cat out.txt
echo "--- err.txt"; cat err.txt
first line
second line
--- out.txt
/etc/hostname
--- err.txt
ls: cannot access '/nope': No such file or directory
The ls command succeeded for one file and failed for the other, and the two results went to different places. That separation is what lets scripts log errors without mixing them into normal output.
You'll see three more redirection patterns constantly:
command > /dev/null # throw away normal output
command 2> /dev/null # throw away errors
command > all.log 2>&1 # send errors to the same place as output
2>&1 reads as "send stream 2 to wherever stream 1 is going." The order matters: the redirect to all.log has to come first, so stream 1 already points at the file when stream 2 copies it.
Warning:
>overwrites without asking.> important.confwith a typo in the command in front of it will empty the file. If that scares you, runset -o noclobberin your shell, and>will refuse to overwrite existing files.
Pipes
A pipe, |, sends one command's stdout into the next command's stdin. This is where the shell gets its power, because you can chain small tools into something bigger.
printf "web01 nginx\ndb01 postgres\nweb02 nginx\nweb03 apache\n" > servers.txt
cat servers.txt | grep web | wc -l
cut -d' ' -f2 servers.txt | sort | uniq -c
3
1 apache
2 nginx
1 postgres
The first line counts web servers. The second takes the second column, sorts it, and counts each unique value. Neither command knows anything about servers; each one does one small job and passes its output along.
Part 2: Your First Script
To create your first script:
- Open a new file called
hello.shin any editor, for examplenano hello.sh. - Add these three lines and save the file:
#!/usr/bin/env bash
# hello.sh - my first script
echo "Hello from $(hostname)"
The first line is the shebang. It tells Linux which program should run the file, and /usr/bin/env bash means "find bash wherever it's installed." Lines starting with # are comments, which Bash ignores.
Run it:
./hello.sh
bash: ./hello.sh: Permission denied
New files aren't executable by default. Mark it as executable once, then run it again:
chmod +x hello.sh
./hello.sh
Hello from demo-vm-01
The ./ tells Bash to run the file from the current folder. Without it, Bash only looks in the folders listed in your PATH variable. You can also skip chmod and run bash hello.sh, which is handy while you're testing.
If you write a lot of scripts, keep them in one folder and add it to your PATH so you can run them from anywhere:
mkdir -p ~/bin
mv hello.sh ~/bin/
echo 'export PATH="$HOME/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
hello.sh
Part 3: Variables and Quoting
Setting and using variables
name="Parveen"
echo "Hello $name"
There's one rule here that catches almost everyone: no spaces around =. If you write name = "Parveen", Bash thinks name is a command and fails with name: command not found.
Variable names are case-sensitive. A common convention is UPPERCASE for environment variables and constants, and lowercase for everything else, so your variables don't accidentally clash with system ones like PATH or HOME.
Single quotes, double quotes and escaping
Quoting is the single most important thing to get right in Bash.
name="Parveen"
echo "Hello $name"
echo 'Hello $name'
echo "Cost: \$5 for $name"
echo "Today is $(date +%A)"
Hello Parveen
Hello $name
Cost: $5 for Parveen
Today is Friday
- Double quotes keep the text together as one value but still expand variables and commands.
- Single quotes keep everything literally. Nothing inside them is expanded.
- A backslash escapes one character, so
\$prints a literal dollar sign. $(...)is command substitution. It runs the command and drops its output into place.
Without quotes, Bash splits a variable's value on spaces:
file="my report.txt"
touch "$file"
ls $file # unquoted: Bash passes two arguments, "my" and "report.txt"
ls "$file" # quoted: one argument, the real file name
ls: cannot access 'my': No such file or directory
ls: cannot access 'report.txt': No such file or directory
my report.txt
Pro tip: Quote every variable by default, like
"$file", and only leave quotes off when you specifically want splitting. That one habit prevents more Bash bugs than anything else in this guide.
Environment variables and export
A variable you set exists only in your current shell. Programs and scripts you start from that shell can't see it unless you export it:
greeting="hello"
bash -c 'echo "child sees: [$greeting]"'
export greeting
bash -c 'echo "child sees: [$greeting]"'
child sees: []
child sees: [hello]
This trips people up when a script works in their terminal and fails somewhere else: the variable it relied on was set in their shell but never exported. Run env to see every exported variable.
For values that should never change, use readonly:
readonly REGION="canadacentral"
REGION="eastus"
bash: REGION: readonly variable
Parameter expansion
Bash can transform variables without calling any other program. These are worth memorizing because you'll use them constantly:
file="/var/log/nginx/access.log"
echo "${#file}" # length of the string
echo "${file##*/}" # strip everything up to the last /
echo "${file%/*}" # strip the last / and everything after it
echo "${file%.log}.bak" # swap the extension
echo "${file/nginx/apache}" # replace the first match
env="Production"
echo "${env,,}" # lowercase
echo "${env^^}" # uppercase
echo "${env:0:4}" # substring: start at 0, take 4
unset region
echo "${region:-canadacentral}" # default if empty or unset
echo "${region:?region is required}" # fail with a message if empty or unset
25
access.log
/var/log/nginx
/var/log/nginx/access.bak
/var/log/apache/access.log
production
PRODUCTION
Prod
canadacentral
bash: region: region is required
The # and % forms look cryptic at first. One way to remember them: on a US keyboard, # is to the left of $ and strips from the left, and % is to the right and strips from the right. One symbol removes the shortest match and two removes the longest.
${var:-default} is the one you'll use most. It's how scripts accept an optional value: THRESHOLD="${1:-80}" means "use the first argument, or 80 if there isn't one."
Arithmetic
Bash does integer math inside $(( )):
disks=3
size=128
echo $(( disks * size ))
echo $(( 10 / 3 ))
echo $(( 10 % 3 ))
count=0
(( count += 5 ))
echo "$count"
384
3
1
5
Notice that 10 / 3 gives 3. Bash only does whole numbers, and it drops the remainder. When you need decimals, hand the math to awk:
awk 'BEGIN { printf "%.2f\n", 10/3 }'
3.33
Part 4: Making Decisions
Exit codes are the foundation
Every command returns an exit code when it finishes. 0 means success, and anything from 1 to 255 means something went wrong. The variable $? holds the exit code of the last command:
ls /does-not-exist > /dev/null 2>&1
echo "ls exit code: $?"
ls exit code: 2
if doesn't need brackets. It runs any command and checks its exit code. The brackets are just one command among many.
if grep -q "nobody" /etc/passwd; then
echo "user exists"
fi
grep -q prints nothing and returns 0 if it finds a match. That's all if needs.
You can also chain commands on exit codes with && (run the next one only if this succeeded) and || (run the next one only if this failed):
mkdir -p /tmp/demo && echo "created /tmp/demo"
grep -q nobody /etc/passwd || echo "never printed"
created /tmp/demo
Tests with [[ ]]
[[ ]] is Bash's test command. It returns 0 when its condition is true, which makes it work with if:
config="/etc/hostname"
if [[ -f "$config" ]]; then echo "$config exists"; fi
if [[ -d /tmp && -w /tmp ]]; then echo "/tmp is a writable directory"; fi
name=""
if [[ -z "$name" ]]; then echo "name is empty"; fi
if [[ "web01" == web* ]]; then echo "web01 matches web*"; fi
if [[ "v2.14.1" =~ ^v([0-9]+)\.([0-9]+) ]]; then
echo "major=${BASH_REMATCH[1]} minor=${BASH_REMATCH[2]}"
fi
/etc/hostname exists
/tmp is a writable directory
name is empty
web01 matches web*
major=2 minor=14
These are the tests you'll actually use:
| Test | True when |
|---|---|
-f path | the path is a regular file |
-d path | the path is a directory |
-e path | the path exists at all |
-r / -w / -x path | you can read / write / execute it |
-s path | the file exists and isn't empty |
-z "$var" | the string is empty |
-n "$var" | the string isn't empty |
"$a" == "$b" | the strings match (the right side can be a pattern like web*) |
"$a" =~ regex | the string matches a regular expression |
$a -eq $b | the numbers are equal (also -ne, -lt, -le, -gt, -ge) |
Gotcha: Use
-gtand friends for numbers, not>. Inside[[ ]],>compares strings alphabetically, so[[ 9 > 10 ]]is true because "9" sorts after "1". For numeric comparisons you can also use(( a > b )), which does real math.
You'll also see single brackets, [ ], in older scripts. They're the original POSIX test command and they work in any shell, but they're less forgiving about quoting and don't support patterns or &&. In a Bash script, use [[ ]].
if, elif and else
usage=85
if (( usage >= 90 )); then
echo "critical"
elif (( usage >= 80 )); then
echo "warning"
else
echo "ok"
fi
warning
case for multiple choices
When you're comparing one value against several options, case is much cleaner than a chain of elif:
for env in dev prod staging qa; do
case "$env" in
dev|qa) size="Standard_B2s" ;;
staging) size="Standard_D2s_v5" ;;
prod) size="Standard_D4s_v5" ;;
*) size="unknown" ;;
esac
echo "$env -> $size"
done
dev -> Standard_B2s
prod -> Standard_D4s_v5
staging -> Standard_D2s_v5
qa -> Standard_B2s
| means "or," *) catches everything else, and each branch ends with ;;. Patterns work here too, so web*) would match any value starting with "web."
Part 5: Loops
for loops
There are several ways to give a for loop its list:
for i in {1..3}; do echo "attempt $i"; done
for (( i = 0; i < 6; i += 2 )); do printf "%s " "$i"; done; echo
for f in /etc/host*; do echo "found $f"; done
attempt 1
attempt 2
attempt 3
0 2 4
found /etc/host.conf
found /etc/hostname
found /etc/hosts
The first uses a brace range, the second is a C-style loop with a counter, and the third uses a glob, which Bash expands to matching file names. Globs are the right way to loop over files.
Warning: Never loop over
$(ls). File names with spaces get split into pieces, andlsoutput isn't designed to be parsed.for f in /path/*.logdoes the same job safely.
while and until
while keeps going as long as its command succeeds, and until keeps going as long as it fails:
n=3
while (( n > 0 )); do
echo "countdown $n"
n=$((n - 1))
done
tries=0
until [[ -f /tmp/ready ]] || (( tries >= 2 )); do
tries=$((tries + 1))
echo "waiting ($tries)"
done
countdown 3
countdown 2
countdown 1
waiting (1)
waiting (2)
The until loop is a pattern you'll use for "wait until something is ready," with a limit so it can't run forever.
break and continue
continue skips to the next item, and break leaves the loop entirely:
for s in web01 db01 skip-me web02; do
[[ "$s" == skip-* ]] && continue
[[ "$s" == db* ]] && { echo "stopping at $s"; break; }
echo "deploying $s"
done
deploying web01
stopping at db01
Reading a file line by line
This is one of the most common things you'll do in Bash, and there's one correct way to do it:
printf "web01,10.0.1.4,nginx\ndb01,10.0.2.4,postgres\n" > inventory.csv
while IFS=, read -r host ip role; do
echo "$host ($role) at $ip"
done < inventory.csv
web01 (nginx) at 10.0.1.4
db01 (postgres) at 10.0.2.4
IFS=, tells read to split on commas for this command only, -r stops it from treating backslashes as escape characters, and < inventory.csv feeds the file into the loop. Always use read -r unless you have a specific reason not to.
To loop over a command's output instead of a file, use < <(command):
while read -r usage mount; do
echo "$mount is at $usage"
done < <(df --output=pcent,target | tail -n +2)
You might be tempted to write df | while read ... instead. It works, but the loop then runs in a subshell, so any variable you change inside it is lost when the loop ends. < <(...) keeps the loop in your current shell.
Part 6: Functions
Once a script grows past about 30 lines, you'll want functions. They let you name a piece of logic and reuse it.
log() {
local level="$1"; shift
printf '%s [%s] %s\n' "$(date +%H:%M:%S)" "$level" "$*"
}
is_port_open() {
local port="$1"
[[ "$port" -eq 22 || "$port" -eq 443 ]]
}
disk_percent() {
df --output=pcent / | tail -1 | tr -dc '0-9'
}
log INFO "starting checks"
if is_port_open 443; then log INFO "443 is open"; fi
is_port_open 8080 || log WARN "8080 is closed (exit code $?)"
usage="$(disk_percent)"
log INFO "root disk at ${usage}%"
18:45:45 [INFO] starting checks
18:45:45 [INFO] 443 is open
18:45:45 [WARN] 8080 is closed (exit code 1)
18:45:45 [INFO] root disk at 43%
There are four things in that example worth understanding properly.
Arguments work just like they do for scripts. Inside a function, $1 is the first argument, $* is all of them, and shift drops the first one. log uses shift to take the level off the front and treat the rest as the message.
Functions return a status, not a value. return only accepts an exit code from 0 to 255. is_port_open doesn't even use return; a function's exit code is the exit code of its last command, so the [[ ]] test decides it. That's what lets you write if is_port_open 443.
To get data out, print it. disk_percent prints a number, and $(disk_percent) captures it. This is the normal way to "return" a value in Bash.
Use local for every variable inside a function. Without it, variables are global and leak out:
counter=1
bump() { counter=$((counter + 1)); local temp="gone"; }
bump
echo "counter=$counter temp=${temp:-unset}"
counter=2 temp=unset
counter was changed globally because it wasn't declared local, and temp disappeared when the function ended because it was. In bigger scripts, a missing local causes some of the most confusing bugs you'll ever chase.
Part 7: Arrays
Indexed arrays
An array holds a list of values. You'll use them for lists of servers, files or anything else you want to loop over.
servers=("web01" "web02" "db01")
servers+=("cache01")
echo "count: ${#servers[@]}"
echo "first: ${servers[0]}"
echo "last: ${servers[-1]}"
for s in "${servers[@]}"; do printf "%s " "$s"; done; echo
count: 4
first: web01
last: cache01
web01 web02 db01 cache01
Always loop with "${servers[@]}", with the quotes and the @. That form keeps each item intact even if it contains spaces.
Associative arrays
Associative arrays use names instead of numbers as keys, like a dictionary. You have to declare them with declare -A:
declare -A owner
owner[web01]="platform"
owner[db01]="data"
for host in "${!owner[@]}"; do
echo "$host is owned by ${owner[$host]}"
done | sort
db01 is owned by data
web01 is owned by platform
"${!owner[@]}" gives you the keys. Associative arrays don't keep insertion order, which is why the output is piped through sort.
To load a file into an array, one line per item, use mapfile:
mapfile -t hosts < hosts.txt
echo "loaded ${#hosts[@]} hosts"
Part 8: Script Arguments
Positional parameters
When you run ./script.sh deploy "my app" prod, Bash gives the script these variables:
#!/usr/bin/env bash
echo "script: $0"
echo "count: $#"
echo "all: $*"
for arg in "$@"; do echo " arg: [$arg]"; done
shift
echo "after shift: $*"
script: ./args-demo.sh
count: 3
all: deploy my app prod
arg: [deploy]
arg: [my app]
arg: [prod]
after shift: my app prod
"my app" stayed one argument because it was quoted on the command line and looped over with "$@". Use "$@" whenever you pass arguments along, and save $* for when you want them joined into one string.
Proper options with getopts
For anything you'll share with other people, use flags like -e prod instead of relying on argument order. getopts is built into Bash and handles the parsing:
#!/usr/bin/env bash
set -euo pipefail
usage() { echo "Usage: $0 -e env [-r region] [-v]" >&2; exit 2; }
env=""; region="canadacentral"; verbose=false
while getopts ":e:r:vh" opt; do
case "$opt" in
e) env="$OPTARG" ;;
r) region="$OPTARG" ;;
v) verbose=true ;;
h) usage ;;
:) echo "Option -$OPTARG needs a value" >&2; usage ;;
*) echo "Unknown option -$OPTARG" >&2; usage ;;
esac
done
shift $((OPTIND - 1))
[[ -z "$env" ]] && usage
echo "env=$env region=$region verbose=$verbose extra=$*"
$ ./deploy.sh -e prod -v
env=prod region=canadacentral verbose=true extra=
$ ./deploy.sh -e dev -r eastus web01 web02
env=dev region=eastus verbose=false extra=web01 web02
$ ./deploy.sh -x
Unknown option -x
Usage: ./deploy.sh -e env [-r region] [-v]
The string ":e:r:vh" defines the options. A letter followed by : takes a value, which lands in $OPTARG. The leading : tells getopts to let your case handle errors instead of printing its own messages. After the loop, shift $((OPTIND - 1)) removes the options, so whatever is left in $@ is the extra arguments.
Pro tip: Send error and usage messages to stderr with
>&2, and exit with a non-zero code. That way, someone piping your script's output into another tool doesn't get usage text mixed into their data.
Part 9: Input and Output
printf over echo
echo is fine for simple messages, but printf gives you formatting control and behaves the same on every system:
printf "%-10s %5s\n" "HOST" "CPU" "web01" "12%" "db01" "87%"
HOST CPU
web01 12%
db01 87%
%-10s means a string, left-aligned, 10 characters wide. printf reuses the format for each set of arguments, which is how one line printed a whole table.
Here-documents
A here-document feeds a block of text into a command. It's the cleanest way to write config files or multi-line messages:
app="billing"; port=8080
cat > app.conf <<EOF
name=$app
port=$port
EOF
cat app.conf
cat <<'EOF'
This keeps $app literally
EOF
name=billing
port=8080
This keeps $app literally
With a plain EOF, variables expand. Quote it as 'EOF' and everything is kept literally, which is what you want when you're writing out another script.
read and here-strings
read takes input from the user or from text you give it. <<< is a here-string, which feeds a single string to a command:
read -rp "Environment: " env # prompt the user
read -r first rest <<< "alpha beta gamma"
echo "first=$first rest=$rest"
first=alpha rest=beta gamma
When read is given more words than variables, the last variable gets everything left over.
Part 10: The Text Processing Toolkit
Bash itself isn't great at processing text, but it's excellent at gluing together the tools that are. In most real scripts, the heavy lifting is done by grep, awk, sed, cut, sort and uniq, with Bash connecting them. Save this small web server log as access.log to practice on:
10.0.0.5 - - [02/Oct/2026:07:01:12] "GET /api/health HTTP/1.1" 200 12
10.0.0.9 - - [02/Oct/2026:07:01:15] "GET /login HTTP/1.1" 200 512
10.0.0.5 - - [02/Oct/2026:07:02:01] "POST /api/orders HTTP/1.1" 500 87
10.0.0.7 - - [02/Oct/2026:07:02:44] "GET /api/orders HTTP/1.1" 200 1024
10.0.0.5 - - [02/Oct/2026:07:03:09] "POST /api/orders HTTP/1.1" 500 91
grep finds lines. -c counts them, -i ignores case, -v inverts the match, and -r searches folders recursively.
grep -c ' 500 ' access.log
2
sort, uniq and cut summarize. This is the "top talkers" pattern, and you'll use it on every log you ever investigate:
awk '{print $1}' access.log | sort | uniq -c | sort -rn
3 10.0.0.5
1 10.0.0.9
1 10.0.0.7
uniq -c only counts lines that are next to each other, which is why the data has to be sorted first. sort -rn then sorts numerically, highest first.
awk works with columns. It splits each line on whitespace into $1, $2 and so on, and it can filter and add up as it goes:
awk '$8 >= 500 {print $1, $6}' access.log
awk '{bytes += $9} END {print bytes " bytes"}' access.log
10.0.0.5 /api/orders
10.0.0.5 /api/orders
1726 bytes
The first command prints the IP and path for every server error. The second adds up the last column and prints the total once, at the end. You can go a long way with just those two patterns.
sed edits text as it streams past. Its most common job is find and replace:
sed -n '2p' access.log | sed 's/GET/FETCH/'
sed -i 's/old-hostname/new-hostname/g' app.conf # edit a file in place
10.0.0.9 - - [02/Oct/2026:07:01:15] "FETCH /login HTTP/1.1" 200 512
-n '2p' prints only line 2. s/find/replace/ replaces the first match on each line, and adding g replaces every match. -i writes the change back to the file, so test your expression without -i first.
cut splits on a delimiter. It's simpler than awk when the data has a clear separator:
cut -d'"' -f2 access.log | head -2
GET /api/health HTTP/1.1
GET /login HTTP/1.1
Part 11: Error Handling
This is the section that separates scripts you run by hand from scripts you trust to run on their own.
set -euo pipefail
By default, Bash keeps going after errors. A failed cd followed by rm -rf * is the classic disaster story, and it happens because nothing stopped the script after the first command failed. Put this at the top of every script:
set -euo pipefail
-eexits the script as soon as a command fails.-uexits when you use a variable that was never set, which catches typos.-o pipefailmakes a pipeline fail if any command in it fails. Without it, only the last command's exit code counts.
Each option changes the result:
bash -c 'false | true; echo "without pipefail: $?"'
bash -c 'set -o pipefail; false | true; echo "with pipefail: $?"'
bash -c 'set -u; echo "$undefined_var"'
without pipefail: 0
with pipefail: 1
bash: line 1: undefined_var: unbound variable
Gotcha:
set -ehas a few traps. The most common is((count++)), which returns a failure exit code whencountis 0, so your script exits on the very first increment. Usecount=$((count + 1))instead. Also, commands inside anifcondition or before&&and||don't trigger-e, because Bash assumes you're already handling their result.
trap for cleanup
trap runs a command when your script exits or hits an error. The EXIT trap runs no matter how the script ends, which makes it perfect for cleaning up temporary files:
#!/usr/bin/env bash
set -euo pipefail
workdir="$(mktemp -d)"
cleanup() { rm -rf "$workdir"; echo "cleaned up $workdir"; }
trap cleanup EXIT
trap 'echo "failed on line $LINENO" >&2' ERR
echo "working in $workdir"
touch "$workdir/data.txt"
cp /does/not/exist "$workdir/"
echo "never reached"
working in /tmp/tmp.XXXX
cp: cannot stat '/does/not/exist': No such file or directory
failed on line 11
cleaned up /tmp/tmp.XXXX
The cp failed, set -e stopped the script, the ERR trap reported the line number, and the EXIT trap still cleaned up. mktemp -d creates a unique temporary folder, so two copies of the script running at once can't overwrite each other's files.
Retrying flaky commands
Network calls fail sometimes. Rather than letting one blip kill your script, wrap them in a retry function:
retry() {
local attempts="$1"; shift
local n=1
until "$@"; do
if (( n >= attempts )); then
echo "gave up after $n attempts: $*" >&2
return 1
fi
echo "attempt $n failed, retrying in ${n}s" >&2
sleep "$n"
n=$((n + 1))
done
}
retry 3 curl -fsS https://example.com/health
When the command keeps failing, you get this:
attempt 1 failed, retrying in 1s
attempt 2 failed, retrying in 2s
gave up after 3 attempts: curl -fsS https://example.com/health
"$@" runs whatever command you passed in, with its arguments intact. The wait gets longer after each failure, which gives a struggling service a moment to recover.
Part 12: Debugging
When a script doesn't behave, use these tools in this order.
Trace with bash -x
bash -x prints every command, with variables already filled in, just before it runs:
bash -x ./health-check.sh -t 90
+ set -euo pipefail
+ DISK_THRESHOLD=80
+ MEM_THRESHOLD=90
+ SERVICES=(sshd cron)
...
When a variable is empty or holds something unexpected, you'll see it right there. To trace only part of a script, wrap that part in set -x and set +x.
Lint with ShellCheck
ShellCheck reads your script and points out bugs before you run it. To try it:
- Install it with
sudo apt install shellcheck. - Save this broken script as
bad.shand runshellcheck bad.sh:
#!/usr/bin/env bash
logdir="/var/log/my app"
for f in $(ls $logdir/*.log); do
echo Processing $f
done
In bad.sh line 3:
for f in $(ls $logdir/*.log); do
^-----------------^ SC2045 (error): Iterating over ls output is fragile. Use globs.
^-----^ SC2086 (info): Double quote to prevent globbing and word splitting.
It caught the ls loop and the missing quotes, which are two of the mistakes this guide has already warned you about. Every warning has a code like SC2086 with a page explaining the fix. I run ShellCheck on every script before it goes anywhere near a server, and most editors, including VS Code, have an extension that runs it as you type.
Watch for Windows line endings
If you write a script in a Windows editor and copy it to Linux, you might see this:
/usr/bin/env: 'bash\r': No such file or directory
That \r is a Windows carriage return hiding at the end of each line. To fix it:
- Remove the carriage returns with
sed -i 's/\r$//' script.sh. - Set your editor to use LF line endings for
.shfiles, so it doesn't happen again.
Part 13: Project, a Server Health Check
This project puts everything together. The script checks disk usage, memory and running processes, accepts options, logs to a file, and exits with a status code other tools can act on. Every piece of it is something you've seen above.
#!/usr/bin/env bash
# health-check.sh - check disk, memory and services on a Linux server
#
# Usage: health-check.sh [-t disk_threshold] [-m mem_threshold] [-s "svc1 svc2"] [-l logfile]
# Exit codes: 0 = healthy, 1 = problems found, 2 = bad usage
set -euo pipefail
# ---------- defaults ----------
DISK_THRESHOLD=80
MEM_THRESHOLD=90
SERVICES=(sshd cron)
LOGFILE=""
PROBLEMS=0
# ---------- helpers ----------
usage() {
cat <`<USAGE >`&2
Usage: $(basename "$0") [-t disk%] [-m mem%] [-s "svc1 svc2"] [-l logfile]
-t disk usage warning threshold (default: 80)
-m memory usage warning threshold (default: 90)
-s space-separated list of processes to check (default: sshd cron)
-l also append output to this log file
USAGE
exit "${1:-2}"
}
log() {
local level="$1"; shift
local line
line="$(printf '%s %-5s %s' "$(date '+%Y-%m-%d %H:%M:%S')" "$level" "$*")"
echo "$line"
if [[ -n "$LOGFILE" ]]; then echo "$line" >> "$LOGFILE"; fi
}
warn() {
log WARN "$@"
PROBLEMS=$((PROBLEMS + 1))
}
is_number() {
[[ "$1" =~ ^[0-9]+$ ]]
}
# ---------- checks ----------
check_disks() {
local usage mount percent
while read -r usage mount; do
percent="${usage%\%}"
if (( percent >= DISK_THRESHOLD )); then
warn "disk ${mount} at ${usage} (threshold ${DISK_THRESHOLD}%)"
else
log OK "disk ${mount} at ${usage}"
fi
done < <(df --output=pcent,target -x tmpfs -x devtmpfs -x overlay -x squashfs | tail -n +2)
}
check_memory() {
local total available used_pct
read -r total available < <(awk '/^MemTotal:/ {t=$2} /^MemAvailable:/ {a=$2} END {print t, a}' /proc/meminfo)
used_pct=$(( (total - available) * 100 / total ))
if (( used_pct >= MEM_THRESHOLD )); then
warn "memory at ${used_pct}% (threshold ${MEM_THRESHOLD}%)"
else
log OK "memory at ${used_pct}%"
fi
}
check_services() {
local service
for service in "${SERVICES[@]}"; do
if pgrep -x "$service" > /dev/null; then
log OK "process ${service} is running"
else
warn "process ${service} is not running"
fi
done
}
# ---------- main ----------
main() {
local opt
while getopts ":t:m:s:l:h" opt; do
case "$opt" in
t) DISK_THRESHOLD="$OPTARG" ;;
m) MEM_THRESHOLD="$OPTARG" ;;
s) read -r -a SERVICES <<< "$OPTARG" ;;
l) LOGFILE="$OPTARG" ;;
h) usage 0 ;;
:) echo "Option -${OPTARG} needs a value" >&2; usage ;;
*) echo "Unknown option -${OPTARG}" >&2; usage ;;
esac
done
shift $((OPTIND - 1))
is_number "$DISK_THRESHOLD" || { echo "-t must be a number" >&2; usage; }
is_number "$MEM_THRESHOLD" || { echo "-m must be a number" >&2; usage; }
log INFO "health check on $(hostname)"
check_disks
check_memory
check_services
if (( PROBLEMS > 0 )); then
log FAIL "${PROBLEMS} problem(s) found"
exit 1
fi
log INFO "all checks passed"
}
main "$@"
A few design choices are worth pointing out, because they're habits worth copying into your own scripts.
- Everything lives in functions, and
main "$@"is the last line. Bash reads a script as it runs it, so callingmainon the last line guarantees every function is defined before anything executes. It also makes the script read top to bottom like a table of contents. - Memory comes from
/proc/meminfo, not fromfree. The output offreeis meant for people and its layout varies between versions./proc/meminfois the sourcefreereads from, and its format is stable. - Input is validated before it's used.
is_numberstops-t abcbefore it can cause a confusing arithmetic error halfway through. - The exit codes mean something. 0 is healthy, 1 is a problem with the server, and 2 is a problem with how you called the script. A monitoring tool can tell those apart.
Output on a healthy server:
./health-check.sh
2026-10-02 07:00:01 INFO health check on demo-vm-01
2026-10-02 07:00:01 OK disk / at 46%
2026-10-02 07:00:01 OK memory at 18%
2026-10-02 07:00:01 OK process sshd is running
2026-10-02 07:00:01 OK process cron is running
2026-10-02 07:00:01 INFO all checks passed
And here it is with a strict disk threshold and a process that isn't running, logging to a file:
./health-check.sh -t 40 -s "sshd nginx" -l /tmp/health.log
echo "exit code: $?"
2026-10-02 07:00:05 INFO health check on demo-vm-01
2026-10-02 07:00:05 WARN disk / at 46% (threshold 40%)
2026-10-02 07:00:05 OK memory at 18%
2026-10-02 07:00:05 OK process sshd is running
2026-10-02 07:00:05 WARN process nginx is not running
2026-10-02 07:00:05 FAIL 2 problem(s) found
exit code: 1
The script also passes ShellCheck with no warnings, which is a good bar to hold your own scripts to.
Part 14: Running It on a Schedule with cron
A health check you have to remember to run isn't really doing its job. On Linux, cron runs commands on a schedule. To schedule the script:
- Run
crontab -eto open your schedule in an editor. - Add this line, which runs the script every morning at 7:00, then save and close the editor:
0 7 * * * /home/azureuser/bin/health-check.sh -l /home/azureuser/health.log > /dev/null 2>&1
The five fields before the command are minute, hour, day of month, month and day of week:
| Schedule | Meaning |
|---|---|
0 7 * * * | every day at 07:00 |
*/15 * * * * | every 15 minutes |
0 9 * * 1-5 | 09:00 on weekdays |
0 2 1 * * | 02:00 on the first of every month |
The script writes its own log with -l, so the screen output is thrown away with > /dev/null 2>&1. Run crontab -l to check what's scheduled.
Warning: cron runs with a very small
PATHand none of your login settings. A script that works fine when you run it by hand can fail under cron because it can't find a command. Use full paths in the crontab line, and check the log the morning after you set it up.
Gotcha: Don't schedule anything in Azure Cloud Shell. The session shuts down after 20 minutes without activity, and your cron job goes with it. Cloud Shell is for learning and one-off commands. Use a VM or WSL for scheduled jobs.
Quick Reference
| Task | Syntax |
|---|---|
| Set a variable | name="value" (no spaces around =) |
| Default value | "${var:-default}" |
| Command output into a variable | result="$(command)" |
| Math | $(( a + b )) |
| Last exit code | $? |
| File exists | [[ -f "$path" ]] |
| String empty | [[ -z "$var" ]] |
| Number compare | (( a > b )) or [[ $a -gt $b ]] |
| Loop over files | for f in /path/*.log; do ...; done |
| Read a file | while IFS= read -r line; do ...; done < file |
| All arguments | "$@" |
| Array, all items | "${arr[@]}" |
| Redirect errors to output | command > file 2>&1 |
| Error message | echo "message" >&2 |
| Safe script header | set -euo pipefail |
| Cleanup on exit | trap cleanup EXIT |
| Trace a script | bash -x script.sh |
Where to Go from Here
The fastest way to get better at Bash is to automate something you do by hand. Here are a few ideas, roughly in order of difficulty:
- Add a check to the health-check script for certificate expiry with
openssl, or for a website returning 200 withcurl -o /dev/null -w '%{http_code}'. - Write a backup script that compresses a folder with
tar, names the archive with today's date, and deletes archives older than seven days withfind -mtime +7 -delete. - Write a script that reads a CSV of server names and runs a command on each one over SSH.
- Use the Azure CLI from a Bash script to list every VM in a subscription and flag the ones that are stopped but still allocated.
If you come from Windows, a lot of this probably felt familiar. Variables, conditions, loops, functions and exit codes all work the same way in PowerShell, with different spelling. I wrote a PowerShell basics guide a while back, and if you read it next to this post you'll see the overlap straight away. Once you know one shell well, you'll find the second one much easier to learn.
What matters more than the language is the habit. When you've typed the same commands three times, it's time to save them in a file. In the cloud you'll spend a lot of time in Bash, whether you're running the Azure CLI, connecting to a VM over SSH or writing a pipeline step, and every one of those jobs gets easier once Bash stops feeling foreign.



