Linux · 40 min

Automate a Task with a Shell Script

Turn the commands you learned into a real, reusable Bash script: arguments, conditionals, loops, and a safe header — building a log-summarizer that turns a raw logfile into a report. The hands-on companion to Terminal & Linux Essentials.

Problem

Typing the same three commands every day is how you know a task should be a script. The gap between "I can use the terminal" and "I automate my work" is small — a shebang, a few variables, an `if` and a `for` — but it's the leap that turns you from a tool user into a tool builder. In this lab you'll write a real Bash script from scratch: a **log summarizer** that takes a logfile, counts errors and warnings, finds the top offending sources, and prints a clean report. Along the way you'll handle arguments, guard against bad input, loop over lines, and add the safety header every serious script needs. You'll finish with a script you could actually drop into a server's crontab.

Objectives

Prerequisites

What you will build

A single script, logstats.sh, that you'll grow step by step. Given a logfile, it will:

Each step adds one shell concept and one feature to the script, so you finish with something real
rather than toy snippets.

Why scripting, not just commands

A command you type is gone the moment it scrolls off screen. A script is that knowledge captured
— versioned, shareable, schedulable. The same grep | sort | uniq -c you'd run by hand becomes,
wrapped in a script, a report anyone (or cron) can run identically at 2 a.m. That reproducibility is
the entire value of automation.

A sample logfile to work on

You'll generate a realistic logfile in Step 1 so everyone works on the same data. Real logs are
messier, but the shape — a timestamp, a level, a source, a message per line — is universal.

How to work through this lab

Build the script incrementally and run it after every step (./logstats.sh sample.log), reading
the real output. Don't paste the whole thing at the end — the point is to see each piece work.

Steps

  1. From command to script: shebang and chmod

    First, create sample data and your first runnable script.

    Generate a sample logfile

    Paste this to create sample.log with a realistic mix of levels and sources:

    cat > sample.log <<'EOF'
    2026-07-10 09:01:12 INFO  auth      user login ok
    2026-07-10 09:02:03 WARN  auth      slow response
    2026-07-10 09:02:44 ERROR payments  gateway timeout
    2026-07-10 09:03:10 INFO  web       GET /home 200
    2026-07-10 09:04:55 ERROR payments  card declined
    2026-07-10 09:05:01 WARN  web       deprecated endpoint
    2026-07-10 09:06:22 ERROR auth      token expired
    2026-07-10 09:07:30 INFO  web       GET /pricing 200
    2026-07-10 09:08:14 ERROR payments  gateway timeout
    EOF
    wc -l sample.log      # 9 lines
    

    Write the first script

    Create logstats.sh:

    #!/usr/bin/env bash
    echo "logstats starting"
    wc -l "sample.log"
    
    • The first line is the shebang: #!/usr/bin/env bash tells the OS to run the file with
      Bash. Without it, the file is just text.
    • Using env bash finds Bash on the user's PATH, which is more portable than hard-coding /bin/bash.

    Make it executable and run it

    chmod +x logstats.sh    # add the execute permission
    ./logstats.sh           # the ./ says "run the file here", not a command on PATH
    

    A file is not runnable just because it contains commands — it needs the execute bit. That's
    what chmod +x grants; ls -l logstats.sh now shows an x in the permissions.

    Deliverable for this step: ls -l logstats.sh output showing the x bit, and the script's output.

  2. Arguments and validation

    Hard-coding sample.log makes a one-trick script. Take the filename as an argument.

    Read the argument

    Replace the body of logstats.sh:

    #!/usr/bin/env bash
    
    logfile="$1"        # first argument passed on the command line
    echo "Analyzing: $logfile"
    wc -l "$logfile"
    
    • $1 is the first argument, $2 the second, and so on. $@ is all arguments, $# is how
      many
      there are.
    • Always quote variables ("$logfile") so filenames with spaces don't break the script.

    Run it two ways:

    ./logstats.sh sample.log
    ./logstats.sh              # no argument — what happens?
    

    With no argument, $1 is empty and wc errors out ugly. A real script validates first.

    Guard the input

    Add a check at the top, before any work:

    #!/usr/bin/env bash
    
    if [ "$#" -lt 1 ]; then
      echo "Usage: $0 <logfile>" >&2
      exit 1
    fi
    
    logfile="$1"
    
    if [ ! -f "$logfile" ]; then
      echo "Error: '$logfile' is not a readable file" >&2
      exit 2
    fi
    
    echo "Analyzing: $logfile"
    

    What each piece does:

    • [ "$#" -lt 1 ] → "fewer than 1 argument". If so, print usage to stderr (>&2) and exit non-zero.
    • $0 is the script's own name — a self-documenting usage message.
    • [ ! -f "$logfile" ] → "not a regular file". Fail early with a distinct exit code.
    • Different exit codes (1 for usage, 2 for bad file) let other scripts react to why you failed.

    Deliverable for this step: the output of running with a missing arg and with a non-existent file, showing the error messages and (via echo $?) the exit codes.

  3. Conditionals and counting

    Now the real work: count what's in the log and branch on the result.

    Count levels

    Add below the validation:

    total=$(wc -l < "$logfile")
    errors=$(grep -c 'ERROR' "$logfile")
    warnings=$(grep -c 'WARN' "$logfile")
    
    echo "Total lines: $total"
    echo "Errors:      $errors"
    echo "Warnings:    $warnings"
    
    • $( ... ) is command substitution: it runs the command and captures its output into the variable.
    • wc -l < "$logfile" feeds the file via stdin so wc prints just the number (no filename).
    • grep -c PATTERN counts matching lines directly — no need to pipe to wc.

    Branch on the counts

    Turn numbers into a verdict with if/elif/else:

    if [ "$errors" -eq 0 ]; then
      echo "Status: healthy ✅"
    elif [ "$errors" -lt 3 ]; then
      echo "Status: warning ⚠️  ($errors errors)"
    else
      echo "Status: critical ❌ ($errors errors)"
    fi
    

    Number comparisons use word operators: -eq (=), -lt (<), -gt (>), -le, -ge. (String
    comparisons use = and != instead — a classic gotcha.)

    Run ./logstats.sh sample.log — with 4 errors in the sample, you should land in the critical branch.

    Deliverable for this step: the script's output showing the three counts and the status verdict.

  4. Loops and aggregation: top error sources

    A count is useful; where the errors come from is actionable. Aggregate by source.

    The pipeline that finds top offenders

    The third column of each log line is the source (auth, payments, web). Extract, group and rank:

    echo "Top error sources:"
    grep 'ERROR' "$logfile" | awk '{ print $4 }' | sort | uniq -c | sort -rn
    

    Read the pipeline left to right — this is the essence of Unix:

    • grep 'ERROR' → keep only error lines.
    • awk '{ print $4 }' → print the 4th whitespace field (the source). awk splits on whitespace automatically.
    • sort → group identical sources together (required before uniq).
    • uniq -c → collapse duplicates and prefix each with its count.
    • sort -rn → sort by that count, numerically (-n), highest first (-r).

    On the sample, payments should top the list with 3 errors.

    Loop over the results

    To format each source line yourself, loop with while read:

    echo "Breakdown:"
    grep 'ERROR' "$logfile" | awk '{ print $4 }' | sort | uniq -c | sort -rn |
    while read -r count source; do
      echo "  - $source caused $count error(s)"
    done
    
    • while read -r count source reads each line, splitting it into two variables at whitespace.
    • -r stops backslashes being interpreted — always use it with read.
    • The loop body runs once per source, letting you format, threshold, or alert per line.

    Deliverable for this step: the "Top error sources" and "Breakdown" output for the sample log.

  5. Make it robust: set -euo pipefail and a report function

    A script that keeps going after an error can do real damage. Add the safety header and structure.

    The safety header

    Put this right after the shebang:

    #!/usr/bin/env bash
    set -euo pipefail
    

    Each flag prevents a class of silent bug:

    • -e → exit immediately if any command fails, instead of blundering on.
    • -u → error on undefined variables (a typo'd $logfyle becomes a crash, not empty string).
    • -o pipefail → a pipeline fails if any stage fails, not just the last one.

    This trio is the single most important habit for reliable scripts. Test it: temporarily
    reference an undefined variable and watch the script stop instead of producing garbage.

    Wrap the report in a function

    Group the output into a function for readability and reuse:

    print_report() {
      local file="$1"
      echo "=============================="
      echo " Log report: $file"
      echo " Generated:  $(date '+%Y-%m-%d %H:%M')"
      echo "=============================="
      echo "Total lines: $(wc -l < "$file")"
      echo "Errors:      $(grep -c 'ERROR' "$file")"
      echo "Warnings:    $(grep -c 'WARN'  "$file")"
    }
    
    print_report "$logfile"
    
    • local file="$1" scopes the variable to the function — it won't leak into the rest of the script.
    • Functions turn a long script into named, testable pieces, exactly like functions in any language.

    Run the full script once more and confirm the report renders cleanly with the header and counts.

    Deliverable for this step: the full report output, and a note on what happened when you triggered set -u with an undefined variable.

  6. Submit: your script and a sample run

    Package the script, the sample log and a captured run, and submit.

    Assemble the repository

    Your logstats.sh should now contain, in order: the shebang, set -euo pipefail, argument
    validation, the print_report function, the top-error-sources pipeline/loop, and the status
    verdict. Create a SUBMISSION.md capturing a real run:

    # Shell Scripting Lab — <your name>
    
    ## The script
    See logstats.sh (final version).
    
    ## Sample run
    $ ./logstats.sh sample.log
    <paste the full real output: report header, counts, status, top sources>
    
    ## Validation behavior
    $ ./logstats.sh            # (usage error, exit 1)
    $ ./logstats.sh nope.log   # (bad file, exit 2)
    <paste both, with `echo $?` after each>
    
    ## Reflection (2-3 sentences)
    In my own words: what `set -euo pipefail` protects against, and one task
    from my own life I could automate with a script like this.
    

    Submit

    git init
    git add logstats.sh sample.log SUBMISSION.md
    git commit -m "Shell scripting lab — log summarizer"
    # push to a public repo or gist, then submit that URL
    

    Submission criteria (self-check)

    • logstats.sh starts with a shebang and set -euo pipefail
    • It validates argument count and file existence, with distinct non-zero exit codes
    • It uses command substitution $(...) to count lines/errors/warnings
    • It uses an if/elif/else to print a status verdict
    • It aggregates top error sources with grep | awk | sort | uniq -c | sort -rn
    • A while read loop or a function is used to structure output
    • The submission shows a real run plus the two validation-failure cases with their exit codes

    What's next

    You can now turn any repetitive terminal task into a safe, reusable script. A natural next step
    is scheduling it (cron) and, in the Backend Developer Path, using scripts like this to
    automate parts of building and deploying the API you'll create.