sed vs awk: When to Use Which (With Real Examples)

sed and awk solve different problems, even though most “comparison” posts treat them as interchangeable. sed is a line-based stream editor: it reads input one line at a time and applies text transformations to it, with no concept of columns or data types. awk is a field-based scripting language: it automatically splits each line into fields, gives you variables to track state across the whole file, and can do arithmetic. If your problem is “change this pattern to that pattern on each line,” reach for sed. If your problem is “do something with column 3, or count something across every line,” reach for awk.

sed: when and how

sed is built around one idea: read a line into a buffer (the “pattern space”), run your script against it, print the result, move to the next line. There’s no persistent memory of previous lines unless you explicitly ask for it, and there’s no native arithmetic. That narrow scope is exactly why it’s fast and why one-liners rarely get longer than a single s/// command.

The core syntax is sed 'script' file, where the most common script is a substitution: s/regexp/replacement/flags. Without a flag, sed replaces only the first match per line. Add the g flag to replace every match on the line:

sed 's/foo/bar/' file.txt      # replaces the first "foo" on each line
sed 's/foo/bar/g' file.txt     # replaces every "foo" on each line

Neither of those commands changes file.txt — by default sed writes the result to standard output and leaves the original file untouched. That’s useful for previewing a change before you commit to it.

Deleting lines that match a pattern is just as common as substituting text inside them. The d command deletes the current pattern space and starts the next cycle, so pairing it with an address pattern strips any line that matches:

sed '/^#/d' nginx.conf   # drop every comment line

To actually modify the file instead of streaming the result to stdout, you need -i for in-place editing. This is where sed scripts break the first time someone runs them on a different machine, and it’s worth getting right.

Important: GNU sed (the default on every mainstream Linux distro) treats the backup suffix after -i as optional — sed -i 's/foo/bar/' file.txt edits the file in place with no backup. BSD sed, which is what ships on macOS, requires that suffix argument. Run the exact same command on macOS and sed consumes 's/foo/bar/' as the backup extension instead of as your script, and the actual script argument is now missing — the command fails or behaves nothing like you expect. The fix is to always pass an explicit extension on macOS, using an empty string if you don’t want a backup file:

# GNU sed (Linux) — suffix is optional, no backup by default
sed -i 's/foo/bar/' file.txt

# BSD sed (macOS) — suffix is required; '' means "no backup"
sed -i '' 's/foo/bar/' file.txt

A script that works fine on a Linux CI runner can fail outright on a teammate’s Mac for exactly this reason. If a shell script needs to run on both, check for GNU sed with sed --version (GNU sed prints a version banner; BSD sed errors on the unrecognized flag) and branch the -i call, or install GNU sed on macOS via Homebrew as gsed and call that directly.

The last piece worth knowing before you move to awk is the & backreference. Inside a replacement, an unescaped & stands for the entire text that the regexp matched, so you can wrap or decorate a match without retyping it:

sed 's/[0-9]\+/[&]/g' access.log   # wraps every run of digits in brackets

That one line turns every number in a log file into [1234] without you having to write out what the number was. It’s the sed equivalent of a capture group, except it always refers to the whole match rather than a specific parenthesized group.

awk: when and how

awk treats every input line as a record made of fields, splits on whitespace by default, and gives you built-in variables to work with those fields directly — no regex needed just to grab “the third column.” That’s the entire reason to reach for it over sed: the moment your problem involves a specific column, a count, a sum, or state that has to persist from one line to the next, awk’s structure fits and sed’s doesn’t.

The basic program shape is pattern { action }, where either half can be omitted but not both. BEGIN runs once before any input is read, and END runs once after the last line. Fields inside a record are referenced as $1, $2, and so on, with $0 standing for the whole line:

awk -F: '{ print $1 }' /etc/passwd   # print just the username field

The -F flag sets the field separator for that run. Here it’s a colon, because /etc/passwd is colon-delimited. Without -F, awk defaults FS to a single space — and that default isn’t a literal space character, it’s a special case that means “any run of whitespace,” collapsing multiple spaces or tabs into one separator and ignoring leading and trailing whitespace entirely. That’s genuinely different behavior from setting FS to any other single character.

Important: set FS to a literal tab (-F'\t') or any other single character that isn’t a space, and awk stops collapsing runs of it — two tabs in a row now produce an empty field between them, as does a leading or trailing one. That trips people up on output that mixes tabs and spaces inconsistently (copy-pasted terminal output is the usual culprit): leaving FS at its default handles irregular whitespace correctly, while pinning it to a non-space character doesn’t forgive irregular spacing. Need a literal single space without the collapsing magic? Use a one-character regexp instead of a bare space: FS="[ ]".

Two built-ins you’ll use constantly are NR (the number of the current record, i.e. the running line count) and NF (the number of fields in the current record). They’re handy for sanity-checking input before you trust it:

awk '{ print NR, NF, $0 }' file.txt       # line number, field count, full line
awk 'NF != 5 { print NR": expected 5 fields, got "NF }' data.csv

That second command flags every malformed row in a file you expect to be consistently 5 columns wide — a real sanity check you’d run before feeding a file into something stricter.

Where awk earns its keep over sed is arithmetic across lines. You can accumulate a running total in a variable and print it once, at the end, with END:

ls -l | awk '{ x += $5 } END { print "total bytes:", x }'

Nothing in sed can do that — there’s no persistent numeric variable to add into, and no concept of “after the last line.” That gap is the clearest signal of which tool you actually need: the instant you’re summing, counting, or comparing across rows instead of transforming each row independently, you’re writing awk, not sed.

Combining sed, awk, and grep

In practice you rarely pick just one of these tools — you chain them, each doing the part it’s actually good at. grep filters which lines matter, awk pulls out and reshapes the fields you need, and sed cleans up what’s left. Here’s a realistic one: pulling the message out of pipe-delimited application log lines, but only for ERROR rows, with leading whitespace stripped from the result:

# log line format: 2026-10-08 14:32:01 | ERROR | OOMKilled: container exceeded memory limit
grep "ERROR" app.log | awk -F'|' '{ print $3 }' | sed 's/^[[:space:]]*//'

grep narrows the stream so awk never field-splits a line that’s getting thrown away anyway. awk’s -F'|' treats the pipe as the field separator and isolates the third field — the message itself. sed’s job at the end is deliberately small: strip the leftover leading space using the POSIX [[:space:]] class rather than a literal space, so it also catches a stray tab if the log format shifts.

That division of labor is the real skill here. Each tool does the one thing it’s built for, and the pipeline reads top to bottom as “filter, then extract, then tidy” — which is also how you should think through any sed/awk/grep problem before you start typing.

Frequently Asked Questions

Is sed faster than awk?

For simple line-based substitution or deletion, yes — sed has a smaller execution model and less per-line overhead than awk’s field-splitting and variable handling. The gap mostly disappears once your awk script is doing something sed can’t do at all, like summing a column, because at that point you’d need sed plus extra tools to match it.

Can awk do what sed does?

Mostly. awk’s sub() and gsub() functions handle regex substitution, and $0 gives you the whole line to rewrite, so you can replicate most sed one-liners in awk. The syntax is heavier for that specific job, though — if all you need is a substitution, sed says it in fewer characters.

Can sed do math like awk can?

No. sed has no arithmetic operators and no numeric variables — it operates purely on text patterns within a line. If your task involves counting, summing, or comparing numbers across lines, that’s awk’s job, not sed’s.

Which one should I learn first?

sed, because its scope is small enough to pick up from a handful of real examples in an afternoon. awk is a full scripting language with its own control flow, and it’s easier to learn once you already understand regular expressions from using sed and grep.

Do sed and awk behave the same on Linux and macOS?

awk’s core behavior is consistent across GNU and BSD/macOS versions for everything covered here. sed is the one to watch — macOS ships BSD sed, and its -i flag requires an explicit backup-suffix argument where GNU sed’s is optional. That’s the single most common reason a sed script that works on a Linux CI runner fails on a contributor’s Mac.

Quick Summary

  • sed is a line-based stream editor for pattern-based substitution and deletion; it has no fields, no arithmetic, and no cross-line state.
  • awk is a field-based scripting language; use it the moment you need a specific column, a running count or sum, or logic that spans multiple lines.
  • macOS/BSD sed’s -i requires a backup-suffix argument (use -i '' for none); GNU sed’s is optional — this is the classic cross-platform script breaker.
  • awk’s default field separator isn’t a literal space — it’s a special case that collapses whitespace runs and ignores leading/trailing space. Setting FS to any other single character loses that forgiveness.
  • Real pipelines combine all three: grep filters, awk extracts and reshapes fields, sed does the final text cleanup.