check-http-status

Find every broken link and redirect on your website.

Crawl a site or check a list of URLs. See status codes, full redirect chains, and the exact pages each broken link was found on. Export to Excel, CSV, JSON or an interactive HTML report.

$npx check-http-status https://your-site.com -o report.xlsx
$ npx check-http-status acme-outdoor.example -e /docs

301 /summer-sale → /new-arrivals
      found on / ("Summer sale")
301 /products/sleeping-bags ⇒ /collections/
      sleeping-bags (404, 2 hops)
      found on /products/ ("Sleeping bags")
404 partner-site.example/deleted-article
      found on /blog/best-hiking-trails
410 /blog/2019/old-news
      found on /blog/ ("Old news")
500 /products/stoves
      found on /products/ ("Stoves")
      found on /blog/packing-list ("Stoves")

36 URLs (17 pages crawled, 2s): 22 OK,
7 redirects, 4 not found, 2 5xx, 1 error
3 orphan page(s): in the sitemap but not linked.

Everything you need for a link audit

It runs on your machine (macOS, Windows or Linux), so it can also check staging sites, password-protected sites and sites behind a VPN.

⟳

Full-site crawl

Follows every internal link from a start page, breadth-first, with page and depth limits.

⌖

Where it was found

Every broken link and redirect lists the pages that link to it, plus the link text.

↪

Redirect chains

Each hop is reported, with the final URL and status. Redirect loops are detected.

⇥

External links

External links are checked but never crawled, so they're listed without leaving your site.

⊘

Include / exclude rules

Skip sections like /docs. Excluded URLs are still listed and checked, just not crawled. Supports wildcards and regex.

≡

Sitemaps & orphans

Reads sitemap indexes and .xml.gz files, and flags pages in the sitemap that nothing links to.

▦

Excel, CSV, JSON, HTML

Every URL found, listed once with all the pages it was found on in one cell. Excel sheets: Summary, All URLs, Issues and External Links.

🔒

Staging & auth

HTTP Basic auth (sent only to your site, never to external links) and custom headers.

☁

Google Drive & OneDrive

Add --drive or --onedrive to upload the report. Add --public to get a link anyone can open.

{}

Config file & TypeScript

Save settings in check-http-status.config.json (with a JSON Schema for autocomplete). Type definitions included.

✓

CI friendly

--fail exits with code 1 when broken links are found, so you can block a deploy.

See a real report

A crawl of a sample store with 40 URLs, including 404s, a 410, 5xx errors, a two-hop redirect chain, a redirect loop, a DNS failure, a broken image and orphan pages. Filter it, search it, and click any row to see where the link was found.

Usage

Requires Node.js 18.17 or newer. Run it with npx, install it globally, or use it as a library.

# Crawl a site, skip /docs (still listed), save an Excel report
npx check-http-status https://www.example.com --exclude /docs -o report.xlsx

# Optional: add a sitemap to also find orphan pages; HTML report
npx check-http-status crawl example.com -s https://example.com/sitemap_index.xml -o report.html

# Only check specific URLs, or every URL in a sitemap
npx check-http-status check https://example.com/old-page https://example.com/pricing
npx check-http-status check -s https://example.com/sitemap.xml --issues-only

# Password-protected staging site
CHS_AUTH=user:pass npx check-http-status https://staging.example.com

# Install globally
npm install -g check-http-status
check-http-status --help

Options

The CLI flags and Node.js options are equivalent. Run check-http-status --help for the full list.

CLINode.jsDefaultDescription
crawl <url>crawlWebsite URL. Pages are found by following internal links.
check <url…>urlsCheck only these URLs (and their redirects).
-s, --sitemapsitemapsOptional. Sitemap or sitemap index URL(s). Not needed for crawling. Gzip supported.
-i, --includeincludeOnly crawl URLs matching one of these rules.
-e, --excludeexcludeDon't crawl matching URLs. They're still listed and checked.
--max-pagesmaxPages5000Maximum pages to crawl.
--max-depthmaxDepth∞Maximum link depth from the start URL.
--no-externalcheckExternal: falsecheckedDon't check external links. Excluded pages and assets are still checked.
--assetscheckAssetsoffAlso check images, scripts, CSS and iframes.
--subdomainssubdomainsoffTreat subdomains as internal.
--respect-robotsrespectRobotsoffDon't request URLs your robots.txt disallows (they're listed as blocked); honour Crawl-delay.
--configloadConfig()Load settings from a JSON file. check-http-status.config.json in the current folder is used automatically.
--ignore-queryignoreQueryoffTreat /a?x=1 and /a as the same URL.
-c, --concurrencyconcurrency10Parallel requests.
-t, --timeoutoptions.timeout15 sPer-request timeout (seconds on the CLI, milliseconds in Node.js).
--delaydelay0Wait between requests to the same site (seconds on the CLI, milliseconds in Node.js).
--retriesretries2Retry 429/502/503/504 and dropped connections, honouring Retry-After.
--user-agentuserAgentCustom User-Agent header.
-a, --authoptions.authHTTP Basic auth, sent only to the site being checked.
-H, --headeroptions.headersExtra request headers.
-o, --outputexport.xlsx, .csv, .json or .html report.
--issues-only / --allskip200Hide or show 200 OK URLs in the terminal. --issues-only also drops them from saved reports, which otherwise list every URL.
--drivegoogleDriveUpload the report(s) to Google Drive. Excel/CSV become Google Sheets. Run check-http-status drive-login once first.
--onedriveoneDriveUpload the report(s) to OneDrive. Run check-http-status onedrive-login once first.
--publicpublic: trueAnyone with the link can view the uploaded files.
--failExit with code 1 if any 4xx, 5xx or connection errors are found.

Why run it locally?

Online checkers can't reach staging servers, intranet sites or VPN-only sites, and many sites block their datacenter IPs. Running on your own machine gets past all of that, and you can crawl as many pages as you need.

Upgrading from v1

v1 configs (urls, sitemaps, skip200, export, options) still work. v2 adds crawling and a CLI, and now returns the results. It also throws errors instead of exiting the process, and requires Node.js 18.17+.