Crawl a site or check a list of URLs. See status codes, full redirect chains, and the exact pages each broken link was found on. Export to Excel, CSV, JSON or an interactive HTML report.
npx check-http-status https://your-site.com -o report.xlsx$ npx check-http-status acme-outdoor.example -e /docs 301 /summer-sale → /new-arrivals found on / ("Summer sale") 301 /products/sleeping-bags ⇒ /collections/ sleeping-bags (404, 2 hops) found on /products/ ("Sleeping bags") 404 partner-site.example/deleted-article found on /blog/best-hiking-trails 410 /blog/2019/old-news found on /blog/ ("Old news") 500 /products/stoves found on /products/ ("Stoves") found on /blog/packing-list ("Stoves") 36 URLs (17 pages crawled, 2s): 22 OK, 7 redirects, 4 not found, 2 5xx, 1 error 3 orphan page(s): in the sitemap but not linked.
It runs on your machine (macOS, Windows or Linux), so it can also check staging sites, password-protected sites and sites behind a VPN.
Follows every internal link from a start page, breadth-first, with page and depth limits.
Every broken link and redirect lists the pages that link to it, plus the link text.
Each hop is reported, with the final URL and status. Redirect loops are detected.
External links are checked but never crawled, so they're listed without leaving your site.
Skip sections like /docs. Excluded URLs are still listed and checked, just not crawled. Supports wildcards and regex.
Reads sitemap indexes and .xml.gz files, and flags pages in the sitemap that nothing links to.
Every URL found, listed once with all the pages it was found on in one cell. Excel sheets: Summary, All URLs, Issues and External Links.
HTTP Basic auth (sent only to your site, never to external links) and custom headers.
Add --drive or --onedrive to upload the report. Add --public to get a link anyone can open.
Save settings in check-http-status.config.json (with a JSON Schema for autocomplete). Type definitions included.
--fail exits with code 1 when broken links are found, so you can block a deploy.
A crawl of a sample store with 40 URLs, including 404s, a 410, 5xx errors, a two-hop redirect chain, a redirect loop, a DNS failure, a broken image and orphan pages. Filter it, search it, and click any row to see where the link was found.
Requires Node.js 18.17 or newer. Run it with npx, install it globally, or use it as a library.
# Crawl a site, skip /docs (still listed), save an Excel report npx check-http-status https://www.example.com --exclude /docs -o report.xlsx # Optional: add a sitemap to also find orphan pages; HTML report npx check-http-status crawl example.com -s https://example.com/sitemap_index.xml -o report.html # Only check specific URLs, or every URL in a sitemap npx check-http-status check https://example.com/old-page https://example.com/pricing npx check-http-status check -s https://example.com/sitemap.xml --issues-only # Password-protected staging site CHS_AUTH=user:pass npx check-http-status https://staging.example.com # Install globally npm install -g check-http-status check-http-status --help
const checkHttpStatus = require('check-http-status');
// Just the website URL: pages are found by crawling, no sitemap needed
const report = await checkHttpStatus({
crawl: 'https://www.example.com',
exclude: ['/docs'],
maxPages: 2000,
export: [
{ format: 'xlsx', location: './reports/' },
{ format: 'html', file: './reports/report.html' }
],
options: {
auth: { username: 'user', password: 'pass' },
headers: { 'Accept-Language': 'en' },
timeout: 15000
}
});
// report.summary.counts → { OK: 412, Redirect: 9, 'Not Found': 3, ... }
for (const r of report.results.filter((r) => r.category === 'Not Found')) {
console.log(r.url, 'found on', r.foundOn.map((s) => s.page));
}# .github/workflows/links.yml
name: Check links
on:
schedule: [{ cron: '0 6 * * 1' }]
workflow_dispatch:
jobs:
links:
runs-on: ubuntu-latest
steps:
- uses: actions/setup-node@v4
with: { node-version: 20 }
- run: npx check-http-status https://www.example.com --fail -o report.html
- uses: actions/upload-artifact@v4
if: always()
with: { name: link-report, path: report.html }The CLI flags and Node.js options are equivalent. Run check-http-status --help for the full list.
| CLI | Node.js | Default | Description |
|---|---|---|---|
crawl <url> | crawl | Website URL. Pages are found by following internal links. | |
check <url…> | urls | Check only these URLs (and their redirects). | |
-s, --sitemap | sitemaps | Optional. Sitemap or sitemap index URL(s). Not needed for crawling. Gzip supported. | |
-i, --include | include | Only crawl URLs matching one of these rules. | |
-e, --exclude | exclude | Don't crawl matching URLs. They're still listed and checked. | |
--max-pages | maxPages | 5000 | Maximum pages to crawl. |
--max-depth | maxDepth | ∞ | Maximum link depth from the start URL. |
--no-external | checkExternal: false | checked | Don't check external links. Excluded pages and assets are still checked. |
--assets | checkAssets | off | Also check images, scripts, CSS and iframes. |
--subdomains | subdomains | off | Treat subdomains as internal. |
--respect-robots | respectRobots | off | Don't request URLs your robots.txt disallows (they're listed as blocked); honour Crawl-delay. |
--config | loadConfig() | Load settings from a JSON file. check-http-status.config.json in the current folder is used automatically. | |
--ignore-query | ignoreQuery | off | Treat /a?x=1 and /a as the same URL. |
-c, --concurrency | concurrency | 10 | Parallel requests. |
-t, --timeout | options.timeout | 15 s | Per-request timeout (seconds on the CLI, milliseconds in Node.js). |
--delay | delay | 0 | Wait between requests to the same site (seconds on the CLI, milliseconds in Node.js). |
--retries | retries | 2 | Retry 429/502/503/504 and dropped connections, honouring Retry-After. |
--user-agent | userAgent | Custom User-Agent header. | |
-a, --auth | options.auth | HTTP Basic auth, sent only to the site being checked. | |
-H, --header | options.headers | Extra request headers. | |
-o, --output | export | .xlsx, .csv, .json or .html report. | |
--issues-only / --all | skip200 | Hide or show 200 OK URLs in the terminal. --issues-only also drops them from saved reports, which otherwise list every URL. | |
--drive | googleDrive | Upload the report(s) to Google Drive. Excel/CSV become Google Sheets. Run check-http-status drive-login once first. | |
--onedrive | oneDrive | Upload the report(s) to OneDrive. Run check-http-status onedrive-login once first. | |
--public | public: true | Anyone with the link can view the uploaded files. | |
--fail | Exit with code 1 if any 4xx, 5xx or connection errors are found. |
Online checkers can't reach staging servers, intranet sites or VPN-only sites, and many sites block their datacenter IPs. Running on your own machine gets past all of that, and you can crawl as many pages as you need.
v1 configs (urls, sitemaps, skip200, export, options) still work. v2 adds crawling and a CLI, and now returns the results. It also throws errors instead of exiting the process, and requires Node.js 18.17+.