Broken link checker for whole websites
Crawls the websites you give it and checks every link on every crawled page. Each row is one link on one page with the HTTP status, the full redirect chain, whether it is broken and the SEO basics of the page it sits on.
The crawl starts at each start URL and follows links breadth first, up to maxPages HTML pages. Only the sites you list are crawled. Links to other websites are checked with a single request and never followed further. robots.txt is read and followed for every host, and the default rate is at most 2 requests per second per domain.
Redirects are followed one hop at a time, so the whole chain is recorded. Servers that refuse automated checks (401, 403, 429 and LinkedIn's 999) are reported as restricted, not as broken. Each row also carries the page title, meta description, number of H1 headings, canonical URL and a short list of page issues such as a missing meta description. Check only sites you own or have permission to check.
Sample from a real run
5 of 143 rows (selected fields) from run KqYlkqUj6i7VgLM4E on 6 October 2026, exactly as the tool returned them.
pageUrl | linkUrl | anchorText | type | statusCode | linkStatus | isBroken | pageIssues |
|---|---|---|---|---|---|---|---|
| butik.nightwave.se/ | butik.nightwave.se/ | Hoppa till innehållet | internal | 200 | ok | false | title-too-long, meta-description-too-long |
| butik.nightwave.se/ | butik.nightwave.se/guider | Guider | internal | 200 | ok | false | title-too-long, meta-description-too-long |
| butik.nightwave.se/ | butik.nightwave.se/tillganglighet | AccessAudit | internal | 200 | ok | false | title-too-long, meta-description-too-long |
| butik.nightwave.se/ | butik.nightwave.se/anbud | BidBrief | internal | 200 | ok | false | title-too-long, meta-description-too-long |
| butik.nightwave.se/ | butik.nightwave.se/growth | Growth | internal | 200 | ok | false | title-too-long, meta-description-too-long |
Fields
Every row has these fields. Field names are stable between versions.
| Field | What it holds |
|---|---|
pageUrl | The crawled page the link was found on (after redirects) |
linkUrl | The link, resolved to an absolute URL without #fragment. null on the single row of a page without links |
anchorText | Link text, or the aria-label or image alt text when the link has no text |
type | internal (same site) or external |
rel | The link's rel attribute, for example nofollow |
statusCode | Final HTTP status code after redirects. null when there was no answer |
linkStatus | ok, redirect (works after one or more redirects), broken, restricted (401, 403, 429 or 999: the server refuses automated checks, the page may work in a browser), timeout (no answer in 15 seconds) or blocked-by-robots (not checked because robots.txt disallows it) |
isBroken | true for 4xx and 5xx answers (except the restricted codes), DNS errors, refused connections, TLS errors, redirect loops and more than 10 redirects |
redirectChain | Every redirect hop as { url, statusCode }. Empty when the link answered directly |
finalUrl | Where the link ends up after redirects |
error | Short code when there was no HTTP answer: dns-not-found, connection-refused, connection-reset, tls-error, timeout, redirect-loop, too-many-redirects, robots |
title | <title> of the page |
metaDescription | Meta description of the page |
h1Count | Number of <h1> elements on the page |
canonical | Canonical URL of the page (<link rel="canonical">), absolute |
pageIssues | SEO findings for the page: missing-title, title-too-long (over 60 characters), missing-meta-description, meta-description-too-long (over 160), missing-h1, multiple-h1, missing-canonical, noindex |
startUrl | The start URL whose crawl found the page |
checkedAt | When the page's links were checked (ISO 8601, UTC) |
Input example
Paste it into the JSON tab of the actor in Apify Console, or send it to the Apify API.
{
"onlyNew": false,
"maxPages": 10,
"startUrls": [
{
"url": "https://butik.nightwave.se/"
}
],
"brokenOnly": false,
"checkExternal": true,
"sameDomainOnly": true,
"requestsPerSecond": 2
}What people use it for
- Site migrations. Run before and after the move, compare the redirect chains and find old URLs that now answer 404.
- SEO audits. Broken internal links, long redirect chains and pages without a meta description or with several H1 headings, in one table you can filter and export to CSV or Excel.
- Agency maintenance. One scheduled run per client site with
onlyNewset to true reports only broken links that are new since the last run.
Price
20 USD per 1 000 pages, plus Apify platform usage.
The price is per crawled page, with all its links included. A run with the default input crawls at most 10 pages, so it costs at most 0.20 USD. The run on 6 October 2026 shown above crawled 10 pages of butik.nightwave.se and returned 143 link rows in 29 seconds, which is 0.20 USD. With onlyNew, a crawled page without a new broken link costs 0.005 USD.
Run it on a schedule
Set onlyNew to true and add the actor to a schedule in Apify Console (for example the cron expression 0 7 * * * for 07:00 every day). The actor remembers what it has already delivered for the same input, in a named key-value store in your own Apify account (nightwave-state-website-link-checker), and each run returns and charges only what is new. The first run returns everything in the selection. A run with nothing new finishes with 0 rows and costs nothing beyond platform usage.
Source and license
The data is what the websites you enter answer, read live during the run. There is no third-party database behind it. The tool follows robots.txt as described in RFC 9309 and identifies itself with the User-Agent NightwaveLinkChecker, so site owners can address it in robots.txt.
No data license applies: the output describes the websites you check and is yours to use. Under Apify's General Terms and Conditions you are responsible for having the right to crawl the sites you enter.
Try it
The tool runs on Apify. Apify handles your account, the runs and the payment, and you can try it with the input above before you schedule anything.
Built and maintained by Nightwave AB. Questions or bugs: kontakt@nightwave.se