01Before you start
The API lets your own software start crawls and read results, with no browser involved. It is available on Pro+.
Create a key in Settings under API keys. The key is shown once, at creation, and is never recoverable: we store only a hash of it. If you lose it, revoke it and create another.
A key authenticates as you and can start crawls against your account, so treat it like a password. Keep it in an environment variable or a secret manager, never in a repository. If one leaks, revoke it: revoking takes effect on the next request.
02Authentication
Send the key as a bearer token on every request. The base URL is https://deadlinkcrawler.com/api/v1.
curl https://deadlinkcrawler.com/api/v1/crawls/CRAWL_ID \
-H "Authorization: Bearer $DLC_API_KEY"A missing, malformed or unrecognised key returns 401 unauthorized. A valid key on a plan without API access returns 403 plan_limit_exceeded, so the two cases are easy to tell apart: the first means check your key, the second means check your plan.
03Start a crawl
POST /crawls returns immediately with a queued crawl. It never blocks: a large crawl can run for a long time, so you start it here and poll for the result.
curl -X POST https://deadlinkcrawler.com/api/v1/crawls \
-H "Authorization: Bearer $DLC_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}'{
"id": "3bfa25d2-c9e8-49b0-baf8-daab0d7213aa",
"status": "queued",
"target_url": "https://example.com/",
"created_at": "2026-10-04T21:05:28.744Z",
"started_at": null,
"finished_at": null,
"scan_id": "cad99d80-708b-4f01-b6ce-bbd69e472b3b",
"stats": { "pages_scanned": 0, "links_checked": 0, "broken_count": 0 },
"stopped_reason": null
}Body
| Field | Type | Default | Meaning |
|---|---|---|---|
| url | string, required | The page to start from. Must be http:// or https://. | |
| check_assets | boolean | true | Also check images, stylesheets and scripts, not only links. |
| strict_status_codes | boolean | false | Treat 401, 403 and 405 as broken. Off by default because they usually mean bot-blocking rather than a dead link. |
| single_page | boolean | false | Check the links on the starting page only, without following them. |
Your plan sets how many pages one crawl will scan (25,000 on Pro+) and how many crawls can run at once (4). Starting one too many returns 409 concurrency_limit rather than queueing it, so your script decides whether to wait or to cancel something.
04Check on a crawl
GET /crawls/{id} returns the same shape, updated. Poll it until status is terminal.
curl https://deadlinkcrawler.com/api/v1/crawls/$CRAWL_ID \
-H "Authorization: Bearer $DLC_API_KEY"{
"id": "3bfa25d2-c9e8-49b0-baf8-daab0d7213aa",
"status": "complete",
"target_url": "https://example.com/",
"created_at": "2026-10-04T21:05:28.744Z",
"started_at": "2026-10-04T21:05:29.421Z",
"finished_at": "2026-10-04T21:05:30.329Z",
"scan_id": "cad99d80-708b-4f01-b6ce-bbd69e472b3b",
"stats": { "pages_scanned": 1, "links_checked": 1, "broken_count": 0 },
"stopped_reason": null
}| status | Terminal | Meaning |
|---|---|---|
| queued | no | Accepted, not started yet. |
| running | no | In progress. stats update as it goes. |
| complete | yes | Finished. Results are ready. |
| failed | yes | Could not complete. See stopped_reason. |
| cancelled | yes | You cancelled it. Whatever it found is kept. |
stopped_reason
Null on an ordinary completion. When set, it says why a crawl ended early. A crawl that reaches your plan’s page limit is still complete, not failed: it scanned everything your plan allows and its results are real, so it carries page_limit_reached rather than an error status.
"status": "complete",
"stopped_reason": {
"code": "page_limit_reached",
"message": "Stopped at the time ceiling after scanning 25,000 of 25,000 pages…"
}05Read the results
GET /crawls/{id}/links returns the broken and redirected links, a page at a time. You can start reading while the crawl is still running.
curl "https://deadlinkcrawler.com/api/v1/crawls/$CRAWL_ID/links?limit=100" \
-H "Authorization: Bearer $DLC_API_KEY"{
"links": [
{
"seq": 3719,
"type": "broken",
"url": "https://example.com/gone",
"status_code": 404,
"error_type": "Not Found",
"link_text": "Our old guide",
"found_on_page": "https://example.com/blog",
"resource_type": "link",
"redirects_to": null
}
],
"next_cursor": "3723.0",
"has_more": true
}| Query | Default | Meaning |
|---|---|---|
| limit | 100 | Links per page, up to 500. |
| cursor | (start) | Pass back the next_cursor from the previous page. |
| type | all | Filter to broken or redirected. |
Paging
Keep calling with the next_cursor you were given until has_more is false. Treat the cursor as opaque: it encodes a position in more than one table and its format may change. An unrecognised cursor starts from the beginning rather than failing, so a replayed cursor returns data instead of an error.
06Cancel a crawl
DELETE /crawls/{id} stops a crawl in progress and returns its final state. Anything it found before stopping is kept.
curl -X DELETE https://deadlinkcrawler.com/api/v1/crawls/$CRAWL_ID \
-H "Authorization: Bearer $DLC_API_KEY"Cancelling is idempotent, and a crawl that already finished is left alone rather than treated as an error: by the time your cancel lands the crawl may have completed on its own, which is a race you cannot avoid. Check the status in the response to see which happened.
07Rate limits
Requests are limited per key. Every response carries your current budget, so you never have to guess.
| Header | Meaning |
|---|---|
| X-RateLimit-Limit | Requests allowed per window. |
| X-RateLimit-Remaining | Requests left in the current window. |
| X-RateLimit-Reset | When the window resets, as a Unix timestamp in seconds. |
| Retry-After | Seconds to wait. Only on a 429. |
retry-after: 20
x-ratelimit-limit: 60
x-ratelimit-remaining: 0
x-ratelimit-reset: 1791148140
{ "error": { "code": "rate_limited", "message": "Rate limit exceeded. Retry in 20 seconds." } }When polling, wait a few seconds between checks rather than looping as fast as you can. A crawl of any size takes minutes, so polling every 5 to 10 seconds costs you nothing and keeps you well inside the limit.
08Errors
Every error has the same shape, so one handler covers all of them.
{ "error": { "code": "not_found", "message": "No crawl with that id." } }Branch on code, never on the message. The codes below are stable; the messages may be reworded.
| Code | HTTP | What to do |
|---|---|---|
| unauthorized | 401 | Check the key and the Authorization header. |
| plan_limit_exceeded | 403 | The key is fine; the plan does not include the API. |
| forbidden | 403 | The key may not do this. |
| not_found | 404 | No crawl with that id on your account. |
| invalid_request | 400 | Fix the body or query parameters; the message says which. |
| concurrency_limit | 409 | Wait for a crawl to finish, or cancel one. |
| rate_limited | 429 | Wait Retry-After seconds, then retry. |
| internal | 500 | Our fault. Retry; if it persists, get in touch. |
09A complete example
Start a crawl, wait for it, and print every broken link. Paste it, set your key, and run it. Needs curl and jq, both of which are a package manager away on any platform.
#!/usr/bin/env bash
set -euo pipefail
API="https://deadlinkcrawler.com/api/v1"
KEY="${DLC_API_KEY:?Set DLC_API_KEY to your API key}"
TARGET="${1:?Usage: ./crawl.sh https://example.com}"
# Every call goes through here, so a revoked key, a rate limit or an outage
# stops the script with the API's own message instead of looping on an empty
# result.
api() {
local method=$1 path=$2 body=${3:-}
local args=(-sS -X "$method" "$API$path" -H "Authorization: Bearer $KEY")
if [ -n "$body" ]; then args+=(-H "Content-Type: application/json" -d "$body"); fi
local out code
out=$(curl "${args[@]}" -w '\n%{http_code}')
code=${out##*$'\n'}
out=${out%$'\n'*}
if [ "$code" -ge 300 ]; then
echo "$method $path failed ($code): $(jq -r '.error.message // "unexpected error"' <<<"$out")" >&2
exit 1
fi
printf '%s' "$out"
}
# 1. Start it. The response is immediate; the crawl is not.
crawl=$(api POST /crawls "{\"url\": \"$TARGET\"}")
id=$(jq -r '.id' <<<"$crawl")
echo "crawl $id started"
# 2. Poll until it reaches a terminal status.
while :; do
crawl=$(api GET "/crawls/$id")
status=$(jq -r '.status' <<<"$crawl")
pages=$(jq -r '.stats.pages_scanned' <<<"$crawl")
broken=$(jq -r '.stats.broken_count' <<<"$crawl")
echo " $status: $pages pages, $broken broken"
case "$status" in
complete|failed|cancelled) break ;;
esac
sleep 5
done
reason=$(jq -r '.stopped_reason.message // empty' <<<"$crawl")
if [ -n "$reason" ]; then echo "note: $reason"; fi
if [ "$status" != "complete" ]; then echo "crawl $status"; exit 1; fi
# 3. Page through the links. Stop when the server says there are no more.
cursor=""
while :; do
path="/crawls/$id/links?limit=100&type=broken"
if [ -n "$cursor" ]; then path="$path&cursor=$cursor"; fi
page=$(api GET "$path")
jq -r '.links[] | "\(.status_code) \(.url) (on \(.found_on_page))"' <<<"$page"
cursor=$(jq -r '.next_cursor // empty' <<<"$page")
[ -z "$cursor" ] && break
doneexport DLC_API_KEY="dlc_live_…"
./crawl.sh https://example.com10Stability
This is v1, and the version is in the path because the shape is a promise. Fields may be added, and existing fields will not be renamed, removed, or quietly given a new meaning. Anything that cannot be done that way ships as v2, with v1 left working.
Deliberately not in v1: webhooks, managing monitors, and listing past scans. If you need one of them, tell us and say what you are building. A small API is one we can keep promises about.