api_reference

API Reference

Last updated: October 4, 2026See also: Pricing

01Before you start

The API lets your own software start crawls and read results, with no browser involved. It is available on Pro+.

Create a key in Settings under API keys. The key is shown once, at creation, and is never recoverable: we store only a hash of it. If you lose it, revoke it and create another.

A key authenticates as you and can start crawls against your account, so treat it like a password. Keep it in an environment variable or a secret manager, never in a repository. If one leaks, revoke it: revoking takes effect on the next request.

02Authentication

Send the key as a bearer token on every request. The base URL is https://deadlinkcrawler.com/api/v1.

Every request
curl https://deadlinkcrawler.com/api/v1/crawls/CRAWL_ID \
  -H "Authorization: Bearer $DLC_API_KEY"

A missing, malformed or unrecognised key returns 401 unauthorized. A valid key on a plan without API access returns 403 plan_limit_exceeded, so the two cases are easy to tell apart: the first means check your key, the second means check your plan.

03Start a crawl

POST /crawls returns immediately with a queued crawl. It never blocks: a large crawl can run for a long time, so you start it here and poll for the result.

Request
curl -X POST https://deadlinkcrawler.com/api/v1/crawls \
  -H "Authorization: Bearer $DLC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'
202 Accepted
{
  "id": "3bfa25d2-c9e8-49b0-baf8-daab0d7213aa",
  "status": "queued",
  "target_url": "https://example.com/",
  "created_at": "2026-10-04T21:05:28.744Z",
  "started_at": null,
  "finished_at": null,
  "scan_id": "cad99d80-708b-4f01-b6ce-bbd69e472b3b",
  "stats": { "pages_scanned": 0, "links_checked": 0, "broken_count": 0 },
  "stopped_reason": null
}

Body

FieldTypeDefaultMeaning
urlstring, requiredThe page to start from. Must be http:// or https://.
check_assetsbooleantrueAlso check images, stylesheets and scripts, not only links.
strict_status_codesbooleanfalseTreat 401, 403 and 405 as broken. Off by default because they usually mean bot-blocking rather than a dead link.
single_pagebooleanfalseCheck the links on the starting page only, without following them.

Your plan sets how many pages one crawl will scan (25,000 on Pro+) and how many crawls can run at once (4). Starting one too many returns 409 concurrency_limit rather than queueing it, so your script decides whether to wait or to cancel something.

04Check on a crawl

GET /crawls/{id} returns the same shape, updated. Poll it until status is terminal.

Request
curl https://deadlinkcrawler.com/api/v1/crawls/$CRAWL_ID \
  -H "Authorization: Bearer $DLC_API_KEY"
200 OK
{
  "id": "3bfa25d2-c9e8-49b0-baf8-daab0d7213aa",
  "status": "complete",
  "target_url": "https://example.com/",
  "created_at": "2026-10-04T21:05:28.744Z",
  "started_at": "2026-10-04T21:05:29.421Z",
  "finished_at": "2026-10-04T21:05:30.329Z",
  "scan_id": "cad99d80-708b-4f01-b6ce-bbd69e472b3b",
  "stats": { "pages_scanned": 1, "links_checked": 1, "broken_count": 0 },
  "stopped_reason": null
}
statusTerminalMeaning
queuednoAccepted, not started yet.
runningnoIn progress. stats update as it goes.
completeyesFinished. Results are ready.
failedyesCould not complete. See stopped_reason.
cancelledyesYou cancelled it. Whatever it found is kept.

stopped_reason

Null on an ordinary completion. When set, it says why a crawl ended early. A crawl that reaches your plan’s page limit is still complete, not failed: it scanned everything your plan allows and its results are real, so it carries page_limit_reached rather than an error status.

Stopped at the page limit
"status": "complete",
"stopped_reason": {
  "code": "page_limit_reached",
  "message": "Stopped at the time ceiling after scanning 25,000 of 25,000 pages…"
}

05Read the results

GET /crawls/{id}/links returns the broken and redirected links, a page at a time. You can start reading while the crawl is still running.

Request
curl "https://deadlinkcrawler.com/api/v1/crawls/$CRAWL_ID/links?limit=100" \
  -H "Authorization: Bearer $DLC_API_KEY"
200 OK
{
  "links": [
    {
      "seq": 3719,
      "type": "broken",
      "url": "https://example.com/gone",
      "status_code": 404,
      "error_type": "Not Found",
      "link_text": "Our old guide",
      "found_on_page": "https://example.com/blog",
      "resource_type": "link",
      "redirects_to": null
    }
  ],
  "next_cursor": "3723.0",
  "has_more": true
}
QueryDefaultMeaning
limit100Links per page, up to 500.
cursor(start)Pass back the next_cursor from the previous page.
typeallFilter to broken or redirected.

Paging

Keep calling with the next_cursor you were given until has_more is false. Treat the cursor as opaque: it encodes a position in more than one table and its format may change. An unrecognised cursor starts from the beginning rather than failing, so a replayed cursor returns data instead of an error.

06Cancel a crawl

DELETE /crawls/{id} stops a crawl in progress and returns its final state. Anything it found before stopping is kept.

Request
curl -X DELETE https://deadlinkcrawler.com/api/v1/crawls/$CRAWL_ID \
  -H "Authorization: Bearer $DLC_API_KEY"

Cancelling is idempotent, and a crawl that already finished is left alone rather than treated as an error: by the time your cancel lands the crawl may have completed on its own, which is a race you cannot avoid. Check the status in the response to see which happened.

07Rate limits

Requests are limited per key. Every response carries your current budget, so you never have to guess.

HeaderMeaning
X-RateLimit-LimitRequests allowed per window.
X-RateLimit-RemainingRequests left in the current window.
X-RateLimit-ResetWhen the window resets, as a Unix timestamp in seconds.
Retry-AfterSeconds to wait. Only on a 429.
429 Too Many Requests
retry-after: 20
x-ratelimit-limit: 60
x-ratelimit-remaining: 0
x-ratelimit-reset: 1791148140

{ "error": { "code": "rate_limited", "message": "Rate limit exceeded. Retry in 20 seconds." } }

When polling, wait a few seconds between checks rather than looping as fast as you can. A crawl of any size takes minutes, so polling every 5 to 10 seconds costs you nothing and keeps you well inside the limit.

08Errors

Every error has the same shape, so one handler covers all of them.

{ "error": { "code": "not_found", "message": "No crawl with that id." } }

Branch on code, never on the message. The codes below are stable; the messages may be reworded.

CodeHTTPWhat to do
unauthorized401Check the key and the Authorization header.
plan_limit_exceeded403The key is fine; the plan does not include the API.
forbidden403The key may not do this.
not_found404No crawl with that id on your account.
invalid_request400Fix the body or query parameters; the message says which.
concurrency_limit409Wait for a crawl to finish, or cancel one.
rate_limited429Wait Retry-After seconds, then retry.
internal500Our fault. Retry; if it persists, get in touch.

09A complete example

Start a crawl, wait for it, and print every broken link. Paste it, set your key, and run it. Needs curl and jq, both of which are a package manager away on any platform.

crawl.sh
#!/usr/bin/env bash
set -euo pipefail

API="https://deadlinkcrawler.com/api/v1"
KEY="${DLC_API_KEY:?Set DLC_API_KEY to your API key}"
TARGET="${1:?Usage: ./crawl.sh https://example.com}"
# Every call goes through here, so a revoked key, a rate limit or an outage
# stops the script with the API's own message instead of looping on an empty
# result.
api() {
  local method=$1 path=$2 body=${3:-}
  local args=(-sS -X "$method" "$API$path" -H "Authorization: Bearer $KEY")
  if [ -n "$body" ]; then args+=(-H "Content-Type: application/json" -d "$body"); fi

  local out code
  out=$(curl "${args[@]}" -w '\n%{http_code}')
  code=${out##*$'\n'}
  out=${out%$'\n'*}

  if [ "$code" -ge 300 ]; then
    echo "$method $path failed ($code): $(jq -r '.error.message // "unexpected error"' <<<"$out")" >&2
    exit 1
  fi
  printf '%s' "$out"
}

# 1. Start it. The response is immediate; the crawl is not.
crawl=$(api POST /crawls "{\"url\": \"$TARGET\"}")
id=$(jq -r '.id' <<<"$crawl")
echo "crawl $id started"

# 2. Poll until it reaches a terminal status.
while :; do
  crawl=$(api GET "/crawls/$id")
  status=$(jq -r '.status' <<<"$crawl")
  pages=$(jq -r '.stats.pages_scanned' <<<"$crawl")
  broken=$(jq -r '.stats.broken_count' <<<"$crawl")
  echo "  $status: $pages pages, $broken broken"
  case "$status" in
    complete|failed|cancelled) break ;;
  esac
  sleep 5
done

reason=$(jq -r '.stopped_reason.message // empty' <<<"$crawl")
if [ -n "$reason" ]; then echo "note: $reason"; fi
if [ "$status" != "complete" ]; then echo "crawl $status"; exit 1; fi

# 3. Page through the links. Stop when the server says there are no more.
cursor=""
while :; do
  path="/crawls/$id/links?limit=100&type=broken"
  if [ -n "$cursor" ]; then path="$path&cursor=$cursor"; fi
  page=$(api GET "$path")
  jq -r '.links[] | "\(.status_code)  \(.url)  (on \(.found_on_page))"' <<<"$page"
  cursor=$(jq -r '.next_cursor // empty' <<<"$page")
  [ -z "$cursor" ] && break
done
Run it
export DLC_API_KEY="dlc_live_…"
./crawl.sh https://example.com

10Stability

This is v1, and the version is in the path because the shape is a promise. Fields may be added, and existing fields will not be renamed, removed, or quietly given a new meaning. Anything that cannot be done that way ships as v2, with v1 left working.

Deliberately not in v1: webhooks, managing monitors, and listing past scans. If you need one of them, tell us and say what you are building. A small API is one we can keep promises about.