Live scrape
SensitiveLive backend health and capacity
/data-api/v2/scrape/status- Scope
- scrape:live
- Freshness
- catalog read
- Group
- Live scrape
- Platforms
- any
What this endpoint answers
Whether the live-scrape backends are up, how much capacity they have, and what the batch limits are. Poll this before a large job and back off on `degraded` rather than discovering it through 502s.
This is a catalog read: it is served from the CRM Solid database in milliseconds, costs one budget unit, and reports how old the reading is in meta.cache_age_s. It needs the scrape:live scope (Live scrape): Read from the platform itself rather than from the catalog: a profile or channel on any of the seven platforms, and on X the whole content layer - timelines, posts, replies, quotes, threads, followers, search, lists and communities.
Good to know
- This endpoint requires scrape:live but does NOT itself hit a scraper - it is two health probes and is not metered as a live call.
- state: degraded means at least one backend is reachable but not serving; expect 502s from that platform.
- error_rate is cumulative since the lookup service last restarted, not a rolling window. A step change matters more than the absolute value.
- accounts is the size of the authenticated X pool. It is small by design - treat it as the real concurrency ceiling for everything you run.
- budgets is the per-source ceiling that the account live check draws against, shared across every caller of this instance. It is set far below what each source would tolerate because two of them - the Instagram proxy pool and the X scraper pool - are the same pools our own crawlers run on. `refused` climbing is the signal to slow down.
- Only the backends with a health probe are listed. A working Instagram lookup is confirmed by a real call to /scrape/instagram/user, not by this endpoint - reporting a health state we cannot actually observe would be worse than leaving it out.
Parameters
This endpoint takes no path or query parameters.
Call it
Authenticate with a bearer token or the x-api-key header. Keys are server-to-server credentials. Never embed one in front-end code - call the API from your own backend and forward the result.
curl "https://crmsolid.com/data-api/v2/scrape/status" \
-H "Authorization: Bearer psk_live_..."
const res = await fetch("https://crmsolid.com/data-api/v2/scrape/status", {
headers: {
Authorization: `Bearer ${process.env.CRM_SOLID_DATA_API_KEY}`,
},
});
if (!res.ok) {
const { error } = await res.json();
throw new Error(`${error.code}: ${error.message} (${error.request_id})`);
}
const { data, meta } = await res.json();
import os
import requests
res = requests.get(
"https://crmsolid.com/data-api/v2/scrape/status",
headers={"Authorization": f"Bearer {os.environ['CRM_SOLID_DATA_API_KEY']}"},
timeout=30,
)
res.raise_for_status()
payload = res.json()
data, meta = payload["data"], payload["meta"]
Keys look like psk_live_... for production keys, psk_test_... for test keys and are minted in the panel.
What comes back
Success is { data, meta }. Failure is { error: { code, message, request_id } }. The body below is the spec's own example: the values in it are illustrative readings, not live numbers.
- State
- operational
- Backends
- 2 items
- Budgets
- 2 items
- Limits batch max handles
- 25
- Limits batch concurrency
- 5
- Limits x cache ttl s
- 60
- Limits x pool concurrency
- 4
{
"data": {
"state": "operational",
"backends": [
{
"name": "xlookup",
"state": "up",
"detail": "Serving from 4 authenticated account(s).",
"metrics": {
"ready": true,
"accounts": 4,
"uptime_s": 82140,
"served": 19442,
"errors": 118,
"error_rate": 0.006,
"cached_profiles": 812
}
},
{
"name": "telegram_bot",
"state": "up",
"detail": "Bot @playersells_bot is answering.",
"metrics": {
"bot": "playersells_bot"
}
}
],
"budgets": [
{
"platform": "instagram",
"per_minute": 10,
"concurrency": 2,
"available": 8,
"in_flight": 1,
"granted": 1204,
"refused": 17
},
{
"platform": "x",
"per_minute": 20,
"concurrency": 2,
"available": 20,
"in_flight": 0,
"granted": 4881,
"refused": 3
}
],
"limits": {
"batch_max_handles": 25,
"batch_concurrency": 5,
"x_cache_ttl_s": 60,
"x_pool_concurrency": 4
}
},
"meta": {
"request_id": "req_9f2c41a8b3d5",
"generated_at": "2026-08-23T09:14:02.317Z",
"took_ms": 42
}
}
The meta block
request_idstringrequiredUnique id for this request. Quote it in a support ticket.
generated_atstringrequiredServer time the response was produced.
took_msintegerrequiredMilliseconds spent server-side.
pagePageoptionalsourcestringoptionalWhich backend served the payload, for endpoints with more than one.
cache_age_sintegeroptionalAge of the underlying data in seconds. 0 for live reads.
When it fails
GET /scrape/status documents 7 failure statuses. Branch on error.code, which is stable and enumerated; message is prose and may change.
- 401Unauthorized
unauthorized | invalid_keyNo key was presented, or the key is unknown, revoked or expired.
- 402PaymentRequired
payment_required | subscription_inactiveThe key is valid but the plan behind it cannot serve the call: the included requests are spent and overage is switched off, capped or unfunded (payment_required), or the billing period lapsed and was not renewed (subscription_inactive). Retrying does not help; paying does. The X-Plan-* headers on this response say how far past the line you are.
Carries X-Plan, X-Plan-Limit, X-Plan-Overage, X-Plan-Period-End, X-Plan-Remaining.
- 403Forbidden
forbidden_scope | forbidden_ipThe key is valid but not allowed to make this call: it lacks the scope, or the request came from an address outside the key's allowlist.
- 422InvalidRequest
invalid_requestA parameter is malformed, out of range or mutually exclusive with another. `details` names the offending fields.
- 429RateLimited
rate_limited | quota_exceededEither the burst ceiling for the current minute or the daily quota is spent. Distinguish with the code: rate_limited clears within the minute, quota_exceeded does not clear until 00:00 UTC.
Carries Retry-After, X-Quota-Limit, X-Quota-Remaining, X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, X-Request-Id.
- 500InternalError
internal_errorSomething failed on our side. Internals are never leaked; quote the request id.
- 503Unavailable
upstream_timeout | upstream_error | not_configuredThe request could not be served right now. BRANCH ON error.code, not on the status: 'upstream_timeout' means a source was too slow (this is what a catalog query hitting its 15-second statement timeout returns, so it is reachable from any endpoint that reads the corpus, not only the live-scrape ones) and the same call is worth retrying with backoff - narrowing it with a smaller limit, a filtered scope or a less popular account makes it far less likely; 'upstream_error' means a source was unreachable, so back off further; 'not_configured' means the capability has no backing service in this deployment, and retrying will never help.
Every response carries X-Request-Id and meta.request_id. Quote it in support requests.
Coverage and limits
This endpoint is not platform-specific. Rate limits come from the tier on your key.
| Tier | Requests a minute | Requests a day | Live reads a minute |
|---|---|---|---|
| free | 30 | 1,000 | 5 |
| standard | 120 | 25,000 | 20 |
| pro | 600 | 250,000 | 60 |
| unlimited | 6,000 | 10,000,000 | 600 |
This call only draws on the ordinary per-minute and per-day columns. Every response reports where you stand in X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset and the X-Plan headers.
Start calling it
A key takes a minute to mint in the panel, no card. The reference covers authentication, the envelope, scopes, rate limits and every error code in one page.
Reference path: /data-api/reference/scrape-status