Rate limits and quotas
Constaia API limits per plan (requests per second, concurrent analyses and pages per minute), RateLimit headers, 429 responses with Retry-After, the monthly spend cap, and how to retry with backoff.
Constaia limits how many requests each API key can make and how much work each account can have in flight, to protect the service and share it fairly. This page explains the limits of each plan, the monthly spend cap and how to design your integration so you don't run into them.
Limits per plan
| Tier | Requests/s | Concurrent analyses | Pages/min | When |
|---|---|---|---|---|
| Free | 2 | 2 | 60 | No purchases |
| Paid | 10 | 10 | 600 | Any pack purchased |
| Enterprise | Custom | Custom | Custom | Contract |
- The tier is computed automatically: your account moves to Paid as soon as you buy a pack. Enterprise and accounts
with raised limits have their own values. Your key's actual value is always in
RateLimit-Limit. - Requests per second, per API key: each key has its own counter (1 s sliding window), and any authenticated request
to
/v1counts: analyze, classify, list, fetch an analysis, etc. If you exceed it:429 rate_limited. - Concurrent analyses, per account: how many synchronous live-mode analyses the account can have running at
once. If you exceed it:
429 concurrency_limit. Analyses withasync: trueand batches don't count here: they are queued. - Pages per minute, per account: pages analysed in live mode (synchronous and asynchronous, batches included) in a
60 s sliding window. If you exceed it:
429 pages_rate_limitedwithRetry-After. - Test mode:
ck_test_keys have the same requests-per-second limit as your tier, with no credit quota and no page or concurrency limit.
Headers
Every authenticated response includes the limit state:
| Header | Meaning |
|---|---|
RateLimit-Limit | Requests allowed in the window. |
RateLimit-Remaining | Requests left in the current window. |
RateLimit-Reset | Seconds until the window resets. |
RateLimit-Policy | Policy applied, e.g. 10;w=1 (10 requests, 1 s window) on the paid plan or 2;w=1 on the free one. |
Retry-After | Only on 429: seconds you must wait. |
HTTP/1.1 200 OK
Content-Type: application/json
X-Request-Id: req_01J9Z8Q3K4M5N6P7Q8R9S0T1V2
RateLimit-Limit: 10
RateLimit-Remaining: 7
RateLimit-Reset: 1
RateLimit-Policy: 10;w=1The RateLimit-* headers are exposed through CORS, although your key must never be used from a browser.
The SDKs keep the ones from the last response: constaia.lastRateLimit in JavaScript, $constaia->lastRateLimit in
PHP and client.last_rate_limit in Python, with limit, remaining, reset, policy and retryAfter
(retry_after in Python).
429 response
If you exceed the limit you get 429 with Retry-After:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 1
RateLimit-Limit: 10
RateLimit-Remaining: 0
RateLimit-Reset: 1
RateLimit-Policy: 10;w=1
{
"error": {
"type": "rate_limited",
"code": "rate_limited",
"message": "Demasiadas peticiones. Reintenta en un momento.",
"request_id": "req_01J9Z8Q3K4M5N6P7Q8R9S0T1V3"
}
}A request rejected with 429 is not processed and not charged.
The other two limits answer the same way, with their own code and message:
code | Limit | What to do |
|---|---|---|
rate_limited | The key's requests per second. | Wait Retry-After and retry. |
concurrency_limit | The account's concurrent synchronous analyses. | Wait for the running ones to finish, or use async: true or batches. |
pages_rate_limited | The account's pages per minute. | Wait Retry-After (until pages leave the 60 s window). |
Other limits
| Limit | Value | Error when exceeded |
|---|---|---|
| File size | 20 MB | 413 file_too_large |
| PDF pages, synchronous analysis | 30 | 422 too_many_pages |
PDF pages, async: true or batches | 200 | 422 too_many_pages |
| Documents per batch | 1–100 | 400 batch_too_large / 400 empty_batch |
Types in expect | 1–20 | 422 invalid_parameter |
metadata | 20 keys; key ≤ 40 characters; value ≤ 500 | 422 invalid_parameter |
file_url download | 15 s, https, public IP | 422 file_url_unreachable / 400 invalid_file_url |
Idempotency-Key | 255 characters | 400 invalid_idempotency_key |
| Results per page when listing | 1–100 (default 10) | 422 invalid_parameter |
| Synchronous wait | 30 s; after that, 202 and the result by webhook | — |
How to retry
Rules:
- Retry only what is retryable:
429,409 idempotency_in_progress, 5xx (except501) and network errors. See errors. - On
429, wait exactly whatRetry-Aftersays. - Otherwise, use exponential backoff with jitter: 0.5 s, 1 s, 2 s, 4 s… with a cap, plus a random component so your processes don't all retry at once.
- Send an
Idempotency-Keyon every POST and reuse it across retries, so you are never charged twice. See idempotency. - Cap the number of attempts (3–5) and log the
X-Request-Id.
The JavaScript, PHP and Python SDKs already do this: they retry 429, 408, 409 idempotency_in_progress, 5xx
(except 501) and network errors, honour Retry-After (JavaScript and Python wait 60 s at most), use exponential
backoff with jitter and reuse the Idempotency-Key. By default they retry twice (maxRetries in JavaScript,
max_retries in PHP and Python).
With the SDK you only need to set maxRetries:
import { Constaia } from "@constaia/sdk";
export const constaia = new Constaia({ maxRetries: 4, timeout: 90_000 });If you call the API with fetch directly:
const RETRYABLE = new Set([408, 429, 500, 502, 503, 504]);
export async function postWithRetry(url: string, body: FormData | string, headers: Record<string, string>, maxRetries = 4) {
const idempotencyKey = crypto.randomUUID();
for (let attempt = 0; ; attempt++) {
try {
const res = await fetch(url, {
method: "POST",
body,
headers: { ...headers, Authorization: `Bearer ${process.env.CONSTAIA_API_KEY}`, "Idempotency-Key": idempotencyKey },
});
const payload = await res.json();
const inProgress = res.status === 409 && payload?.error?.code === "idempotency_in_progress";
if ((!RETRYABLE.has(res.status) && !inProgress) || attempt >= maxRetries) {
if (!res.ok) throw Object.assign(new Error(payload.error?.message), { status: res.status, error: payload.error });
return payload;
}
const retryAfter = Number(res.headers.get("retry-after"));
await sleep(Number.isFinite(retryAfter) && retryAfter > 0 ? retryAfter * 1000 : backoff(attempt));
} catch (err) {
if ((err as { status?: number }).status || attempt >= maxRetries) throw err;
await sleep(backoff(attempt)); // network error
}
}
}
const backoff = (attempt: number) => Math.min(8000, 500 * 2 ** attempt) * (0.5 + Math.random() / 2);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));Limit concurrency in your client
Retrying well is not enough if you fire hundreds of analyses at once: most will end in 429. Limit how many requests
you have in flight. All three SDKs do it on the client (4 by default); extra requests wait their turn:
| SDK | Option |
|---|---|
| JavaScript | new Constaia({ maxConcurrency: 4 }) |
| PHP | new Client(null, ['max_concurrency' => 4]), used by analyzeMany() to analyse in parallel |
| Python | Constaia(max_concurrency=4) o AsyncConstaia(max_concurrency=4), with threads or asyncio tasks |
The limit is per client (per process). With several workers, the key's per-second limit is shared: keep
concurrency × workers in line with it.
import { readdir } from "node:fs/promises";
import path from "node:path";
import { Constaia } from "@constaia/sdk";
import { fromPath } from "@constaia/sdk/node";
const constaia = new Constaia({ maxConcurrency: 4, maxRetries: 4 }); // at most 4 requests in flight
const dir = process.argv[2] ?? "./docs";
const files = (await readdir(dir)).filter((f) => /\.(jpe?g|png|webp|heic|pdf)$/i.test(f));
const results = await Promise.all(
files.map(async (name) => {
const analysis = await constaia.analyze(await fromPath(path.join(dir, name)), { expect: "es_dni" });
return { name, status: analysis.verdict?.status };
}),
);
console.table(results);Without an SDK, limit concurrency yourself: p-limit in JavaScript or an asyncio.Semaphore in Python.
High volume: use batches
To process hundreds of documents, a batch is more efficient than hundreds of calls: a single
POST /v1/batches request takes up to 100 documents, counts as one request for the per-second limit, takes no
concurrent-analysis slots, answers 202 right away and notifies you with the batch.completed webhook. Its pages do
count towards the pages-per-minute limit: on the free plan (60 pages/min) a batch with more pages is rejected as a whole
with 429 pages_rate_limited, so split it or wait Retry-After. Full guide:
bulk batches.
Controlling spend
Monthly spend cap
You can set a monthly credit cap for the account (monthly_credit_cap). It counts the credits charged in live mode
during the calendar month (UTC) and resets on the 1st:
- If an analysis would exceed the cap, it is rejected with
402 monthly_cap_reachedbefore processing and is not charged. Batches and asynchronous analyses are checked the same way. - The account's owners and admins get an email at 80 % and 100 % of the cap (once per level and month). Changing the cap resets the month's alerts.
- Test mode doesn't count.
{
"error": {
"type": "insufficient_credits",
"code": "monthly_cap_reached",
"message": "Se ha alcanzado el tope de gasto mensual de la cuenta (1000 créditos). Súbelo en el panel o espera al día 1.",
"request_id": "req_01J9Z8Q3K4M5N6P7Q8R9S0T1V4"
}
}Don't retry a 402: raise the cap or wait for the next month. The cap can't be set from the dashboard yet: write to
hola@constaia.com with the value you want.
Also:
- Set the low-balance threshold in the dashboard and subscribe to the
credits.lowwebhook (live mode only, sent once until you top up). See webhooks. - Check
GET /v1/balancebefore large jobs. What you can spend iscredits_spendable(credits_available+free_tier_remaining). - Develop and test with
ck_test_keys: they never charge. See test mode.
More in credits and billing.
Asking for higher limits
If you need more requests per second, concurrent analyses or pages per minute, write to hola@constaia.com with your account, the expected volume (documents per day and peaks) and the use case. Limits can be raised per account. The Enterprise plan includes custom limits.
Next steps
Errors
Constaia API error format, every error code by HTTP status, the JavaScript and PHP SDK error classes and which errors are safe to retry.
Idempotency
Use the Idempotency-Key header to retry analyses, classifications and batches without processing or paying for them twice, with examples in several languages.