> ## Documentation Index
> Fetch the complete documentation index at: https://developers.kardinal.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Limits and quotas

> Rate limits, maximum payload size, computation time, and SLA.

These are current platform defaults, not fixed constants — contact [customer.success@kardinal.ai](mailto:customer.success@kardinal.ai) to raise any of them for your account.

| Item | Sandbox | Production |
| - | - | - |
| Rate limit (optimizations) | 1,500 per day, 50,000 per month | 1,500 per day, 50,000 per month |
| Max simultaneous running plans | 10 | 10 |
| Max payload size | 3,000 orders, 3,000 stops, 250 resources per plan | 3,000 orders, 3,000 stops, 250 resources per plan |
| Max computation time (`maxOptimizationDuration`) | 5 seconds (default 5 seconds) | 6 hours (default 15 minutes) |
| Max stops per order without `successiveStops` | 4 | 4 |

## Reaching a limit (403)

A `403` with `code: NOT_ALLOWED` currently always signals that an agency limit was reached, not a missing permission. There are three causes, and the `properties.details` field of the error states which one applies:

* **Optimization quota.** Once an agency's daily or monthly optimization quota is reached, further requests fail.
* **Plan size limit.** A plan that exceeds one of the per-plan limits in the preceding table (resources, orders, stops, or stops per order without `successiveStops`) is rejected.
* **Insufficient credits.** In production, a request that would consume more credits than your balance has left is rejected. See [Pricing and credits](/reference/pricing-and-credits#insufficient-credits).

The error comes in the same `EnvelopedErrors` envelope as any other business error. See the [`Error`](/api-reference/plan/create-a-plan) schema.

A plan size limit returns one `INVALID_VALUE` error per exceeded limit, each naming the field in `properties.path` (for example `resources` or `orders.stops`) and the exceeded limit in `properties.details`, followed by a single `NOT_ALLOWED` stating that the plan violates the agency's quota configuration. Read every error in the response, not only the first:

```json theme={null}
{
  "errors": [
    {
      "code": "INVALID_VALUE",
      "message": "The field value is not valid.",
      "properties": {
        "details": "The plan's number of resources (251) is greater than the allowed value (250) for agency BLD1234567_sandbox",
        "path": "resources"
      }
    },
    {
      "code": "NOT_ALLOWED",
      "message": "The requested action is not allowed.",
      "properties": {
        "details": "The resulting plan violates the quotas configuration of the agency"
      }
    }
  ]
}
```

## Temporary overload (429)

A `429` response signals a temporary overload rather than a quota: no server was available to take the request. It comes with an empty body and no `Retry-After` header. Retry after a pause, with an increasing delay between attempts, and avoid sending large bursts of simultaneous requests.

## Max computation time

You control how long a plan can optimize with the plan-level `maxOptimizationDuration` field (see [How the optimization engine works](/concepts/how-the-optimization-engine-works#the-quality-vs-computation-time-trade-off)), within your agency's maximum: 5 seconds for the sandbox, 6 hours for production by default. A longer value is clamped to that maximum, and the response to the `POST` or `PUT` carries a warning saying so — warnings are only returned in that response, never by a later `GET`, so read them there. If the field is omitted, the default applies: 5 seconds for the sandbox, 15 minutes for production. The engine stops earlier on its own once it stops finding improvements. Contact [customer.success@kardinal.ai](mailto:customer.success@kardinal.ai) if your production agency needs a longer maximum.

The one indirect limit on computation time is throughput, not duration: if your agency already has its maximum number of simultaneous running plans in progress (10 by default), a new plan waits in the **waiting room** (`status.waitingRoom`) until a slot frees up, before its own `maxOptimizationDuration` clock effectively starts mattering. Contact [customer.success@kardinal.ai](mailto:customer.success@kardinal.ai) if your account needs a higher threshold.

## Payload size in practice

There's no explicit byte-size limit on a submitted plan, but it's bounded in practice by the \~16 MB storage limit of the document it's stored as internally — an order of magnitude, not an exact JSON byte cap. The per-plan object counts in the preceding table (orders, stops, resources) are the limits to design against.

## SLA

Kardinal targets 99.9% uptime as an internal objective, but this is not currently a contractual uptime commitment — no formal availability guarantee is offered for API access today.

Support response times, on the other hand, are defined by severity:

| Severity | First response |
| - | - |
| P1 — Critical (service down, or major operational impact) | 1 hour |
| P2 — Major (significant functional degradation) | 4 hours |
| P3 — Minor (limited impact, workaround available) | 1 business day |
| P4 — Low (cosmetic or informational) | 2 business days |

Report an issue to [customer.success@kardinal.ai](mailto:customer.success@kardinal.ai) and state its severity to get the applicable response time.

## See also

* [How the optimization engine works](/concepts/how-the-optimization-engine-works) — the quality-vs-time trade-off and the `status` lifecycle (`waitingRoom`, `creation`, `optimization`, `waitingTraffic`).
* [Handling large volumes](/guides/handling-large-volumes) — practical guidance for sizing `maxOptimizationDuration` and paginating large result sets.
* [Pricing and credits](/reference/pricing-and-credits) — credit consumption, counted separately from the optimization quota, although insufficient credits return the same `403`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.