Reaching a limit (403)
A403 with code: NOT_ALLOWED currently always signals that an agency limit was reached, not a missing permission. There are three causes, and the properties.details field of the error states which one applies:
- Optimization quota. Once an agency’s daily or monthly optimization quota is reached, further requests fail.
- Plan size limit. A plan that exceeds one of the per-plan limits in the preceding table (resources, orders, stops, or stops per order without
successiveStops) is rejected. - Insufficient credits. In production, a request that would consume more credits than your balance has left is rejected. See Pricing and credits.
EnvelopedErrors envelope as any other business error. See the Error schema.
A plan size limit returns one INVALID_VALUE error per exceeded limit, each naming the field in properties.path (for example resources or orders.stops) and the exceeded limit in properties.details, followed by a single NOT_ALLOWED stating that the plan violates the agency’s quota configuration. Read every error in the response, not only the first:
Temporary overload (429)
A429 response signals a temporary overload rather than a quota: no server was available to take the request. It comes with an empty body and no Retry-After header. Retry after a pause, with an increasing delay between attempts, and avoid sending large bursts of simultaneous requests.
Max computation time
You control how long a plan can optimize with the plan-levelmaxOptimizationDuration field (see How the optimization engine works), within your agency’s maximum: 5 seconds for the sandbox, 6 hours for production by default. A longer value is clamped to that maximum, and the response to the POST or PUT carries a warning saying so — warnings are only returned in that response, never by a later GET, so read them there. If the field is omitted, the default applies: 5 seconds for the sandbox, 15 minutes for production. The engine stops earlier on its own once it stops finding improvements. Contact customer.success@kardinal.ai if your production agency needs a longer maximum.
The one indirect limit on computation time is throughput, not duration: if your agency already has its maximum number of simultaneous running plans in progress (10 by default), a new plan waits in the waiting room (status.waitingRoom) until a slot frees up, before its own maxOptimizationDuration clock effectively starts mattering. Contact customer.success@kardinal.ai if your account needs a higher threshold.
Payload size in practice
There’s no explicit byte-size limit on a submitted plan, but it’s bounded in practice by the ~16 MB storage limit of the document it’s stored as internally — an order of magnitude, not an exact JSON byte cap. The per-plan object counts in the preceding table (orders, stops, resources) are the limits to design against.SLA
Kardinal targets 99.9% uptime as an internal objective, but this is not currently a contractual uptime commitment — no formal availability guarantee is offered for API access today. Support response times, on the other hand, are defined by severity:
Report an issue to customer.success@kardinal.ai and state its severity to get the applicable response time.
See also
- How the optimization engine works — the quality-vs-time trade-off and the
statuslifecycle (waitingRoom,creation,optimization,waitingTraffic). - Handling large volumes — practical guidance for sizing
maxOptimizationDurationand paginating large result sets. - Pricing and credits — credit consumption, counted separately from the optimization quota, although insufficient credits return the same
403.

