This organization's rate limits and concurrent-call quotas
Two budgets that fail in different ways for different reasons. Exceeding the request budget is a 429 rate_limited with a Retry-After — slow down. Exceeding a concurrency pool is a 429 concurrent_call_limit_reached, which slowing down does not fix: you have to wait for calls to end. Live and test have separate budgets, and this answers for whichever key you asked with.
Requires the read scope. A key with less gets 403 insufficient_scope.
Authorization: Bearer tone_live_… or tone_test_…. The prefix IS the environment: a test key reaches only the sandbox, and no request field bridges the two.
In: header
Response Body
application/json
application/json
application/json
application/json
application/json
curl -X GET "https://example.com/v1/limits"{ "data": { "concurrency": { "agentCalls": { "inUse": 0, "limit": 0 }, "byoCalls": { "inUse": 0, "limit": 0 } }, "environment": "live", "requests": { "limitPerMinute": 0, "remaining": 0 } }}Readiness — are dependencies reachable?
Checks Postgres. Answers 503 when a dependency is down, which is the signal to stop sending this instance traffic without restarting it. Cache status rides along in the payload but never fails the probe — it is an optimisation, not a dependency. Queue depths ride along the same way: the API only ever produces, so a broker outage defers work rather than breaking a request. The process that cannot run without a broker is tone-worker, which refuses to boot instead of reporting here.
Create an agent (always as a draft)
A new agent is always a draft, whatever you send: a draft can be dialled outbound but will NOT answer inbound calls, so publishing is the deliberate step that puts it on a phone line. The (ttsModel, ttsVoice) pair is validated here — a voice belonging to a different model version is rejected now rather than failing mid-call.