APIRate limits & quotas

Rate limits & quotas

Per-plan request caps, the headers that report them, and the monthly word quota.

Two independent limits apply to API use.

LimitCountsEnforced as
Rate limitRequests, per plan, in three rolling windows429 before the endpoint runs
Word quotaWords generated, per month, per account422 on article creation

Subscription gate

Every /v1 route is behind a subscription check that runs first. An account with no plan gets:

{
  "error": "No active subscription",
  "message": "This account has no active subscription. Subscribe to a plan to access the API.",
  "code": "NO_ACTIVE_SUBSCRIPTION"
}

with status 402. No rate limit is consumed and nothing is created.

Rate limits

Requests are counted per account against three windows at once — minute, hour, and day — and against the endpoint family being called. Whichever window fills first rejects the request.

Endpoint familyMatches
articlesAny path containing articles
workflowsAny path containing workflows
generalEverything else

Caps by plan

PlanWindowGeneralArticlesWorkflows
StarterMinute301020
StarterHour30060200
StarterDay2,0002001,000
ProfessionalMinute602040
ProfessionalHour1,000200600
ProfessionalDay8,0001,0004,000
TeamMinute1204080
TeamHour3,0006001,500
TeamDay20,0003,00010,000
AgencyMinute24080160
AgencyHour8,0001,5004,000
AgencyDay50,0008,00025,000

Headers

Every successful response reports the minute window:

HeaderMeaning
X-RateLimit-LimitRequests allowed in the current minute
X-RateLimit-RemainingRequests left in it
X-RateLimit-ResetUnix timestamp when it resets

A rejection adds Retry-After (seconds) and returns 429:

{
  "error": "Rate limit exceeded",
  "message": "You have exceeded the hour rate limit for this endpoint.",
  "code": "RATE_LIMIT_EXCEEDED",
  "retry_after": 1284
}

The message names which window was hit, and retry_after is that window's reset — so an hourly rejection can ask you to wait far longer than a minute. Honour the value rather than retrying on a fixed interval.

Windows are fixed buckets aligned to the clock, not sliding. The minute counter resets at the top of each minute, the hour counter on the hour. X-RateLimit-Reset tells you exactly when.

Word quota

Article creation also spends the account's monthly word allowance, which is set by the plan.

Serpon reserves 120% of target_word_count at creation time, because generation routinely overshoots the target before the revision steps trim it. A 2,000-word article needs 2,400 words of remaining quota to start.

If the reservation does not fit, creation fails with 422 and nothing is queued:

{
  "message": "Insufficient word quota. You have 500 words remaining this month. This article would require approximately 2400 words. Please reduce the target word count or upgrade your plan.",
  "errors": {
    "target_word_count": [
      "Insufficient word quota. You have 500 words remaining this month. This article would require approximately 2400 words. Please reduce the target word count or upgrade your plan."
    ]
  }
}

The message states remaining and required words, so it can be surfaced to a user or logged as-is.

Designing around the limits

  • Read the headers, not the clock. Plans change; X-RateLimit-Remaining is always current.
  • Back off on 429 using Retry-After. Retrying sooner just consumes another rejection.
  • Do not poll hard. Article generation takes minutes, and every poll spends the articles budget. Use a webhook instead, or poll every 30–60 seconds.
  • Fail fast on 402. It will not resolve on retry — the account needs a plan.
  • Size batches against the daily cap, not the minute cap. On Starter, 200 articles a day is the ceiling no matter how the requests are spaced.