Skip to main content

Help us keep the Simkl API at 100% uptime — optimise your app’s requests

The limits below are generous and most apps never hit them. The ones that do are almost always parallelizing calls to endpoints that should be sequential, or polling /sync/all-items without gating behind /sync/activities. The patterns on this page keep your app fast and keep the API healthy for every other developer building on Simkl.

Request limits

Two limits apply at the same time, and you have to stay under both. One caps how fast you send requests; the other caps how many you send in a day.

Per-second rate

10 GET per second and 1 POST per second, counted against the same bucket as your daily quota: per user for AUTH V2. Going over on a GET returns 429 — back off and retry. Going over on a POST triggers a temporary throttling block on the offending token or client_id, and repeated overages extend the block. The Cloudflare-cached endpoints are exempt, because they are served from the edge and never reach Simkl’s origin. That is exactly why parallel requests are allowed on those endpoints and nowhere else.
One POST per second is less restrictive than it sounds, because every write endpoint accepts arrays. Fifty items sent as one batched POST is a single request against the rate; the same fifty sent individually is fifty requests and will throttle you. See Batch writes below.

Daily quota

Which model applies depends on which auth version your app uses. These are two different systems — read the one that matches your app, not both.

AUTH V2 — per user

The allowance follows the user’s plan, not yours, and is counted per user account. Because V2 requires a user token on every user-scoped request, this is the only limit a V2 app deals with — there is no app-wide quota to manage or exhaust. Your capacity grows with your user base rather than being a fixed ceiling you share between everyone.
The allowance belongs to the user, not to your app. Counting is against the Simkl account, so it is shared across every app that user has connected — yours, plus any other integration they use. Two grants for your own app share it too, so authorizing repeatedly buys nothing.In practice this rarely bites, because the allowance is generous and most users run one or two integrations. But it does mean a user running a busy media-server sync alongside your app has less headroom for you than a user who only runs yours. The way to get headroom is fewer, better-batched requests — not more tokens.

What the user can see

You cannot see another app’s usage, but the user can see all of it. Connected Apps shows them, in one screen:
  • Today’s requests against their allowance, what is left, and how close they are — the V2 allowance, so V1 traffic is not in this number
  • A breakdown by app — every integration they have connected, V1 and V2 alike, sortable by today, yesterday or all-time
  • Which auth version each app uses, labelled AUTH V1 or AUTH V2
  • What each app is permitted to do, in plain words: “Can read and update your media library” or “Read only”
  • Which connections are expiring soon
The two halves answer different questions, which is worth knowing before you send someone there: the allowance is what can actually refuse a request, and only V2 apps draw on it. The per-app breakdown is every app’s traffic regardless of version, so a V1 integration can top that list while contributing nothing to the limit. That page is still the fastest route through “your app stopped syncing”. Ask the user to open it and sort by today’s requests: if another V2 app is eating the allowance, it is visible immediately, and the page already prompts them to take those numbers to the relevant developer. It is worth linking directly from your own error state — see showing the reset time.

AUTH V1 — per client_id

A V1 app has one daily allowance covering everything it does, whether or not a request carries a user token. Every user of your app draws on the same pool, so the busiest user can consume the budget the rest of them needed. It also means a V1 client_id is worth something to a stranger: it works on its own across the whole API, so anyone who finds yours can spend your allowance with it. A V2 client_id is not worth stealing — on its own it reaches only public catalog data, which costs nobody any quota.
This is the strongest practical reason to migrate. Under V1 your ceiling is fixed no matter how many users you have; under V2 each user brings their own allowance. An app that keeps growing on V1 eventually runs into a wall that V2 does not have.V1 is being retired around April 2027 in any case — see Migrating from V1 to V2. If your V1 app is hitting its ceiling before you can migrate, talk to us rather than rationing.
Treat the daily number as approximate, under either model. Simkl runs in more than one region and each counts independently, so a client spread across regions can get somewhat further than the published number before being cut off. Build against the published limit rather than the slack — the slack is an implementation detail and will tighten.

Headers

Read X-RateLimit-Remaining rather than counting requests yourself — it accounts for the exemptions below, which your own counter will not.
The two X-RateLimit-* headers are sent for AUTH V2 bearer tokens only. A V1 token has no per-user allowance to report, so it gets neither header — if you are testing with a V1 token and seeing nothing, that is why, and it is not a bug. You can see them on demand with ?debug_limit=user.V1 requests do not consume the user’s V2 allowance. The per-user daily counter tracks AUTH V2 traffic only, so a user running a busy V1 integration alongside your V2 app is not quietly spending your headroom. V1 requests are still counted per app — they appear in your Analytics and in the user’s per-app breakdown — they simply do not draw down the V2 budget.
The very first request with a brand-new access token carries no X-RateLimit-* headers. They appear from the second request onward.This catches people out because it happens at exactly the wrong moment: you finish authorizing, make your first call, and the headers you just read about are missing. Nothing is wrong. A newly minted token is not in the edge cache yet, so that one request is served behind it and never reaches the counter; the token is cached as a side effect, and every later request has both headers.Do not treat their absence as an error. Read them when present and fall back to your own pacing when they are not — which is also the right shape for the exempt endpoints, since those never carry them either.Verified on production, 2026-09-18:
Every quota resets at midnight US Eastern — not at a rolling 24 hours from your first call, and not at your own local midnight. The same instant for every user, every app, every day.That is the whole reset rule, so you already know when your window ends without asking. Schedule a deferred sync against the next New York midnight and you will be right every time.

Showing the reset time in your UI

Until you are actually blocked, no response tells you when the quota resets — so compute it. Three states, and only the last one carries a reset: The middle row is the one that catches people: Remaining: 0 arrives on a successful response, with no Retry-After. The counter is compared after it increments, so the request that takes you to zero still goes through — it is the next one that fails. If your UI wants to say “resets in 6h” at the moment the user hits zero, you are computing it yourself. Everything you need is in hand: X-RateLimit-Remaining gives the count, and the reset is always the next midnight in New York. That is enough to show the user something useful before they run out:
120 requests left for the next 8 hours.
or, if you prefer the absolute form, “375 of 500 left — resets in 6h 12m”. Either beats a silent failure at request 501. The reset instant, precisely: It is the same instant for every user and every app, worldwide. Nothing about your own server’s timezone changes it.
Do not hardcode either UTC offset. New York shifts between UTC-5 and UTC-4 twice a year, so a fixed +5h is an hour wrong for roughly eight months of the year — and the error is silent. Use a timezone-aware date library and name the zone America/New_York; let it resolve the offset for the date in question.
Show it as a relative time (“resets in 6 hours”) rather than an absolute clock time. The reset is in New York’s timezone, so “resets at midnight” is wrong for most of your users and “resets at 05:00” invites them to wonder why.
Retry-After is a duration, not a timestamp. It is the number of seconds from now until the reset — Retry-After: 45000 means “wait 45,000 seconds”, not “reset at epoch 45000”. Feeding it to code that expects a Unix timestamp will schedule your retry for 1970.It also appears only on a 429, so it is for recovering, not for planning. To plan, use the midnight rule above.
The quota headers are only sent for AUTH V2 tokens. AUTH V1 requests are not subject to the per-user allowance and do not draw it down, so they carry no X-RateLimit-* headers at all — an empty header is not a quota of zero. Anonymous and catalog requests do not carry them either.Treat “no header” as “not applicable” rather than as a number, or a V1 app will read a missing header as 0 remaining and throttle itself for no reason.

What does not count

Detail lookups for movies, TV and anime — GET /movies/{id}, GET /tv/{id}, GET /anime/{id}, and the episode-list endpoints — do not count against your daily quota and are not subject to the per-second rate, along with /redirect and the trending and calendar data files. They are served from the Cloudflare edge and never reach Simkl’s origin, so they are free on both limits. This is why X-RateLimit-Remaining is the number to trust: a client that increments a local counter on every call will think it is far closer to the limit than it is.

Limit signals

Read the body before you decide how long to wait, because the three 429s want completely different responses. rate_limit is a burst you can retry in a second. user_limit_exceeded means one user was noisy and is done for the day. app_limit_exceeded means your whole app is out of budget and every user is affected at once.
400 with {"error": "RATE_LIMIT"} is not a rate limit. Despite the name, it means “your previous write for this user is still running”. It is a 20-second lock held per user on /scrobble and POST /sync/history so that two writes for the same user cannot interleave.Do not treat it as a quota error and back off for minutes. Serialise your writes for that user and retry shortly — the lock clears as soon as the in-flight write finishes. Backing off exponentially on this one makes your app slower for no reason, and retrying in parallel will just hit it again.It is a different thing from 429 app_limit_exceeded in every respect: different status, different cause, different fix. They share a name only because of an old internal label.

Trigger a 429 on demand

Quota errors are hard to test precisely because the honest way to reach one is to spend a real day’s allowance. Add ?debug_limit=user or ?debug_limit=app to any request and you get that refusal immediately:
The response is byte-for-byte what a genuine quota hit returns — same body, same status, same headers — because both come from the same code path. So it exercises your back-off logic, your retry timer and your “you’ve hit today’s limit” screen for real.
It costs nothing and affects nobody. The parameter turns your own request into an error and stops there: it is answered before any counter is touched, so it spends none of your quota and none of the user’s, writes no cache, and never reaches another caller. Leaving it in a request by accident wastes the call, nothing more.There is no flag to switch it on — it works on any request, on production, today.
Remove it before you ship. It is a testing aid, not a feature: an app that sends it in normal traffic simply fails every request it is attached to.

Parallel requests — when allowed

Do not parallelize requests unless necessary. A single sequential client is the default — it stays well under the limits and never triggers a throttle. Parallel requests are explicitly allowed on endpoints cached at the Cloudflare edge, because parallel hits there are served from cache and don’t pressure Simkl’s origin: Everything else stays sequential. Sync endpoints, user-state endpoints, and search endpoints serve per-user or dynamic data that bypasses the edge cache — every parallel call hits the origin directly and counts individually against the cap.

Why parallel on uncached endpoints hurts

It doesn’t increase throughput. The ceiling is 10 GET/sec, 1 POST/sec. Ten parallel GETs at t=0 all land in the same one-second window — you have spent the whole per-second budget at once, and request 11 gets 429. The same ten calls spread across that second land you in exactly the same place with zero 429s, and cost the identical daily quota either way. Parallelism buys you nothing here and costs you the errors. Servers are commodity hardware, not supercomputers. Imagine opening 10 copies of Photoshop on your laptop at the same time — each one loads the full app into RAM independently, the fans spin up, and everything else slows to a crawl. A web server handling 10 parallel uncached requests does the same thing: each request loads its own slice of the application — framework boot, autoloader, config, model code, request context — and holds it until the response is sent. Ten parallel requests means 10 full copies of that boot state in RAM at once, plus 10 database connections, 10 worker threads, 10 response buffers, and 10 fan-out calls to downstream services (read replicas, search index, image proxy). Connection pools, CPU schedulers, GC pauses, and disk I/O all have finite headroom; when a burst exceeds capacity, in-flight requests slow down or fail and the next caller waits for the queue to drain. Sequential traffic reuses the same worker over and over — boot cost paid once per worker, not once per request — so no single resource ever spikes above safe levels. A 429 has user-visible cost. A throttled call breaks the user’s session — the page stalls, the sync stalls, a mid-flight write fails. Your client then has to back off and retry, doubling the time-to-success. Retry storms also re-pressurize the origin and can extend the throttle window. “Just parallelize and let the API tell us to slow down” is not a viable strategy; the failed requests have already cost user-visible time. Sustained overage suspends your client_id. Per the /sync/all-items warning: apps that hammer the API — particularly write endpoints or polling without /sync/activities gating — get their client_id suspended. No warning, no appeal. Burst patterns look like abuse. Edge and origin monitoring can’t tell a well-meaning client firing 50 parallel sync calls apart from a botnet or scraper — the signature is the same: one client_id or IP sending many concurrent requests to non-cacheable URLs in a short window. Automated mitigation kicks in (temporary block, IP throttle, longer back-off) without anyone making a judgment call about intent. The correct pattern for uncached endpoints: one in-flight request at a time per token; batch into one call instead of N parallel calls whenever the endpoint accepts arrays (every write endpoint does — see Batch writes below); gate polling behind /sync/activities so most /sync/all-items calls never happen.
Cache invalidation on the cached endpoints is automatic. When Simkl updates the underlying record (admin edits, metadata refresh, image swap, new episode airing, etc.), the corresponding Cloudflare cache entry is purged server-side. The next call returns fresh data — no TTL to wait out, no ?nocache=… trick needed. Same applies to the trending and calendar JSON files. Your own app-level cache, if you have one, still has to be invalidated by your client.

Best practices

Most metadata never changes. Cache responses on the device for the longest sensible time — minutes for user data, hours for catalog data, indefinitely for images.Cache catalog data by URL. Cache user data by URL and account. Anything behind a user token — the library, ratings, playback, custom lists — can return completely different content, or different permissions, for the same URL depending on who is signed in. A cache keyed on URL alone will happily serve one user’s private list to the next person who signs in on that device.Include the account identity in the cache key, and drop those entries on sign-out or account switch. Catalog responses are the same for everyone and need none of this.
Always check /sync/activities before pulling watchlists. Skip watchlists whose timestamp didn’t change. See the Sync guide for the two-phase model.
Every write endpoint accepts arrays. POST /sync/history, POST /sync/history/remove, POST /sync/ratings, and POST /sync/add-to-list will happily process 50 items in one request.Sending 50 separate single-item POSTs immediately blows past the 1-POST-per-second cap and gets your token or client_id throttled — sending one batched array stays well under. It also costs one request of daily quota instead of 50, and avoids the 50 calls queueing behind each other on the short per-user write lock.One batched array is faster, cheaper, and the only version that does not throttle.
On 500, 502, 503: wait, retry, double the wait, cap it. Give up after 5 attempts. A common starting pattern is 1s → 2s → 4s → 8s → 16s, but tune to your application’s tolerance.On a 429, read the body before choosing a wait — the two kinds are hours apart:
  • A bare 429 is the per-second rate. It clears in about a second, so pause briefly and carry on. Exponential backoff here is overkill and makes your app feel broken.
  • user_limit_exceeded or app_limit_exceeded is the daily quota. Five retries over thirty seconds cannot possibly clear it — you will burn the attempts and fail anyway. Read Retry-After, wait for the window, or stop and surface the problem to the user.
And do not back off on 400 RATE_LIMIT at all: that is the short per-user write lock, which clears as soon as the in-flight write finishes. Serialise your writes for that user and retry shortly.Four failures, four different waits: seconds for a server error, about a second for a burst, moments for the write lock, until reset for the daily quota.
If you have an IMDb / TMDB / MAL / etc. ID and need the matching Simkl ID, use GET /redirect — it returns a tiny 301 redirect with the Simkl ID in the Location header (no JSON body), and the follow-up GET /movies/{id} / GET /tv/{id} / GET /anime/{id} call is Cloudflare-cached. Much cheaper than GET /search/id.

Hide your client_id in CI

If you’re using a CI system, store client_id and client_secret in your provider’s encrypted secrets / environment variables — not in repository files. A leaked credential burns your quota and is hard to recover.