Help us keep the Simkl API at 100% uptime — optimise your app’s requests
The limits below are generous and most apps never hit them. The ones that do are almost always parallelizing calls to endpoints that should be sequential, or polling/sync/all-items without gating behind /sync/activities. The patterns on this page keep your app fast and keep the API healthy for every other developer building on Simkl.Request limits
Two limits apply at the same time, and you have to stay under both. One caps how fast you send requests; the other caps how many you send in a day.Per-second rate
10 GET per second and 1 POST per second, counted against the same bucket as your daily quota: per user for AUTH V2. Going over on aGET returns 429 — back off and retry. Going over on a POST triggers a temporary throttling block on the offending token or client_id, and repeated overages extend the block.
The Cloudflare-cached endpoints are exempt, because they are served from the edge and never reach Simkl’s origin. That is exactly why parallel requests are allowed on those endpoints and nowhere else.
One POST per second is less restrictive than it sounds, because every write endpoint accepts arrays. Fifty items sent as one batched
POST is a single request against the rate; the same fifty sent individually is fifty requests and will throttle you. See Batch writes below.Daily quota
Which model applies depends on which auth version your app uses. These are two different systems — read the one that matches your app, not both.AUTH V2 — per user
The allowance follows the user’s plan, not yours, and is counted per user account.
Because V2 requires a user token on every user-scoped request, this is the only limit a V2 app deals with — there is no app-wide quota to manage or exhaust. Your capacity grows with your user base rather than being a fixed ceiling you share between everyone.
The allowance belongs to the user, not to your app. Counting is against the Simkl account, so it is shared across every app that user has connected — yours, plus any other integration they use. Two grants for your own app share it too, so authorizing repeatedly buys nothing.In practice this rarely bites, because the allowance is generous and most users run one or two integrations. But it does mean a user running a busy media-server sync alongside your app has less headroom for you than a user who only runs yours. The way to get headroom is fewer, better-batched requests — not more tokens.
What the user can see
You cannot see another app’s usage, but the user can see all of it. Connected Apps shows them, in one screen:- Today’s requests against their allowance, what is left, and how close they are — the V2 allowance, so V1 traffic is not in this number
- A breakdown by app — every integration they have connected, V1 and V2 alike, sortable by today, yesterday or all-time
- Which auth version each app uses, labelled
AUTH V1orAUTH V2 - What each app is permitted to do, in plain words: “Can read and update your media library” or “Read only”
- Which connections are expiring soon
AUTH V1 — per client_id
A V1 app has one daily allowance covering everything it does, whether or not a request carries a user token. Every user of your app draws on the same pool, so the busiest user can consume the budget the rest of them needed.
It also means a V1 client_id is worth something to a stranger: it works on its own across the whole API, so anyone who finds yours can spend your allowance with it. A V2 client_id is not worth stealing — on its own it reaches only public catalog data, which costs nobody any quota.
This is the strongest practical reason to migrate. Under V1 your ceiling is fixed no matter how many users you have; under V2 each user brings their own allowance. An app that keeps growing on V1 eventually runs into a wall that V2 does not have.V1 is being retired around April 2027 in any case — see Migrating from V1 to V2. If your V1 app is hitting its ceiling before you can migrate, talk to us rather than rationing.
Treat the daily number as approximate, under either model. Simkl runs in more than one region and each counts independently, so a client spread across regions can get somewhat further than the published number before being cut off. Build against the published limit rather than the slack — the slack is an implementation detail and will tighten.
Headers
Read
X-RateLimit-Remaining rather than counting requests yourself — it accounts for the exemptions below, which your own counter will not.
The two
X-RateLimit-* headers are sent for AUTH V2 bearer tokens only. A V1 token has no per-user allowance to report, so it gets neither header — if you are testing with a V1 token and seeing nothing, that is why, and it is not a bug. You can see them on demand with ?debug_limit=user.V1 requests do not consume the user’s V2 allowance. The per-user daily counter tracks AUTH V2 traffic only, so a user running a busy V1 integration alongside your V2 app is not quietly spending your headroom. V1 requests are still counted per app — they appear in your Analytics and in the user’s per-app breakdown — they simply do not draw down the V2 budget.Every quota resets at midnight US Eastern — not at a rolling 24 hours from your first call, and not at your own local midnight. The same instant for every user, every app, every day.That is the whole reset rule, so you already know when your window ends without asking. Schedule a deferred sync against the next New York midnight and you will be right every time.
Showing the reset time in your UI
Until you are actually blocked, no response tells you when the quota resets — so compute it. Three states, and only the last one carries a reset:
The middle row is the one that catches people:
Remaining: 0 arrives on a successful response, with no Retry-After. The counter is compared after it increments, so the request that takes you to zero still goes through — it is the next one that fails. If your UI wants to say “resets in 6h” at the moment the user hits zero, you are computing it yourself.
Everything you need is in hand: X-RateLimit-Remaining gives the count, and the reset is always the next midnight in New York. That is enough to show the user something useful before they run out:
120 requests left for the next 8 hours.or, if you prefer the absolute form, “375 of 500 left — resets in 6h 12m”. Either beats a silent failure at request 501. The reset instant, precisely:
It is the same instant for every user and every app, worldwide. Nothing about your own server’s timezone changes it.
What does not count
Detail lookups for movies, TV and anime —GET /movies/{id}, GET /tv/{id}, GET /anime/{id}, and the episode-list endpoints — do not count against your daily quota and are not subject to the per-second rate, along with /redirect and the trending and calendar data files. They are served from the Cloudflare edge and never reach Simkl’s origin, so they are free on both limits.
This is why X-RateLimit-Remaining is the number to trust: a client that increments a local counter on every call will think it is far closer to the limit than it is.
Limit signals
Read the body before you decide how long to wait, because the three
429s want completely different responses. rate_limit is a burst you can retry in a second. user_limit_exceeded means one user was noisy and is done for the day. app_limit_exceeded means your whole app is out of budget and every user is affected at once.
Trigger a 429 on demand
Quota errors are hard to test precisely because the honest way to reach one is to spend a real day’s allowance. Add?debug_limit=user or ?debug_limit=app to any request and you get that refusal immediately:
The response is byte-for-byte what a genuine quota hit returns — same body, same status, same headers — because both come from the same code path. So it exercises your back-off logic, your retry timer and your “you’ve hit today’s limit” screen for real.
It costs nothing and affects nobody. The parameter turns your own request into an error and stops there: it is answered before any counter is touched, so it spends none of your quota and none of the user’s, writes no cache, and never reaches another caller. Leaving it in a request by accident wastes the call, nothing more.There is no flag to switch it on — it works on any request, on production, today.
Parallel requests — when allowed
Do not parallelize requests unless necessary. A single sequential client is the default — it stays well under the limits and never triggers a throttle. Parallel requests are explicitly allowed on endpoints cached at the Cloudflare edge, because parallel hits there are served from cache and don’t pressure Simkl’s origin:
Everything else stays sequential. Sync endpoints, user-state endpoints, and search endpoints serve per-user or dynamic data that bypasses the edge cache — every parallel call hits the origin directly and counts individually against the cap.
Why parallel on uncached endpoints hurts
It doesn’t increase throughput. The ceiling is 10 GET/sec, 1 POST/sec. Ten parallel GETs att=0 all land in the same one-second window — you have spent the whole per-second budget at once, and request 11 gets 429. The same ten calls spread across that second land you in exactly the same place with zero 429s, and cost the identical daily quota either way. Parallelism buys you nothing here and costs you the errors.
Servers are commodity hardware, not supercomputers. Imagine opening 10 copies of Photoshop on your laptop at the same time — each one loads the full app into RAM independently, the fans spin up, and everything else slows to a crawl. A web server handling 10 parallel uncached requests does the same thing: each request loads its own slice of the application — framework boot, autoloader, config, model code, request context — and holds it until the response is sent. Ten parallel requests means 10 full copies of that boot state in RAM at once, plus 10 database connections, 10 worker threads, 10 response buffers, and 10 fan-out calls to downstream services (read replicas, search index, image proxy). Connection pools, CPU schedulers, GC pauses, and disk I/O all have finite headroom; when a burst exceeds capacity, in-flight requests slow down or fail and the next caller waits for the queue to drain. Sequential traffic reuses the same worker over and over — boot cost paid once per worker, not once per request — so no single resource ever spikes above safe levels.
A 429 has user-visible cost. A throttled call breaks the user’s session — the page stalls, the sync stalls, a mid-flight write fails. Your client then has to back off and retry, doubling the time-to-success. Retry storms also re-pressurize the origin and can extend the throttle window. “Just parallelize and let the API tell us to slow down” is not a viable strategy; the failed requests have already cost user-visible time.
Sustained overage suspends your client_id. Per the /sync/all-items warning: apps that hammer the API — particularly write endpoints or polling without /sync/activities gating — get their client_id suspended. No warning, no appeal.
Burst patterns look like abuse. Edge and origin monitoring can’t tell a well-meaning client firing 50 parallel sync calls apart from a botnet or scraper — the signature is the same: one client_id or IP sending many concurrent requests to non-cacheable URLs in a short window. Automated mitigation kicks in (temporary block, IP throttle, longer back-off) without anyone making a judgment call about intent.
The correct pattern for uncached endpoints: one in-flight request at a time per token; batch into one call instead of N parallel calls whenever the endpoint accepts arrays (every write endpoint does — see Batch writes below); gate polling behind /sync/activities so most /sync/all-items calls never happen.
Cache invalidation on the cached endpoints is automatic. When Simkl updates the underlying record (admin edits, metadata refresh, image swap, new episode airing, etc.), the corresponding Cloudflare cache entry is purged server-side. The next call returns fresh data — no TTL to wait out, no
?nocache=… trick needed. Same applies to the trending and calendar JSON files. Your own app-level cache, if you have one, still has to be invalidated by your client.Best practices
Cache aggressively — but key user data by account
Cache aggressively — but key user data by account
Most metadata never changes. Cache responses on the device for the longest sensible time — minutes for user data, hours for catalog data, indefinitely for images.Cache catalog data by URL. Cache user data by URL and account. Anything behind a user token — the library, ratings, playback, custom lists — can return completely different content, or different permissions, for the same URL depending on who is signed in. A cache keyed on URL alone will happily serve one user’s private list to the next person who signs in on that device.Include the account identity in the cache key, and drop those entries on sign-out or account switch. Catalog responses are the same for everyone and need none of this.
Use the CDN files for trending and calendars
Use the CDN files for trending and calendars
Sync incrementally
Sync incrementally
Always check
/sync/activities before pulling watchlists. Skip watchlists whose timestamp didn’t change. See the Sync guide for the two-phase model.Batch writes — one POST per second is enough for arrays of 50+ items
Batch writes — one POST per second is enough for arrays of 50+ items
Every write endpoint accepts arrays.
POST /sync/history, POST /sync/history/remove, POST /sync/ratings, and POST /sync/add-to-list will happily process 50 items in one request.Sending 50 separate single-item POSTs immediately blows past the 1-POST-per-second cap and gets your token or client_id throttled — sending one batched array stays well under. It also costs one request of daily quota instead of 50, and avoids the 50 calls queueing behind each other on the short per-user write lock.One batched array is faster, cheaper, and the only version that does not throttle.Retry transient errors with exponential backoff — but check what kind of 429 you got
Retry transient errors with exponential backoff — but check what kind of 429 you got
On
500, 502, 503: wait, retry, double the wait, cap it. Give up after 5 attempts. A common starting pattern is 1s → 2s → 4s → 8s → 16s, but tune to your application’s tolerance.On a 429, read the body before choosing a wait — the two kinds are hours apart:- A bare
429is the per-second rate. It clears in about a second, so pause briefly and carry on. Exponential backoff here is overkill and makes your app feel broken. user_limit_exceededorapp_limit_exceededis the daily quota. Five retries over thirty seconds cannot possibly clear it — you will burn the attempts and fail anyway. ReadRetry-After, wait for the window, or stop and surface the problem to the user.
400 RATE_LIMIT at all: that is the short per-user write lock, which clears as soon as the in-flight write finishes. Serialise your writes for that user and retry shortly.Four failures, four different waits: seconds for a server error, about a second for a burst, moments for the write lock, until reset for the daily quota.Resolve external IDs via /redirect rather than search
Resolve external IDs via /redirect rather than search
If you have an IMDb / TMDB / MAL / etc. ID and need the matching Simkl ID, use
GET /redirect — it returns a tiny 301 redirect with the Simkl ID in the Location header (no JSON body), and the follow-up GET /movies/{id} / GET /tv/{id} / GET /anime/{id} call is Cloudflare-cached. Much cheaper than GET /search/id.Hide your client_id in CI
If you’re using a CI system, storeclient_id and client_secret in your provider’s encrypted secrets / environment variables — not in repository files. A leaked credential burns your quota and is hard to recover.