API rate limiting controls how much request traffic a client, user, API key, tenant, IP address, or service can generate within a defined period. The goal is not simply to block traffic. Good rate limiting protects backend capacity, login flows, expensive endpoints, third-party dependencies, and fair access while giving legitimate clients a predictable way to recover when they reach a limit.
For Raff Technologies workloads, rate limiting is usually an application or reverse-proxy decision layered on top of the compute, network, security, and observability controls around the workload. The best policy depends on who owns the traffic, how expensive the endpoint is, whether the caller is authenticated, and whether rejected work could instead be queued or slowed down.
A useful rule is:
Rate-limit the identity that owns the traffic, protect expensive paths first, and make 429 responses actionable for legitimate clients.
API rate limiting: quick decision framework
| API surface | Main risk | Better limit key | Typical control |
|---|---|---|---|
| Public unauthenticated endpoint | Bots, scraping, accidental bursts | IP/source + route | Conservative burst + sustained limits |
| Login | Brute force, credential stuffing | Account + source + route | Tight throttling, progressive delay |
| Signup/password reset | Spam and abuse | Account/email/source | Frequency limits + abuse checks |
| Authenticated API | Tenant or integration overuse | User, tenant, API key | Identity-aware quota |
| Expensive search/report | CPU/database pressure | Identity + endpoint | Lower rate + concurrency limit |
| File upload | Bandwidth, storage, processing | User/tenant + size + route | Size + frequency + concurrency |
| Webhook receiver | Retry storms, bursts | Provider/account/event | Queue + backpressure |
| Internal API | Service overload | Service identity + route | Service-aware rate/concurrency limits |
| Admin action | High-impact misuse | Admin identity + action | Conservative limits + audit evidence |
Do not begin with “100 requests per minute for everything.” Begin with which resource or user experience needs protection.
What is API rate limiting?
API rate limiting is a policy that restricts request activity according to a key, quota, time window, concurrency level, cost unit, or combination of those controls.
Examples include:
- 60 requests per minute per API key;
- 5 login attempts per account in a short window;
- 2 report-generation jobs running concurrently per tenant;
- 100 uploads per day per workspace;
- a short burst allowance followed by a lower sustained rate;
- different limits for free and paid plans.
The limit key matters as much as the number. A limit applied to the wrong identity can block legitimate users while failing to constrain the actor causing pressure.
Why APIs need rate limits
Rate limiting protects more than security.
A production API may need protection from:
- malicious scraping or brute force;
- buggy clients retrying too aggressively;
- batch jobs that accidentally loop;
- one tenant consuming shared capacity;
- expensive search/report endpoints;
- file-processing spikes;
- webhook retry storms;
- external API or email-provider quotas;
- sudden load on a database or queue;
- plan or contract usage limits.
OWASP API Security identifies unrestricted resource consumption as an API risk because requests can consume compute, memory, storage, bandwidth, database capacity, and paid third-party services.
That makes rate limiting partly a reliability and cost-control mechanism, not just an abuse filter.
Rate limiting vs throttling vs quotas
These terms overlap but are useful to separate operationally.
| Control | Main purpose |
|---|---|
| Rate limit | Restrict request activity over time |
| Throttling | Slow or reject traffic when policy/capacity is exceeded |
| Quota | Control a larger usage allowance, often per plan or billing period |
| Concurrency limit | Restrict work executing at the same time |
| Queue/backpressure | Absorb valuable work and process it at a safe rate |
A product can use several at once.
For example, an API may allow a short burst, enforce a per-minute request rate, cap daily usage, and permit only two heavy exports to run concurrently.
429 Too Many Requests: what it means
HTTP 429 Too Many Requests means the client has exceeded a rate policy and should reduce request activity. RFC 6585 defines the status code and allows the response to include Retry-After so the client knows when to retry.
A useful 429 response should help a legitimate developer recover.
Include, where appropriate:
- a clear error code/message;
Retry-Afteror equivalent retry guidance;- the scope of the limit, such as user, token, tenant, or route;
- safe quota information;
- a request/correlation ID;
- documentation for public APIs.
A minimal example:
HTTP/1.1 429 Too Many Requests Retry-After: 30 Content-Type: application/json { "error": "rate_limit_exceeded", "retry_after_seconds": 30 }
Do not expose internal enforcement details that make abuse easier merely to make the error verbose.
Rate-limit headers are useful, but the standard is still evolving
Public APIs often expose quota information through provider-specific headers such as X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset.
The IETF HTTPAPI working group is also developing standardized RateLimit-Policy and RateLimit fields. As of May 2026, that work remains an Internet-Draft, not a finalized RFC, so treat it as work in progress rather than a settled standard.
The durable design principle is simpler: if customers depend on your API, give them enough information to avoid unnecessary throttling and implement backoff correctly.
Choose the right rate-limit key
The key identifies who or what is consuming the quota.
| Key | Good fit | Main limitation |
|---|---|---|
| IP address | Unauthenticated traffic and coarse abuse controls | NAT, VPNs, offices, mobile networks share IPs |
| User ID | Authenticated user fairness | One user may act for a larger organization |
| Tenant/account | B2B SaaS fairness | Large tenants may need higher policy limits |
| API key | Integrations and automation | Shared keys reduce attribution |
| Route/action | Protecting expensive operations | Usually needs another identity dimension |
| Service identity | Internal APIs | Requires reliable service authentication |
| Plan | Contract or product quota | Not enough for burst protection |
For authenticated APIs, user, tenant, API key, or organization is usually more meaningful than IP alone.
IP-based controls still help at the public edge, but they should not be the only fairness mechanism for logged-in users.
Do not use one global limit for every endpoint
Endpoints have different costs and risks.
A health check, login request, database search, export, upload, and billing action should not automatically share one policy.
Classify endpoints by:
- cost — CPU, database, storage, bandwidth, external API usage;
- security sensitivity — login, reset, admin, account changes;
- business impact — writes versus reads, destructive actions;
- expected burst behavior — UI page load, batch process, webhook;
- caller identity — anonymous, user, tenant, API client, internal service.
Then apply the narrowest policy that protects the actual bottleneck.
Burst limits and sustained limits solve different problems
Normal applications can create legitimate bursts.
A dashboard may load several resources at once. A mobile client may synchronize after reconnecting. A batch workflow may submit several related operations quickly.
A good policy often separates:
- burst capacity for short legitimate spikes;
- sustained rate for ongoing usage;
- quota for longer-term allowance;
- concurrency for expensive active work.
This is why token-bucket and similar designs are popular: they can allow controlled bursts while still limiting average throughput.
Rate limiting algorithms: fixed window, sliding window, token bucket, leaky bucket
The algorithm matters, but only after the policy is clear.
| Algorithm | Strength | Main trade-off |
|---|---|---|
| Fixed window | Simple to implement and explain | Boundary spikes near window reset |
| Sliding window | Smoother recent-time enforcement | More state/calculation |
| Token bucket | Allows bursts while controlling average rate | Bucket size/refill need tuning |
| Leaky bucket | Smooth output rate | Bursts may be delayed or rejected |
| Concurrency limiter | Protects expensive in-flight work | Does not control total request volume |
Fixed window
A counter resets on a fixed boundary such as each minute. It is easy to operate but can permit traffic spikes around the boundary.
Sliding window
The server evaluates activity across the most recent interval rather than a shared reset point. It gives smoother enforcement at higher implementation cost.
Token bucket
Tokens refill at a defined rate. Requests consume tokens. A larger bucket permits short bursts without allowing unlimited sustained traffic.
Leaky bucket
Work leaves the bucket at a controlled rate, which is useful when smoothing traffic matters more than allowing immediate bursts.
For most small teams, the best algorithm is the one that matches expected client behavior and can be observed and tuned reliably.
Rate limiting should protect expensive endpoints first
If implementation time is limited, prioritize endpoints where one request can trigger disproportionate work.
Start with:
- login and password reset;
- signup and verification;
- expensive search;
- reports and exports;
- file uploads;
- AI/compute-heavy operations;
- external API calls;
- email/SMS actions;
- webhook receivers;
- administrative writes.
A cheap cached read endpoint may tolerate much higher request rates than a report that performs several database joins and generates a file.
Rate limiting should work with queues and backpressure
Some valid work should be delayed rather than rejected.
Queues are useful for:
- exports;
- image/file processing;
- webhook processing;
- email delivery;
- background reports;
- batch jobs;
- long-running tasks.
A useful pattern is:
Request ↓ validate + authorize Accept job ↓ Queue ↓ controlled concurrency Workers ↓ Result / callback / status
Rate limits still matter because a queue can also be overwhelmed. The difference is that valuable work can be absorbed and processed safely instead of failing immediately.
Read Cron Jobs vs Queues vs Workflow Automation for the wider background-work decision.
Login rate limiting needs special treatment
Authentication endpoints are both security-sensitive and user-sensitive.
Avoid simplistic policies that permanently lock accounts after a small number of failures. Attackers can intentionally trigger those failures and create denial of service for real users.
A stronger design can combine:
- account-aware throttling;
- source/IP signals;
- progressive delay;
- device/session context where available;
- temporary challenge or verification;
- monitoring for distributed attempts;
- clear recovery paths.
The exact policy depends on the authentication system, but the goal is to slow abuse without giving an attacker an easy way to lock legitimate users out.
API keys and rate limits should be designed together
API keys are useful rate-limit identities because they can map traffic to a known integration.
Prefer:
- one key per integration/customer/environment where practical;
- an owner for every key;
- scoped permissions;
- separate development and production keys;
- rotation and revocation procedures;
- usage monitoring by key.
A shared key across several applications makes both rate limiting and incident investigation harder.
Use API Keys for Automation for the credential-lifecycle side of the design.
Rate limiting is not DDoS protection
Rate limiting and DDoS protection can both support availability, but they solve different layers of the problem.
| Control | Primary job |
|---|---|
| Firewall | Restrict ports, protocols, and network sources |
| DDoS protection | Detect/filter high-volume or hostile traffic patterns |
| WAF/application filtering | Inspect HTTP requests and web attack patterns |
| API rate limiting | Allocate request capacity by endpoint and identity |
| Concurrency limit | Protect active workers/dependencies |
| Quota | Enforce longer-term usage policy |
Application-aware rate limiting knows things network-level controls often do not: user ID, tenant, API key, endpoint cost, plan, or business action.
Use Cloud Security Fundamentals and Cloud Firewall Best Practices for the surrounding security layers.
Observe rate-limit behavior after launch
A rate limit should not be “set and forgotten.”
Track signals such as:
- 429 responses by endpoint;
- 429s by user/tenant/API key;
- top limited clients;
- request rate before and after limiting;
- backend CPU/memory/database load;
- queue depth;
- retry behavior;
- support tickets about throttling;
- error rates after limit changes;
- cost/resource trends on expensive endpoints.
Interpret the signals carefully.
If only known abusive clients hit limits, the control may be working as intended. If normal users repeatedly hit them, the limit may be too strict or the application may be generating unnecessary requests.
Read Observability for Small Teams for the monitoring layer.
API rate limiting best practices
A practical production baseline is:
- Define the resource being protected. CPU, database, login flow, third-party API, storage, or fair tenant capacity.
- Choose the responsible identity. IP, user, tenant, API key, service, or action.
- Protect high-cost endpoints first. Do not spend the same engineering effort on every route.
- Allow normal bursts deliberately. Avoid punishing legitimate page loads or synchronization.
- Use concurrency limits for expensive work. Request count alone may not protect active workers.
- Queue valuable asynchronous work. Rejecting every burst is not always the best UX.
- Return useful 429 responses. Include retry guidance where possible.
- Document public limits. Customers should not discover quotas only through production failures.
- Monitor limit hits. Rate policy needs evidence to improve.
- Review policy after traffic or architecture changes. New endpoints and larger customers can change the correct threshold.
How API rate limiting fits Raff
Rate limiting is generally implemented close to the application because the application knows the identity and business cost of each request.
A Raff-hosted API can apply controls at several layers:
Client traffic ↓ Network/security controls ↓ Reverse proxy or gateway ↓ Application middleware ↓ Queue / workers / database
Raff VM and Linux VM can host reverse proxies, API services, containers, queues, caches, workers, and application-level limiters. Raff VPC can keep supporting services on private paths where appropriate, while Raff Security covers the current infrastructure security surface.
The important ownership boundary is that Raff provides the infrastructure layer; your application still decides which user, tenant, route, API key, or business action receives which rate policy.
Use the live pricing page for current commercial details rather than encoding mutable plan, bandwidth, or deployment claims inside an evergreen rate-limiting guide.
