Posted on Oct 8
A cache can absorb a burst of requests for one popular product page. But a crawler visiting thousands of different product pages creates a different problem: every URL may need its own render.
Your application can be busy generating responses that nobody else will reuse. Meanwhile, ordinary visitors are competing for the same rendering capacity.
Tanod is an open-source Rust reverse proxy for controlling origin work in server-side rendering (SSR) and other expensive HTTP applications. It combines concurrency limits, bounded queues, request coalescing, and short-lived caching, with a dedicated integration for self-hosted Next.js.
Its goal is to stop traffic spikes from becoming render spikes.
The problem with counting requests alone
Two HTTP requests can place different demands on your server. One returns a static file. Another renders a page, queries a database, and calls a content management system.
A requests-per-second limit does not express how much of that work may be active at once. Longer responses can leave the origin with more concurrent work even when the arrival rate stays the same.
For an illustrative steady-state workload, 40 requests per second with a 500 ms response time means about 20 requests in flight. If response time grows to 1.5 seconds at the same arrival rate, that becomes about 60. Those numbers describe the relationship, not a benchmark or a recommended server capacity.
Caching and client rate limits still matter. They address different parts of the problem:
| Mechanism | What it controls |
|---|---|
| Client rate limiting | How much traffic a client may send over time |
| Response caching | How much repeated work can be avoided |
| Request coalescing | How many identical concurrent requests need separate origin responses |
| Origin admission control | How much new origin work may run at once |
Tanod concentrates on the last three. You still need an appropriate public edge and client rate limiting.
What Tanod does
Tanod runs in front of your application server, also called the origin. Place it behind a content delivery network (CDN), an edge proxy, or a load balancer:
Browser / crawler
↓
CDN / Caddy / NGINX / load balancer
↓
Tanod
↓
Next.js or another HTTP application
The edge handles responsibilities such as certificate management and client rate limits. Tanod handles origin capacity and eligible response reuse.
Tanod is built on Pingora, Cloudflare's Rust proxy framework. Its application-specific layer adds route classification, response-sharing policy, and bounded origin admission.
The name comes from the Tagalog tanod: a barangay watchman who keeps order at the gate.
Reuse first, then spend origin capacity
The ordering matters because a cache hit should not wait for capacity that it does not need.
For governed requests, Tanod follows this sequence:
- Match the route and classify the request.
- Reuse an eligible cached response or attach to a safely shareable in-flight response.
- Acquire the configured route and global capacity permits for remaining origin work.
- Wait within a bounded queue if capacity is busy.
- Forward admitted work, or reject it when the queue is full or its deadline expires.
- Check the origin response before sharing or storing it.
Request coalescing means equivalent concurrent requests can receive one origin response. Microcaching means an eligible response is stored briefly so later requests can reuse it.
Neither helps when every request is a distinct, uncacheable miss. The concurrency ceiling still applies to that work, including private dynamic requests.
You configure the limits and optional route weights. Tanod does not automatically measure the CPU or memory cost of a render, and it cannot infer your origin's safe capacity.
Overload needs a defined response
A queue is useful only when it has an end. Allowing unlimited requests to wait can move the overload from the application into the proxy.
Tanod bounds both queue size and waiting time. When it cannot admit a request, the default response is 503 Service Unavailable with a Retry-After header.
For a browser loading a document, Tanod can return a configurable busy page that refreshes automatically. API clients and Next.js payload requests receive the error status without that HTML page.
The busy page is a retry mechanism, not a reserved place in line. Each refresh is a new request.
Eligible stale content can be served during background revalidation or an upstream failure. It is not currently served as a fallback for overload shedding. That distinction matters when choosing how a public page should behave under load.
Response sharing requires a privacy decision
A popular URL is not necessarily a public response. A product page might include account information, a customer-specific price, or a session cookie.
Tanod checks both request eligibility and response shareability. Its rules include:
- Requests with
Authorizationand unsafe methods bypass sharing - Cookie-bearing requests are private by default
- Responses containing
Set-Cookiecannot be shared through a route override - Unsupported response
Varyheaders prevent sharing - Origin
private,no-cache, andno-storedirectives are respected unless a validated public route explicitly overrides them
A public-route override is an application privacy assertion. It requires reviewing what the route returns, including for visitors who send cookies. It does not make a personalized response safe to share.
This is why a first deployment can enable origin protection while leaving caching and coalescing disabled. You can review public routes separately after measuring the origin.
Next.js needs more than a URL-based cache key
With Next.js, the same URL can return an HTML document or a React Server Components (RSC) payload. Prefetch requests introduce additional variants, while Server Actions and draft mode need their own handling.
Tanod's Next.js adapter accounts for those distinctions. Cache keys separate relevant framework variants, and Server Actions and draft-mode content bypass response sharing.
The companion @tanod/next package reads production build manifests to generate route configuration. Dynamic pages and Route Handlers remain private unless you approve them through a checked-in policy. Prerendered routes can receive public policies based on build evidence.
There is also an invalidation boundary to understand: calling Next.js revalidateTag() or revalidatePath() alone does not invalidate a response held by Tanod. The package provides integration for coordinating both caches.
Tanod can proxy other HTTP applications with manual route configuration. Next.js currently has the dedicated build integration; equivalent integrations for other frameworks remain future work.
Try origin protection locally
This walkthrough starts with sharing disabled. It requires Linux x64, a Rust toolchain, and a Next.js application you can build and run locally. Tanod declares Rust 1.88 as its minimum; the repository pins its build toolchain in rust-toolchain.toml.
In a terminal inside your Next.js project, build and start the origin on loopback:
npm run build
npx next start --hostname 127.0.0.1 --port 3000
In a second terminal, clone Tanod and build its binary:
git clone https://github.com/akosidencio/tanod.git
cd tanod
cargo build --release --locked
Generate a local configuration that forwards to your running application:
./target/release/tanod init \
--config tanod.local.yaml \
--upstream 127.0.0.1:3000 \
--listen 127.0.0.1:8080 \
--concurrency 16
The generated configuration admits up to 16 concurrent origin work units and allows up to 32 queued requests, with a two-second queue deadline. Its private catch-all disables caching and coalescing.
The value 16 is a starting point for experimentation, not a recommendation for your application. Measure representative requests, origin memory use, latency, and queue waits before choosing a deployment limit.
Validate the configuration, then start the proxy:
./target/release/tanod check --config tanod.local.yaml
./target/release/tanod run --config tanod.local.yaml
In another terminal, send a request through Tanod and inspect its separate admin listener:
curl -i http://127.0.0.1:8080/
curl -fsS http://127.0.0.1:9091/health/live
curl -fsS http://127.0.0.1:9091/status
Application traffic enters on port 8080; administrative checks use port 9091. Keep the admin listener private in a public deployment.
For a server installation with systemd and a public domain, follow the standalone guide.
Test the case where caching cannot help
The repository includes a focused admission-control demonstration. From the Tanod directory, run it separately from your application traffic:
./bench/demo.sh 60 1000
The script starts a fixture origin and sends 60 concurrent requests for 60 distinct URLs. Each fixture response simulates one second of work. The benchmark configuration sets the origin ceiling to 10 and disables caching.
It compares direct traffic with traffic through Tanod. The origin reports its own peak concurrency, and the script asserts that the direct run exceeds the configured ceiling while the governed run stays within it.
The fixture sleeps rather than performing a real React render. This checks admission behavior, not Next.js throughput. It also demonstrates the trade-off: serializing work into smaller groups can take longer when the origin could have handled the entire batch.
The repository has separate Next.js HTTP, browser, and replica checks. Those exercise framework behavior, response isolation, invalidation, and deployment boundaries.
What to measure before tuning a limit
The useful question is how your origin behaves as concurrent work grows. Request totals alone cannot answer it.
Tanod exports Prometheus metrics, structured logs, and optional OpenTelemetry data. A Grafana dashboard and alerting rules help inspect:
- Origin work in flight compared with the configured ceiling
- Queue waiting time and total request duration
- Responses shed by Tanod
- Origin latency and backend health
- Cache reuse and memory occupancy
Raising a concurrency ceiling can reduce queueing or increase overload. Check the origin's behavior before treating either outcome as an improvement.
Current limits and the next steps
The version covered here is 0.3.0, and Tanod remains pre-1.0. Repository tests and bounded staging tests support its mechanisms, but sustained production validation and an independent cache-safety review remain open.
There are deployment boundaries to account for:
- Cache and coalescing are local to each process
- Several proxies sharing one origin pool need statically partitioned capacity budgets
- Purges must reach every instance that can hold a cached response
- Exact-path purges do not expand dynamic route patterns
- Slow readers can keep origin permits occupied; optional bounded spooling sacrifices progressive rendering to release capacity earlier
- Static and long-lived streaming classes do not use render permits; WebSocket upgrades have separate limits
Tanod can also supervise one local application through origin.command, allowing the app and proxy to run within one container or service. That simplifies the process arrangement, but each instance still owns its own cache and capacity boundary.
The roadmap focuses on deployment validation, independent review, overload behavior, and additional integrations. Tanod aims to make origin overload controlled and observable; it does not promise that every application becomes faster or cheaper.
Try Tanod with a workload you understand
If you operate an expensive self-hosted origin, start by measuring concurrent work and response latency. Try Tanod with private defaults, set a capacity budget, and inspect what happens when that budget is reached. Enable response reuse only after reviewing the routes that can safely share it.
The project is available under the Apache 2.0 license:
Feedback on self-hosting workloads, reproducible failures, and cache-safety assumptions can help determine the next priorities. When reporting behavior, include a redacted configuration, the request pattern, and the Tanod version.
Top comments (1)
-
Email
-
LocationToronto, ON, Canada
-
EducationUniversity of Toronto
-
PronounsHe/Him
-
WorkSoftware Development Manager at Purolator
-
JoinedJul 15, 2018
I'd look at the two-second queue deadline first, because under sustained overload it becomes a standing queue. If arrivals stay above the ceiling, the queue never empties, so nearly every admitted request waits close to the full two seconds before its render even starts.
Facebook's "Fail at Scale" (ACM Queue, 2015) describes a fix: keep the normal timeout while the queue drains regularly, but once it hasn't been empty for N ms, cap the wait at M ms. They found M = 5 ms and N = 100 ms worked across a wide set of services. They pair it with adaptive LIFO, switching to newest first once a queue starts to form, so the requests you do admit belong to clients that are probably still waiting. For Tanod that could sit as a per-route option beside the deadline, and the queue waiting time you already export would show whether it helps.
For further actions, you may consider blocking this person and/or reporting abuse
