Embedding Tollgate on a request path
For a service putting Tollgate's staged admission pipeline in front of its own work. It states the contract: the request order, what you implement, what you cannot, and how to shut down without losing usage.
It is not the rationale. docs/DESIGN.md § "Staged admission interface (GL-96)"
says why the seam has this shape and is the document an interface change
amends first. It is also not the provisioning API — creating accounts, setting
budgets and issuing credentials is docs/ACCOUNT_ADMINISTRATION.md.
examples/pricing-api is the worked example. Read it alongside this; the last
section says which of its choices are the contract and which are its own.
Depending on Tollgate
The crates are published to crates.io and release as one unit: every crate
carries the same version, and they depend on each other at exactly that
version, so Cargo resolves a consistent set. Depend on the same 0.x series
for each:
[dependencies]
tollgate-core = "0.30"
tollgate-admission = "0.30"
# Only if you want the managed runtime — see "Two ways in" below.
tollgate-client = "0.30"
Under 0.x, a breaking change moves the minor version, so a caret requirement
never crosses one. Pin the exact version your conformance run was recorded
against ("=0.30.1"), and move it deliberately.
Two ways in
tollgate_admission::AdmissionEngine is the request path itself. You own
snapshot publication, lease funding and usage export.
tollgate_client::InstanceRuntime owns those for you — snapshot
distribution, lease acquisition and renewal, usage batching, readiness and
shutdown — and hands back a RuntimeHandle whose begin delegates straight to
the engine. This is what the example uses and what most services want.
Everything below is identical either way: RuntimeHandle::begin and
AdmissionEngine::begin have the same signature and return the same
RequestContext.
The supported request order
This is the contract, from docs/DESIGN.md:
The supported request order is authenticate,
begin, reserve the usage slot, read/decode underctx.limits(), then callctx.admit.
Each position earns its place, and reordering them loses something specific:
- Authenticate from bounded transport metadata — a header, not a body.
beginresolves the account snapshot once and checks the route permission. Everything downstream reads that one pinned generation, so a snapshot refresh mid-request cannot change the answer under you.- Reserve the usage slot. After
begin, so an unknown or unauthorized credential never occupies the usage queue. Before the body, so backpressure sheds before expensive input work. - Read and decode under
ctx.limits(). A body-read failure drops the context and the slot and creates no funding reservation. admitwith the compiled workload. This is the stage that can create pending funding — everything before it is free to fail.
Then acquire_capacity, commit at the moment execution actually starts, and
drop the Committed guard when the work is done.
The stages
#![allow(unused)] fn main() { // 2. begin — one snapshot lookup plus the route permission. pub fn begin(&self, principal: Principal, required: PermissionBits, now: Timestamp) -> Result<RequestContext, DenyReason>; // 5. admit — shape, rate, concurrency and funding. pub fn admit<O: OpIndex, S: UsageSlot>(self, workload: &[(O, u64)], slot: S, now: Timestamp) -> Result<Pending<S>, DenyReason>; // 6. capacity — a startup-selected policy, not caller input. pub fn acquire_capacity<G: CapacityGate>(self, gate: &G) -> Result<ReadyToStart<S, G::Permit>, (DenyReason, Released)>; // 7. commit — at actual execution start, not at admission. pub fn commit(self, request_id: RequestId, now: Timestamp) -> Result<Committed<S, P>, (CommitError, Released)>; }
RequestContext is Send + Sync + 'static and cheap to hold, so it crosses the
body-read await. It carries snapshot(), limits(), generation() and
policy_revision() — read your effective limits from it rather than from
anywhere else, because it is the generation this request was authorized under.
Commit can fail. CommitError at execution start means funding lapsed
between reserving and starting. The Result is what forces the executor to
check before running the kernel; the #[must_use] on commit is deliberate.
The workload is borrowed. &[(O, u64)] — a stack array or your own
fixed-capacity buffer. Tollgate owns no Vec, box, hash table, string or
product identifier, and quoting is O(distinct classes), not O(policy records).
Dropping Committed is what bills. The usage event is emitted from Drop,
so the guard must live exactly as long as the work. cancel() before commit
releases everything and charges nothing.
What you implement
OpIndex — dense indices for your operation classes.
#![allow(unused)] fn main() { impl OpIndex for Op { fn index(&self) -> usize { *self as usize } } }
UsageSlot — pre-reserved capacity for exactly one event.
#![allow(unused)] fn main() { pub trait UsageSlot: Send + 'static { fn record(self, event: UsageEvent); } }
record must not perform fallible I/O. Obtaining the slot is the backpressure
decision; consuming it is the committed-charge path, and it runs from Drop.
Take tollgate_client::UsagePermit from the runtime's recorder unless you have
a reason not to.
Authentication — verify before the body. tollgate_auth::CredentialVerifier
is the scheme seam and HmacRegistry the implementation in the box. Strip any
transport prefix before the library sees the bytes, so what you cache and what
you verify are the same bytes by construction.
A panic boundary around your kernel. Tollgate's guard is drop-safe and
allocation-free, but a panic that escapes a worker thread can leak the guard
rather than drop it, and a leaked guard emits nothing. catch_unwind (or your
executor's equivalent) is yours to place. Under panic = "abort" this does not
apply, and INVARIANTS GL-13 states the process-loss boundary instead.
Shutdown ordering — see below.
What you cannot implement
CapacityGate and CapacityPermit are sealed. An application's own compute
permit is not a Tollgate capacity decision, and a type outside the crate
satisfying the permit would be a way to start work the gate refused. Choose a
policy at startup: NoGate (zero-sized, disabled) or ExecutionCapacityGate
with Uniform or Reserved.
Your own compute admission stays yours and stays separate. Weights you use to schedule work are not cost units.
Adopting in stages
You do not have to wire everything at once. Two zero-cost stand-ins exist for exactly this, and both are honest about being stand-ins:
NoGate— admission without execution-capacity limiting.DiscardedUsageSlot— admission without usage export. It counts what it discarded, deliberately: a silently dropping slot makes "usage is not wired up" indistinguishable from "no usage happened", and one of those means you are admitting billable work and losing the record. Not for production billing.
Shutting down
One safe order, because usage events must land while their lease is still live:
- Stop admitting — stop your listener first.
- Quiesce request tasks still holding permits or
Committedguards, bounded by the deadline the runtime returns. - Await the usage writer's shutdown; it refuses new reservations, drains outstanding permits, and reports anything unresolved.
- Only then release leases.
InstanceRuntime does steps 2–4 under one deadline. Step 1 is yours, and so is
bounding your own tasks by the deadline it gives back.
What is contract, and what is pricing-api's own choice
Contract:
- the request order, and every signature above;
- reading limits from the pinned context;
- commit at execution start, and checking its
Result; - the
Committedguard living as long as the work; - the shutdown order.
pricing-api's own choices, which you should not copy without deciding:
- Doing admission inside an Axum
FromRequestextractor. It is a tidy place to enforce "before the body" in that framework, and nothing requires it. MemoryStorein-process. Pointing the same stack at atollgate-serveris a swap toHttpStore.NoGateby default, with the capacity gate behind configuration.- A per-connection credential cache. Sound for long-lived connections; irrelevant if yours are short.
- Its error mapping. RFC-7807 shapes and which
DenyReasonbecomes which status are product decisions.
Keeping an account on one core
Admission writes a handful of per-account cache lines — the rate bucket, the concurrency gauges, the lease counter, and the reference counts of the account state and lease — and each is fast only while it stays in the cache of the core that last wrote it. A work-stealing runtime moves a connection's task between worker threads, so an account's next request often runs on a different core and pulls every one of those lines across first.
On the controlled host, eight threads admitting for eight distinct accounts
cost about 115 ns per admission when each account stays on one thread and
about 1 µs when the threads cycle through the accounts
(admission/full_check_contended_8_distinct_accounts and its _rotating
twin). In the example service at ten connections the in-handler admit grew
from about 200 ns to about 500 ns for this reason alone (GL-138). Nothing in the
library is shared between those accounts; the cost is the account's own lines
moving.
The library does not choose where your tasks run, so this is yours to decide. It matters when admission is a meaningful share of your request cost:
- A thread-per-core runtime — one single-threaded runtime per core, each
accepting its own connections (for example through
SO_REUSEPORT) — keeps a connection's requests on one core for its whole life. - Connection affinity keeps an account's traffic on few connections, so its requests land on few cores. Many connections for one account spread it back out whatever the runtime does.
- A default multi-threaded runtime is correct and usually fine; it pays this migration on some fraction of requests, which is what the example measures.
Related
docs/DESIGN.md§ "Staged admission interface (GL-96)" — the rationale, and the document an interface change amends first.docs/ACCOUNT_ADMINISTRATION.md— provisioning accounts, budgets and credentials over HTTP, with a conformance list for that surface.docs/CREDENTIAL_PROJECTION.md— authenticating customer keys in an HTTP-backed deployment.docs/USAGE_ACCOUNTING.md— what happens to the events you emit.docs/DESIGN.md§ "Instance-local admission sharding (GL-3)" — the opt-in layout, once same-account contention warrants it. It is off by default and costs nothing until you enable it.INVARIANTS.md— the testable contract. Three of them are reachable from outside, which is why they appear above: GL-8 accounting backpressure sheds (why the slot is reserved before the body), GL-12 no commit outside the usability window (why leases are released last), and GL-13 a committed charge is always emitted (whose process-loss boundary is what your panic boundary keeps you inside).
Funding refusal and retry advice
Two reasons say the account itself cannot fund a request. Both are
Retry::Never under the current funding, both charge nothing, and both are
reached only after local lease credit and any elastic fallback have failed:
DenyReason::BalanceExhausted: the allocator confirmed that no account funding remains, including units held in leases. Callers need a top-up, a changed funding policy, or the next budget period.DenyReason::BalanceInsufficient { remaining }: the account still hasremainingunits of funding, counting units held in leases, and this request's quote is larger. Retrying this quote needs new funding or the next period; a request that quotes at mostremainingmay still succeed.
The pricing example returns HTTP 402 with balance-exhausted or
balance-insufficient for these, and 429/503 for the lease refusals. A service
that exposes a reset time can use the budget period end.
remaining is an upper bound on what the account can spend, never an estimate
to show as a balance: unreported usage can only lower it. A quote within it is
not a promise of admission. Such a quote keeps the lease refusal's transient
advice (LeaseUnavailable, LeaseExpired, LeaseExhausted), because funds
held by another instance can return when that instance releases its lease. Usage
is batched, so evidence can lag consumption until billing has recorded it and
the background manager has talked to the allocator again. Do not turn an
estimated remaining balance or a snapshot's old budget view into a funding
refusal.
Evidence comes from the allocator in two ways. Every grant carries the ledger's
remaining funding as of the transaction that made it. A refusal with nothing
allocatable carries it as AllocateError::BalanceExhausted (zero remaining) or
AllocateError::BalanceInsufficient (held in other leases).
InsufficientBalance remains the refusal that attests nothing. A consolidation
whose rolled-back settlement would have recorded loss or expired allowance
returns it. So do custom allocators that cannot establish the ledger fact.
Evidence is shared by all principals and localities of an account. A new grant
replaces it with that grant's own evidence. Accepted snapshots clear it when the
budget view or enforcement mode changes. So does a grant whose allocator call
overlapped such a change. Snapshot maps order generations separately for each
principal: an accepted change clears account evidence even if its generation is
below another principal's. Rejected replays and unchanged funding do not clear
it. A stored period end expires the evidence using caller-supplied admission
time, even if the allocator response arrives after rollover. Unscheduled
balances retain evidence until a funding observation replaces or invalidates it.
Top-ups propagate through the normal grant and snapshot loops, not
synchronously. A funding refusal still rings the lease's refill doorbell, so the
next consolidation carries fresh evidence. Refundable pending elastic
reservations and an in-flight overage commit retain their transient and
AfterInFlight advice, because they can recover without new funding.
Contract changes (GL-130)
LeaseAllocator::acquireandconsolidatereturntollgate_store::Allocation { grant, funding }instead of a bareLeaseGrant.funding: Nonemeans the allocator attested nothing. Custom allocators may returnNoneuntil they can read committed ledger state inside the grant's own transaction.- New enum variants:
AllocateError::BalanceInsufficient(BalanceShortfall)andDenyReason::BalanceInsufficient { remaining }. Exhaustive matches need them. - Dense counters append
balance_insufficient: deny slot 23 and allocator slot 10. Existing slots do not move. FundingAttempt::exhaustedis nowshortfall, and there is a newgranted.LeaseSlot::balance_exhaustedis nowfunding_evidence, which returns the evidenced remaining.ProblemandApiErrorgain an optionalbalance_shortfallfield, so Rust struct literals need it.
The wire change is backward compatible in both directions and needs no ordering or migration:
- Grant responses flatten the grant and add an optional
fundingobject. Old clients ignore it; old servers omit it, and new clients then have no evidence. - A confirmed shortfall keeps the
insufficient-balancecode and adds abalance_shortfallextension, whoseperiod_endis required and may be null. Old clients read the unattested refusal they always did. - New clients discard malformed or zero-remaining extensions, and grant evidence below the grant's own units, rather than letting them refuse fundable quotes.
The exhaustion code from GL-128, balance-exhausted, still wants clients updated
before servers: a client older than that code reads it as an unknown storage
error.
Local evidence publication uses a mutex only in the control plane. Admission reads one atomic word on a stable funding refusal when no evidence is live. Live evidence is read as a seqlock, four more loads, with no allocation, lock, I/O or clock read.
Grant sizing and large quotes
GrantPolicy halves grants near exhaustion by default (shrink_divisor = 2),
so one instance cannot hoard a small balance ahead of demand. A quote larger
than a lease the policy would grant is still reachable. The lease records the
largest quote it refused for want of units. The refill plane consolidates the
lease, and the allocator grows the replacement to that quote when the account's
restored balance can fund it. The first such request is refused with transient
LeaseExhausted advice, and a retry after the consolidation is funded. Growth
is at most one refused quote. A quote the account cannot fund does not grow the
grant; GL-130's evidence answers it instead.
Consolidation is the safety net, not the steady state. A rising consolidated
count means target_grant is undersized against the largest quote the service
prices.
Contract changes (GL-131)
LeaseAllocator::consolidatetakesneeded: CostUnitsafterrequested: the largest quote the returned lease refused. Custom allocators should size withGrantPolicy::consolidation_grant; passing zero keeps the earlier sizing.LocalLease::largest_refused_quoteexposes the recorded demand, which is exact at quiescence.ConsolidateRequestgains an optionalneededfield, omitted when zero. Old servers ignore it and keep the earlier sizing; old clients omit it. No rollout order or migration is needed.
Crashes and forfeited leases
Only a release returns a lease's unspent units. An instance that dies without
a graceful shutdown (SIGKILL, host loss, OOM, preemption without drain) never
releases. Its committed-but-unflushed usage also dies with it, so nobody can
prove any of its units unspent. After expires_at + reclaim_grace, server
maintenance sweeps such a lease and forfeits its whole unaccounted remainder
as settlement loss: nothing is credited back, and executed work cannot become
spendable again (INVARIANTS.md GL-9, GL-136). The same applies to a lease a
graceful shutdown abandoned because its release deadline lapsed.
- Cost: a hard kill forfeits the unspent remainder of every lease the
instance held, parked ones included. Size
target_grantandlease_ttlwith that in mind: smaller, shorter grants forfeit less and rotate more. - Late usage is still billed. An instance that outlives a control-plane outage past its lease's grace flushes its backlog afterwards, and those events bill against the forfeit instead of being rejected. Usage beyond what was forfeited is rejected and reported.
- Graceful drain matters.
InstanceRuntime::shutdownreleases every quiesced lease, which returns its units. Giveshutdown_release_deadlineenough room for your control plane.
Contract changes (GL-136)
ReclaimedLease.reclaimedis renamedforfeited, in Rust and in the JSON thatPOST /v1/leases/reclaimreturns, because its meaning changed: units recorded as loss, not units credited.- The server's sweep event reports
forfeited_units(wasunits, orreclaimed_unitson failure) and iswarnwhen it is non-zero. - No schema change or migration is needed. Leases swept before the upgrade keep their recorded credit and still reject stragglers, which is correct because they were credited.