Control-plane security
Usage ingestion also derives non-authoritative credential activity from accepted
events. Its instance role and secured transport are unchanged. A publisher's
optional key ID must match the snapshot principal/account or publication returns
422 invalid-credential-binding. Unknown attribution at ingestion preserves the
bill and is reported; activity grants no lifecycle or authorization capability.
See Credential activity for coverage and mixed-version
reporting. There is no additional activity HTTP endpoint.
tollgate-server authenticates every lease, snapshot, credential, usage, and administrative
route. /livez and /readyz accept unauthenticated probes. Readiness checks
storage reachability and the server-owned maintenance task's latest reclaim and
rollover outcomes. It does not certify that every configured identity is usable.
Monitor security reload warnings separately.
The independently owned reloader reports unexpected task exit at error with
operation=security-reload and reason=unexpected-exit; restart its owner to
restore refresh. The last valid configuration and signing-key expiry retain
their existing authority. Deliberately dropping the owner is not a failure alert.
Server configuration
Lease TTL compatibility
Both /v1/leases/acquire and /v1/leases/consolidate accept positive whole
seconds as before ("ttl_seconds": 60). A fractional or wider duration uses
an exact Jiff duration string with a zero compatibility sentinel, for example
"ttl_seconds": 0, "ttl": "PT0.5S". A non-null ttl requires the sentinel
to be zero; two nonzero declarations are refused with 422 invalid-ttl, even
when they express the same duration. Malformed fields return invalid-json.
The server's allocator continues to apply its configured maximum TTL and
supplies the authoritative expiry in its response.
Upgrade every control-plane server before enabling fractional or greater-than-
u32::MAX-second TTLs in HTTP clients. Whole-second clients remain compatible
throughout the rollout. An older server ignores ttl and rejects the zero
sentinel with invalid-ttl, so a mixed deployment cannot silently grant a
different lifetime. Before rolling servers back, stop clients from issuing
the precise form; otherwise those lease operations refuse until compatible
servers return. Existing grant responses are unchanged. PostgreSQL's durable
nanosecond timing requires the coordinated migration 0017 upgrade described in
Lease timing; that storage upgrade has its own compatibility
fence and must precede resuming lease operations on the new backend.
Rust callers constructing AcquireRequest or ConsolidateRequest must replace
the ttl_seconds member with ttl: LeaseTtl::try_from(duration)?. This is a Rust
DTO source break. HttpStore's public allocator signatures and valid runtime
configuration remain unchanged; nonpositive durations return InvalidTtl
before HTTP I/O. The invalid-ttl code remains stable; its title says that a
TTL must specify one positive duration.
Security manifest
Set TOLLGATE_SECURITY_CONFIG to a JSON manifest. There is no default credential,
including on loopback. Relative paths resolve against the manifest's directory.
Unknown fields, duplicate credential mappings, malformed credentials, and invalid
TLS material reject the initial configuration. Failed reloads preserve the last
valid generation.
{
"tls": {
"certificate": "server-chain.pem",
"private_key": "server-key.pem",
"client_ca": "instance-ca.pem"
},
"bearers": [
{"identity": "deployment-operator", "role": "operator", "token_file": "operator.token"},
{"identity": "signup-service", "role": "provisioner", "token_file": "signup.token",
"max_budget_allowance": 100000, "allowed_policy_templates": ["standard"]}
],
"certificates": [
{"identity": "instance-service", "role": "instance", "certificate": "instance-leaf.pem"}
],
"google": {
"audience": "https://tollgate.example.net",
"subjects": [
{"subject": "123456789012345678901", "identity": "pricing-api", "role": "instance"}
]
},
"issuer": {"secret_file": "issuer.secret"},
"policy_templates": {
"standard": {
"cost_table": {"fixed_request": 1, "minimum_charge": 1, "weights": [1]},
"limits": {"max_items_per_request": 100, "rate_units_per_second": 1000,
"rate_burst_units": 1000},
"permissions": 1
}
}
}
The subject above is illustrative; replace it with the service account's numeric
unique ID. google, tls, client_ca and issuer are optional; bearers and
certificates default to empty lists. An empty role map intentionally denies all
protected operations. Keep the TLS block when withdrawing all identities from
an encrypted listener.
Use separate instance, operator and provisioner identities. Roles are disjoint:
| Role | Routes under /v1 |
|---|---|
instance | POST /leases/{acquire,release,consolidate,reclaim}, GET /snapshots, GET /snapshots/{principal}, GET /keys, POST /usage/ingest |
operator | POST /admin/accounts, GET /admin/accounts/{id}, POST /admin/accounts/{id}/{deposit,status,capacity-class}, PUT /admin/accounts/{id}/budget, POST/GET /admin/accounts/{id}/keys, DELETE /admin/accounts/{id}/keys/{key}, PUT/DELETE /admin/accounts/{id}/keys/{key}/snapshot, PUT /admin/snapshots/{principal}, DELETE /admin/snapshots/{principal} |
provisioner | every operator route except POST /admin/accounts/{id}/deposit and /admin/snapshots/{principal}, within the provisioner scope |
Give an internet-facing account service a provisioner, never an
operator. A provisioner creates only unfunded, suspended, best-effort
accounts; activates them; sets budgets up to its max_budget_allowance; and
manages credentials and strict key snapshots — all on accounts a provisioner
created. It cannot deposit, suspend, close, grant Assured, publish a
principal snapshot, reach an account an operator created, or lift a status an
operator set. Compromising it therefore cannot fund an account or undo an abuse
suspension. Keep operator for people and trusted tooling.
max_budget_allowance and a nonempty allowed_policy_templates list are
required on every provisioner bearer, certificate or Google subject entry.
Both fields are refused on other roles. Names resolve in the manifest's
policy_templates map; a missing name, missing/empty allowlist or invalid
template rejects startup. Failed reloads retain the last valid generation.
A provisioner's key snapshot must match one complete approved template's
cost_table, limits, permissions and policy_revision. Revision defaults
to the unstated value when omitted. Equality compares the typed values,
including every operation weight/permission and legacy rate fallback; JSON
object order is irrelevant. Templates cannot be combined field by field.
Strict enforcement remains mandatory. A mismatched snapshot receives audited
403 scope-forbidden before any store read or write. Operators remain
unrestricted by these templates.
Templates exclude account/key binding, status, capacity class, valid_until,
generation and budget. Existing store checks own binding/status/class;
provisioner generations and budget views are assigned by the store. The
provisioner chooses snapshot validity. Template changes are part of the normal
atomic security-generation reload and constrain requests authenticated after
installation. Already authenticated requests retain their captured generation;
already published snapshots remain valid until expiry, replacement or
withdrawal. Revoke or replace existing snapshots separately when withdrawing
an entitlement; changing an allowlist is not retroactive revocation.
For an existing provisioner deployment, stage the templates and identity
allowlists before starting the upgraded server; there is no allow-all fallback
or database migration. Replace/drain every old server before resuming
provisioner traffic. An old binary rejects the new manifest and can retain its
previous permissive generation on reload, so a mixed fleet does not enforce
this restriction uniformly. Instance/operator-only manifests are unchanged.
Library embedders pass the approved policies to the now-fallible
ProvisionerLimits::new; an empty set cannot create an identity.
A provisioner reaches only accounts whose origin is Provisioner. Every
account that existed before the role did is an operator's, so a signup service
moved from an operator credential to a provisioner cannot administer the
accounts it created under the old credential. Keep administering those with
operator tooling, or re-create them through the provisioner.
GET /keys serves customer credential digests from the store's KeyDirectory.
That key space is separate from this server's control-plane bearer credentials,
which are loaded from the security manifest into their own HmacRegistry. The
route sends neither customer secrets nor the HMAC secret already held by a
customer verifier. See credential projection for
paging, clock, freshness and rollout contracts.
Static credential files contain 32–16,377 visible ASCII bytes; a trailing newline
is accepted. Generate at least 256 random bits and store them using a secret
manager or protected mounted file. Only HMAC digests remain in the loaded static
verifier. Identities are non-secret audit names, limited to 128 ASCII identifier
characters. Raw tokens, authorization headers, private keys, and JWT bodies are
not logged. Snapshot publication logs the affected principal at debug, so do
not run RUST_LOG=debug in production: principals are digest material that no
administrative response or audit event discloses.
Credential issuer
The optional issuer entry lets this server mint customer credentials:
POST /admin/accounts/{id}/keys answers 501 issuance-unsupported without it.
See account administration for the routes.
- Format. The file holds exactly 64 lowercase hexadecimal characters, with
one optional trailing newline (
openssl rand -hex 32). Anything else, including uppercase, rejects the configuration. - The HMAC key is those 64 characters as bytes, not their decoded value.
Give every verifier (
tollgate_client::KeyManager::spawn) the same 64 characters without the newline, so one stored value configures both sides. Uppercase is refused because under this rule it would be a different key. - Distinct authority. The issuer secret must differ from every bearer token in the manifest; a match rejects the configuration. It is unrelated to the per-start registry that verifies control-plane bearers.
- Fixed for the process lifetime. The issuer is read when the server starts.
A reload whose manifest names a different issuer — changed, added or removed —
still applies TLS, bearer, certificate and Google changes, keeps the running
issuer, and logs
reason="issuer-change-requires-restart"once per distinct staged value. An invalid issuer file fails the reload like any other malformed file, and the previous generation stays live. - Neither the secret nor anything derived from it is logged, returned or audited.
Rotating the issuer secret invalidates every credential issued under the old one, because a credential verifies only under the secret that minted it and a verifier holds exactly one (see credential projection). There is no multi-secret overlap. To rotate: write the new secret, restart the server, restart every verifier with the same new value, then re-issue each account's credentials and revoke the old ones. Plan the rotation as a credential reissue, not a transparent reload.
The server accepts a single Authorization: Bearer … header, bounded to 16 KiB.
Bearer verification uses tollgate-auth::CredentialVerifier. The configured role
map authorizes the verified principal. Invalid, missing, expired, or conflicting
credentials return RFC-7807 401 authentication-required with a Bearer challenge;
valid identities without the required role return 403 scope-forbidden, and
the refusal is written to the audit log. These checks precede path/body
decoding and store calls. Existing JSON errors and wire
DTOs retain their meanings.
TLS uses rustls with safe protocol defaults. TOLLGATE_BIND defaults to
127.0.0.1:8080; a non-loopback bound address requires TLS and fails startup
without it. Network restriction alone does not permit remote plaintext. The
library serve function enforces the same rule. Embedders must use this listener
entry point; the standalone router is useful for in-process tests and cannot
validate an external listener it does not own. Its /readyz returns 503 because
it has no maintenance owner; /livez and authenticated API handlers still work.
An optional client CA requests and verifies client certificates while allowing
bearer-only callers and probes. A trusted certificate also needs an exact leaf
SHA-256 fingerprint mapping to gain authority; the manifest derives that
fingerprint from the configured leaf PEM. HTTP headers such as
X-Forwarded-Client-Cert never create an identity. If bearer and certificate
credentials are both supplied, they must resolve to the same name, role and
provisioner ceiling.
TLS handshakes run concurrently, with at most 128 pending tasks and a five-second
deadline each. Excess connections wait in the OS backlog. Library embedders may
set these bounds with TlsConfig::with_handshake_limits; the binary uses the
defaults. Dropping the listener aborts pending handshakes. SIGTERM and Ctrl-C
start graceful HTTP shutdown; cancelling the server also aborts its maintenance
task. This does not introduce a total deadline for all server HTTP requests.
The caller's shutdown future remains owned by serve. Cancellation or task
failure also releases its internal signal waiter and tells existing connections
to shut down.
Maintenance readiness and recovery
serve starts maintenance immediately. /readyz stays 503 until both the reclaim
and budget-rollover passes reach a partial batch successfully, and returns 200
only while their latest outcomes are successful, the task is alive and the store
answers ping. A failed pass withdraws readiness immediately, even if the other
pass succeeds. Successful recovery of the failed operation restores its own
health; /livez remains 200 while the service can run. The handlers continue to
return their normal domain results while maintenance retries on its existing
cadence. No partial batch or failed call is treated as proof of rollback.
Each operation logs its first two consecutive failures at warn, the third and
subsequent failures at error. Monitor operation (reclaim or budget-rollover)
and consecutive_failures; one info event reports after_failures on recovery.
The three-attempt threshold is an alerting policy, not proof that a particular
lease TTL has expired. Inspect backend permissions, connectivity and transaction
health; successful ping alone does not prove that maintenance writes can run.
Completed progress fields survive a later failed batch. No backend details are
formatted into these events.
A successful sweep that settled leases logs leases, forfeited_units and
batches. A swept lease is one its holder never released, so its remainder is
forfeited as settlement loss rather than returned to the account (GL-136). The
event is warn when forfeited_units is non-zero: an instance crashed, or its
shutdown's release deadline lapsed. Late usage for a forfeited lease is still
billed against the forfeit. The failure events carry forfeited_units for the
batches that committed.
An unexpected maintenance return, cancellation or unwinding panic emits an
error with operation=maintenance and a static reason, then serve returns
an I/O error and stops listening. Process supervision should restart it. A build
using panic=abort terminates the process immediately on panic; it cannot emit
the unwind supervisor's event. Graceful shutdown withdraws readiness before
cancelling maintenance, and expected cancellation is not a task-failure alert.
Pending store calls may have committed effects even when cancelled; reconcile
with backend records rather than inferring rollback from shutdown.
This health policy observes completed outcomes and task liveness. It adds no
timeout or cancellation deadline to a backend call that remains pending, and it
does not certify that SKIP LOCKED left no rows with another replica. Backend
futures must yield to the executor for task supervision and cancellation to run.
There are no Rust signature, database or wire-schema changes. Readiness is
intentionally stricter: deployments should allow initial maintenance to complete
before routing traffic. Embedders using the bare router for a listener must move
to serve; their probe now stays 503 instead of claiming health without an
owned worker. Existing TOLLGATE_RECLAIM_INTERVAL_SECS retains its scheduling
meaning. Zero or a duration that cannot fit the monotonic clock is rejected
before starting tasks. Rolling back restores the old false-ready behavior.
Instance shutdown accounting
The instance runtime has the same allocation boundary. A successful
InstanceRuntime::shutdown can leave a grant whose acquire or consolidation
result never reached the manager. Its known leases can all be released while
that unanswered capability still holds units in the backend. Retain a
RuntimeHandle and inspect report().uncertain_acquires and
account_reports(now) after the join, alongside the shutdown report's abandoned
leases, task failures and usage-drain counters. Uncertainty is a count of
possible grants, not a unit amount or proof that each call committed. Such units
remain in active grants until server maintenance sweeps them after expiry and
grace, when they are forfeited as settlement loss rather than returned (GL-136).
Do not credit them manually: nobody can prove them unspent. Nor should you
classify the liquidity difference alone as lost billing, because late usage for
the lease still bills against the forfeit. Reconcile recorded usage and ledger conservation both before
and after reclamation. Expiry does not erase the runtime's historical counters.
Instance clients and Google service-account identity
HttpStore::with_config validates its URL, credentials, TLS roots and deadlines
before construction. HTTPS verifies the server certificate and host name. A
custom root_ca_pem replaces public trust roots; identity_pem contains the
client certificate chain followed by its private key. StaticBearer supplies a
rotatable token. HttpStoreConfig defaults to a two-second connection deadline
and a ten-second total request deadline, including credential retrieval and
response reading. Redirects and environment proxies are disabled. Plain HTTP
is accepted only for literal loopback addresses or localhost, which is pinned
to loopback rather than resolved through DNS. URL userinfo, queries and fragments
are refused.
Instances running on Google Cloud can use their attached service account, through the platform's metadata server, without distributing per-replica secrets:
use tollgate_client::{GoogleIdentity, HttpStore, HttpStoreConfig};
let identity = GoogleIdentity::new("https://tollgate.example.net")?;
let store = HttpStore::with_config("https://tollgate.example.net", HttpStoreConfig {
bearer: Some(identity),
..Default::default()
})?;
The provider obtains a Google ID token from the fixed metadata identity endpoint
with the configured audience and Metadata-Flavor: Google. It caches tokens for
60 seconds and bounds metadata requests to five seconds and 16 KiB. The server
accepts only RS256 signatures under Google's published keys, Google issuer names,
the exact configured audience, a nonempty subject, and unexpired verified evidence.
Authorization uses sub, not a mutable email claim. This is service-account
identity support; customer login, arbitrary OIDC issuers and user directories
are outside this API. Replicas sharing a service account share one identity and
revocation scope.
The server fetches Google signing keys off the handler path. Their usability is
bounded by the issuer's Cache-Control/Age policy and at most one hour after
fetch; absent cache metadata defaults to five minutes. Refresh is attempted at
half the remaining lifetime, between five seconds and five minutes. Fetch or
validation failures retain the preceding keys with their original expiry and
never extend identity validity. Expired or unknown keys fail closed until a
valid refresh succeeds. Internet access to Google's signing-key endpoint and
accurate server time are deployment prerequisites for this mode.
The supported topology has clients connecting over TLS to a directly encrypted server endpoint, with PostgreSQL behind the server. The server must therefore run where it can expose its own TLS listener. A platform front end that terminates TLS and forwards plaintext is not an exemption from the non-loopback plaintext rule; this feature does not provide a forwarded-identity or trusted-proxy bypass.
References: Google ID tokens for service accounts, Google ID token validation and key caching.
Rotation and revocation
- Write new versioned credential/certificate files and validate them. For CA changes, allow both old and new CAs and certificate identities during overlap.
- Atomically replace the manifest only after all referenced files exist. The
owned
SecurityReloaderchecks every five seconds, stages the complete set, then publishes verification, role mapping and TLS as one immutable generation. - Rotate client bearers with
StaticBearer::replace, or roots, mTLS identity and bearer provider together withHttpStore::reconfigure. Existing managers keep the sameArc<HttpStore>; new calls use the replacement. Invalid replacements preserve the working transport. Embedders own client file watching. - Remove retired mappings/roots after clients have moved. Monitor rejected calls
and reload warnings.
StaticBearer::revokeprevents subsequent client calls from silently becoming anonymous.
New requests, including those on keep-alive connections, use the current role map and recheck a presented certificate against the current CA and verification time. Requests already authorized retain their pinned generation. New TLS handshakes use the configuration current at connection acceptance. Existing TLS connections do not renegotiate the server certificate; clients needing immediate server-trust withdrawal must replace their transport. Enabling or disabling TLS requires a listener restart, so a reload cannot expose an encrypted service as plaintext. Changing the credential issuer requires a server restart; a reload defers it without blocking the rest (see credential issuer).
Usage ingestion treats authentication failures as retryable. Rotation can pause billing delivery but must not turn a valid usage batch into a terminal refusal. Existing writer buffering, backpressure and shutdown bounds still apply; a long credential outage eventually denies admissions rather than losing usage silently.
Backend failures and diagnostics
Backend errors carry arbitrary text. The server does not expose that text in
HTTP bodies, ApiError debug output or its backend-failure logs. Public storage
failures use 503 / storage / backend unavailable; permanent usage refusals use
422 / usage-refused / usage batch refused. Existing domain codes, generation
responses and retry decisions are preserved. Credential-page failures retain
their existing credential-source-unavailable code and title.
HTTP problem responses with 5xx or usage-refused add an optional error_id containing 32
lowercase hexadecimal digits. Find the matching warning on tollgate::diagnostics
to identify the route template, status and code. The server generates this ID;
request headers cannot choose it. It is diagnostic context, not authorization
evidence. If system entropy is unavailable, the failure keeps its original
status, omits the ID and logs error_id_unavailable=true.
Retain this target at warn or above, for example with
RUST_LOG=info,tollgate::diagnostics=warn. Library embedders own subscriber
installation and log delivery. Correlation is limited by that delivery; it is
not a durable record across process death. Problem's public Rust shape is
unchanged, and existing clients ignore the additive JSON field. HTTP consumers
that need the correlation ID can read error_id from the response object.
Readiness retains its empty 200/503 response and logs failures by operation. Maintenance logs retain static operation codes and completed progress counters; PostgreSQL startup logs identify initialization failure and the configuration to check. Neither connection strings nor driver text are logged. Inspect connectivity, backend health, migration status and appropriately protected backend operational records using the incident's time and operation. An error does not prove rollback: administrative audit receipts remain the authority for confirmed writes, and ambiguous failures need reconciliation.
This change requires no schema or configuration migration and preserves the public Rust error types. Deploying the server updates the public titles and adds the optional field. Consumers must classify errors by status/code, not by parsing the old backend-specific title. A rollback to an older server restores the disclosure defect.
Administrative audit
Every HTTP administrative operation reaching the store emits structured
tollgate::audit events with a random operation ID, stable actor, role, action, resource and server time.
started precedes the store call. confirmed carries the backend's typed
AdminReceipt: before/after values captured under the memory lock or inside the
PostgreSQL transaction that serialized the mutation. It never substitutes a
separate read that could describe somebody else's concurrent write.
Receipts identify changed fields: creation balance/status/class/origin,
deposited and top-up totals, status and the authority that set it, capacity
class, or snapshot generation plus revocation state. The resource identifies the account or principal. Snapshot receipts name
the immutable publication generation; they do not copy its full policy graph.
No-op operations report equal states. failed carries a stable code/status;
cancelled_unknown marks an interrupted operation. Neither invents a before/after
pair or asserts that a storage error ruled out a commit. Unmatched started
events after a process crash also require reconciliation.
A request refused for its scope emits one refused event and no started,
because it never reaches the store. It carries the actor, role, action,
resource, time and code. A credential on a route its role does not reach is
refused by the router: its action is the method and route template
(POST /v1/admin/accounts/{account}/deposit) and its code scope-forbidden. A
provisioner refused for an argument or an account is refused by the handler:
its action is the operation name (create_account) and its code
scope-forbidden or account-not-provisioned, with a reason. An operator
hold is found inside the store transaction, so it is a failed event with code
operator-hold. Alert on refused events from a provisioner: a signup
service that asks for a deposit is not behaving like one.
The binary keeps audit events enabled even with RUST_LOG=error; library
embedders must install a subscriber that retains this target. Route these events
to the deployment's retained, access-controlled audit log and alert on delivery
failures. This is structured audit logging, not a durable transaction outbox:
process death or logging infrastructure failure can lose delivery after a commit.
No database migration or audit table is introduced. Direct AdminStore callers
receive receipts but remain responsible for attaching identity and persisting
their own audit record.
Rollout and local development
This is an intentional Rust/configuration API break in unpublished crates:
ServerState requires security, HttpStore constructors return Result, and six
HTTP-facing AdminStore mutations return AdminReceipt<T> (read .outcome for
the previous result). Custom backends must capture receipts at their serialization
point. Wire payloads, existing domain error codes, and database schemas are
unchanged; 401/403 authentication codes are additive.
Bring up a secure endpoint against the existing PostgreSQL backend, configure instance identity and TLS, then deploy clients using the new configuration. Keep overlapping identities during rotation and remove the old endpoint once clients have moved. Rollback must preserve an authenticated TLS endpoint; reverting to an older unauthenticated server on a public listener is not a compatible rollback. No schema rollback is required.
For a disposable loopback demonstration only:
mkdir -m 700 .local-control
printf '%s\n' 'demo-only-operator-token-do-not-deploy-98' > .local-control/operator.token
printf '%s\n' '{"bearers":[{"identity":"demo-operator","role":"operator","token_file":"operator.token"}]}' > .local-control/security.json
chmod 600 .local-control/*
TOLLGATE_SECURITY_CONFIG="$PWD/.local-control/security.json" cargo run -p tollgate-server
This uses the ephemeral memory backend and grants no instance role. Supply an
independent instance credential to exercise HttpStore, or run the generated
TLS/bearer/mTLS fixtures in cargo test -p tollgate-server. Keep local security
material out of version control. --help and --version work without secrets or
database configuration; -- ends option processing.