My metrics service used to ask my authentication service for permission to accept metrics. Including metrics about the authentication service. If Auth was having a bad day, the machinery recording that bad day needed its cooperation. Very orderly. Not especially helpful.
I've written about that circular dependency before. The replacement is now deployed: machine publishing no longer calls Auth. A separate credential, checked by Metrics itself, does the job.
Apparently the department of counting things can be trusted with its own key.
A Small Division of Responsibilities
Each host runs one local agent that collects events from its applications and ships them to Metrics over HTTPS. The agent holds a randomly generated publishing credential. The applications don't need that credential, or even the address of the collector.
Metrics owns the credential registry and checks it locally. The keys can be rotated or revoked independently of user accounts, and they grant ingestion access only: no reading dashboards or administering the service. The old anonymous publishing route is gone.
This trusts the enrolled hosts. Possession of a key does not prove which application produced a point, and authentication doesn't limit how much work a caller can create. Batch-size, rate, and distinct-series limits are separate controls.
The regression test makes Auth unavailable before Metrics starts, then submits a batch using a valid machine credential. It checks that the batch is accepted without an authorization callback. This runs in isolation, not by sabotaging my own login service for the sake of a test.
When Does "Sent" Mean Sent?
The delivery side needed a similarly specific contract. Calling emit() does
not mean the collector has received anything. It means the application has
handed a point to its local producer, which queues it in memory. A background
thread passes it to the agent.
That first queue holds at most 256 points or 256 KiB. Under pressure it evicts its oldest entries. Before the agent commits an event to its persistent queue, crashes can lose it. I'd rather lose some telemetry than let the machinery counting requests eat the application serving them.
After that commit, the agent retries within explicit capacity limits and a one-hour lateness window. Each event keeps the same ID across retries. Metrics commits the event and its receipt together before acknowledging it, so a lost reply doesn't mean counting the same event again on retry.
That is a useful guarantee. It is not a promise that every emitted point will arrive. This is telemetry, not the place I'd keep a billing ledger.
Is It Alive, or Just Catching Up?
There is one more distinction I wanted the monitoring to make. A recovering queue delivering old authorization events tells me something about delivery. It doesn't tell me whether Auth is working now. Neither does a quiet graph: perhaps nobody has asked it to do anything.
Auth's availability alarm now uses a direct readiness check instead of missing authorization metrics. It checks whether the service responds and can read its expected database schema. Historical events cannot satisfy it. The pipeline separately tracks publisher receipts, queue age, and aggregation progress.
Human dashboards still require Auth, and alert email still has an auth dependency. External monitoring remains necessary. I haven't made the whole system independent; I've removed the authorization callback from ingestion and stopped treating request activity as the availability check.
The machines can now file their complaints without getting the subject's signature.