I'd recently written my own authentication service and my own metrics service, rather than using anything off the shelf. Every login and every data point went through code I'd written and understood completely.
Which is how I managed to make them wait on each other.
Revised September 10, 2026. This started as an August 2025 debugging story. The later workarounds and eventual replacement are dated below; they were not all part of the original fix.
The Guard Who Locked His Keys in the Office
I wanted my shiny new Authentication Service to emit metrics. How many login attempts per minute? How long does token verification take? Simple questions. My Metrics Service was ready and waiting.
The intended flow was straightforward: Auth handles a login, sends a metric to Metrics, and responds to the user. But Metrics didn't accept writes from just anyone. It called Auth to check the incoming request's token.
You see the loop:
- Auth handles a login and calls Metrics before responding.
- Metrics calls Auth to authorize the write.
- That authorization request is itself instrumented, so Auth tries to send another metric. Back to step two.
With synchronous calls, that can tie up the available request workers waiting for work that needs those same workers. More workers can postpone the jam; they don't remove the feedback loop.
I had built a security guard who needs to swipe his ID to get into the building, and whose ID card is on his desk, inside the building.
Two Ways to Let the Guard In
The Trusted Fortress
My first idea was to let requests from the private network bypass the auth check. That would remove the callback, but it would make network membership the permission to publish. I didn't want moving a service to another machine or network to mean redesigning that permission rule.
A network allowlist can be useful. It just wasn't the identity boundary I wanted for this service.
The Skeleton Key
The next idea was a separate secret that Metrics could check for itself, without consulting Auth. I called it a skeleton key, which probably tells you how warmly I received the idea. Another credential to distribute, another place to revoke it, another exception to my nice centralized system.
But "separate" does not have to mean "all-powerful." A credential that can only publish metrics need not grant access to dashboards, user accounts, or anything else. The operational work is real; it isn't a reason to treat local credential verification as inherently worse than a live authorization call.
That distinction took rather longer to settle than the nickname did.
Let Him Sign the Logbook Later
The part I could change immediately was the waiting.
Auth didn't need confirmation from Metrics before answering a login request. It could record the event in memory and let a background thread deliver it. Change the guard's job description: he signs a logbook, and someone collects it later. He does not stand at the door waiting for a receipt.
The first shared publisher flushed periodically, or when enough metrics had accumulated. The HTTP work moved off the request thread.
That removed the synchronous wait for delivery. It did not remove the need to authorize the eventual write, or the metrics generated by doing so. During the early work I also temporarily disabled protection on the publishing APIs. That was a workaround, not a property that asynchrony supplied.
Walk the feedback loop with buffering in place. Auth handles a login and buffers a metric. Ten seconds later, the publisher sends it to Metrics. Metrics calls Auth to authorize the batch. Auth records having served that authorization. Ten seconds later, the publisher wakes up again.
I would have traded an infinite loop for an infinite loop with better manners. Two services could still spend their time discussing the fact that they were talking to each other.
Moving a call off the request thread and removing a dependency are different changes. I had done the first.
A Buffer Is Not a Delivery Guarantee
The original publisher also did less for outages than I gave it credit for.
Its MAX_BUFFER_SIZE was 500, but that was a trigger to flush, not a hard limit
on how much the process could hold. New events could accumulate while the sender
was busy. Giving a setting an imposing name does not make it enforce anything.
At flush time, the publisher removed the batch from memory before attempting the HTTP request. A network exception meant losing that batch; the code didn't put it back. It didn't inspect unsuccessful HTTP responses either. A process crash could lose whatever hadn't been sent yet.
So the useful claim was that a login no longer waited for the metrics HTTP call. It was not "zero impact during a metrics outage," and it was not "buffer until recovery." Memory pressure, delivery failure, and shared machine resources still existed. This was best-effort telemetry, not an audit log.
May 2026: The Mail Slot
A later workaround gave Auth its own publisher and a dedicated ingestion route.
That route skipped the authorization callback, accepted only names beginning
with AuthService., and had no request-metrics decorator of its own. Recording
an auth metric no longer required another auth decision or generated another
metric about recording it.
That cut the recursive path. It also accepted writes without a credential.
I described the namespace restriction as a mail slot rather than a back door. There is something to that: it didn't grant dashboard access or let callers write under unrelated metric names. But it did not prove that the caller was Auth. The route was reachable through a public hostname; "internal" described a path around the CDN, not a private network boundary.
And a prefix is not a resource budget. A caller could submit nonsense under that prefix, generate many distinct metric names or attribute combinations, or send requests fast enough to consume storage and processing time. Restricting where a write lands does not bound how much can be written.
So "my login-count graph might be wrong" understated the risk. The route solved the recursion by removing a check. It did not solve ingestion security, and it was not safer merely because there was no secret to leak.
September 2026: A Key With One Job
The replacement, deployed on September 10, uses independently generated machine credentials after all. Not a universal skeleton key: one ingestion credential per enrolled host, verified against a registry owned by Metrics. Those credentials cannot read or administer metrics.
Applications, including Auth, hand their events to a shared local agent. The agent owns the publishing credential and sends batches over HTTPS. Metrics checks the credential locally; accepting a metric no longer requires an Auth authorization request. The old anonymous route is retired.
The replacement also has explicit queue budgets, delivery acknowledgments, and separate resource controls. It remains best-effort telemetry before durable storage, and dashboards and alert email still depend on Auth. I've written a short follow-up on what the new pipeline promises, including the regression test that starts Metrics with Auth unavailable.
The guard still needs his ID to enter the building. Reporting that he's locked out no longer requires him to get inside first.