2026-10-05 · 8 min read

The Breach Was Authorized: Stop Auditing Logins, Audit the Shape of Access

Flat isometric vault of indigo record blocks with one narrow violet path entering a green-marked gate and a wide grey fan-shaped arrow sweeping out, illustrating disproportionate access through a legitimate opening.

Denmark's national person registry, CPR, disclosed in early October that unauthorized parties had obtained names, addresses and personal ID numbers for roughly 8.8 million registered citizens. No vulnerability was exploited. No firewall fell. According to the registry's own statement, they got in by misusing a Danish company's lawful access to the system. They logged in exactly the way the system was built to let them log in.

Sit with that for a second, because it is the most common shape of modern data loss and almost nobody is instrumented for it. Your audit log will show a valid account, during business hours, from an expected address, hitting a supported endpoint. Every individual row looks clean. You could read a million of them and learn nothing.

A record tells you who was touched. It never tells you whether the access was proportionate. That is a different question, and it needs a different instrument.

Your audit log answers the wrong question

Most teams log access like this: actor X read record Y at time T. That is an accountability ledger, and it is genuinely valuable — after the fact, it is how you tell 8.8 million people what happened to them. As a detector it is close to useless, because it describes events one at a time in a world where the attack only exists in aggregate. One lookup is a lookup. Eight million lookups is a breach.

The signal is never in a row. It is in the set.

That has teeth. Is a bulk read by an authorized internal tool suspicious? It depends entirely on what that tool has ever needed before. So the first thing you need is not a threshold. It is a baseline of your own behavior.

Four signals you can compute this week

None of these require new infrastructure. All of them require deciding what "normal" means for each access path.

  1. Subjects per session. Count distinct records touched inside a single authenticated session — not per request, per session. A support agent looking up one customer touches one subject. A tool that suddenly touches 40,000 in one session is doing something its interface was never designed for.
  2. Ratio to history. Compare this session's subject count to the 90-day maximum for that same tool. Not to a global constant. A generic "more than 1,000 records" rule will either miss everything or page you constantly.
  3. Purpose binding. If your request carries a reason code, log the reason the caller claimed next to the reason the data implies. "Customer support ticket 8812" is a reason. Touching every record in a postal code is not that reason.
  4. The derivative, not the level. Alert on the change in rate. A registry that normally serves 2,000 lookups an hour and is now serving 90,000 did not gradually become more popular. Watch the slope.

Give every access path a budget, not a rule

Here is a framework that survives contact with real traffic. For each tool or service that can reach personal data, pull three numbers from the last 90 days: median subjects per session, 99th percentile, and all-time maximum. Then set the budget at roughly three times the maximum.

  • Support lookup tool: median 1, p99 4, all-time max 11, so budget 33. Crossing it stops the session and asks a human.
  • Fraud investigation console: median 12, p99 60, all-time max 300, so budget 900. Same mechanism, completely different number, because its job is different.
  • Annual reconciliation job: legitimately touches 400,000 subjects per run. It gets its own budget, a named owner, a scheduled window, and its own identity — not the credential a support agent uses.

Notice what this does. It stops pretending one number can govern everything. The tool that has never needed more than eleven records in a session now cannot quietly fetch eleven thousand. The reconciliation job is untouched, because someone wrote down that it is supposed to be huge.

And notice the enforcement point: the session, not the request. Per-request limits are easy to stay under and impossible to reason about. Per-session budgets match how a human or an automated job actually behaves.

The trade-off you have to accept

Hard blocks break legitimate work, and broken legitimate work creates pressure to delete the block. The person who gets stopped at 2 a.m. mid-incident will find the exception list, or the read replica nobody instrumented.

So make the control noisy but passable. When a session crosses its budget, do not silently kill it and do not silently allow it. Stop it, log it, and offer a break-glass path that requires a name, a reason string and a one-time approval — then tags the rest of that session as elevated in the ledger. You keep the work moving, you keep the signal, and you convert an invisible bypass into a visible, attributable event.

The honest cost: you will get false positives and someone will be annoyed. That is the price of a control that can actually see the shape of access. If your tool has a static threshold nobody has tuned in a year, you either have no false positives, or nobody is watching the alerts.

Design the interface so the attack cannot be expressed

The strongest version of this is not a monitor. It is an API that only knows how to answer narrow questions.

If a support tool genuinely needs one customer's record, ship it a lookup that takes one identifier and returns one record. Do not hand it a "search" or "export" surface and then try to detect abuse of that surface. Interfaces that cannot express a full-registry scan do not need alerting, because the risky operation does not exist.

This runs into the wall you would expect: analysts and data teams need bulk views, and they will build them somewhere. The answer is to give bulk work a separate identity, a separate budget, a scheduled window and a named owner, instead of letting it ride along inside the everyday tool. You are not banning the capability. You are making it a decision with a signature on it.

What to ask in an interview (and how to answer)

If you are hiring for backend, platform or security roles, "how would you detect an authorized account doing unauthorized volume?" separates people who have operated systems from people who have read about them. Weak answers reach for MFA and quarterly role reviews — both of which would have done nothing in the Danish case, because the access was legitimate. Strong answers ask what the tool's baseline is, and what happens at three times that.

If you are the one answering: say "shape, not rows." Then say what you would measure, what you would set the budget to, and what happens when it is crossed. Three sentences. It is a genuinely hard problem, and three specific sentences are worth more than a paragraph of general.

The uncomfortable part

That same ledger is your legal and moral answer sheet. When the CPR administration found the problem, it cut the company's access and handed the matter to the data protection authority and the police. The first question any regulator, journalist or affected citizen asks is always the same: which records, exactly, and whose?

If you only keep aggregated metrics — "audit events: 4,112,908" — you cannot answer it. A second decision hides inside the first: retain per-subject access rows, at least for personal data, for longer than you retain application logs. They are small. They are boring. And they are the difference between "we think about 8 million people" and a letter with a name on every page.

Key Takeaways

  • Most large data losses ride on legitimate access. If your detector only fires on bad credentials or odd IP addresses, you are blind to the common case.
  • Log per-subject access events, then compute subjects per session. The anomaly lives in the set, not in any single row.
  • Set a budget per access path from that path's own history — roughly 3x its 90-day maximum. One global threshold is either noise or nothing.
  • Enforce at the session level with a break-glass path that requires a name and a reason. A control people can route around is a control you do not have.
  • Prefer narrow interfaces. A lookup that takes one ID and returns one record cannot be abused into a full-registry sweep.
  • Keep the per-subject ledger for personal data longer than your application logs. It is how you answer "who was affected" in a week rather than a quarter.

You can read the CPR administration's own statement here. It is four short paragraphs, and not one of them describes a technical failure.

[ Call to action ]

Ready to scale your engineering team?

Tell us the roles you need to fill and we'll get back within 24 hours.