Break-glass that never closes: Why emergency access becomes your real permission model

Break-glass access starts as a kindness to the future. Someone invents a fast path for the night when production is on fire and the usual approvals are asleep. A shared admin role, a standing exception, and a temporary vault that never quite expires. Then the outage ends, the glass stays open, and the exception quietly becomes how work gets done on ordinary Tuesdays. Mid-size teams invent emergency elevated access for real pain, then forget to treat closing it as part of the product.
This post is about owning break-glass as an operational surface. Named owners, time bounds, evidence, review, and a habit that closes the glass before emergency privilege becomes your real permission model.
Emergency access that outlives the emergency
Break-glass is supposed to be rare, loud, and short. In practice it becomes ambient. The on-call who needed production write access keeps the role "just in case." A vendor bridge account stays active after the ticket closes. A cloud organization's emergency group grows by one person every quarter and sheds none. None of those paths feel dramatic on the day they are created. Together they rewrite who can actually change the system.
In practice, the danger is not only abuse. It is unowned power. When elevated paths linger without a named owner, an expiry, and a review date, nobody can answer who may use them, for what, or how you would revoke them under stress. The formal model still looks tidy in the wiki. The real model lives in stale groups, sticky roles, and chat history.
That's why inventing another emergency role rarely helps. More glass means more edges that never close. Fewer, owned break-glass paths beat a maze of half-forgotten keys.
Why the glass stays open
Teams leave break-glass open for reasons that sound responsible. Closing it feels like inviting the next outage to hurt longer. Re-requesting access feels bureaucratic. Documenting the path feels optional once the pager is quiet. Meanwhile the path keeps working, so the incentive to tidy it is weak.
However, the social drift is worse than the technical one. When engineers learn that production privilege is one Slack message away and never expires, they stop investing in ordinary least privilege. When managers learn that break-glass is how releases land on Fridays, they stop funding better approval paths. Culture follows the path of least resistance. An open glass trains everyone to lean on it.
Still, you can name the usual leftovers without needing a forensic lab. Shared admin identities, standing "break-glass" groups with no membership review, long-lived personal access tokens minted during an incident, and cloud roles that grant far more than the outage required all tend to linger once the pager is quiet. If those leftovers have no owner and no expiry, they are not readiness. They are your permission model wearing an emergency badge.
Who owns break-glass as a product surface
Someone has to own the break-glass contract. Not "security in general," and not "whoever last approved an exception." A named owner decides which elevated paths may exist, who may request them, how long they last, what evidence must be recorded, and when unused glass must be retired.
In reality, mid-size org charts often leave that ownership floating. Platform wants fewer tickets during outages. Security wants locked-down production and a softer story for emergencies. Product wants velocity. On-call wants to sleep after the fix. Nobody wants to be the person who says the emergency role expires in four hours even if the postmortem is unfinished.
Write the contract in operational language. Name the systems covered, the roles that qualify as break-glass, who may open the glass, who must be notified, the default time bound, the evidence required, and who reviews membership on a cadence. If that list is empty, you do not have emergency access. You have permanent privilege with better storytelling.
Meanwhile, pair the glass with authority that matches the night. A junior engineer who can open break-glass but cannot roll back, page a peer team, or revoke a peer's lingering role is not safer. They are a human with a louder key. Ownership includes the decisions the role may make and the escalation when those decisions are not enough.
Evidence that the glass still closes
An unused policy page is not evidence that break-glass is under control. Evidence is specific. It includes a current inventory of emergency paths, recent openings with recorded intent, proven expiry or revocation, and a review that actually removed stale members. Prefer a short record a tired engineer can trust over a dashboard that only proves SSO is configured.
Finally, treat "temporary" as a timed contract, not a label. Temporary should mean a start, an end, a reason, and a human accountable for closing it. It should not mean "until someone remembers." If your tooling cannot show who holds elevated access right now, when it ends, and who approved the open, you do not have break-glass. You have standing privilege with an incident narrative.
Yet keep the bar humane. You do not need a novel for every exception. You need proof the common emergency paths have owners, time bounds, and a practiced close. Paths that cannot meet that bar should be fixed, time-boxed with an owner, or deleted. Leaving them open trains the team to treat the next outage as permission to mint more.
A close-the-glass habit mid-size teams can keep
You do not need perfect zero-standing-privilege. You need a repeatable habit that keeps emergency access from becoming the default way to ship and fix.
In practice, a workable pattern looks like this.
- Inventory every path that grants production-elevated power outside normal roles, and require a named owner, a purpose, and a review date for each one in plain language.
- Default every break-glass grant to a short time bound with automatic expiry, and make extension an explicit decision with the same evidence bar as opening.
- Record who opened the glass, for which incident or change, what they could do, and when it closed, in a place the next on-call and the next auditor can both find.
- Review membership and unused paths on a cadence, including shared accounts and long-lived tokens, and remove what the last quarter did not need.
- Rehearse opening and closing the glass before the outage, including notification and revocation, so the first real night is not a discovery exercise.
- After every messy incident, ask which emergency grants are still live, assign owners to close them, and refuse to call the incident done while glass remains open without a dated plan.
Instead of minting another standing admin after a bad night, mint a time-boxed grant, evidence, and a close. Privilege volume is easy to grow. Closing the glass is the scarce discipline.
When emergency privilege becomes expensive
Unclosed break-glass feels like resilience until the first mistake or confused deploy that only standing elevated access made possible. Then you pay in blast radius, audit pain, and a culture that cannot tell emergency from everyday work. Mid-size teams feel it faster because the same few people own the glass, the outage, and the apology.
In the end, treat break-glass as a product surface for operators and security together. Name who owns it, bound every open, keep evidence, and practice the close. Ask the uncomfortable question while the next outage is still optional.
If production went dark in the next hour, would your emergency paths open with a named owner, a clock, and a rehearsed close, or would the night mint privilege that quietly becomes how the team works forever after?
TL;DR
- Emergency elevated access that never expires quietly becomes the real permission model, even when the wiki still describes least privilege.
- Break-glass should be rare, loud, and short. Lingering roles, shared admins, and sticky tokens are unowned power, not readiness.
- Name an owner for the break-glass contract, including who may open it, default time bounds, evidence, and review cadence.
- Treat temporary as a timed contract with start, end, reason, and accountable close, not as a label on standing privilege.
- Keep a simple close-the-glass habit: inventory, short defaults, recorded openings, membership review, and rehearsal of open plus close.
- Ask whether tomorrow's outage would leave the glass closed afterward, or mint another path that never quite ends.