Safer AI needs two things at once: mature engineering and new science
Frontier AI arguments often turn into stories about institutions. When a safety researcher at OpenAI resigned and criticized frontier safety practice, coverage mostly followed the person, the employer, and the optics of the departure. That is the least durable part of the story. Remove the one lab and the one exit, and a better argument remains: the field needs mature safety engineering, and it needs new science for problems mature safety engineering has never encountered.
Those two tasks usually get split into rival camps, engineers versus researchers, "ship it with guardrails" versus "we do not understand what we are building." That framing is wrong, and it is expensive, because it lets each side ignore the other's best points. The practical question for anyone building or running these systems is simpler:
> Are we trying to make increasingly powerful AI with an engineering discipline that has not caught up?
For most organizations, the answer is "somewhat, unevenly, and faster than the discipline is catching up." So we do not have to choose between AI safety through better engineering and AI safety through new research. We need both.
Existing safety knowledge + AI-specific scientific research = More reliable AI systems
Change one: borrow wisdom
Aviation, nuclear power, aerospace, medical devices, chemical process engineering, and critical infrastructure operators have spent a century learning what to do when a bug is not just a bug. These fields share one structural fact: failure can be much worse than the immediate defect, because the systems operate in the physical world, at speed, with tired people, and with consequences that cannot be reverted by a redeploy.
The practices those fields accumulated are not exotic. They are a checklist of organizational maturity.
Defense in depth uses multiple independent controls, so one failure is not enough to cause disaster. Independent verification means the people who build the system are not the people who certify it. Containment assumes a component will fail and designs so the failure stays local. Redundancy means critical functions survive the loss of any one element. Monitoring means instrumented telemetry reflects real state, not marketing state. Auditability means enough recorded evidence exists to reconstruct what actually happened. Incident reporting means cheap, protected channels let bad news move upward. Rollback means a rehearsed path exists back to a known-good configuration. Risk ownership means a named human is accountable for accepting each residual risk. Operational discipline means the boring habits that keep controls working at 2 a.m.
For AI, defense in depth looks less like a poster and more like a control stack:
Model -> Policy -> Tool permissions -> Sandbox -> Data boundaries -> Human approval -> Monitoring -> Audit logs -> Incident response
The design goal is not zero failures. The goal is that a failure in one layer should not automatically become a system-wide failure. A model that produces a bad action is a different problem from a system that executes that action against production data with no rate limit, no approval gate, and no record. The first is a model defect. The second is an architecture choice. Most AI outages that look like model failures are, on inspection, stack failures: permission boundaries too wide, approval gates absent, logging retrofitted, rollback untested.
The annoying version of this argument is "AI companies should use aviation checklists," and that version deserves to be dismissed. Checklists are the visible artifact of something harder to copy. The real aviation lesson is a set of institutional power relations, and most AI organizations lack those. The useful questions are structural:
- Who is allowed to stop a deployment, and is that power real or nominal? - Who independently evaluates the system before it touches customers? - Who owns the residual risk after the mitigations run out? - How are incidents reported, and what happens to the person who reports them? - Are safety findings visible to leadership in the same format as latency dashboards? - Are safety engineers rewarded for finding problems, or treated as friction? - Can a team ship despite unresolved safety concerns, and if so, who signs off? - Are failures treated as learning opportunities or as reputational threats to manage? - Can you prove after the fact which controls were actually active?
That last question trips up more teams than it should. "We have guardrails" is not the same claim as "we can demonstrate, from immutable logs, which guardrail evaluated which request, what it decided, and what happened next."
High-reliability organizations are often misunderstood. They do not eliminate human error. They build systems that remain safe when humans make mistakes.
An engineer who forgets a step, a reviewer who approves too quickly, a leader who wants a launch date: these are not deviations from the model of how organizations work. They are the model. Aviation assumes people will err and then makes erring survivable. Most AI development still implicitly assumes the on-call person will be sharp, the reviewer will be thorough, and the model will be well behaved enough that nobody has to check. That assumption does not survive contact with anything at scale. AI development needs the mindset shift, not the paperwork.
Change two: discover what we do not know
Here borrowing runs into a wall. An airframe, a reactor, and an infusion pump are all specified by their designers. Their dangerous behaviors are discoverable through physics and testing. A learned system is different in kind. Its behavior emerges from optimization over data nobody fully read, inside representations we cannot yet read, and the designer's intent enters only indirectly through objectives, gradients, and fine-tuning.
The list of AI-specific problems is long, and every item weakens the analogy to older industries: learned behavior rather than written behavior; emergent capabilities that appear at scale without being trained for; opaque internal states; strategic behavior; deception that adapts to being tested; goal misgeneralization, where the trained objective points somewhere the designer did not mean; distribution shift between evaluation and deployment; multi-turn and multi-agent interactions nobody scripted; autonomous planning; tool use that touches real systems; long-horizon reasoning that compounds small errors into large ones.
This exposes a distinction worth internalizing, because it determines which tool you reach for.
An engineering unknown is a question like this:
"We do not know whether this system will fail under condition X."
You can address that with testing, simulation, redundancy, monitoring, isolation, better architecture, and a bigger sample size.
A scientific unknown is harder:
"We do not yet understand the mechanism that determines whether this class of AI behavior emerges."
More testing does not address that. More testing gives you more correlated data about a phenomenon whose generating process you still cannot describe. You need new research: mechanistic interpretability, evaluation science, behavioral characterization, robustness theory, and alignment methods with actual guarantees.
Some of the most safety-critical questions sit in the second category:
- How do we know what an advanced model is actually doing on a given inference, as opposed to what the trace suggests? - How do we know a model is genuinely aligned rather than merely passing the tests we happen to run? - How do we detect deceptive or strategically adaptive behavior that is, by construction, selected to survive detection? - How do we evaluate increasingly autonomous agents whose behavior depends on environments we cannot enumerate? - How do we maintain control as capability increases, in a way that does not degrade gracefully, or degrade at all?
None of these have answers mature enough to certify. "Our eval scores look fine" is not a scientific account of why the model behaves the way it does. It has the epistemic status it should have: a sample of behavior, not an explanation.
Why one without the other fails
Hold the two changes against each other with an aircraft thought experiment.
Imagine an aircraft built with exemplary engineering discipline: checklists, redundant flight controls, structured inspections, emergency protocols, a blame-free reporting culture, and independent certification. Now imagine it was designed by people who have no theory of aerodynamics. Lift is a correlation they noticed in test flights. Stall behavior is a vibe. Every control surface is well maintained, and nobody can say why the thing stays up. It will fly right until it encounters a regime that is not in the test corpus, and then it will crash for reasons that cannot be diagnosed, predicted, or fixed, only noticed statistically after enough hull losses.
Now imagine the reverse: perfect scientific understanding of aerodynamics, terrible operational discipline. Someone skips the inspection, ignores a warning, misreads a procedure, or deploys an untested component because the launch date was immovable. The physics was never the problem. The aircraft goes down anyway.
Applied to AI organizations, that gives two recognizable failure modes.
Organization A has excellent engineering discipline and poor scientific understanding. It has a beautiful control stack: permissions, sandboxing, approval gates, immutable audit logs, rollback rehearsed quarterly, and a safety reviewer with real stop authority. But it has no mechanistic or behavioral account of why the model does what it does. Every control is empirical, keyed to behaviors observed in testing. If behavior generalizes past the eval distribution, or adapts to the guardrails, the organization does not see the failure coming. It also may not see it after the fact, because its telemetry records actions rather than intent. It has a fortified castle whose foundations nobody has surveyed. Its confidence is calibrated to the sample, not to the phenomenon.
Organization B has excellent AI research and poor operational discipline. It does genuinely interesting work on interpretability, builds novel evaluations, and is honest about what the model does not yet understand. Then it ships to production with shared credentials, no rate limits, no approval step, no incident process, and a review culture that treats safety findings as schedule risk. The research output is real, and it dies on the way to a control surface. A model behaves badly at 3 a.m., the action executes against a customer database, nobody knows for two days, and the post-mortem has no logs.
The safest organization is the one that is not confused about either side. The three layers have to be present at the same time:
Science What makes these systems behave this way?
Engineering How do we make failures survivable?
Organization How do we make people consistently apply both?
Science tells you what to protect against. Engineering makes being wrong survivable. Organization is the mechanism that keeps the first two from quietly decaying under delivery pressure. It is also the layer that most often fails, because it is the only one made of incentives.
The gap that actually matters
The concern is not that safety is getting worse. In many organizations it is getting better, and credibly so. The concern is relative rate of change.
Capability: steep growth in benchmarks, autonomy, tool use, context length, planning horizons.
Safety understanding: slow growth in evals, interpretability, incident data, control mechanisms. Gaps open at thresholds nobody can name precisely in advance.
The dangerous variable is not capability alone. It is capability crossing a threshold that existing safety science was not designed to handle. No one can tell you which threshold that is or when it arrives. That is not a failure of the analysis. The uncertainty is part of the problem statement. A field that cannot explain why its systems behave as they do also cannot certify that its controls are sufficient for behavior it has not seen. Aviation could name its failure modes before it needed to prevent them. We are operating the reverse order.
Agentic systems make this concrete rather than philosophical. A chatbot is:
Human -> AI -> Answer
The blast radius is text. The human is the actuator. An agent is:
Human -> Goal -> AI -> Plan -> Tools -> External systems -> Real-world consequences
The difference is not one of degree. An agent interprets an ambiguous goal, decomposes it into a plan, executes commands, reads back live state, encounters conditions outside its eval set, revises its strategy, and continues for minutes, hours, or across a thousand tool calls. A bad answer can be deleted. A bad action can be committed, compounded, and then covered by the next plausible-looking action. Error rates that are tolerable per decision become near-certain over long horizons. This is precisely the case where both changes are mandatory: containment, permissions, approval gates, and observability, because single-step failure is no longer contained by the human; and new science, because the behavior we are containing is planned, adaptive, and not well characterized by anything we currently publish.
The new AI safety stack
Layer 5: Governance
Layer 4: Organizational Culture
Layer 3: Safety Engineering
Layer 2: AI Evaluation and Monitoring
Layer 1: AI Safety Science
Layer 1 produces explanations: what causes behavior, what is generalizable, and what the evals actually measure. Layer 2 turns partial knowledge into signal, with continuous evaluation, behavioral drift detection, and adversarial probing against production traffic. Layer 3 turns signal into survivability through the control stack described earlier. Layer 4 is whether the controls hold under deadline pressure, when the reviewer is new, and when the finding is inconvenient. Layer 5 decides who is accountable when layers 1 through 4 disagree with the roadmap.
No single layer is sufficient, and the failure modes are predictable. Science without engineering is a paper. Evaluation without enforcement is a dashboard nobody reads. Engineering without culture is controls that erode quietly, and the erosion is invisible because nobody built the instrumentation for the process. Culture without science is confidence with no basis. Governance without the other four is plausible deniability.
What this means for AI engineers
If you build or operate these systems, the practical agenda is not abstract.
Permissions should use least privilege for tools, be scoped per task, and be revocable mid-run. Treat "the agent can call anything" as a security incident already in progress.
Isolation should mean ephemeral sandboxes, network egress limits, and no shared credentials across trust domains.
Observability should mean full traces of plans, actions, arguments, and results, with retention long enough to matter.
Evaluation should not stop at offline benchmarks; offline benchmarks are theater. You need online evaluation against real traffic and real failure classes.
Rollback should mean state you can actually revert, tested by someone who was not in the room.
Failure containment should put budgets and ceilings on tool calls, spend, writes, and blast radius per unit time.
Human approval should gate irreversible actions, designed so the approver can meaningfully approve. A wall of "Are you sure?" prompts just trains rubber stamps.
Auditability should produce evidence that survives the incident, immutable and reconstructible.
Threat modeling should treat misuse, prompt injection, tool-chain compromise, and reward hacking as first-class adversarial scenarios.
Adversarial testing needs red teams with clear authority, protected reporting, and findings that block.
Run that list at your organization and count what is genuinely implemented rather than intended. The number will be lower than your roadmap implies.
But resist the satisfying conclusion that this list is the answer. Every control above is a bet about behavior you cannot fully explain. Sandboxing constrains where a model acts; it says nothing about why it chose that action. Approval gates add a human check; they do not tell you whether the human is reviewing something a deceptive planner has curated for their expectations. Evaluation detects divergence from known-good; it cannot detect a failure mode that has not been characterized. These controls are worth having. They convert unbounded risk into bounded risk, which is the whole game in every mature industry. They are not a substitute for understanding, and they are not a way to defer the research agenda.
The leadership question
> Are we building organizations that can safely operate the systems we are racing to create?
For most of us, honestly, the answer is not yet. That means the right move is not to slow capability or abandon it. It is to make leadership interrogate the gap directly:
- What do we actually know about this system's behavior? - What do we not know, stated plainly rather than as "ongoing work"? - Which assumptions is this deployment riding on? - Which failures are recoverable, and which are not? - How will we know when an assumption is wrong, and how fast? - Who has real authority to stop deployment, and when was the last time they used it? - Which scientific questions remain unanswered, and who is funded to answer them?
Underneath all of that is one principle, and it does not depend on anyone's risk philosophy.
> The greater the potential consequence of failure, the less comfortable we should be relying on assumptions we have not validated.
You can believe transformative AI risk is likely, unlikely, or somewhere in between and still accept that principle. It applies equally to an agent with database write access and to a system whose behavior nobody can explain. What it asks of us is not fear and not paralysis. It is the same thing every high-reliability industry eventually had to learn: understand the failure modes you can, engineer around the ones you cannot, and remember that the organization layer is made of incentives, which is why it so often fails.
