Try Your Ideas logo

Try Your Ideas

The culture change we need in the AI times.

The culture change we need in the AI times.

Aviation does not depend on one pilot never making a mistake. Chemical plants do not rely on one engineer never opening the wrong valve. Financial infrastructure does not trust a single transaction. The shared principle is defense in depth. For AI, the goal is not to make the Agentic flow perfect. The goal is to make failure survivable.

starstarstarstarstar
starstarstarstarstar
No ratings yet

When the software stops waiting for you

Some former members of OpenAI's safety team recently raised a concern worth taking seriously beyond the headlines. They said safety is hard to maintain inside a culture optimized for rapid development and deployment. This is not a claim that safety researchers dislike shipping. It is a claim about incentives. Safety work behaves badly in organizations whose reward systems treat velocity as the main virtue.

That leaves a larger question. What kind of engineering culture do we need when the software we build becomes increasingly autonomous, powerful, and capable of affecting the real world?

For decades, good software culture has rewarded shipping. We build, measure, learn from production, and iterate. If something breaks, we revert. If latency spikes, we optimize. If users complain, we fix it before the next release. This culture produced enormous innovation. It also assumed something that no longer holds cleanly: the software will wait for us. It waits for a request, a diff review, or an operator noticing the bad deploy.

The culture that helped us build AI may not be sufficient to safely deploy AI.

The most important safety technology may not be another model, benchmark, guardrail, monitoring dashboard, or safety classifier. It may be the culture in which AI systems are designed, tested, deployed, and operated.

What changes when code starts acting

A chatbot answer, a code completion, and a recommendation are relatively easy to reason about. They produce output. A human consumes or rejects it. If the output is wrong, the failure is often bounded: edit, ignore, ask again.

Agentic software is different. An agent can accept a goal, plan, call tools, read files, write code, install dependencies, run commands, browse the web, send messages, touch databases, invoke APIs, and delegate subtasks to other agents. It can run for thousands of steps and continue while you sleep.

That changes the failure model. A wrong answer is usually recoverable. A wrong action may not be. If a coding assistant suggests dropping a production column, review catches it. If an agent is permitted to apply the migration itself, you may discover the failure only after the data is gone and restore takes hours.

We can argue about model intelligence, but the engineering problem is narrower. Capability, authority, autonomy, and access together produce real-world effect.

The problem is often incentives, not intent

The easy story is bad actors. The truer story is ordinary pressure. Most AI companies are not careless. They are optimized. Teams are rewarded for demos, releases, latency improvements, adoption, eval scores, and launch dates. Safety work moves slower. It looks like a regression in success rate. It sounds like "not yet." It asks for budget before anyone has been burned.

None of that requires malice. It only requires structure.

A release engineer who blocks a launch becomes the blocker. A researcher who reports unexpected behavior becomes a distraction. An infra team that removes ambient database access becomes annoying. Smart, ethical people adapt to those incentives.

That concern generalizes. Safety is hard inside a fast-moving company not because safety is impossible, but because velocity can turn every safety question into a delay to be optimized away.

Two engineering cultures

You can see the tension as two cultures.

Culture A is the move-fast culture: ship quickly, experiment in production, iterate, fix failures, optimize velocity, trust experienced engineers, and treat safety as a feature, compliance checkpoint, or checklist in the final week.

Culture B is the high-reliability culture. It assumes systems will fail, humans will make mistakes, and models will behave unexpectedly. It designs for containment, independently verifies critical decisions, separates incentives where appropriate, red-teams continuously, makes failures visible, learns systematically from incidents, and treats safety as an architectural property.

Culture A is excellent at invention. It is often mediocre at operating a system that can invent its own actions.

Culture B is not bureaucracy. Aviation is not bureaucracy because it requires a second check before a critical action. Nuclear operations are not bureaucracy because they assume loss of power, loss of cooling, human error, and sensor failure. They are cultures that respect consequences.

What other industries learned

Aviation does not depend on one pilot never making a mistake. It uses checklists, standardized communication, crew resource management, maintenance logs, duplicate systems, automation envelopes, and reporting cultures where near misses are valuable data.

Chemical plants do not rely on one engineer never opening the wrong valve. They use interlocks, relief systems, containment, automatic shutdowns, and procedures designed to make common errors harmless.

Financial infrastructure does not trust a single transaction. It uses reconciliation, idempotency, limits, settlement windows, fraud rules, and audit trails.

The shared principle is defense in depth. Do not build a system whose safety depends on every component being correct. Build layers so failure in one layer is caught by the next.

For AI, the goal is not to make the model perfect. The goal is to make failure survivable.

A layered architecture for autonomous AI

A serious agent system needs layers, not because the model is untrusted, but because trust is not a control.

The model layer covers training behavior, refusal boundaries, robustness to prompt injection, and output constraints. The agent policy sets permitted goals, maximum runtime, allowed delegations, and approval thresholds. Tool permissions should expose least-privilege APIs. Instead of giving an agent any tool, give it scoped capabilities: read this bucket, write this log, send this message, run this script. The sandbox provides isolated execution boundaries for files, network, processes, packages, and dependencies. Data access controls need tenant and column scope, PII redaction, time-limited tokens, query budgets, and audited paths. Human approval adds friction at the right places. Not every action needs review; irreversible actions do. Monitoring gives observability into tool calls, resource consumption, unexpected branches, permission requests, and behavioral drift. The audit trail keeps an immutable record of what happened, why it was allowed, and what authorized it. Incident response needs pre-agreed kill switches, rollback plans, containment playbooks, communication templates, and blameless reviews.

Teams skip audit trails when the dashboard looks cooler. They loosen permissions because debugging is annoying. They postpone sandboxing because the demo needs one more integration. Safety survives when it is part of the architecture, not a promise in the release notes.

The culture around the controls

Culture is not a poster. It is what people do when the launch date slips or an eval looks odd.

A strong AI engineering culture should make these behaviors normal. People can stop the system when it is not safe enough to ship without being treated as a blocker. Bad news surfaces early when evaluations reveal unexpected behavior, instead of being buried by metric tuning. Prevention gets rewarded, not only shipping. Removing an ambient credential before an incident should count. Incident reviews ask, "What allowed this failure to happen?" rather than "Who made the mistake?" Safety is treated as architectural. Security, permissions, observability, sandboxing, evaluation, rollback, and incident response are designed in, not added immediately before launch. The team shipping should not always be the only group deciding whether it is safe to ship. Uncertainty is acceptable. "We don't know" should be valuable engineering information, not a sign of weakness.

These are not ethics slogans. They are operational mechanics. A system where people are punished for raising concerns will eventually ship concerns into production.

The developer's new responsibility

Developers increasingly work alongside agents that generate code, modify repositories, run tests, install dependencies, access APIs, operate terminals, make architectural decisions, and deploy applications. The developer's job changes.

The question is no longer only, "Can the AI produce correct code?"

It becomes, "What happens when the AI is wrong, and what authority does the system give it when it is wrong?"

That is deceptively deep. Generating code is relatively benign. Executing code with credentials is not. An agent writing `DROP TABLE users` in a diff is a review comment. An agent running that migration with `db-admin` access is an outage.

Developers need to think in terms of authority. A tool call carries authority. A token is a leash. An automation trigger grants permission. A long-running agent is a process that may do something you did not intend, at a scale you did not expect, while you are doing something else.

The dangerous case is silent authority: inherited credentials, broad scopes granted for convenience, network access "for now," and destructive operations protected only by the model's judgment.

Capability must be matched by control. An AI agent with more capability should not automatically receive more authority.

Autonomy demands maturity

Autonomy is not a switch. It is a ladder: chatbot, coding assistant, tool-using agent, autonomous agent, multi-agent system, AI operating critical infrastructure.

As you move up, every control must become more sophisticated. Testing needs to cover adversarial inputs, tool combinations, and stateful behavior. Observability needs to explain why an agent took a sequence of actions. Isolation needs to reduce blast radius. Human oversight needs to happen at meaningful decision points, not in a dashboard no one reads. Rollback needs to be more than reverting a commit. Incident response needs to be rehearsed.

The more autonomy you give an AI system, the stronger the surrounding engineering culture needs to become.

What managers owe the work

Managers do not need to become safety philosophers. They need to change what gets rewarded.

Ask what gets rewarded: shipped features, eval wins, or systems that survive hard cases. Ask who can stop a release, and whether they are protected when they do. Ask who owns safety: a team with authority or a person with influence. Ask who challenges assumptions before they become architecture. Ask what happens when someone reports a serious problem. Ask whether failures are hidden or shared.

If you want a safety culture, pay for it. Budget for evaluations that slow success metrics. Reserve time for incident drills. Fund permissioning work no user notices. Create space for "not yet." Let people stop things without damaging their careers.

Developers have a matching responsibility. Do not grant an agent more permissions than you would grant a contractor. Do not let convenience become architecture. Build the kill switch before the autonomous loop. Assume the model will one day behave in a way your happy path does not cover.

The real shift

The move is not from reckless to paralyzed. That would be a false binary, and it would probably lose.

The shift is from "move fast and fix it" to "move fast where failure is recoverable, slow down where failure isn't."

Risk tolerance should be proportional to recoverability and autonomy. A shadow-mode suggestion engine can experiment aggressively. An agent with write access to a production database needs different humility. A prototype can be messy. A system that moves money, executes shell commands, changes infrastructure, or sends messages needs containment.

Powerful technology requires institutions capable of handling that power responsibly. Not abstractions, but ordinary engineering institutions: reviews, permission models, monitoring, audit trails, blameless postmortems, independent challenge, and leadership that treats a raised concern as a contribution rather than a delay.

If we get those properties right, we can keep moving fast where speed helps us learn. If we do not, we are not building a future with autonomous software. We are outsourcing incidents to it.

Was this article accurate and helpful?

Keep reading to unlock rating this article.

Subscribe to News

Get the latest articles delivered to your inbox.