Smart Contract Audit & Certification
All Insights

The AI Challenge

Why machine-speed analysis changes the security assumptions behind the digital economy.

Christian Krumrey8 min read

Lawyer & Founder, SentaTrust Technologies LLC, Dubai

A wartime Enigma cipher machine, its rotors and keyboard in close focus

Earlier this year, Google disclosed something that would have sounded speculative not long ago. Its Threat Intelligence Group identified a threat actor using a zero-day exploit that Google believes was developed with the assistance of artificial intelligence. The attacker intended to use it in a wide-scale campaign, but Google says its own proactive discovery may have disrupted that plan before it could unfold.

The significance of the incident goes well beyond one vulnerability. It points to a change in the economics of software security.

Artificial intelligence is already transforming the way software is built. Developers use it to understand unfamiliar code, generate functions, write tests, troubleshoot errors and move from an idea to working software far more quickly than before. Most discussion around this shift has understandably focused on productivity: more software, produced faster, with fewer barriers between an idea and its execution.

But the same capabilities work in the opposite direction. Software that can be understood more quickly can also be examined more quickly. Patterns can be compared across huge codebases, potential weaknesses identified, attack paths explored and known exploit techniques tested with increasing levels of automation. Google's own threat research now describes a transition from relatively experimental uses of AI toward what it calls the industrial-scale application of generative models within adversarial workflows.

That creates an uncomfortable symmetry. AI will help build more of the digital economy, but it will also make that economy easier to probe, test and challenge.

The economics of finding a weakness

Software security has always involved a race. Developers build systems, security specialists examine them, and attackers look for what everyone else missed. Until recently, all three sides shared the same fundamental limitation: human attention.

A skilled security researcher can study only so much code in a day. An attacker has to decide which targets justify the time required to understand them. Even large organizations have finite numbers of engineers, analysts and penetration testers. That scarcity of attention creates its own kind of friction. Not every vulnerability will be discovered simply because there are not enough people available to look for every vulnerability in every application.

Machine-scale analysis begins to weaken that assumption.

An attacker does not need an AI system capable of independently carrying out a perfect cyberattack for the economics to change. Even an imperfect system can become powerful if it dramatically reduces the cost of searching. A tool that can review thousands of functions and identify the few that deserve deeper investigation has already extended the reach of the person using it. So has one that can compare implementations, explore unusual inputs, reproduce known exploit patterns or continuously monitor newly published code.

The same is true for defenders. AI can help security teams search larger codebases, reproduce failures, identify suspicious behaviour and develop fixes more quickly. Google has explicitly argued that highly capable AI models are increasingly able to find vulnerabilities and assist with exploit generation, while defensive teams are simultaneously beginning to use AI to harden software more aggressively.

The important point is therefore not that AI inherently favours the attacker. It does not. The same technology can strengthen defence enormously. What changes is the scale at which both sides can operate.

And this is happening while the amount of software itself is expanding.

AI-assisted development lowers the cost of creating code. Small teams can build products that once required much larger engineering organizations. Specialists can automate repetitive work and concentrate on harder problems. People who could never previously write software can increasingly turn ideas into functioning applications.

That is a remarkable productivity gain. It also means more code is entering the economy more quickly, while the tools available to inspect that code are becoming more capable at the same time.

The real challenge is therefore not that AI necessarily writes bad code; human beings have always managed that perfectly well on their own. The more consequential change is that both sides of the security equation are accelerating simultaneously. We are producing code faster at precisely the moment when we are becoming better at searching it for weaknesses.

Traditional security processes cannot simply assume that the old balance will remain intact.

When the code controls the money

This challenge applies across software, but smart contracts expose it particularly clearly.

A conventional software vulnerability may interrupt a service, expose confidential information or allow unauthorized access. Those consequences can be extremely serious. A smart contract can add another dimension because the software itself may directly control economic value.

Smart contracts can determine who owns an asset, when money moves, whether tokens can be created, whether transactions can be paused, who holds privileged permissions and under what conditions a financial agreement executes. In that environment, the code is not merely describing the rules. It may be enforcing them.

A weakness can therefore become something more than a defect waiting to be patched. Under the wrong circumstances, it can become an executable financial opportunity.

Blockchain also introduces an unusual transparency. Much smart-contract code is publicly visible, allowing investors, developers and researchers to inspect what a protocol actually does. That transparency is one of the technology's great strengths, but it is symmetrical. What defenders can examine, attackers can examine as well.

That has always been true. What AI changes is the scale at which examination can occur.

The industry is already preparing for that reality. In February, OpenAI and Paradigm introduced EVMbench, a benchmark specifically designed to evaluate whether AI agents can detect, patch and exploit high-severity smart-contract vulnerabilities. It draws on 117 curated vulnerabilities from 40 audits and tests AI systems across both defensive and offensive tasks.

The existence of a benchmark like EVMbench is revealing in itself. We are no longer asking whether AI can participate meaningfully in smart-contract security. Researchers are already building standardized environments to measure how effectively AI agents can find vulnerabilities, repair them and exploit them.

That should not be interpreted as evidence that autonomous AI hackers are about to empty every blockchain protocol. Current systems still make mistakes. They misunderstand context, generate false positives and sometimes identify problems that disappear under closer examination. Human expertise remains essential.

But perfection is not required for the technology to matter. If an AI system can narrow a large codebase to a handful of promising attack surfaces, it has already increased an attacker's reach. If a defensive system can do the same for an auditor, it has increased the defender's reach as well.

The security contest increasingly becomes one of who can combine automation, judgment and speed most effectively.

Why humans still matter

Whenever AI enters a profession, the discussion quickly turns into a replacement debate. Will it replace programmers? Security researchers? Auditors?

That is probably the wrong way to think about this problem.

Machines and human experts have different strengths. Automated analysis is exceptionally good at repetition. It can examine enormous amounts of code without fatigue, compare patterns across thousands of examples and repeatedly test assumptions that would be tedious for a human reviewer to explore manually.

Human judgment remains essential where the meaning of those findings has to be understood.

A smart contract may contain an administrator role with extensive permissions. Whether that represents a serious risk depends partly on what the contract is meant to do, how those permissions are governed and what users have been told to expect. An automated system might identify an unusual economic interaction, but determining whether it is genuinely exploitable may require understanding market behaviour, incentives, governance arrangements and the surrounding protocol.

A piece of code may function precisely as written and still produce a result that nobody intended.

Those are not simply code-quality problems. They are questions of context.

The most credible model for verification is therefore unlikely to be purely human or purely automated. Machines can expand what can be examined; people still have to determine what the findings mean, whether they matter and what should be done about them.

That combination becomes more important, not less, as automated analysis improves.

Raising the standard

The deeper consequence of AI is that it raises the standard against which security itself has to operate. If software can be created at unprecedented speed, security cannot remain a process that happens occasionally and only at the end. If potential weaknesses can increasingly be searched for automatically, defensive verification has to become more systematic, reproducible and continuous as well.

That does not mean promising perfect security. No serious audit or certification can guarantee that software will remain immune to every future attack. New techniques appear, new interactions emerge and new vulnerabilities are discovered.

Verification is not about claiming perfection. It is about reducing uncertainty through evidence.

At SentaTrust, this is one of the assumptions behind the verification infrastructure we are building. Automated analysis allows code to be examined systematically and repeatedly at a scale that manual review alone cannot easily sustain. Human interpretation remains essential where technical findings intersect with economic logic, governance, privileged control and real-world consequences.

The objective is not to use AI because AI is fashionable. It is to recognize that the environment itself is changing.

The technology available to those searching for weaknesses is becoming faster and more capable. The technology available to those defending software must evolve accordingly. There is reason for optimism in that. The same advances that make it easier to challenge software also give us better tools to strengthen it, catch mistakes earlier and allow security professionals to spend more of their time on the questions that genuinely require judgment.

But the balance will not maintain itself.

As more money, ownership and commercial activity move through code, society will increasingly depend on software behaving as intended. And as the tools capable of interrogating that software improve, assuming that a system is secure simply because it has worked so far will become an increasingly weak form of assurance.

AI will help build an extraordinary share of the digital economy. It will also subject that economy to a level of scrutiny that human teams alone could never sustain.

That should not make us pessimistic about AI.

It should make us more ambitious about verification.

The standard has changed. Security now has to change with it.

Read more Insights — and follow us on LinkedIn and X for new articles as they are published.