
A line of text should not be able to take over a machine, but that is exactly the uncomfortable direction enterprise AI security is heading in, after Microsoft published research showing how prompt injection can be chained to remote code execution (RCE) vulnerabilities in AI agent frameworks, including work involving Semantic Kernel.
The significance is bigger than a single framework or a single vulnerability class. It is a warning that the attack surface has shifted again; from the network perimeter, to applications, to models, and now to the instruction layer that governs what agents do.
“This is not a new type of threat,” said Francois van der Merwe, founder & CEO of Otinga.io. “It is the scale and the automatability of social engineering that is going to go through the roof, as well as the personalisation. The same hyper-personalisation enterprises use to deliver better services is equally available to attackers. Here, the better service is a social engineering attack.”
The key difference with prompt injection is that the “payload” does not always look like malware. It looks like instructions. In the real world, that can mean an agent being manipulated into exfiltrating data, changing permissions, approving actions it should not, or coaxing a user into running a command they would never execute if they understood the consequences. The mechanism is as old as phishing; the delivery is what is new.
Van der Merwe says executives are still thinking about AI risk in the wrong frame, focusing on chatbot outputs rather than agent behaviour. “You need to think about AI agents the way you think about users in your environment. Sometimes a bad user, but more often a dumb user. You need two completely different threat models for each.”
A bad AI or a dumb AI
His first model is the uninformed agent: a well-intentioned system that crosses boundaries due to a lack of context. That is a governance and permissions problem, and the answer is stronger access control, clear guardrails and zero-trust thinking designed for agents, not just people. The second model is the compromised agent: an agent that has been weaponised and now operates persistently with the permissions granted by the organisation. This is where prompt injection starts to resemble something far more dangerous than phishing, because it can become an execution pathway.
Microsoft’s RCE research makes the escalation risk concrete: prompt injection is not just about tricking a model into saying the wrong thing. It can, under the wrong conditions, become a route into systems where agents have access to tools and can trigger real actions.
The business pressure to move fast is making the risk harder to manage. Van der Merwe describes this as FOLO: Fear of Losing Out. “FOLO is driving organisations to prioritise velocity at the cost of security and sometimes that cost is leaked credentials, or a well-intentioned agent purchasing expensive training courses because the user said ‘do it at all costs.’”
Build the secure environment you need
His advice to CISOs is blunt. Do not assume you can block this away. “Instead of focusing on containing and blocking, focus on building a safe environment to work within. You are not going to contain this. People will always find a smart way around whatever obstruction is put in their way.”
In practice, he argues the goal should be visibility; a monitored environment where organisations can see what agents are doing, what data they touch, what instructions they execute and where information goes, before an incident becomes a headline.
“Agentic AI” changes productivity, but it also changes threat models. If prompts can become shells, then security programmes must treat agents as privileged actors, not just software anymore, and apply the same rigour to instruction inputs, tool permissions, and third-party integrations that they already apply to identity and endpoint protection.
Find out more at www.otinga.io
© Technews Publishing (Pty) Ltd. | All Rights Reserved.