Seven items from the week of 31 July - 6 August 2026, each with a comment on what it changes in your systems.
This week three independent communities - a standards body, a research firm and an enterprise software vendor - described the same phenomenon from three different angles. Nobody coordinated it. OWASP moved excessive model autonomy up to third place on its risk list. Red Hat counted how many companies actually oversee that autonomy and arrived at thirty-one percent. SAP wrote that decisions about agent permissions belong to the board.
The question that emerges is, at heart, an old one and has nothing to do with artificial intelligence. It reads: who decided what this tool is allowed to do. The only new part is that the tool has started acting on its own, while the answer to that question is rarely assigned to a named person.
Add to that two items on what agents are actually starting to do in automation, one Polish data breach and one attack for which there is no good answer today.
1. OWASP Top 10 for LLM applications: the new edition stopped describing fears

On 4 August the 2026 edition of the list of the ten most serious risks in applications built on large language models was published. It is the first edition in which the ranking was produced by colliding two sources: practitioner votes and the record of real incidents. The authors name the difference outright - it is the gap between what the industry fears and what it has already been burned by.
Three moves are worth knowing by heart.
Excessive Agency moved up to third place. Not because anyone changed their mind, but because that is where the damage lands. The model was given tools and started using them more broadly than anyone had planned.
Unbounded Consumption climbed four positions. Uncontrolled token cost has stopped being a budget line and become a security risk. In many companies this shift is not yet reflected in the division of responsibility - the model invoice and the incident report usually land on two different desks.
Improper Output Handling dropped from fifth to tenth, but with a broader scope - it now covers unsafe code generated at scale by assistants. A drop in the ranking does not mean the problem went away.
There is one more distinction, more important than the order itself. This list describes the risk when the model is a component inside an application. The moment it becomes an actor - with tools it calls on its own, memory carried across sessions and effects that propagate down the process - the risk moves to a separate list, the OWASP Top 10 for Agentic Applications. The authors state plainly that many incidents sit exactly on the boundary and neither list covers it alone.
The practical takeaway: if someone in an agentic project cites only the LLM list, they are looking at half the picture.
In parallel, AISVS 1.0 was released - the first testable verification standard for AI system security, with maturity levels and chapters on agent orchestration and attack resilience. The difference is fundamental. A risk list tells you what to fear. A standard tells you what to check and how to prove it was checked.
Sources: OWASP GenAI, document dated 4.08.2026, CC BY-SA 4.0 licence; coverage in Help Net Security, 6.08.2026.
2. Thirty-one percent

Red Hat published a study with three numbers that are worth reading together, because it is the spread between them that says something.
92% of organisations declare they know where the data used by AI is stored. 49% have full control over that data. 31% have implemented mature oversight mechanisms for agentic AI.
The drop from ninety-two to forty-nine is the difference between “we know” and “we can do something about it”. The drop from forty-nine to thirty-one is something else - the difference between technical control and organisational governance. The second gap is the harder one, because it does not come bundled with a licence.
As a maturity test, the study’s authors point to something that cannot be faked: the ability to switch model or platform vendors without disrupting the business. 63% of companies have a formal exit strategy, while 30% have not prepared one at all. Within that second group, 14% assume a migration would be easy - which is itself an interesting measurement of optimism.
Until the first agentic deployment, oversight is a sentence in a policy. Our observation from projects is that the real test arrives on the day an agent is granted access to a production system and someone has to decide what it may do without asking.
To be fair: the study comes from a vendor whose platform is meant to solve this very problem. That does not invalidate the numbers, but it does mean reading them as a measurement taken by an interested party.
The study came out one day after the first threshold of AI Act obligations. A coincidence, but a telling one.
Source: Red Hat study as reported by ITwiz, 3.08.2026. We could not locate the original Red Hat report in open access - the figures are quoted from the coverage.
3. UiPath Maestro Case: numbers worth knowing

Early adopters of the Maestro Case module report average case handling time cut by 60-80% and three to five times more cases closed without human intervention.
One caveat before these numbers travel any further: they describe early deployments selected and written up by the vendor. That is the upper bound of what a process well matched to the tool can achieve. Not an average and not a promise.
More interesting than the percentages is what they apply to. Maestro Case manages a case, not a task. The difference is practical. A task has a beginning, an end and a robot. A case can live for weeks, pass through several systems and several pairs of hands, and stop along the way at every step that requires someone’s approval. Classic automation handles the task brilliantly and stalls exactly at the moments where the case is waiting.
The processes where numbers like these repeat share one profile: high case volume, repeatable decisions, data scattered across several systems. Claims handling. HR helpdesk. Approvals in SAP flows.
And the inverse, which needs saying: if a company has no automated tasks yet, the orchestration layer has nothing to manage. Maestro does not replace the first step.
Source: UiPath investor announcement.
4. Autopilot stopped suggesting and started building

Autopilot in UiPath Studio Desktop is presented as a full coding agent: it is meant to plan, build, run, diagnose, explain and rework automations. Public preview, available from the STS S195 line upwards. We do not yet know the quality of its output on a real client project, and we say so plainly.
The list of announced capabilities is longer, but three items show the scale of the change. It takes a thirty-page specification of an employee onboarding process and builds a complete automation from it, interface steps included. It takes a deployed job that reported an access-denied error at three in the morning and traces the failure to missing robot permissions. It takes an action that stopped working after an application change, names the cause and repairs the selector.
That last one is precisely what separates it from a general-purpose agent. General-purpose agents do not know what a selector is. Autopilot does, because it works inside Studio, on the same skills as the rest of the platform, with the object repository within reach. UI automation - the core of classic RPA - works for it from day one, while for external agents it tends to be the weakest link.
The second difference concerns oversight and will matter more in the conversation with your security team than the feature list. Within the monthly limit there is no separate subscription and no token billing. There is also no account to open with an external provider and no extra keys to guard. Destructive actions are gated, the autonomy level is set by an administrator, and events go to the audit log. Studio remains the visual layer - every change can be opened, inspected and debugged.
Three caveats without which this item would be marketing material. It is a public preview, not general availability - you do not build a project schedule on it. The monthly usage limit has not been quantified, so before anyone says “no extra cost”, it needs to be established for the specific licence. And the simplest thing of all: customers on the LTS line - the long-term support releases - will not see this feature.
Source: UiPath forum, announcement of 7.07.2026.
5. Żabka: in through a vendor, out with a map

On 4 August Żabka Polska - Poland’s largest convenience store chain - confirmed unauthorised access to selected technical resources through an external service provider’s account. The incident was detected at the end of the previous week. Independently, an offer appeared on a criminal forum to sell the allegedly stolen set for EUR 5,000: Jira tickets, code repositories, user information. The company did not disclose the scope, and the description of the set comes from the attacker, not the victim.
The most interesting thing in this story is the asking price. Five thousand euros is not much, which suggests the attacker himself does not consider the haul spectacular. If his description is accurate, he reached for material that most companies classify as technical rather than sensitive.
It is worth pausing on what actually sits in a ticketing system. Configuration dumps pasted into a comment so a colleague can see the error. System addresses. Names and roles of the people who know each integration. A description of what exactly keeps breaking and since when. Together this is a map of the environment written by the people who know it best. Repositories add integration details and, in the worse variant, credentials left in the change history.
Formally, two regimes may apply here: GDPR obligations, including the seventy-two-hour notification deadline, and - if the entity qualifies as an important one - obligations under Poland’s national cybersecurity act, which implements NIS2.
Two questions are worth asking before someone else asks them. Which of our resources can each vendor reach today. And what happens if one of them calls tomorrow morning about their own incident. A data processing agreement answers the first question only if someone has read it since it was signed.
Sources: Sekurak, CRN, ITwiz, The Record.
6. SAP called agent sprawl a board-level problem

On 3 August SAP published an article on a phenomenon it named agent sprawl: organisations deploy agents faster than they build the mechanisms to oversee them. A sentence from the piece: “As adoption of AI agents accelerates, governance is struggling to keep pace”.
The thesis goes further than a technical diagnosis. Decisions about agent permissions are decisions about business risk, so they belong to the board, not to the operations team. The vendor positions its own layers as the answer along the way, which is natural and worth saying out loud right away.
In an ERP system the stakes look different than in a marketing tool. There, an agent’s mistake costs a campaign. Here, the agent works on the data that the month-end close, supplier settlements and tax filings stand on. A change made by an agent can trigger an accounting effect - and all the more it must leave a trace in the audit log. We wrote about this at length in the context of SAP agent security.
Then there is the thing specific to this environment. Permissions in SAP are built over years, layer upon layer, and rarely can anyone say from memory what exactly a technical user - the one whose permissions an integration runs on - is able to do. An agent granted permissions “same as the existing integration” inherits all that baggage, along with everything everyone forgot about.
A question for your next review: how many agents run in your SAP landscape today, who approved their permissions, and can that list be reconstructed without polling three people. If gathering the answer takes a week, that is exactly the sprawl SAP is writing about.
Source: news.sap.com, 3.08.2026.
7. The page plants an instruction, the agent executes it. Can we defend against this?

We saved for last the item where the answer is “not in a way that closes the subject”.
Independent research teams have been showing for over a year that browsers with a built-in assistant can be taken over without a single click from the user. The mechanism is banal, and that is exactly why it is dangerous.
The assistant reads the page through the same text stream it reads your command through. All the attacker needs is to hide a few sentences phrased like an instruction in the content - white text on a white background, or in a paragraph the agent was going to summarise anyway. The assistant takes those sentences for a command and executes them. Nobody clicks anything. It is enough to let the agent onto a prepared page.
The consequences described in the research go beyond data theft. One scenario serves the agent a page pretending to require a login, to harvest credentials. Another steers what information the agent collects and what conclusion it reaches - an attack not on data but on the recommendation the agent will hand you. Vendors keep adding safeguards, researchers keep bypassing them one by one and state plainly that no perfect fix is in sight.
A note on a number making the rounds. A figure of 86% attack success is in circulation. We checked it, and it does not hold up in the form it is being repeated: it comes from an earlier study (the WASP benchmark, 2025) and describes partial success of a prompt injection, not the share of successful fake logins. That is why we do not present it here as fact. The mechanism is documented and is enough on its own - the number is not needed.
Why there is no patch here
The reasons are structural, not an oversight by any particular vendor.
First, the agent has one door for content and for commands. A single input channel for text to read and orders to execute, with no reliable way to tell them apart, because both are plain text.
Second, the human has left the loop. It was the human who usually caught the odd command before executing it. An agent acting autonomously removes that moment. What remains is the text and the action - and gone with them is the instant in which anyone could say: hold on, I did not order this.
Hence a conclusion worth remembering: input-side control will not catch this, because the attack looks exactly like the content the agent was meant to read. It only becomes visible in what the agent does.
What can be done today
Four things. None of them is a solution; together they limit the blast radius.
Watch behaviour instead of filtering input. Since the attack is indistinguishable at the input, the budget goes into an oversight layer and detecting deviations from the intended trajectory, not into another command classifier. Classifiers get bypassed - we have been through that with earlier vulnerabilities of this family.
Put a gate on irreversible actions. Payment, dispatch, permission changes, data deletion - everything that cannot be undone requires human confirmation. This is where the removed moment gets put back in.
Give the agent narrower permissions than the user. If the answer to “what can the agent reach after a takeover” is “everything the user has access to”, that is the wrong answer.
Isolate the environment. An agent driving a browser works in a segregated environment, not on a workstation logged into production systems.
Finally, something it is only right to say about ourselves: we use agent-driven browser tools. The same class of risk applies to us, not just to our clients, and we treat it accordingly.
What follows from this
Seven items, one common denominator. The risk list moves excessive autonomy up. The study shows that fewer than one company in three has mature oversight. The ERP vendor says it is a board matter. Two automation items show how much that autonomy actually delivers when it is properly fenced. The breach reminds us that a company’s perimeter stopped running along its own infrastructure long ago. And the last item says outright that in one place we have no good answer today.
If we were to leave you with one question for Monday, it would be this: who in your company signs off on an agent’s permission scope - by name, not by a job title in a policy. Everything else is achievable once that answer exists. Without it, every other item on this list is just industry news.
If any of these topics touches you directly - oversight of agentic AI, SAP security, orchestration and automation with agents - let’s talk.
The weekly review is our selection from several hundred radar and RSS items each week. Sources are linked at every item. Informational material; it does not constitute legal advice.
