Skip to content

Weekly review W40: enforcing AI agent permissions outside the prompt

SNOK's review of 25 September-1 October 2026: authorisation and containment for agents in SAP Joule Studio, SAP ERP 6.0 support dates, UiPath Cartographer, AG-UI 1.0, open decision models and a SNOK Press book on SAP security.

SNOK’s review of 25 September - 1 October 2026: SAP, AI agent security, UiPath and decision models.

This week’s releases share a design choice: the limits on an AI agent are enforced somewhere other than the prompt. SAP and NVIDIA split business authorisation from execution containment, AG-UI 1.0 lets an agent pause for human approval and resume from the same point, and Google AX gives every agent task its own sandbox.

A prompt can tell an agent what to do. Enforcement has to sit in permissions, process, runtime and protocol, none of which care what the model has just read. SAP ERP 6.0 offers a parallel from the contract side: how long a system stays supported is governed by eligibility terms and technical prerequisites, not by the year quoted in the headlines.

1. SAP ERP 6.0: support timelines and eligibility

Mainstream maintenance for SAP Business Suite 7, SAP ERP 6.0 included, ends with 2027 and covers the three most recent enhancement packages. SAP prices extended maintenance for 2028-2030 openly: two percentage points on top of the maintenance base. Systems left outside it fall back to customer-specific maintenance.

For 2031-2033 the only route is a subscription, “SAP ERP, private edition, transition option”, and it covers SAP ERP alone rather than the whole of SAP Business Suite 7. It requires a move to SAP HANA in private edition before 31 December 2030, a system of at least 2 TB and the max success plan. According to SAP’s August 2025 announcement, contracts signed in 2026 carry a 20 percent uplift, and pricing for contracts from 2027 will be published with the offer in 2028.

The backdrop is the European Commission decision of 9 July 2026, which made SAP’s commitments legally binding for ten years, worldwide. They include clearer terms for splitting a system landscape between different support providers or leaving parts unsupported, the removal of reinstatement fees and lower back-maintenance fees. We summarise the announcement here; what it means for a given contract is a question for a lawyer.

So 2033 is available to some ERP systems, on stricter terms and at a higher price. Every SAP Business Suite 7 system now needs a costed comparison of three options: conversion to SAP S/4HANA before mainstream maintenance ends, extended maintenance to 2030, or - for SAP ERP - the transition option to 2033. Each needs a named decision-maker, and a contract signed this year carries the 20 percent uplift.

Sources: SAP Support Portal, “Maintenance strategy for SAP Business Suite 7”; SAP News Center, Stefan Steinle, “Navigating Your RISE with SAP Journey: Updates for SAP ERP, Private Edition, Transition Option”, 4 August 2025; European Commission, press release IP/26/1554, 9 July 2026. Status: confirmed - read on 1 October 2026. Detailed SAP Notes require a login and are not quoted.

Three hourglasses of different sizes in aluminium frames on a steel lab bench, the sand in each at a different stage

2. SAP Joule Studio and NVIDIA OpenShell: authorise, then contain

On 28 September NVIDIA announced the Open Agent Safety Platform. Its foundation is OpenShell, an Apache 2.0 runtime that sets operating boundaries for agents running on CPUs, whatever model or harness sits behind them; NVIDIA describes it as broadly available. The second component, Sentry, an out-of-band watchdog on BlueField-4 DPUs, exists for now as a reference design.

SAP is working to embed OpenShell in SAP Joule Studio runtime, part of SAP Business AI Platform. The division of labour is clean: the SAP runtime applies business authorisation, role-based policy and process context before a request reaches execution, and OpenShell governs the execution itself. SAP engineers contribute to OpenShell - separating the supervisor from agent execution, Kubernetes support and observability. The title of SAP’s own post says “working toward”, and FedRAMP and FIPS sit on the roadmap. NVIDIA’s list of infrastructure partners includes Lenovo and SUSE, among others.

The result is two independent checkpoints: one admits the action, the other constrains how it runs. Neither depends on what the model happens to read in its input, which is precisely what separates them from limits written into a system prompt.

Sources: NVIDIA Newsroom, “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment”, 28 September 2026; SAP News Center, Andre Lamego, “SAP and NVIDIA OpenShell Working Toward Governance and Security for Auditable AI Agents in Enterprise Systems”, 28 September 2026; the NVIDIA/OpenShell repository on GitHub. Status: confirmed - read on 1 October 2026. SAP sources give two different end dates for free access to SAP Joule Studio runtime, so we quote neither.

A compact robot arm working inside a sealed glass-and-steel isolator, with two round glove ports on the front

3. Validating tickets before the agent runs in UiPath Maestro

Panagiotis Drakopoulos, a UiPath practitioner, has described on LinkedIn an IT ticketing process built in UiPath Maestro BPMN. Every 15 minutes it reads a mailbox and stores each new message - sender, subject and body - as a separate record. It works through the records one at a time and invokes the agent only once a record has passed validation; his reasoning is cost, since an agent call on meaningless input is money spent for nothing.

The agent classifies the ticket and sets its priority against the supplier’s SLA. The process opens a case in Jira Service Management, notifies the customer and the supplier, and loops until the queue is empty. One note of our own, from the UiPath documentation: the Jira connector in UiPath Integration Service supports Jira Software Cloud, not Server or Data Center. We have not checked whether it handles Jira Service Management projects.

The example shows clearly where control sits in such a design. The process decides when the agent runs and what it receives, and validating first cuts both the bill and the attack surface. Organisations running Jira on their own infrastructure need a different integration path before any demo is built.

Sources: Panagiotis Drakopoulos, “The Evolution of ITSM: Moving from Manual to Agentic Automation!”, LinkedIn, September 2026; UiPath Docs, “About the Jira connector” (Integration Service). Status: confirmed - one practitioner’s design, not a UiPath reference architecture.

A pneumatic tube station on a graphite wall, a clear capsule with an amber light waiting on a steel tray next to the tube diverter

4. UiPath Cartographer and process modelling: a model, not a diagram

Andrzej Sobczak has written on LinkedIn about moving from generating BPMN diagrams out of interview transcripts to building a process model from them. He draws the line precisely: a diagram is a visual depiction of a flow, while a model carries a process hierarchy, a glossary, roles, data, rules, metrics and assumptions, together with an explicit list of unknowns - what the interview could not establish. His toolset records these elements, imports them into Sparx EA 17.1 and links them.

UiPath Cartographer, generally available since UiPath FUSION 2026 on 23 September, does comparable work inside the UiPath stack. From documents, procedures, recordings and interviews it builds a sourced, versioned map of how the process runs today, reconciles conflicting accounts and asks follow-up questions about the gaps. It then designs the target state, assigning each step a mode of work: fully automated, assisted, agent, human review or manual. Finally it generates a PDD as a Word document and an SDD in Markdown for development teams and coding agents. UiPath stresses that the PDD and SDD are outputs of the map, not the source of truth. Cartographer runs on UiPath Delegate; pricing has not been published.

Both approaches move the output of discovery from a picture to a model with its unknowns made explicit. An explicit list of unknowns tells the team what it still has to find out before designing the process, and an agent asked to build that process needs exactly this: hierarchy, rules and a record of what is not yet known.

Sources: Andrzej Sobczak, LinkedIn post, September 2026; UiPath Newsroom, “UiPath Launches UiPath Cartographer Map of Work”, 23 September 2026; UiPath Cartographer product page; UiPath Docs, “Cartographer overview”. Status: confirmed - read on 1 October 2026. Pairing the two approaches is our reading - the author does not mention UiPath.

A surveying total station on a tripod in a new production hall, aimed down an aisle of machines, a gloved hand adjusting the instrument

5. A read-only compromise assessment for Microsoft 365

M365 Investigation Toolkit is an MIT-licensed set of PowerShell 7.5+ scripts for a first-pass compromise assessment of Microsoft 365 and Entra ID. According to its documentation it performs read operations only. It signs in with a delegated admin account, interactively or with a device code, with no app registration and no tokens written to disk.

It collects evidence from mailbox forwarding and rules, transport rules, interactive and non-interactive sign-ins, service principal sign-ins, OAuth consents and privileged role assignments. Version 1.0.0 ships 18 evidence collectors and 10 detections, among them OAuth consent abuse, BEC indicators, password spraying, dormant account takeover and service principal backdoors. The report states a confidence level and flags gaps in what it could collect, and the author is clear that the verdict proves neither compromise nor its absence.

A read-only toolkit can speed up initial triage, yet it still needs broad access. The required permissions include Mail.Read, which exposes message content, so using it for a client requires written contractual authorisation and a GDPR assessment. The project has a single release from March 2026, which should weigh on any decision to build it into an incident response procedure.

Source: the securigeek/M365-Investigation-Toolkit repository on GitHub, README, release notes and licence file, read on 1 October 2026. Status: confirmed as to the documentation; we have not run the tool.

An open rugged case with cut foam holding a set of forensic tools, next to a closed laptop on a graphite desk

6. Agent security as a business analyst’s requirement

In part four of her #BA4AI series on business analysis for AI, Dorota Roszkowska turns prompt injection risk into requirements an analyst can write into a specification. Her premise: an agent allowed to cancel an order can also cancel it on an attacker’s instruction, given directly or hidden in an email, document or web page the agent reads. She frames it as four questions.

Who checks a tool call before it runs? The model only proposes; business logic verifies permissions and parameters, such as a refund limit. Whose permissions does the agent use? Those of the user it acts for, with no broad technical accounts. Which actions are irreversible? Cancellations, payments and data deletion need human confirmation, and the confirmation window shows parameters read from the system, not a description generated by the model. How are input and output filtered? By a separate classification layer before and after the agent, and by separators that keep instructions apart from data.

OWASP lists prompt injection among the top risks for applications built on large language models (LLM01:2025) and says it is unclear whether any fool-proof defence exists. Her requirements therefore limit the damage rather than the attempt, which is why the first question belongs in every agent specification.

Sources: Dorota Roszkowska, “#BA4AI 4/5 Agenci AI: 4 pytania bezpieczeństwa” (in Polish), LinkedIn, 28 September 2026; OWASP, “LLM01:2025 Prompt Injection”. Status: confirmed - read on 1 October 2026. The questions are our translation.

A gloved hand attaching a fourth padlock to a lockout hasp on a machine isolation switch, blank tags hanging from the padlocks

7. Autonomous application pentesting within enforced scope

Pentest Swarm AI is an AGPL-3.0 tool for autonomous penetration testing of APIs and web applications, which its authors position as an open alternative to XBOW. It chains reconnaissance into attack sequences - BOLA and IDOR, JWT forgery, mass assignment, SSRF and injection - and backs findings with collected evidence. It runs on Claude, any OpenAI-compatible API or Gemini, or entirely locally through Ollama or LM Studio.

Two design decisions matter more than the swarm itself. Scope is enforced in the tool layer and again by the executor, and a clean-up registry runs on interruption, failure or budget exhaustion. The authors are explicit about maturity: the default five-phase sequential runner is stable, the agent swarm is alpha and the exploit chains are beta. The README requires the system owner’s written permission before any scan.

Attackers have the same tooling, and reconnaissance of applications and APIs will speed up accordingly. In any tool admitted to your own testing, scope and clean-up must be enforced outside the model. The AGPL-3.0 licence also carries obligations if a modified version is offered as a service - a question for a lawyer.

Source: the Armur-Ai/Pentest-Swarm-AI repository on GitHub, README, licence file and releases (v0.2.31 of 29 September 2026), read on 1 October 2026. Status: confirmed as to the documentation; effectiveness is the authors’ claim, and we have not tested the tool.

A swarm of small inspection drones in formation along the glass-and-steel facade of a new data centre at dusk

8. Aikido Altar-1: code security analysis on your own infrastructure

On 21 September Aikido Security released Altar-1, open weights for code security analysis and pentesting; run on the company’s own infrastructure, it keeps source code in-house. It is neither trained from scratch nor fine-tuned for security; it is a pruned GLM-5.3. Expert pruning (REAP) kept 168 of the 256 experts per layer, and their weights were quantised to INT4. The result takes 328 GB, has 504 billion parameters and is served in vLLM on four H200 cards.

Aikido reports results on its own test set: 32 known vulnerabilities across 30 repositories, three runs per case. Altar-1 reaches a mean recall of 60.4 percent, against 65.6 percent for the full GLM-5.3. The GLM-5.3 licence allows commercial use, modification and redistribution on its own terms, with separate requirements for the largest model-as-a-service operators, but it is not an OSI-approved open-source licence.

Keeping code in-house therefore costs a few points of recall and four Hopper-class cards. That trade-off belongs in the infrastructure budget, and the licence should be read before the first test. This is not legal advice.

Sources: Aikido Security, “Aikido Altar” blog post, 21 September 2026; the AikidoSec/altar-1 model card on Hugging Face; the GLM-5.3 licence. Status: confirmed - read on 1 October 2026. The test-set result is the vendor’s measurement.

An X-ray inspection cabinet in a bright security room, a sealed aluminium case in the inspection tunnel, a gloved hand at the panel

9. AG-UI 1.0: pausing an agent for human approval

On 30 September CopilotKit announced AG-UI 1.0, the stable version of the Agent-User Interaction Protocol, which defines the event stream between an agent backend and a user interface. Every event has a JSON Schema from which the TypeScript, Python and .NET SDKs are generated. The 1.0.0 packages have been available since 17 September, and the release is backwards compatible.

The change that matters most lets an agent pause for human approval and resume from exactly the same point. Subagents become first-class stream objects, so it is clear which agent is doing which part of the work; input and tool results go multimodal; and token usage is reported in the final event. The README lists supported integrations including LangGraph, CrewAI, Microsoft Agent Framework, Google ADK and Mastra. CopilotKit’s framing of the stack: MCP connects agents to tools, A2A connects them to each other, and AG-UI connects them to the application.

The moment of human decision thus becomes a standard event the application can record. The control logic and the audit trail are still the application’s job; the protocol standardises only how the decision is signalled. CopilotKit’s statement that Google, Microsoft, Amazon and Oracle have adopted the protocol is a claim; the integrations themselves are confirmed.

Sources: CopilotKit, “Introducing AG-UI 1.0: a stable spec for connecting any agent to any application”, 30 September 2026; the AG-UI 1.0 spec changelog; the ag-ui-protocol/ag-ui repository on GitHub (MIT licence). Status: confirmed - read on 1 October 2026.

A gloved hand plugging a multi-pin industrial connector into a socket on the door of a steel control cabinet

10. Open decision models, including one for Polish

Within two weeks of Jev’s launch on 15 September, five open models appeared that honour the same contract: they choose among given answers and return a probability for each in a single pass, without generating text. basal-1.0 by Remek Kinas, under Apache 2.0, is fine-tuned on the Polish Bielik v3.0 model in 4.5-billion and 1.5-billion parameter versions and is built for Polish. Its server exposes the same interface as Jev, so an application only changes the service address.

Intern-Decision from InternLM accepts images, and Qwen licence terms apply to its weights. CLM-8B adds small heads to a frozen Qwen3-8B. GLiNER2.5-Decide from Fastino and Julia 1 from Supersonic Labs are small models that run on an ordinary CPU. Liquid AI d1 offers the same contract, but only as a hosted service, without open weights.

Structured decisions are becoming a layer you can run locally, Polish data included. The Polish and English results, however, diverge. The author of basal-1.0 reports 0.884 on Polish decisions against 0.780 for Jev, but 0.740 against 0.861 for Jev on a public English benchmark. The authors of Julia 1 report 64 percent on Banking77 against 87 percent for Jev. All of these figures come from the authors; we have not verified them independently.

Sources: the rkinas/basal repository and the Remek/basal-1.0-4.5B and -1.5B cards on Hugging Face; the InternLM/Intern-Decision repository; the Contrastive-LM/CLM repository and CLM-v0.1-8B card; the fastino/GLiNER2.5-Decide and SupersonicLabs/Julia-1 cards; Liquid AI documentation. Status: confirmed as to existence, licences and sizes - read on 1 October 2026; results not verified.

A hand checking a freshly machined steel part with a go/no-go plug gauge on a granite surface plate

11. Google AX: a sandbox for every agent task

Google has published AX on GitHub, an Apache 2.0 declarative orchestrator for agent workloads that runs on the Agent Substrate layer. A Task runs untrusted agent code in its own sandbox with CPU and memory limits, while a Workspace supplies repositories, MCP servers and skills so the agent starts with its context in place. A task can be suspended and resumed exactly where it stopped, and the command syntax will feel familiar to anyone who uses kubectl.

The authors open the documentation with a warning: AX is in heavy development and will introduce breaking changes before a stable release. The API is v1alpha1, and the latest release is v0.3.1 of 25 September. “Billions of tasks per cluster” is the authors’ ambition, not a measurement.

Agents are a new kind of workload. They accumulate state, call external services and can exhaust a budget in a loop before anyone notices. Isolation and limits have to come from the platform; a prompt cannot impose them. We treat AX as a reference point in discussions about agent runtime environments and would not put it into production yet.

Source: the google/ax repository on GitHub, README and releases, read on 1 October 2026. Status: confirmed; an experimental project, not a Google Cloud service.

A gantry crane in an automated container terminal at dawn placing one container into its slot among neat stacks

12. Apple and AI servers: a press report

The Information, cited by Bloomberg on 16 September, reports that Apple is developing an enterprise server on its own silicon, aimed at AI developers, government and business. Two configurations are planned, with two and with four future M8 Ultra chips, and Apple has reportedly discussed linking the chips with NVIDIA’s NVLink Fusion.

The server would reach the market in 2029 at the earliest, and the project could be cancelled or go ahead without NVIDIA’s technology. The backdrop is unexpectedly strong demand for the Mac mini and Mac Studio among AI developers. Apple did not comment, and no pricing was given.

The report suggests Apple is weighing a server built for running models locally. With no product and no price, it changes no purchasing decision for 2026-2027.

Source: Bloomberg, “Apple Is Developing Enterprise Server for AI Age, Report Says”, 16 September 2026, based on The Information. Status: recorded - a press report based on anonymous sources; The Information article is paywalled and we did not open it.

A gloved hand sliding a smooth, unmarked server chassis into an open black rack in a new server room

13. The book “SAP Cybersecurity from A to Z”

On 29 September SNOK Press published “SAP Cybersecurity from A to Z: A Guide for CIOs, CISOs and SAP Teams”, alongside the Polish original. It is written for CIOs, CISOs, SAP project managers, SAP Basis administrators and SOC teams, and runs to 304 pages, 24 chapters in five parts and appendices A-D. The authors are Jacek Bugajski, the SNOK RedTeam and large language models.

The book covers attack paths into SAP systems, authorisations, AI agents as both threat and defence, a red team lab, detection and response, and NIS2. Each chapter carries a “For decision-makers” box, alongside “Control in practice”, “What to watch for” and “Question for the CIO” boxes.

The decision-maker boxes give executives a way into every chapter without stripping the technical detail out. Both language editions can be downloaded at snok.ai with a business email address.

Source: SNOK Press, edition of 29 September 2026; the resource page and launch post on snok.ai. Status: confirmed - read on 1 October 2026.

A thick closed book with a plain dark cover on a steel bench, next to a laptop and safety glasses, a hand resting on the cover

Reviewing agent controls before rollout

Our pre-production review of an agent starts with the limits that exist only as a sentence in the prompt. Such a limit carries no guarantee of enforcement: crafted content can lead the model to break it. Limits written into permissions, process, runtime and protocol hold regardless of what the model reads.

At SNOK we carry out security assessments of AI agents, UiPath Maestro deployments, deployments of language models on the client’s own infrastructure and conversions to SAP S/4HANA with security controls built in. We would be glad to discuss your case.

Found this useful? Please pass it on:

Get in touch