Weekly review W41, 2-8 October 2026: AI agents are reaching SAP, UiPath and Snowflake faster than the controls around them.
Putting this week together, I kept picturing a test cell where the car is already spinning its wheels on the rollers while someone at the bench is still fitting the brake callipers. That is the agent market right now. UiPath has connected to Snowflake without copying data, and SAP is buying a work graph for its HR agents while publishing the official route by which non-SAP agents should enter its systems. In the same week, researchers showed that the Jev decision model gives away hidden personal data through nothing more than its choice of answer, Apple announced stricter consent for full disk access because of agents, and Gartner forecast that most agentic systems built by vendors’ engineers will be abandoned.
The brakes can be fitted - provided someone thinks about them before the car leaves the test cell.
1. UiPath and Snowflake: the data stays put, the agent comes to it
On 30 September UiPath announced a two-way integration with Snowflake. UiPath Data Fabric reads and models data directly in Snowflake, without copying or moving it. UiPath Cartographer and Delegate reach Snowflake Cortex through Integration Service. Snowflake CoCo users can call UiPath agentic skills and trigger automations from inside the data platform. UiPath Maestro orchestrates the lot, and UiPath is now listed on Snowflake Marketplace.
I like the direction because it answers a recurring objection: data teams do not want an agent walking off with a copy of the data and living outside the access policies. Here the data stays where its permissions and lineage live, and the record of what the agent did stays in UiPath. The flip side is that automations can now be triggered from the data platform, so a security review has to cover both directions - what Data Fabric reads from Snowflake, and what UiPath skills called from CoCo are allowed to do. That is why an AI Security review is mandatory in our UiPath projects: these questions need answering before go-live, not after the first incident.
Source: UiPath press release, “UiPath and Snowflake Announce Two-Way Integration Connecting Automation and Governed Data, Now Available on Snowflake Marketplace”, 30 September 2026 (Business Wire reprint). Status: confirmed, read 8 October 2026.

2. Half of organisations run agents in production - according to the UiPath community
In June and July UiPath surveyed 1,226 people who build automation and AI. 53% report agents in production: 30% in a limited number of use cases and 23% broadly. Another 23% are piloting. The most cited barrier to production is data security or privacy (40%), ahead of cost (39%) and difficulty proving ROI (31%). Only 7% see no barrier at all.
The more telling figure sits further in. Among organisations using agents broadly, 72% have fully documented and enforced governance; among those still evaluating, 36%. The organisations that have gone furthest with agents have not relaxed control: twice as many of them report full governance. A survey cannot say which came first, but the two clearly go together. One honest caveat: this is a vendor survey, respondents were recruited mainly through UiPath’s own channels, and 54% are from Asia-Pacific. It describes the UiPath community, not the whole market.
Source: UiPath, “2026 State of AI and Automation Professionals”, pp. 15, 16, 26, 48-50. Status: figures match the report; vendor survey.

3. SAP buys TechWolf: an HR agent needs a graph, not just a model
On 6 October SAP announced an agreement to acquire TechWolf of Ghent. It is buying a “context graph for work” - a model of the tasks inside jobs, people’s skills and the external labour market, built from existing HR systems. TechWolf is to become a core of SAP SuccessFactors and the layer on which SAP Joule bases HR decisions: skills-based hiring, workforce planning, role redesign. Financial terms were not disclosed; closing is expected in Q4.
One sentence from SAP’s Manoj Swaminathan stood out: the graph “makes token usage more efficient, lowers the cost of deploying workforce agents”. That is an open admission that a language model alone is not enough and that an agent’s cost depends on the quality of the context underneath it - the same conclusion we have reached working with knowledge graphs. The less comfortable side: a graph of what employees do and what they can do is personal data about them. Who will see it through SAP Joule, and with which authorisations, is a question for the SAP security team before any new product reaches the market.
Source: SAP News Center, “SAP to Acquire TechWolf, Giving Enterprises Evidence-Based View of Work in the Age of AI”, 6 October 2026. Status: recorded - full release in our notes.

4. SAP Architecture Center: the official route for non-SAP agents
For months, non-SAP agents have also reached SAP systems through unofficial MCP servers built on top of OData. SAP Architecture Center now sets out, in its reference architecture “A2A and MCP for Interoperability”, how SAP wants this done. There are three routes: Agent Gateway using the A2A protocol for traffic to SAP Joule agents; MCP Gateway in SAP Integration Suite for exposing SAP and non-SAP APIs as tools; and A2A through SAP Integration Suite, which takes over authentication, rate limiting and monitoring.
Having the map is useful - it gives every architecture conversation a reference point. But the caveat SAP put at the very top matters: some components are not yet generally available, and bidirectional communication through Agent Gateway with third-party and self-hosted agents “is not yet supported”. Bring Your Own Agent also handles text messages only, with a 60-second limit in synchronous mode. Anyone designing a UiPath agent to work with SAP Joule today should plan for the current state, not the roadmap.
Source: SAP Architecture Center, “A2A and MCP for Interoperability”, updated 27 August 2026. Status: confirmed, read 8 October 2026.

5. Jev versus SAP RPT: who should make the small decisions inside ERP
I recently wrote about Jev, TypeSafe’s decision model that returns a typed decision with a probability instead of generating text. Now SAP researchers have published, on SAP Community, a comparison of Jev with SAP RPT-1.6, SAP’s tabular foundation model. According to the authors, RPT-1.6 scores 76.75% on typed-decision tasks against 72.70% for Jev 1.13.0. On ticket assignment from an internal SAP system the gap widens: 68.9% against 49.4%, with RPT-1.6 taking 44 milliseconds and Jev 322.
Read these numbers with care. RPT’s latency was measured directly on an H200 GPU without a production deployment, and the ticket benchmark cannot be reproduced outside SAP. The direction matters more than the score: SAP is telling customers that the decision layer need not be a separate model bought alongside ERP, because a tabular model inside SAP already speaks that language. For SAP customers this is a genuine architectural choice - decisions inside the platform or next to it - and worth making deliberately, with your own measurements on your own data.
Source: J. von Rueden, M. Boerner, G. Schindler, “Jev vs. RPT: The Decision Layer Is Here, and RPT Already Speaks Its Language”, SAP Community, Technology Blog by SAP, 5 October 2026. Status: confirmed; figures as reported by the authors.

6. HoneySAP 0.2.0: a decoy for a vulnerability rated 9.9
On 6 October OWASP released HoneySAP 0.2.0, a honeypot that emulates SAP services. The new trap impersonates an SAP RFC Gateway: it handles the NWRFC SDK handshake, presents function module catalogues and captures credentials. One feature caught my eye. HoneySAP’s documentation treats calls to /SLOAE/DEPLOY as attempts to exploit CVE-2025-42957 - ABAP code injection in SAP S/4HANA, rated CVSS 9.9 and patched on SAP Security Patch Day in August 2025.
The logic is simple, which is why it works. Nobody calls that module from outside in normal operation, so any call the decoy records is suspicious and points to an attempt to exploit a known flaw. Patching and hardening are the foundation of SAP security, but they say nothing about whether someone is trying right now; a honeypot tells you that cheaply. It is a community tool without vendor support, so in a client environment it needs a repository scan, network isolation and agreement with the SOC on where the events go.
Source: OWASP HoneySAP release v0.2.0, GitHub, 6 October 2026; NVD, CVE-2025-42957. Status: confirmed.

7. Jev’s decisions give away hidden data: a typed answer is not a privacy control
This was the piece that stayed with me this week. On 4 October a team led by Shang Wang posted to arXiv the first systematic study of Jev’s security and privacy when it is used as an application’s decision layer, testing both the official API and NanoJev, a locally run model. The privacy result worried me most. In a support-routing task, an attacker who controls only the text of a request and sees only the resulting decision was able to recover gender, age group, heart disease status and ethnicity held in the hidden application state. All it took was a pair of queries with the condition reversed: if the decision flips along with the condition, the attribute is known. According to the authors, the official Jev gave away all four attributes across all 200 test states. Hiding the probabilities makes no difference, because the choice of answer already carries the information.
The other results point the same way. Two injection attacks designed specifically for Jev steered its decision to the attacker’s chosen answer in at least 97% of cases on each of the three tasks tested. A local model downloaded from a public hub could be backdoored through poisoned training data, with the trigger working in more than 89% of attempts in most settings while accuracy on clean inputs barely moved. And 200 of Jev’s answers were enough to train a surrogate model that matched it on 92% of cases for the same task.
The lesson reaches beyond Jev to any model that makes decisions inside a process: a typed answer constrains the format, not the leakage. The authors recommend separating trusted application context from user-supplied input, validating the decision against trusted data and enforcing authorisation before any action runs, and minimising the sensitive fields the model sees at all. In my view that is the baseline for any agent that decides on client or employee data.
Source: S. Wang, T. Zhu, H. Chen, J. Li, M. Yang, B. Liu, “Hidden Risks of Jev: An Empirical Study of Security, Privacy, and Dual Use”, arXiv:2610.04985, 4 October 2026. Status: confirmed against the full text; figures as reported by the authors, preprint not peer reviewed.

8. Anthropic opens Red Team Access - at the price of data retention
On 6 October Anthropic expanded its Cyber Verification Program to three access tiers. Defense Access covers defensive work: SOC, incident response, malware analysis. Red Team Access lifts safeguards for authorised penetration testing and red teaming, for organisations only - individual researchers are not eligible, and review takes a few weeks. Specialized Access is reserved for a small group testing critical systems, vetted in cooperation with the US government; Project Glasswing members move into it automatically.
For a firm that runs penetration tests, a frontier model without safeguards inside an authorised scope is an interesting offer. One condition changes the calculation, though: data retention is mandatory in the programme. The exception is narrow: until Enterprise Frontier Safeguards launches, organisations with zero-data-retention access to Claude Fable 5.1 or Claude Mythos 5.1 can also use the programme without retention. On client engagements under NDA and within an ISO 27001 regime, any decision to send test data to an outside provider needs a security assessment and the client’s consent. For us that is a decision to weigh carefully, not a reflex.
Source: Anthropic, “Expanding the Cyber Verification Program”, 6 October 2026. Status: confirmed.

9. OWASP FinBot: a range where attacking the agent is the point
OWASP is promoting FinBot again this week - an intentionally vulnerable agentic platform billed as “the Juice Shop for Agentic AI”. It is a multi-agent vendor management system covering onboarding, fraud detection and invoice processing, and the task is to break it through prompt injection, tool misuse, data exfiltration and even remote code execution. The challenges map to the OWASP Top 10 for LLM, the Top 10 for Agentic Applications, CWE and MITRE ATLAS. It runs in the browser, and the code is on GitHub.
The platform is not new; it launched in the spring. It earns its place because it fits this week so well. FinBot gives teams a place to rehearse an attack on an agent without causing harm, before someone tries it on a live system. If your team ships agents and has never tried to break one, this is a good place to start.
Source: OWASP GenAI Security Project, owasp-finbot-ctf.org; GenAI-Security-Project/finbot-ctf, GitHub. Status: confirmed.

10. Apple: agents change the risk of full disk access
On 2 October Apple published a developer note on changes to Full Disk Access in macOS, and stated the reason plainly: “As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially.” Apple plans additional controls so that an app can be granted full disk access only through very explicit user action.
Moves like this shift the debate from statements of intent to the operating system itself. An agent granted access to the whole disk with one click can see mail, documents and keys. Apple is not banning that access; it is making it a deliberate decision. Companies should treat agent permissions the same way - as a decision with an owner, not a default setting. The note gives no macOS version or date yet, so for now it is an announcement rather than a change to roll out.
Source: Apple Developer News, “Updates to Full Disk Access in macOS”, 2 October 2026. Status: confirmed.

11. HyperFrames Studio: the agent edits, the human decides
In early October HeyGen released HyperFrames Studio, a desktop app for macOS and Linux described as a video editor built for agents. You describe the film, and Claude Code or Codex builds a first cut. Then both work on the same timeline: the human edits by hand, draws notes on a frame or comments on a stretch of footage, and the agent applies the changes.
It is a detour from SAP and security, but it shows this week’s theme from another angle. The agent produces a first version at a pace no person can match, and the human keeps the decisions - the same division of roles we try to build into process automation. One caveat before using it at work: the app requires a HeyGen account, and the documentation does not say clearly whether projects and prompts reach the vendor’s servers. With company material, that is the first question to ask.
Source: HyperFrames Studio documentation, HeyGen, read 8 October 2026; heygen-com/hyperframes, GitHub. Status: confirmed.

12. OpenAI loses safety people while the industry argues about pace
Three stories from recent weeks form one picture. On 1 October, reports citing the Wall Street Journal said OpenAI had dismissed three researchers for breaching rules on access to confidential information; on 8 October the three disputed the allegations. On 3 October David Robinson, who spent three and a half years preparing safety reports for OpenAI’s launches, resigned and wrote in The Atlantic that the company’s culture “is broken”. In the background, a consumer class action accuses Anthropic, OpenAI, Google and SpaceXAI of colluding to slow AI development - allegations by the plaintiffs, with the case undecided. And in September Sam Altman endorsed Dario Amodei’s essay on pacing the frontier.
I am not passing judgement on who is right. My conclusion is different: the labs are simultaneously talking about slowing down, being sued for slowing down and losing their safety staff. Agent security inside a company cannot rest on a model provider’s assurances. It has to rest on what you control yourselves: permissions, process, logs and a review before go-live.
Sources: TechCrunch, 3 and 8 October 2026; Decrypt citing the WSJ, 1 October 2026; Bloomberg Law, 18 September 2026 (Buist v. Anthropic PBC, N.D. Cal.); NBC News, 12 September 2026. Status: confirmed by press reports; lawsuit undecided.

13. Gartner: 70% will abandon agentic systems built by vendors’ engineers
Gartner predicts that by 2028, 70% of enterprises will abandon agentic AI built through vendor forward-deployed engineering - systems assembled by a vendor’s engineers working on site. Gartner cites rising costs and the client’s inability to develop such systems further on its own.
I should declare an interest: SNOK is an implementation firm, and we build agents for clients. That is exactly why I take this forecast seriously. An agent that the client cannot maintain, change and understand on its own is a debt that grows with every process change. That is why I believe knowledge transfer and documentation have to be part of the implementation scope from day one. The capability has to stay with the client, or the agent will not survive the next budget review.
Source: The Register, Dan Robinson, “7 in 10 enterprises expected to abandon vendor-built agentic AI by 2028”, 30 September 2026; Gartner press release, 29 September 2026. Status: confirmed.

14. Mistral Large 4 in Polish: sovereignty is not the same as competence
On 6 October Mistral released Mistral Large 4 in public preview, a one-trillion-parameter model with 52 billion active parameters, under the line “Forged in Europe. Built for AI sovereignty.” On PLCC, a Polish benchmark of linguistic and cultural competence, Mistral Large 4 averages 68.17. Bielik-11B-v3.0-Instruct scores 70.67 and PLLuM-12B, in its nc-chat-250715 variant, 69.67. Mistral’s previous version, Large 2512, also scored 70.67.
This matters. A European model is not automatically a good model for a Polish client: on this test a small Polish model outscores the European flagship and matches its predecessor. One caveat: the lead holds for the best Bielik and PLLuM variants, while others score lower. When choosing a model for on-premises deployment, what counts is measurement on your own Polish-language tasks, not the flag on the press release.
Sources: Mistral AI, “Mistral Large 4”, 6 October 2026; PLCC leaderboard (sdadas/plcc, Hugging Face), data as of 6 October 2026. Status: confirmed; averages computed from leaderboard data.

Brakes you can fit today
My practical list from this week. Agent permissions are a decision with an owner. A model’s decision is checked against trusted data before any action runs, and sensitive fields are kept to the minimum. Security reviews cover both directions of an integration. The capability to maintain an agent stays with the client. Models are chosen by measurement on your own data.
SNOK weekly review · W41 · 2-8 October 2026
The whole edition in one file
PDF, seventeen pages, about 2 MB - no form, no details to hand over. A version to pass around the team or read on a phone.
Download edition W41 (PDF)If you are deploying agents in SAP or UiPath and want to see where your brakes are still on order, take a look at our AI Security, SAP security and on-premises LLM services. Questions: office@snok.ai.
Jacek Bugajski, CEO, SNOK
