Context
Databricks Genie Code is an agentic assistant in Databricks that enables users to operate on data in their tenant through natural language instructions.
Genie Code can render displays in the chat for the user; malicious Skills can exploit this functionality to trigger data exfiltration upon tool results being viewed and display phishing attacks to the user.
This attack can occur without any human-in-the-loop approval and without constraint by the controls organizations may be counting on:
- Organization-level Skill governance; Genie Code loads skills from users' personal workspaces, not an organization's governed catalog
- A guardrail agent that assesses commands before they are run; Databricks considers this a best effort productivity feature, not a security control
- Coding-environment egress controls; egress occurs from the user's browser, not the coding environment
- Sandboxing of displays rendered in chat; the Skill builds the display with user data, the display does not need to reach into the tenant to get it
This risk was disclosed to Databricks on August 16, 2026, but it was determined to fall under the “user's responsibility to ensure that uploaded skills do not contain malicious content”. Given the prevalence of malicious Skills being distributed in online marketplaces, and that the controls organizations are likely relying on do not eliminate this risk, we are releasing this report to inform those at risk.
The Attack Chain
Genie is prompted to analyze data using an uploaded Skill
Skills are commonly distributed via online marketplaces, an ecosystem known to be polluted with malicious Skills. Databricks does have a system for Skill governance, but the Skills used by Genie Code are configured from users’ personal workspaces, not an organization’s governed Catalog.
Genie executes code from the malicious Skill
When Genie executes code from the Skill, it is checked by a second agent that flags commands that act outside the user’s intended action, such as those that “send data to third parties”.
The agent-in-the-loop approves running the Skill code, failing to catch the malicious capabilities.
Genie instructs the user to open the full data analysis results
When the results render, a phishing modal is displayed and the user's dataset is exfiltrated
The code in the malicious Skill builds an HTML element that serves two malicious purposes. One effect is that it renders an overlay of an attacker's website that phishes the user for credentials.
The other effect is that when the modal is rendered (e.g., when the user views the results), data from the victim's tenant is exfiltrated. The malicious Skill's code uses its access to the victim's tenant to collect sensitive data (like the victim's datasets) and then builds that data into the HTML display. When the display renders, JavaScript in the malicious display causes the user's browser to issue network requests that exfiltrate the data to an attacker's server.
While the coding environment has network egress controls that would prevent communication with the attacker's server, the display interface does not and can send data to the attacker.
Responsible Disclosure
This risk was reported to Databricks on August 16, 2026. The Databricks team determined that:
“it is ultimately the user's responsibility to ensure that uploaded skills do not contain malicious content.”
Network egress controls dictated that the coding environment was not allowed to contact untrusted external parties, and that was upheld. It just queried the data and used it to build an iframe.
Controls for displays rendered to the user dictated that the display was not allowed to query data from the tenant, and that was upheld. It just processed the data embedded in it by the Skill’s code.
As such, the malicious Skill code can query data in the computing environment, build it into an iframe, and the iframe can make network requests to exfiltrate the data…
We believe this interaction represents a gap in the Databricks threat model, as the purpose of the two guarantees that were technically upheld was to prevent the data exfiltration outcome that ultimately occurred.
With respect to the malicious command being approved by the guardrail agent, the Databricks team noted,
“The auto-allow functionality is not intended as a security boundary, but rather as a control measure to prevent untrusted input from running automatically.”
Documentation supports the statement:
However, as ‘auto-approve’ is documented as both the default and ‘recommended’ command approval mode, we believe it is unlikely that users will, in practice, adhere to the instruction discouraging its use with production data.
No statement was given related to the fact that Genie Code's Skills are loaded from users' personal workspaces, instead of the organization-level governed Catalog.
Timeline
| August 16, 2026 | PromptArmor discloses to Databricks |
| August 17 - September 15, 2026 | PromptArmor and Databricks coordinate |
| September 16, 2026 | PromptArmor informs Databricks of intent to publish |
