barha's picture
aLoRA retrain of genai-attack-vector on granite-4.0-350m (gated on assistant invocation tokens)
85ba633 verified
Raw
History Blame
219 kB
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Description Several filters in the `twig/*` extras packages are registered with `is_safe => ['all']`, which tells Twig's autoescaper to treat their output as safe in every context (`html`, `js`, `css`, `url`, ...). The output of these filters is plain text or HTML markup, neither of which is safe in every escaping context. Affected filters: - `html_to_markdown` (`twig/markdown-extra`) emits plain Markdown text. `league/html-to-markdown` decodes HTML entities when producing code spans and fenced blocks, so an attacker-controlled `<code>&lt;img src=x onerror=alert(1)&gt;</code>` becomes `` `<img src=x onerror=alert(1)>` ``, which renders live when interpolated into an HTML page. - `markdown_to_html` (`twig/markdown-extra`) emits HTML. Safe in an HTML context but not in JS, CSS or URL contexts (e.g. when interpolated into an inline `<script>` block). - `inline_css` (`twig/cssinliner-extra`) emits HTML with inlined styles. Same constraint as `markdown_to_html`. In all three cases, `is_safe => ['all']` causes the autoescaper to emit the output verbatim in any context, even when the developer never wrote `|raw`. In a context such as a JS string or a URL parameter, this produces unescaped HTML and is exploitable as XSS. ### Resolution - `html_to_markdown` no longer claims to be safe in any escaping context; its plain-text output is now autoescaped for the surrounding context. - `markdown_to_html` and `inline_css` are now declared `is_safe => ['html']`, asserting only what\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAn authenticated stored cross-site scripting (XSS) vulnerability in the creation/editing module of Feehi CMS v2.1.1 allows attackers to execute arbitrary web scripts or HTML via injecting a crafted payload into the Content field.\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact In the Prometheus server's legacy web UI (enabled via the command-line flag `--enable-feature=old-ui`), the histogram heatmap chart view does not escape `le` label values when inserting them into the HTML for use as axis tick mark labels. An attacker who can inject crafted metrics (e.g. via a compromised scrape target, remote write, or OTLP receiver endpoint) can execute JavaScript in the browser of any Prometheus user who views the metric in the heatmap chart UI. From the XSS context, an attacker could for example: - Read `/api/v1/status/config` to extract sensitive configuration (although credentials / secrets are redacted by the server) - Call `/-/quit` to shut down Prometheus (only if `--web.enable-lifecycle` is set) - Call `/api/v1/admin/tsdb/delete_series` to delete data (only if `--web.enable-admin-api` is set) - Exfiltrate metric data to an external server Note that this only affects users who have explicitly enabled the legacy Prometheus web UI using the `--enable-feature=old-ui` command-line flag. ### Patches https://github.com/prometheus/prometheus/commit/38f23b9075ced1de2b82d2dad8b2bebb1ecd5b7d ### Workarounds If at all possible, disable the legacy web UI by removing the `--enable-feature=old-ui` command-line flag). If this is not an option, take the following precautions: - If using the remote write receiver (`--web.enable-remote-write-receiver`), ensure it is not exposed to untrusted sources. - If using the OTLP receiver (`--web.enable-otlp-\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Overview On many platforms, a third party can create a Git repository under a name that includes a shell command substitution [^1] string in the syntax `$(<command>)`. These directory names are allowed in macOS and a majority of Linux distributions [^2]. If a user starts `jupyter-lab` in a parent directory of this inappropriately-named Git repository, opens it, and clicks \"Git > Open Git Repository in Terminal\" from the menu bar, then the injected command `<command>` is run in the user's shell without the user's permission. This issue is occurring because when that menu entry is clicked, `jupyterlab-git` opens the terminal and runs `cd <git-repo-path>` through the shell to set the current directory [^3]. Doing so runs any command substitution strings present in the directory name, which leads to the command injection issue described here. A previous patch provided an incomplete fix [^4]. [^1]: https://www.gnu.org/software/bash/manual/html_node/Command-Substitution.html [^2]: https://www.gnu.org/software/libc/manual/html_node/File-Name-Portability.html [^3]: https://github.com/jupyterlab/jupyterlab-git/blob/7eb3b06f0092223bd5494688ec264527bbeb2195/src/commandsAndMenu.tsx#L175-L184 [^4]: https://github.com/jupyterlab/jupyterlab-git/pull/1196 ## Scope of Impact This issue allows for arbitrary code execution via command injection. A wide range of actions are permitted by this issue, including but not limited to: modifying files, exfiltrating data, halting services, or c\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nJohann Rehberger (Embrace The Red) demonstrated that ChatGPT plugins (e.g., WebPilot, Cross-Plugin Request Forgery) could exfiltrate conversation data via Markdown image rendering and indirect prompt injection, including auto-invocation of plugins without user consent.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The `fetch-apify-docs` tool validates URLs against a domain allowlist using `String.startsWith()` instead of proper URL hostname comparison. This allows bypass via attacker-controlled subdomains (e.g., `https://docs.apify.com.evil.com/`), enabling the tool to fetch and return arbitrary web content to the LLM. ### Details #### Vulnerable component `src/tools/common/fetch_apify_docs.ts`, line 51: ```typescript const isAllowedDomain = ALLOWED_DOC_DOMAINS.some((domain) => url.startsWith(domain)); ``` `src/const.ts`, lines 167-170: ```typescript export const ALLOWED_DOC_DOMAINS = [ 'https://docs.apify.com', 'https://crawlee.dev', ] as const; ``` #### How the bypass works `String.startsWith('https://docs.apify.com')` matches any string beginning with that prefix, including: - `https://docs.apify.com.evil.com/payload` - attacker-controlled subdomain - `https://docs.apify.com@evil.com/payload` - userinfo component in URL (browser behavior varies, but `fetch()` in Node.js may follow this) - `https://docs.apify.com.evil.com:8080/path` - custom port on attacker domain All of these pass the `startsWith` check because they begin with the exact string `https://docs.apify.com`. #### The fetched content is returned to the LLM After the allowlist check passes, the tool fetches the URL and returns the full page content as markdown (`fetch_apify_docs.ts:69-103`): ```typescript const response = await fetch(url); // ... const html = await response.text(); markdown\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nLangflow is a tool for building and deploying AI-powered agents and workflows. Prior to 1.9.0, Langflow is vulnerable to Path Traversal in the Knowledge Bases API (DELETE /api/v1/knowledge_bases). This occurs because user-supplied knowledge base names are concatenated directly into file paths without proper sanitization or boundary validation. An authenticated attacker can exploit this flaw to delete arbitrary directories anywhere on the server's filesystem, leading to data loss and potential service disruption. This vulnerability is fixed in 1.9.0.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Grok spews misinformation about Bondi Beach mass shooting. System: Grok. Technology: Generative AI. Purpose: Check facts. Ethical issues: Accuracy/reliability; Mis/disinformation; Transparency. Response: System review/update.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability was detected in JoeCastrom mcp-chat-studio up to 1.5.0. Affected by this issue is some unknown functionality of the file server/routes/llm.js of the component LLM Models API. Performing a manipulation of the argument req.query.base_url results in server-side request forgery. Remote exploitation of the attack is possible. The exploit is now public and may be used. The project was informed of the problem early through an issue report but has not responded yet.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n# Authenticated Local File Inclusion (LFI) via selectobject.php leading to sensitive data disclosure ## Target Dolibarr Core (Tested on version 22.0.4) ## Summary A Local File Inclusion (LFI) vulnerability has been discovered in the core AJAX endpoint `/core/ajax/selectobject.php`. By manipulating the `objectdesc` parameter and exploiting a fail-open logic flaw in the core access control function `restrictedArea()`, an authenticated user with no specific privileges can read the contents of arbitrary non-PHP files on the server (such as `.env`, `.htaccess`, configuration backups, or logs\u2026). ## Vulnerability Details The vulnerability is caused by a critical design flaw in `/core/ajax/selectobject.php` where dynamic file inclusion occurs **before** any access control checks are performed, combined with a fail-open logic in the core ACL function. - **Arbitrary File Inclusion BEFORE Authorization:** The endpoint parses the `objectdesc` parameter into a `$classpath`. If `fetchObjectByElement` fails (e.g., by providing a fake class like `A:conf/.htaccess:0`), the application falls back to `dol_include_once($classpath)` at **line 71**. At this point, the arbitrary file is included and its content is dumped into the HTTP response buffer. This happens *before* the application checks any user permissions. - **Access Control Bypass (Fail-Open):** At **line 102**, the application finally attempts to verify permissions by calling `restrictedArea()`. Because the object creation faile\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary Missing validation logic in the storage volume import logic allows an authenticated user with access to Incus' storage volume feature to cause the Incus daemon to crash. Repeated use of this issue can be used to keep Incus offline causing a denial of service. ### Details The backup restore subsystem contains an out-of-bounds panic vulnerability caused by an invalid bounds check when indexing snapshot metadata arrays. The same flawed pattern also appears in the migration path. When iterating through physical snapshots provided in a backup archive, the loop uses the index i to look up corresponding metadata in the parsed Config.Snapshots and Config.VolumeSnapshots slices. To ensure that the metadata slice is long enough, the code uses the guard condition len(slice) >= i-1. This check is incorrect because it can still evaluate to true when the subsequent slice[i] access is out of bounds, including when i >= len(slice), triggering a runtime panic. An attacker can trigger this by submitting a backup archive that contains physical snapshot directories, which drive the loop variable i, while supplying a tampered index.yaml with an empty or truncated snapshot metadata array. This causes the daemon to index beyond the end of the metadata slice and crash, resulting in immediate denial of service on the node. Affected File: https://github.com/lxc/incus/blob/v6.22.0/internal/server/storage/backend.go Affected Code: ``` func (b *backend) CreateInstanceFromBackup(srcBacku\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Republicans support Trump using AI-generated kitten and duck images. Technology: Generative AI. Purpose: Defend reputation. Ethical issues: Mis/disinformation; Transparency.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary An authentication bypass vulnerability in phpMyFAQ allows any unauthenticated attacker to reset the password of any user account, including SuperAdmin accounts. By sending a PUT request with just a valid username and associated email address to /api/user/password/update, an attacker receives a new plaintext password via email without any token verification, rate limiting, or email confirmation. This enables complete account takeover of any user, including full administrative access. ### Details File: phpmyfaq/src/phpMyFAQ/Controller/Frontend/Api/UnauthorizedUserController.php Lines: 56-130 The updatePassword() method at line 56 accepts PUT requests to /user/password/update with only username and email in the JSON body: #[Route(path: 'user/password/update', name: 'api.private.user.password', methods: ['PUT'])] ```php public function updatePassword(Request $request): JsonResponse { $data = json_decode($request->getContent()); $username = trim((string) Filter::filterVar($data->username, FILTER_SANITIZE_SPECIAL_CHARS)); $email = trim((string) Filter::filterEmail($data->email)); if ($username !== '' && $username !== '0' && ($email !== '' && $email !== '0')) { $user = ($this->currentUserFactory ?? CurrentUser::getCurrentUser(...))($this->configuration); $loginExist = $user->getUserByLogin($username); if ($loginExist && $email === $user->getUserData('email')) { // NO TOKEN CHECK // NO RATE LIMITING\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nn8n-MCP is an MCP server that provides AI assistants access to n8n node documentation, properties, and operations. Prior to version 2.47.11, when n8n-mcp runs in HTTP transport mode, incoming requests to the POST /mcp endpoint had their request metadata written to server logs regardless of the authentication outcome. In deployments where logs are collected, forwarded to external systems, or viewable outside the request trust boundary (shared log storage, SIEM pipelines, support/ops access), this can result in disclosure of: bearer tokens from the Authorization header, per-tenant API keys from the, x-n8n-key header in multi-tenant setups, JSON-RPC request payloads sent to the MCP endpoint. Access control itself was not bypassed \u2014 unauthenticated requests were correctly rejected with 401 Unauthorized \u2014 but sensitive values from those rejected requests could still be persisted in logs. This issue has been patched in version 2.47.11.\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary diesel-async exposes uninitialized stack padding to safe code on every read of a MySQL `DATE`, `TIME`, `DATETIME`, or `TIMESTAMP` column. Reading that buffer is undefined behavior, and the leaked bytes can contain stale heap/stack contents, so this is both a soundness bug and a potential information-disclosure vector. ### Details In `diesel-async/src/mysql/row.rs` (lines 65-103), `MysqlRow::get` builds a `MysqlTime` from the parsed `mysql_async::Value` and then fabricates the byte buffer that downstream `FromSql` impls expect like this: ```rust let date = MysqlTime::new(/* fields from Value::Date / Value::Time */); let buffer = unsafe { let ptr = &date as *const MysqlTime as *const u8; let slice = std::slice::from_raw_parts(ptr, std::mem::size_of::<MysqlTime>()); slice.to_vec() }; ``` `MysqlTime` is `#[repr(C)]` with 3 bytes of padding after `bool neg` (Linux x86_64, offsets 0x21..0x23). The literal construction leaves that padding uninitialized, and `to_vec()` carries it into a `Vec<u8>` that becomes the `MysqlValue`'s backing buffer, reachable from safe code via `MysqlValue::as_bytes() -> &[u8]`. `diesel` itself avoids this by going through `MaybeUninit::<MysqlTime>::zeroed()` + `ptr::copy_nonoverlapping` (see `diesel/src/mysql/value.rs:43-94`); the same pattern would fix this. Alternatively, write the bytes diesel's `FromSql` reads without round-tripping through a `MysqlTime` value. ### PoC `Cargo.toml`: ```toml [dependencies] diesel = { ver\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The `list_files()` tool in `FileTools` validates the `directory` parameter against workspace boundaries via `_validate_path()`, but passes the `pattern` parameter directly to `Path.glob()` without any validation. Since Python's `Path.glob()` supports `..` path segments, an attacker can use relative path traversal in the glob pattern to enumerate arbitrary files outside the workspace, obtaining file metadata (existence, name, size, timestamps) for any path on the filesystem. ## Details The `_validate_path()` method at `file_tools.py:25` correctly prevents path traversal by checking for `..` segments and verifying the resolved path falls within the current workspace. All file operations (`read_file`, `write_file`, `copy_file`, etc.) route through this validation. However, `list_files()` at `file_tools.py:114` only validates the `directory` parameter (line 127), while the `pattern` parameter is passed directly to `Path.glob()` on line 130: ```python @staticmethod def list_files(directory: str, pattern: Optional[str] = None) -> List[Dict[str, Union[str, int]]]: try: safe_dir = FileTools._validate_path(directory) # directory validated path = Path(safe_dir) if pattern: files = path.glob(pattern) # pattern NOT validated \u2014 traversal possible else: files = path.iterdir() result = [] for file in files: if file.is_file(): stat = file.stat() result.ap\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMicrosoft allowed contractors in China to access and review thousands of Skype and Cortana audio recordings, including sensitive conversations, without proper security measures or employee vetting. The lack of safeguards exposed users\u2019 private data to potential criminal or state access, constituting a significant privacy violation linked to AI system use.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary Component state manipulation is possible in `django-unicorn` due to missing access control checks during property updates and method calls. An attacker can bypass the intended `_is_public` protection to modify internal attributes such as `template_name` or trigger protected methods. ## Vulnerability Details: Component Access Control Bypass Security analysis identified that the framework fails to enforce visibility boundaries defined by `_is_public` within the action parsers. Specifically, the logic in `set_property_value()` and `_call_method_name()` utilizes `getattr` and `setattr` directly on component instances without verifying if the target attribute or method is explicitly marked as public. Vulnerability resides in: - `src/django_unicorn/views/action_parsers/call_method.py` - `src/django_unicorn/views/action_parsers/utils.py` While Django's template engine restricts rendering to registered directories, an unauthorized user can still force a component to render sensitive templates (e.g., admin layouts) from other installed applications or reset the component state by invoking the internal `reset()` method. ## Proof of Concept (PoC) Attacker can overwrite the `template_name` attribute by sending a crafted JSON payload to the message endpoint: 1. Construct a payload targeting a protected attribute: ```json { \"actionQueue\": [ { \"type\": \"syncInput\", \"payload\": { \"name\": \"template_name\", \"value\": \"admin/base.html\" } }\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nFacebook&#x27;s and Twitter were not able to sufficiently moderate content of small language groups such as the Balkan languages using AI, allegedly due to the lack of investment in human moderation and difficulty in AI-solution design for the languages.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA previously undisclosed facial recognition initiative operated by Israeli military intelligence units was reportedly deployed across Gaza after the October 7, 2023 attacks. According to multiple intelligence officers, the program uses Corsight technology alongside Google Photos to identify individuals from checkpoints, crowds, and drone footage. The system has allegedly produced misidentifications, including the widely reported detention of Palestinian poet Mosab Abu Toha on November 19, 2023.\n</incident>"}], "ground_truth": "rce", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe Australian government operates a voiceprint authentication system used by Centrelink and the Australian Taxation Office (ATO) that allows people to verify their identity over the phone using voice biometrics combined with customer reference numbers. As of February, 3.8 million Centrelink clients used the voiceprint service, and over 7.1 million people had verified their voice with the ATO. Guardian Australia conducted an investigation following reports of AI voice cloning attacks on overseas banking services. A Guardian journalist successfully created an AI clone of their own voice using just four minutes of audio and used this clone, combined with their customer reference number, to gain access to their own Centrelink self-service account. The voiceprint system is described by Services Australia as 'secure, accurate and reliable' and claims it can detect when someone is using a recording. However, the investigation demonstrated the vulnerability to AI-generated voice synthesis. The self-service phone system allows access to sensitive information including benefit payments and can authorize mailing of replacement cards. The technology is provided by Nuance, which claims its latest systems can detect synthetic voices in 86% to 99% of cases. When contacted about the vulnerability, Services Australia declined to say if the technology would be changed or removed.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The `make:controller` CLI command calls `mkdir(..., recursive: true)` on a path built from the user-supplied controller name, **before** Nette's class-name validation runs. The class-file write is correctly rejected by Nette when the name contains `/`, but the recursive directory creation side effect is already committed \u2014 including directories located outside the project root through `../` traversal. ### Affected code `flight/commands/ControllerCommand.php` (\u2248 63-66): ```php if (is_dir(dirname($controllerPath)) === false) { $io->info('Creating directory ' . dirname($controllerPath), true); mkdir(dirname($controllerPath), 0755, true); // un-normalized, runs before validation } ``` ### Proof of concept ``` $ php vendor/flightphp/runway/runway make:controller '../../../../tmp/CONTROLLER_TRAVERSAL_TEST/pwn' Creating directory .../app/controllers/../../../../tmp/CONTROLLER_TRAVERSAL_TEST Nette\\InvalidArgumentException: Value '../../../../tmp/CONTROLLER_TRAVERSAL_TEST/pwnController' is not valid class name. $ ls /home/user/tmp/CONTROLLER_TRAVERSAL_TEST (directory exists \u2014 created before the exception was thrown) ``` ### Impact - **Arbitrary directory creation outside the project root**, executable by any local actor that can run the Flight CLI (developer machine, shared CI build agent, compromised dev container). - Primes log-file planting for chained LFI exploitation (e.g. creating a directory where an attacker can later drop a `.php` file to be include\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nUS Congresswoman Alexandria Ocasio-Cortez became the target of AI-generated sexual deepfakes circulated on social media, drawing direct legislative action including her co-sponsored DEFIANCE Act creating civil remedies for non-consensual sexual deepfakes.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary Astro server actions have no default request body size limit, which can lead to memory exhaustion DoS. A single large POST to a valid action endpoint can crash the server process on memory-constrained deployments. ## Details On-demand rendered sites built with Astro can define server actions, which automatically parse incoming request bodies (JSON or FormData). The body is buffered entirely into memory with no size limit \u2014 a single oversized request is sufficient to exhaust the process heap and crash the server. Astro's Node adapter (`mode: 'standalone'`) creates an HTTP server with no body size protection. In containerized environments, the crashed process is automatically restarted, and repeated requests cause a persistent crash-restart loop. Action names are discoverable from HTML form attributes on any public page, so no authentication is required. ## PoC <details> ### Setup Create a new Astro project with the following files: `package.json`: ```json { \"name\": \"poc-dos\", \"private\": true, \"scripts\": { \"build\": \"astro build\", \"start:128mb\": \"node --max-old-space-size=128 dist/server/entry.mjs\" }, \"dependencies\": { \"astro\": \"5.17.2\", \"@astrojs/node\": \"9.5.3\" } } ``` `astro.config.mjs`: ```javascript import { defineConfig } from 'astro/config'; import node from '@astrojs/node'; export default defineConfig({ output: 'server', adapter: node({ mode: 'standalone' }), }); ``` `src/actions/index.ts`: ```typescript import { def\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAn AI-powered call center platform in the Middle East reportedly experienced a significant data breach, allegedly exposing over 10 million conversations between consumers, operators, and AI agents. Attackers allegedly accessed the platform\u2019s management dashboard, stealing sensitive data, including national ID documents. The breach poses reported risks such as phishing, identity theft, and social engineering attacks. The stolen data was reportedly listed for sale on the dark web.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary An attacker can cause a denial of service in webtransport-go by preventing or indefinitely delaying WebTransport session closure. A malicious peer can withhold QUIC flow control credit on the CONNECT stream, blocking transmission of the WT_CLOSE_SESSION capsule and causing the close operation to hang. ## Details WebTransport over HTTP/3 signals session termination by sending a WT_CLOSE_SESSION capsule on the CONNECT stream. The capsule is only needed to transmit a reason phrase and an error code to the peer. After the capsule is sent, the CONNECT stream is closed. In affected versions, the closure procedure blocked indefinitely while waiting for sufficient QUIC flow control credit from the peer. A malicious peer can withhold this credit, preventing the capsule from being sent. ## The Fix webtransport-go now attempts to send the WT_CLOSE_SESSION capsule with a short deadline. If the capsule cannot be sent within this deadline, the CONNECT stream is reset instead. This closes the WebTransport session promptly without transmitting the optional error details. This prevents indefinite blocking on session closure.\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe fix for GHSA-p5g2-jm85-8g35 (ClickHouse SQL injection via aggregate query parameters) added column name validation to the `_aggregateBy` method but did not apply the same validation to three other query construction paths in `StatementGenerator`. The `toSortStatement`, `toSelectStatement`, and `toGroupByStatement` methods accept user-controlled object keys from API request bodies and interpolate them as ClickHouse `Identifier` parameters without verifying they correspond to actual model columns. ClickHouse Identifier parameters are substituted directly into queries without escaping, so an attacker who can reach any analytics list or aggregate endpoint can inject arbitrary SQL through crafted `sort`, `select`, or `groupBy` keys. ## Details ### Root cause `StatementGenerator.ts` has four methods that iterate over user-provided object keys to build SQL: | Method | Validates keys? | |--------|----------------| | `toWhereStatement` (line 292) | Yes - calls `this.model.getTableColumn(key)` | | `toSortStatement` (line 467) | **No** | | `toSelectStatement` (line 483) | **No** | | `toGroupByStatement` (line 451) | **No** | In `Statement.ts`, when a value passed to the `SQL` tagged template is a string, it receives the `Identifier` data type (line 40). Per [ClickHouse documentation](https://clickhouse.com/docs/en/sql-reference/syntax#defining-and-using-query-parameters), Identifier parameters are substituted directly into the query without quoting or escaping. This is correct\n</incident>"}], "ground_truth": "deserialization", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact picomatch is vulnerable to a **method injection vulnerability (CWE-1321)** affecting the `POSIX_REGEX_SOURCE` object. Because the object inherits from `Object.prototype`, specially crafted POSIX bracket expressions (e.g., `[[:constructor:]]`) can reference inherited method names. These methods are implicitly converted to strings and injected into the generated regular expression. This leads to **incorrect glob matching behavior (integrity impact)**, where patterns may match unintended filenames. The issue does **not enable remote code execution**, but it can cause security-relevant logic errors in applications that rely on glob matching for filtering, validation, or access control. All users of affected `picomatch` versions that process untrusted or user-controlled glob patterns are potentially impacted. ### Patches This issue is fixed in picomatch 4.0.4, 3.0.2 and 2.3.2. Users should upgrade to one of these versions or later, depending on their supported release line. ### Workarounds If upgrading is not immediately possible, avoid passing untrusted glob patterns to picomatch. Possible mitigations include: - Sanitizing or rejecting untrusted glob patterns, especially those containing POSIX character classes like `[[:...:]]`. - Avoiding the use of POSIX bracket expressions if user input is involved. - Manually patching the library by modifying `POSIX_REGEX_SOURCE` to use a null prototype: ```js const POSIX_REGEX_SOURCE = { __proto__: null, alnum\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary A denial-of-service vulnerability exists in the SM2 public-key encryption (PKE) implementation: the `decrypt()` path performs unchecked `slice::split_at` operations on input buffers derived from untrusted ciphertext. An attacker can submit short/undersized ciphertext or carefully-crafted DER-encoded structures to trigger bounds-check panics (Rust unwinding) which crash the calling thread or process. ### Affected Component / Versions - File: `src/pke/decrypting.rs` - Functions: `DecryptingKey::decrypt_digest/decrypt/decrypt_der`, internal `decrypt()` implementation - Affected releases: - sm2 0.14.0-rc.0 (https://crates.io/crates/sm2/0.14.0-rc.0) - sm2 0.14.0-pre.0 (https://crates.io/crates/sm2/0.14.0-pre.0) ### Details The vulnerability is located in the file `sm2/src/pke/decrypting.rs`. The **fundamental cause** of the vulnerability is that the decryption function **does not strictly check** the ciphertext's format and length information. Consequently, a maliciously crafted ciphertext can trigger Rust's **panic mechanism** instead of the expected error handling (`Error`) mechanism. The Rust function `C.split_at(L)` will trigger a Panic if the length is less than `L`, as shown in the code comment below: the `decrypting` function has **at least three locations** where a slice operation might trigger a Panic. ```rust fn decrypt( secret_scalar: &Scalar, mode: Mode, hasher: &mut dyn DynDigest, cipher: &[u8], ) -> Result<Vec<u8>>\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary A path traversal vulnerability in `IArchive.WriteToDirectory()` allows a malicious archive to create directories outside the intended extraction root. For TAR archives, this can be escalated to arbitrary file writes by chaining with a symlink entry, giving a full write primitive on the target filesystem subject to the permissions of the running process. ### Details The vulnerable code is in the directory-entry branch of `WriteToDirectoryInternal` (sync, `IArchiveExtensions.cs:48\u201361`) and `WriteToDirectoryAsyncInternal` (async, `IAsyncArchiveExtensions.cs:70\u201384`): ```csharp var dirPath = Path.Combine(destinationDirectory, entry.Key); Directory.CreateDirectory(Path.GetDirectoryName(dirPath + \"/\")); ``` No `Path.GetFullPath()` normalisation and no bounds check are applied before the `Directory.CreateDirectory` call. Two .NET `Path.Combine` behaviours make this exploitable: - **Relative traversal**: `Path.Combine(\"/safe/extract\", \"../../evil\")` \u2192 the OS resolves `..` segments on the raw path, placing the directory outside the extraction root. - **Absolute path override**: `Path.Combine(\"/safe/extract\", \"/tmp/evil\")` \u2192 returns `\"/tmp/evil\"` \u2014 the base is discarded entirely for rooted paths. File entries are **not** directly affected \u2014 they route through `ExtractionMethods.WriteEntryToDirectory` which applies the correct guard (`GetFullPath` + `StartsWith`, see `ExtractionMethods.cs:54\u201365`). The directory-entry branch is a separate fast-path that was added withou\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The SignalK Server exposes an unauthenticated HTTP endpoint that allows remote attackers to modify navigation data source priorities. This endpoint, accessible via `PUT /signalk/v1/api/sourcePriorities`, does not enforce authentication or authorization checks and directly assigns user-controlled input to the server configuration. As a result, attackers can influence which GPS, AIS, or other sensor data sources are trusted by the system. The changes are immediately applied and persisted to disk, allowing the manipulation to survive server restarts. ### Affected Component - **File**: `src/serverroutes.ts` - **Endpoint**: `PUT /signalk/v1/api/sourcePriorities` (also accessible at `/skServer/sourcePriorities`) - **Lines**: 1064-1076 - **Function**: Source priorities configuration handler ### Vulnerable Code ```typescript // src/serverroutes.ts - Lines 1064-1076 app.put( `${SERVERROUTESPREFIX}/sourcePriorities`, (req: Request, res: Response) => { app.config.settings.sourcePriorities = req.body app.activateSourcePriorities() writeSettingsFile(app, app.config.settings, (err: any) => { if (err) { res .status(500) .send('Unable to save to sourcePrefences in settings file') } else { res.json({ result: 'ok' }) } }) } ) ``` ## Vulnerability Characteristics **Missing Authentication**: The endpoint has zero authentication middleware, allowing unauthenticated access from any network-adjacent attacke\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nHugging Face Transformers Transformer-XL Model Deserialization of Untrusted Data Remote Code Execution Vulnerability. This vulnerability allows remote attackers to execute arbitrary code on affected installations of Hugging Face Transformers. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the parsing of model files. The issue results from the lack of proper validation of user-supplied data, which can result in deserialization of untrusted data. An attacker can leverage this vulnerability to execute code in the context of the current user. Was ZDI-CAN-25424.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nTikTok&#x27;s &quot;Suggested Accounts&quot; recommendations allegedly reinforced racial bias despite not basing recommendations on race or creators&#x27; profile photo.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nTorchServe gRPC service binds to all interfaces by default without authentication, enabling unauthorized model management requests.\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Met Police retrospective facial recognition system raises privacy concerns. System: NeoFace Watch. Technology: Facial recognition. Purpose: Identify criminal suspects. Ethical issues: Accountability; Fairness; Privacy/surveillance.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The regex-based SVG sanitizer in phpMyFAQ (`SvgSanitizer.php`) can be bypassed using HTML entity encoding in `javascript:` URLs within SVG `<a href>` attributes. Any user with `edit_faq` permission can upload a malicious SVG that executes arbitrary JavaScript when viewed, enabling privilege escalation from editor to full admin takeover. ### Details The file `phpmyfaq/src/phpMyFAQ/Helper/SvgSanitizer.php` (introduced 2026-01-15) uses regex patterns to detect dangerous content in uploaded SVG files. The regex for `javascript:` URL detection is: `/href\\s*=\\s*[\"\\']javascript:[^\"\\']*[\"\\']/i` This pattern matches the literal string `javascript:` but fails when the URL is HTML entity encoded. For example, `&#106;&#97;&#118;&#97;&#115;&#99;&#114;&#105;&#112;&#116;&#58;` decodes to `javascript:` in the browser, but does NOT match the regex. The `isSafe()` method returns `true`, so the SVG is accepted without sanitization. Additionally, the `DANGEROUS_ELEMENTS` blocklist misses `<animate>`, `<set>`, and `<use>` elements which can also be used to execute JavaScript in SVG context. Uploaded SVG files are served with `Content-Type: image/svg+xml` and no `Content-Disposition: attachment` header, so browsers render them inline and execute any JavaScript they contain. The image upload endpoint (`/admin/api/content/images`) only requires the `edit_faq` permission \u2014 not full admin \u2014 so any editor-level user can upload malicious SVGs. ### PoC ### Basic XSS (confirmed working\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nWeKnora is an LLM-powered framework designed for deep document understanding and semantic retrieval. Prior to version 0.3.0, a DNS rebinding vulnerability in the web_fetch tool allows an unauthenticated attacker to bypass URL validation and access internal resources on the server, including private IP addresses (e.g., 127.0.0.1, 192.168.x.x). By crafting a malicious domain that resolves to a public IP during validation and subsequently resolves to a private IP during execution, an attacker can access sensitive local services and potentially exfiltrate data. This issue has been patched in version 0.3.0.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary Commit 49d0bb7 introduced a regression in sanitize-html that bypasses `allowedTags` enforcement for text inside `nonTextTagsArray` elements (`textarea` and `option`). Entity-encoded HTML inside these elements passes through the sanitizer as decoded, unescaped HTML, allowing injection of arbitrary tags including XSS payloads. This affects any application using sanitize-html that includes `option` or `textarea` in its `allowedTags` configuration. ## Details The vulnerable code is at `packages/sanitize-html/index.js:569-573`: ```javascript } else if ((options.disallowedTagsMode === 'discard' || options.disallowedTagsMode === 'completelyDiscard') && (nonTextTagsArray.indexOf(tag) !== -1)) { // htmlparser2 does not decode entities inside raw text elements like // textarea and option. The text is already properly encoded, so pass // it through without additional escaping to avoid double-encoding. result += text; } ``` The comment is factually incorrect. htmlparser2 10.x **does** decode HTML entities inside both `<textarea>` and `<option>` elements before passing text to the `ontext` callback. This can be verified: ```javascript const htmlparser2 = require('htmlparser2'); const parser = new htmlparser2.Parser({ ontext(text) { console.log(JSON.stringify(text)); } }); parser.write('<option>&lt;script&gt;</option>'); // Outputs: \"<\", \"script\", \">\" \u2014 entities are decoded ``` Because the code assumes the text is \"already properly encoded\" and skips `escapeHtm\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nImproper Restriction of XML External Entity Reference vulnerability in Apache Syncope Console. An administrator with adequate entitlements to create or edit Keymaster parameters via Console can construct malicious XML text to launch an XXE attack, thereby causing sensitive data leakage occurs. This issue affects Apache Syncope: from 3.0 through 3.0.15, from 4.0 through 4.0.3. Users are recommended to upgrade to version 3.0.16 / 4.0.4, which fix this issue.\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nFraudsters used AI voice cloning to impersonate a company director and direct a UAE bank manager to authorize roughly USD 35 million in wire transfers, in one of the earliest publicly documented large-scale voice-clone BEC attacks. The recovered funds were minimal and the case became a benchmark for vishing-style AI fraud.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "vishing"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe CEO of the German subsidiary of a UK-based energy company was impersonated by attackers using deepfake audio of his voice on a phone call, with the resulting wire transfer of EUR 220,000 (USD 243,000) routed through Hungary and Mexico. This was widely cited as the first publicly reported AI-voice-clone BEC fraud.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nOpen-source model hub CivitAI was found to host bounty programs paying users for LoRAs and image bounties depicting identifiable real people, including in sexual contexts, providing infrastructure and incentives for non-consensual deepfakes.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMichigan\u2019s MiDAS system falsely accused over 34,000 people of unemployment fraud from 2013 to 2015, which reportedly caused financial ruin for many. The automated system was designed to cut costs, but it adjudicated fraud cases without human oversight. That led to an 85% error rate. Victims faced wage garnishments, some lost homes, and some faced bankruptcy. Despite early warnings, Michigan\u2019s UIA defended MiDAS until lawsuits and federal pressure forced reforms. Legislators have been seeking compensation for those wrongfully accused.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Delhi court orders takedown of AI-doctored content of Sadhguru. Technology: Deepfake, Generative AI; Text-to-speech. Purpose: Defraud. Ethical issues: Accountability; Autonomy/agency; Human rights/civil liberties. Reported consequences: Litigation.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The XML parser can be forced to do an unlimited amount of entity expansion. With a very small XML input, it\u2019s possible to make the parser spend seconds or even minutes processing a single request, effectively freezing the application. ### Details There is a check in `DocTypeReader.js` that tries to prevent entity expansion attacks by rejecting entities that reference other entities (it looks for & inside entity values). This does stop classic \u201cBillion Laughs\u201d payloads. However, it doesn\u2019t stop a much simpler variant. If you define one large entity that contains only raw text (no & characters) and then reference it many times, the parser will happily expand it every time. There is no limit on how large the expanded result can become, or how many replacements are allowed. The problem is in `replaceEntitiesValue()` inside `OrderedObjParser.js`. It repeatedly runs `val.replace()` in a loop, without any checks on total output size or execution cost. As the entity grows or the number of references increases, parsing time explodes. Relevant code: `DocTypeReader.js` (lines 28\u201333): entity registration only checks for & `OrderedObjParser.js` (lines 439\u2013458): entity replacement loop with no limits ### PoC ```js const { XMLParser } = require('fast-xml-parser'); const entity = 'A'.repeat(1000); const refs = '&big;'.repeat(100); const xml = `<!DOCTYPE foo [<!ENTITY big \"${entity}\">]><root>${refs}</root>`; console.time('parse'); new XMLParser().parse(xml); // ~4\u20138 sec\n</incident>"}], "ground_truth": "dos", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary Seven recursive traversals in `lib/dom.js` operate without a depth limit. A sufficiently deeply nested DOM tree causes a `RangeError: Maximum call stack size exceeded`, crashing the application. **Reported operations:** - `Node.prototype.normalize()` \u2014 reported by @praveen-kv (email 2026-04-05) and @KarimTantawey (GHSA-fwmp-8wwc-qhv6, via `DOMParser.parseFromString()`) - `XMLSerializer.serializeToString()` \u2014 reported by @Jvr2022 (GHSA-2v35-w6hq-6mfw) and @KarimTantawey (GHSA-j2hf-fqwf-rrjf) **Additionally, discovered in research:** - `Element.getElementsByTagName()` / `getElementsByTagNameNS()` / `getElementsByClassName()` / `getElementById()` - `Node.cloneNode(true)` - `Document.importNode(node, true)` - `node.textContent` (getter) - `Node.isEqualNode(other)` All seven share the same root cause: pure-JavaScript recursive tree traversal with no depth guard. A single deeply nested document (parsed successfully) triggers any or all of these operations. --- ## Details ### Root cause `lib/dom.js` implements DOM tree traversals as depth-first recursive functions. Each level of element nesting adds one JavaScript call frame. The JS engine's call stack is finite; once exhausted, a `RangeError: Maximum call stack size exceeded` is thrown. This error may not be caught reliably at stack-exhaustion depths because the catch handler itself requires stack frames to execute \u2014 especially in async scenarios, where an uncaught `RangeError` inside a callback or promise chain\n</incident>"}], "ground_truth": "dos", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Buenos Aires government uses live facial recognition to identify child criminals. System: Sistema de Reconocimiento Facial de Pr\u00f3fugos (SNRP). Technology: Facial recognition. Purpose: Identify/track criminals. Ethical issues: Accuracy/reliability, Fairness; Human rights/civil liberties. Reported consequences: Legislative questions/complaints.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability was identified in BigSweetPotatoStudio HyperChat up to 2.0.0-alpha.63. Affected by this issue is the function fetch of the file packages/core/src/http/aiProxyMiddleware.mts of the component AI Proxy Middleware. Such manipulation of the argument baseurl leads to server-side request forgery. The attack can be launched remotely. The exploit is publicly available and might be used. The project was informed of the problem early through an issue report but has not responded yet.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Alonzo Sawyer facial recognition wrongful arrest, jailing. Technology: Facial recognition. Purpose: Identify criminal suspect. Ethical issues: Accountability; Accuracy/reliability; Fairness.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nTrivy Vulnerability Scanner is a VS Code extension that helps find vulnerabilities. In Trivy VSCode Extension version 1.8.12, which was distributed via OpenVSX marketplace was compromised and contained malicious code designed to leverage local AI coding agent to collect and exfiltrate sensitive information. Users using the affected artifact are advised to immediately remove it and rotate environment secrets. The malicious artifact has been removed from the marketplace. No other affected artifacts have been identified.\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary `Executrix.getCommand()` constructs shell commands by substituting temporary file paths directly into a `/bin/sh -c` string with no escaping. The `IN_FILE_ENDING` and `OUT_FILE_ENDING` configuration keys flow into those paths unmodified. A place author who sets either key to a shell metacharacter sequence achieves arbitrary OS command execution in the JVM's security context when the place processes any payload. No runtime privileges beyond place configuration authorship are required, and no API or network access is needed. This is a **framework-level defect** \u2014 `Executrix` provides no escaping mechanism and no validation on file ending values. Downstream implementors have no safe way to use the API as designed. --- ### Root Cause #### Step 1 \u2014 `IN_FILE_ENDING` flows into temp path construction without validation **[`TempFileNames.java:32-36`](src/main/java/emissary/util/shell/TempFileNames.java#L32-L36)** ```java public TempFileNames(String tmpDir, String placeName, String inFileEnding, String outFileEnding) { base = Long.toString(System.nanoTime()); tempDir = FileManipulator.mkTempFile(tmpDir, placeName); in = base + inFileEnding; // no sanitization out = base + outFileEnding; // no sanitization basePath = tempDir + File.separator + base; inputFilename = basePath + inFileEnding; // injected value lands here outputFilename = basePath + outFileEnding; // and here } ``` `inFileEnding` is concatenated direct\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary Vue 3's v-html directive is the framework-documented mechanism for injecting raw HTML, and it intentionally disables the auto-escaping that {{ }} interpolation provides. The PropertyCard.vue component uses v-html for the else branch of the URL check, meaning any non-URL string value flows directly into the DOM as HTML. The isURL() guard only filters values that parse as http: or https: URLs, so any HTML payload not starting with those schemes (e.g., `<img src=x onerror=alert(1)>` padded to exceed 75 chars) bypasses it entirely. The data originates from Kubernetes PolicyReport .results[].properties fields, which are arbitrary string maps populated by policy engines and potentially by any principal with write access to PolicyReport objects in the cluster. No DOMPurify or equivalent HTML sanitization library is present anywhere in the frontend codebase, confirming there is no compensating control between the API response and the sink. This vulnerability was reproduced on the latest policy reporter UI version \u2013 2.5.1. ### PoC Prerequisites: Kubernetes write access to PolicyReport resources in the target cluster (e.g., via a policy engine service account or direct kubectl access) Create a Kubernetes PolicyReport resource with a crafted property value longer than 75 characters. When an authenticated Policy Reporter UI user browses to the affected namespace and expands the result row containing this property, the injected script executes in their browser. ```bash ku\n</incident>"}], "ground_truth": "xss", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nResearch and journalism showed that ChatGPT could be manipulated, both through direct prompts and indirect web content, into harvesting personal identifying information from users and from sources it had memorized in training data.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nPrompt injection vulnerability in 1millionbot Millie chatbot that occurs when a user manages to evade chat restrictions using Boolean prompt injection techniques (formulating a question in such a way that, upon receiving an affirmative response ('true'), the model executes the injected instruction), causing it to return prohibited information and information outside its intended context. Successful exploitation of this vulnerability could allow a malicious remote attacker to abuse the service for purposes other than those originally intended, or even execute out-of-context tasks using 1millionbot's resources and/or OpenAI's API key. This allows the attacker to evade the containment mechanisms implemented during LLM model training and obtain responses or chat behaviors that were originally restricted.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA weakness has been identified in atototo api-lab-mcp up to 0.2.1. This affects the function analyze_api_spec/generate_test_scenarios/test_http_endpoint of the file src/mcp/http-server.ts of the component HTTP Interface. This manipulation of the argument source/url causes server-side request forgery. The attack is possible to be carried out remotely. The exploit has been made available to the public and could be used for attacks. The project was informed of the problem early through an issue report but has not responded yet.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: North Korea uses AI to covertly place IT workers in western tech companies. Technology: Deepfake. Purpose: Impersonate workers. Ethical issues: Authenticity/integrity; Employment/labour; Security; Transparency.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n# Assessment It is believed that the analysis pass works as intended, `REDUCE` and `BUILD` are not at fault here. The few potentially unsafe modules have been added to the blocklist (https://github.com/trailofbits/fickling/commit/0c4558d950daf70e134090573450ddcedaf10400). # Original report ### Summary All 5 of fickling's safety interfaces \u2014 `is_likely_safe()`, `check_safety()`, CLI `--check-safety`, `always_check_safety()`, and the `check_safety()` context manager \u2014 report `LIKELY_SAFE` / raise no exceptions for pickle files that call dangerous top-level stdlib functions (signal handlers, network servers, network connections, file operations) when the REDUCE opcode is followed by a BUILD opcode. Demonstrated impacts include backdoor network listeners (`socketserver.TCPServer`), process persistence (`signal.signal`), outbound data exfiltration (`smtplib.SMTP`), and file creation on disk (`sqlite3.connect`). An attacker can append a trivial BUILD opcode to any payload to eliminate all detection. ## Details The bypass exploits three weaknesses in fickling's static analysis pipeline: 1. **`likely_safe_imports` over-inclusion** (`fickle.py:432-435`): When fickling decompiles a pickle and encounters `from smtplib import SMTP`, it adds `\"SMTP\"` to the `likely_safe_imports` set because `smtplib` is a Python stdlib module. This happens for ALL stdlib modules, including dangerous ones like smtplib, ftplib, sqlite3, etc. 2. **`OvertlyBadEvals` exemption** (`analysis.py:301-310`):\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n# Summary The `webapi` authentication layer trusts a client-controlled `X-lobe-chat-auth` header that is only XOR-obfuscated, not signed or otherwise authenticated. Because the XOR key is hardcoded in the repository, an attacker can forge arbitrary auth payloads and bypass authentication on protected `webapi` routes. Affected routes include: - `POST /webapi/chat/[provider]` - `GET /webapi/models/[provider]` - `POST /webapi/models/[provider]/pull` - `POST /webapi/create-image/comfyui` ## Details The frontend creates `X-lobe-chat-auth` by XOR-obfuscating JSON with the static key `LobeHub \u00b7 LobeHub`, and the backend reverses that operation and treats the decoded JSON as trusted authentication data. The backend then accepts any truthy `apiKey` field in that decoded payload as sufficient authentication. No real API key validation is performed in this path. As a result, an unauthenticated attacker can forge payloads such as: ```json {\"apiKey\":\"x\"} ``` or ``` {\"userId\":\"victim-user-123\",\"apiKey\":\"x\"} ``` and access webapi routes as an authenticated user. Confirmed PoC The following forged header was generated directly from the published XOR key using payload {\"apiKey\":\"x\"}: ``` X-lobe-chat-auth: N00DFSE+B1ngjQI0TR8= ``` That header decodes server-side to: ``` {\"apiKey\":\"x\"}``` A simple request is: ``` curl 'https://TARGET/webapi/models/openai' \\ -H 'X-lobe-chat-auth: N00DFSE+B1ngjQI0TR8=' ``` If the deployment has OPENAI_API_KEY configured, the request should\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: German TV star accuses husband of spreading deepfake porn images of her. Technology: Deepfake. Purpose: Harass ex-wife. Ethical issues: Accountability; Authenticity/integrity; Privacy/surveillance; Safety; Transparency. Reported consequences: Legal complaint/investigation; Legislation introduction/update.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA path traversal vulnerability exists in the awesome-llm-apps project in commit e46690f99c3f08be80a9877fab52acacf7ab8251 (2026-01-19) in the Beifong AI News and Podcast Agent backend in FastAPI backend, stream-audio endpoint, in file routers/podcast_router.py, in function stream_audio. The stream-audio endpoint accepts a user-controlled path parameter that is concatenated into a filesystem path without proper validation or restriction. An unauthenticated remote attacker can exploit this vulnerability to read arbitrary files from the server filesystem, potentially disclosing sensitive information such as configuration files and credentials.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The webhook middleware spawns a goroutine that holds a reference to the request's `echo.Context` after the synchronous handler returns `ErrAsyncProcess` and Echo recycles the context back to its `sync.Pool`. When a concurrent request claims the recycled context, `c.Reset()` clears the store. If the webhook goroutine reaches `hardTimeoutMiddleware` at that moment, an unchecked type assertion on a nil store entry panics outside any `recover()` scope, crashing the Gotenberg process. Any anonymous caller reaches the webhook path (default `webhook-deny-list` filters only the webhook destination, not the submitter). A single-source stress of ~24 webhook requests plus ~60 `GET /version` requests crashes the process in about two seconds. ## Details `pkg/modules/webhook/middleware.go:338-382` starts the async goroutine and immediately returns `api.ErrAsyncProcess` to the caller: ```go w.asyncCount.Add(1) go func() { defer cancel() defer w.asyncCount.Add(-1) err := next(c) // line 343 ... sendOutputFile(sendOutputFileParams{ ctx: ctx, ... }) }() return api.ErrAsyncProcess // line 382 ``` `pkg/modules/api/middlewares.go:356-361` sees the sentinel, responds with `204 No Content`, and lets Echo return `c` to the pool: ```go if errors.Is(err, ErrAsyncProcess) { return c.NoContent(http.StatusNoContent) } ``` Echo's router calls `c.Reset()` before serving the next request from the same goroutine pool slot, wiping `c.s\n</incident>"}], "ground_truth": "dos", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n**Summary** The proxyUi template engine uses Go's text/template (which performs no HTML escaping) instead of html/template. The GitHub OAuth callback handlers in both publicProxy and dynamicProxy embed the attacker-controlled refreshInterval query parameter verbatim into an error message when time.ParseDuration fails, and render that error unescaped into HTML. An attacker can deliver a crafted login URL to a victim; after the victim completes the GitHub OAuth flow, the callback page executes arbitrary JavaScript in the OAuth server's origin. - Attack Vector: Network \u2014 the attack is delivered as a crafted URL over the internet. - Attack Complexity: Low \u2014 no race conditions or special environment prerequisites. - Privileges Required: None \u2014 the attacker needs no account on the zrok instance. - User Interaction: Required \u2014 the victim must click the crafted link and complete the GitHub OAuth flow. - Scope: Changed \u2014 the injected script executes in the OAuth server's origin, not the victim's share origin. - Confidentiality Impact: Low \u2014 the script runs in the OAuth server origin after a failed flow; no session cookie is set at this point, limiting what can be exfiltrated to what is visible in the DOM and what can be requested from the OAuth server. - Integrity Impact: Low \u2014 the script can initiate new OAuth flows or submit forms on behalf of the victim in the OAuth server origin. - Availability Impact: None. **Affected Components** - endpoints/proxyUi/template.go \u2014 init() / Wri\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nOpen WebUI PIP install_frontmatter_requirements Command Injection Remote Code Execution Vulnerability. This vulnerability allows remote attackers to execute arbitrary code on affected installations of Open WebUI. Authentication is required to exploit this vulnerability. The specific flaw exists within the install_frontmatter_requirements function.The issue results from the lack of proper validation of a user-supplied string before using it to execute a system call. An attacker can leverage this vulnerability to execute code in the context of the service account. Was ZDI-CAN-28258.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nIn early March 2024, a network named CopyCop began publishing modified news stories using AI, altering content to spread partisan biases and disinformation. These articles, initially from legitimate sources, were manipulated by AI models, possibly developed by OpenAI, to disseminate Russian propaganda. Over 19,000 articles were published, targeting divisive political issues and creating false narratives.\n</incident>"}], "ground_truth": "rce", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary PraisonAI's call server exposes a network-facing agent control API without authentication when `CALL_SERVER_TOKEN` is not configured. The affected component is the `praisonai.api.agent_invoke` router as mounted by `praisonai.api.call`. The authentication helper `verify_token()` fails open when `CALL_SERVER_TOKEN` is unset. Since every sensitive agent-control endpoint depends on this helper, starting the call server without a token allows any reachable client to list agents, inspect agent metadata and instructions, invoke agents, and unregister agents. This is security-relevant because the bundled call server includes the vulnerable router and binds to `0.0.0.0`. As a result, operators who launch the call server without explicitly setting `CALL_SERVER_TOKEN` may unintentionally expose an unauthenticated remote agent control plane. ### Details The vulnerable behavior is caused by a fail-open authentication default. In `praisonai/api/agent_invoke.py`, `CALL_SERVER_TOKEN` is read from the environment: ```python CALL_SERVER_TOKEN = os.getenv('CALL_SERVER_TOKEN') ``` The authentication dependency then returns successfully when the token is not configured: ```python async def verify_token(request: Request, authorization: Optional[str] = Header(None)) -> None: if not FASTAPI_AVAILABLE or not CALL_SERVER_TOKEN: return # No authentication if FastAPI unavailable or no token set ``` This means that the absence of `CALL_SERVER_TOKEN` disables authenticat\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: UK parliamentarian calls for Deepsukebe nudifier ban. System: Deepsukebe. Technology: Deepfake. Purpose: Undress individuals. Ethical issues: Accountability; Authenticity/integrity; Safety; Privacy/surveillance; Transparency.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The server-side mitigation for the YPTSocket `autoEvalCodeOnHTML` eval sink (prior advisory GHSA-gph2-j4c9-vhhr, commit `c08694bf6`) only strips the payload when it sits under `$json['msg']`, but the relay function `msgToResourceId()` selects the outbound message from `$msg['json']` *before* `$msg['msg']`. An unauthenticated attacker can obtain a WebSocket token from `plugin/YPTSocket/getWebSocket.json.php`, connect to the WebSocket server, and send a message with `autoEvalCodeOnHTML` nested under a top-level `json` field \u2014 the strip branch is skipped, the relay delivers the payload verbatim to any logged-in user identified by `to_users_id`, and the client script runs it through `eval()`. ## Details ### Entry point (unauthenticated) `plugin/YPTSocket/getWebSocket.json.php` (lines 1\u201321) issues a valid WebSocket token to any caller, with no authentication or CSRF check: ```php $obj->webSocketToken = getEncryptedInfo(0); $obj->webSocketURL = YPTSocket::getWebSocketURL(); die(json_encode($obj)); ``` `getEncryptedInfo()` defaults to `sentFrom = 'browser'` and a non-CLI flag (`plugin/YPTSocket/functions.php:3-47`), so a token minted for an anonymous browser client will cause the strip branch below to run \u2014 which is exactly what we want to audit. ### Incomplete strip (the fix from commit c08694bf6) `plugin/YPTSocket/Message.php:236-247`: ```php // Strip eval-able fields from browser/guest messages. if (empty($msgObj->isCommandLineInterface) && ($msgObj->sentFrom\n</incident>"}], "ground_truth": "xss", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nHi Fedify team! \ud83d\udc4b Thank you for your work on Fedify\u2014it's a fantastic library for building federated applications. While reviewing the codebase, I discovered a Regular Expression Denial of Service (ReDoS) vulnerability that I'd like to report. I hope this helps improve the project's security. --- ## Summary A Regular Expression Denial of Service (ReDoS) vulnerability exists in Fedify's document loader. The HTML parsing regex at `packages/fedify/src/runtime/docloader.ts:259` contains nested quantifiers that cause catastrophic backtracking when processing maliciously crafted HTML responses. **An attacker-controlled federated server can respond with a small (~170 bytes) malicious HTML payload that blocks the victim's Node.js event loop for 14+ seconds, causing a Denial of Service.** | Field | Value | |-------|-------| | **CWE** | CWE-1333 (Inefficient Regular Expression Complexity) | --- ## Details ### Vulnerable Code The vulnerability is located in `packages/fedify/src/runtime/docloader.ts`, lines 258-264: ```typescript // Line 258-259: Vulnerable regex with nested quantifiers const p = /<(a|link)((\\s+[a-z][a-z:_-]*=(\"[^\"]*\"|'[^']*'|[^\\s>]+))+)\\s*\\/?>/ig; // Line 261: No size limit on response body const html = await response.text(); // Line 264: Regex execution loop while ((m = p.exec(html)) !== null) rawAttribs.push(m[2]); ``` ### Root Cause Analysis The regex has **nested quantifiers with alternation**, which is a classic ReDoS pattern: ``` /<(a|link)((\\s+[\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Affected openclaw-claude-bridge v1.1.0 ## Issue v1.1.0 spawns the Claude Code CLI subprocess with `--allowed-tools \"\"` and the release notes + README claim this **\"disables all CLI tools\"** for sandboxing. This claim is incorrect. Per the Claude Code CLI documentation, `--allowed-tools` (alias `--allowedTools`) is an **auto-approve allowlist** of tools that execute without permission prompts \u2014 NOT a restriction on which tools are available. The correct flag to restrict the available tool set is `--tools`: > `--tools <tools...>` Specify the list of available tools from the built-in set. **Use `\"\"` to disable all tools**, `\"default\"` to use all tools, or specify tool names (e.g. `\"Bash,Edit,Read\"`). ## Impact - All CLI tools (Read/Write/Bash/WebFetch/...) remain nominally available to the spawned subprocess. - Actual execution behavior in `--print` non-interactive mode depends on undocumented CLI defaults (may auto-deny, may error out, may hang). - Users who deploy the bridge behind any interface that forwards untrusted prompts (e.g., publicly exposed OpenClaw gateway, automated pipelines with web-fetched context, agents that consume tool results from other systems) may be relying on a sandbox that does not exist. The README explicitly makes a security claim the code does not uphold, creating a false sense of safety for downstream operators. If the underlying CLI behavior changes in a future version to auto-allow tools in `--print` mode, prompt-injection attacks cou\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability was found in 1Panel-dev MaxKB up to 2.6.1. The affected element is the function execute of the file apps/application/flow/step_node/mcp_node/impl/base_mcp_node.py of the component Model Context Protocol Node. Performing a manipulation results in os command injection. The attack is possible to be carried out remotely. The exploit has been made public and could be used. You should upgrade the affected component. The vendor was contacted early, responded in a very professional manner and quickly released a fixed version of the affected product.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary A Manager account (`access_all=false`) was able to escalate privileges by directly invoking the **bulk-access API** against collections that were not originally assigned to them. The API allowed changing `assigned=false` to `assigned=true`, resulting in unauthorized access. Additionally, prior to the bulk-access call, the regular single-update API correctly returned **401 Unauthorized** for the same collection. After executing the bulk-access API, the same update API returned **200 OK**, confirming an authorization gap at the HTTP level. --- ## Description * The endpoint accepts `ManagerHeadersLoose` and does not validate access rights for the specified `collectionIds`. src/api/core/organizations.rs:551 ```rust headers: ManagerHeadersLoose, ``` * The received `collection_ids` are processed directly without per-collection authorization checks. src/api/core/organizations.rs:564 ```rust for col_id in data.collection_ids { ``` * Existing group assignments for the collection are deleted. src/api/core/organizations.rs:583 ```rust CollectionGroup::delete_all_by_collection(&col_id, &conn).await?; ``` * Existing user assignments for the collection are deleted. src/api/core/organizations.rs:590 ```rust CollectionUser::delete_all_by_collection(&col_id, &conn).await?; ``` * By comparison, another bulk-processing endpoint performs per-collection validation using `from_loose`. src/api/core/organizations.rs:787 ```rust let heade\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nStarship is a cross-shell prompt. Starting in version 1.0.0 and prior to version 1.20.0, undocumented and unpredictable shell expansion and/or quoting rules make it easily to accidentally cause shell injection when using custom commands with starship in bash. This issue only affects users with custom commands, so the scope is limited, and without knowledge of others' commands, it could be hard to successfully target someone. Version 1.20.0 fixes the vulnerability.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n# Stored XSS via Hostile YouTube Video Title in AVideo YouTubeAPI Gallery Section ## Summary A stored Cross-Site Scripting vulnerability (CWE-79; chained CWE-829, Inclusion of Functionality from Untrusted Control Sphere) in the AVideo YouTubeAPI plugin renders the `snippet.title` field returned by the YouTube Data API into the homepage gallery markup with no HTML encoding. The title is set by the YouTube video uploader (anyone in the world) and is treated by AVideo as trusted content. A YouTube uploader who controls a video matching the operator's configured query injects HTML into the AVideo homepage by setting their video's title to a JavaScript-bearing string; the payload then executes in the browser of every visitor who loads any page that renders the gallery. ## Details `plugin/YouTubeAPI/YouTubeAPI.php::listVideos()` fetches search results from the YouTube Data API and stores the `snippet.title` field unchanged inside a `YPTvideoObject`: ```php // plugin/YouTubeAPI/YouTubeAPI.php (listVideos excerpt) $searchResponse = $youtube->search->listSearch( 'snippet,contentDetails,statistics', $options ); foreach ($searchResponse['items'] as $searchResult) { $vid = new YPTvideoObject( $searchResult[\"id\"][\"videoId\"], $searchResult['snippet'][\"title\"], // uploader-controlled $searchResult['snippet'][\"description\"], $searchResult['snippet'][\"thumbnails\"][\"high\"][\"url\"], $searchResult['snippet'][\"channelTitle\"], \"h\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact The HTML output format in wlc embeds API response data into HTML without escaping, allowing cross-site scripting when the output is rendered in a browser. ### Patches * https://github.com/WeblateOrg/wlc/pull/1327 ### Workarounds The only vulnerable code path is HTML output which is opt-in. ### References Weblate thanks @fg0x0 for reporting this on GitHub.\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nIn versions prior to 2.26.1, the RMI instrumentation registered a custom endpoint that deserialized incoming data without applying serialization filters. An attacker with network access to a JMX or RMI port on an instrumented JVM could exploit this to potentially achieve remote code execution. All three of the following conditions must be true to exploit this vulnerability: 1. Splunk Distribution of OpenTelemetry Java is attached as a Java agent (`-javaagent`) 2. An RMI endpoint is network-reachable (e.g. JMX remote port, an RMI registry, or any application-exported RMI service) 3. A gadget-chain-compatible library is present on the classpath ### Impact Arbitrary remote code execution with the privileges of the user running the instrumented JVM. ### Recommendation Upgrade to version 2.26.1 or later. ### Workarounds Set the following system property to disable the RMI integration: ``` -Dotel.instrumentation.rmi.enabled=false ``` ### References [Advisory in OpenTelemetry Instrumentation for Java](https://github.com/open-telemetry/opentelemetry-java-instrumentation/security/advisories/GHSA-xw7x-h9fj-p2c7)\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Executive Summary: Microsoft is releasing this security advisory to provide information about a vulnerability in System.Security.Cryptography.Xml. This advisory also provides guidance on what developers can do to update their applications to remove this vulnerability. A vulnerability exists in EncryptedXml class where uncontrolled resource consumption can give an attacker to the ability to perform a Denial of Service attack. ## Announcement Announcement for this issue can be found at https://github.com/dotnet/announcements/issues/389 ## CVSS Details - **Version:** 3.1 - **Severity:** High - **Score:** 7.5 - **Vector:** 7.5: AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H E:U/RL:O/RC:C - **Weakness:** CWE-400 CWE-611: Uncontrolled Resource Consumption Improper Restriction of XML External Entity Reference ## Affected Platforms - **Platforms:** All - **Architectures:** All ## <a name=\"affected-packages\"></a>Affected Packages The vulnerability affects any Microsoft .NET project if it uses any of affected packages versions listed below ### <a name=\".NET 10\"></a>.NET 10 Package name | Affected version | Patched version ------------ | ---------------- | ------------------------- [System.Security.Cryptography.xml](https://www.nuget.org/packages/System.Security.Cryptography.Xml) | >=10.0.0, <=10.0.5; | 10.0.6 ### <a name=\".NET 9\"></a>.NET 9 Package name | Affected version | Patched version ------------ | ---------------- | ------------------------- [System.Security.Cr\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Amazon AI publishes false suicide helpline number. System: Rufus. Technology: Generative AI. Purpose: Provide shopping support. Ethical issues: Accuracy/reliability; Mis/disinformation; Safety.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary When an application using Pydantic AI opts a URL into `force_download='allow-local'` (which disables the default block on private/internal IPs), the cloud-metadata blocklist could be bypassed by encoding the metadata IP in an IPv6 transition form (IPv4-mapped IPv6, 6to4, or NAT64). Dual-stack and translated networks route the IPv6 wrapper to the underlying IPv4 endpoint, exposing cloud IAM short-term credentials. This is an incomplete fix of [GHSA-2jrp-274c-jhv3](https://github.com/pydantic/pydantic-ai/security/advisories/GHSA-2jrp-274c-jhv3) / [CVE-2026-25580](https://nvd.nist.gov/vuln/detail/CVE-2026-25580). The parent advisory's remediation guaranteed that \"cloud metadata endpoints are always blocked, even with `allow-local`.\" That guarantee did not hold for IPv6-encoded forms of the metadata IPs. ## Severity Same impact metrics as the parent CVE, but materially narrower attack surface (AC:H instead of AC:L), because exploitation requires the application to have opted into `allow-local` on a URL influenced by untrusted input. ## Who Is Affected Applications are affected **only if** they explicitly opt for `FileUrl` (`ImageUrl`, `AudioUrl`, `VideoUrl`, `DocumentUrl`) into `force_download='allow-local'` on a URL that is, or could be, influenced by untrusted input. Applications are **not** affected if they use any of the bundled integrations to ingest user input, because they do not propagate `force_download` from external data: - `Agent.to_web` / `clai web\n</incident>"}], "ground_truth": "ssrf", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAn issue in ollama v.0.12.10 allows a remote attacker to cause a denial of service via the GGUF decoder\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary `execute_code()` in `praisonai-agents` runs attacker-controlled Python inside a three-layer sandbox that can be fully bypassed by passing a `str` subclass with an overridden `startswith()` method to the `_safe_getattr` wrapper, achieving arbitrary OS command execution on the host. ### Details `python_tools.py:20` (source) -> `python_tools.py:22` (guard bypass) -> `python_tools.py:161` (sink) ```python # source -- _safe_getattr accepts any str subclass def _safe_getattr(obj, name, *default): if isinstance(name, str) and name.startswith('_'): # isinstance passes for subclasses raise AttributeError(...) # hop -- type() is whitelisted in safe_builtins, creates str subclass without class keyword FakeStr = type('FakeStr', (str,), {'startswith': lambda self, *a: False}) # sink -- Popen reached via __subclasses__ walk r = Popen(['id'], stdout=PIPE, stderr=PIPE) ``` ### PoC ```python from praisonaiagents.tools.python_tools import execute_code payload = \"\"\" t = type FakeStr = t('FakeStr', (str,), {'startswith': lambda self, *a: False}) mro_attr = FakeStr(''.join(['_','_','m','r','o','_','_'])) subs_attr = FakeStr(''.join(['_','_','s','u','b','c','l','a','s','s','e','s','_','_'])) mod_attr = FakeStr(''.join(['_','_','m','o','d','u','l','e','_','_'])) name_attr = FakeStr(''.join(['_','_','n','a','m','e','_','_'])) PIPE = -1 obj_class = getattr(type(()), mro_attr)[1] for cls in getattr(obj_class, subs_attr)(): try: m = getattr(cls, mod_attr\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA Python command injection vulnerability exists in the `SagemakerLLM` class's `complete()` method within `./private_gpt/components/llm/custom/sagemaker.py` of the imartinez/privategpt application, versions up to and including 0.3.0. The vulnerability arises due to the use of the `eval()` function to parse a string received from a remote AWS SageMaker LLM endpoint into a dictionary. This method of parsing is unsafe as it can execute arbitrary Python code contained within the response. An attacker can exploit this vulnerability by manipulating the response from the AWS SageMaker LLM endpoint to include malicious Python code, leading to potential execution of arbitrary commands on the system hosting the application. The issue is fixed in version 0.6.0.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Apple AI alert falsely claims Luke Littler has won darts championship. System: Apple Intelligence. Technology: Generative AI. Purpose: Summarise news. Ethical issues: Accuracy/reliability; Mis/disinformation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA Canadian couple were defrauded of approximately USD 21,000 after scammers used AI voice cloning to impersonate their adult son in a distress call demanding bail money.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Faception claim to identify paedophiles from their faces draws controversy. System: Faception. Technology: Computer vision; Behavioural analysis; Emotion recognition; Facial recognition; Personality analysis; Machine learning. Purpose: Identify personality type; Predict behaviour. Ethical issues: Accountability; Accuracy/reliability; Fairness; Transparency.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nUnsafe Deserialization vulnerability in Modular Max Serve before 25.6, specifically when the \"--experimental-enable-kvcache-agent\" feature is used allowing attackers to execute arbitrary code.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nScammers tricked a Hong Kong-based employee of British engineering firm Arup into paying HKD 200 million (USD ~25-26M) via a fake group video call in which every other attendee, including a UK-based 'CFO', was a deepfake recreation of real colleagues built from public video footage.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability was detected in ComfyUI up to 0.13.0. This impacts the function get_model_preview of the file app/model_manager.py of the component Model Preview Endpoint. The manipulation results in path traversal. The attack may be launched remotely. The exploit is now public and may be used. The vendor was contacted early about this disclosure but did not respond in any way.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Executive Summary: The path validation has a critical logic bug: it checks for `..` AFTER `normpath()` has already collapsed all `..` sequences. This makes the check completely useless and allows trivial path traversal to any file on the system. The path validation function also does not resolve the symlink wich could potentially cause path traversal. ### Details: `_validate_path()` calls `os.path.normpath()` first, which collapses `..` sequences, then checks for `'..'` in normalized. Since `..` is already collapsed, the check always passes. **Vulnerable File:** `src/praisonai-agents/praisonaiagents/tools/file_tools.py` **Lines:** 42-49 ```python class FileTools: \"\"\"Tools for file operations including read, write, list, and information.\"\"\" @staticmethod def _validate_path(filepath: str) -> str: # Normalize the path normalized = os.path.normpath(filepath) absolute = os.path.abspath(normalized) # Check for path traversal attempts (.. after normalization) # We check the original input for '..' to catch traversal attempts if '..' in normalized: raise ValueError(f\"Path traversal detected: {filepath}\") return absolute ``` **Severity:** CRITICAL **CVSS v3.1:** 9.2 (CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:H/VI:N/VA:N/SC:H/SI:N/SA:N **CWE:** CWE-22: Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal') ### Proof of concept (PoC) **Prerequisites:** -\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nClaude Code is an agentic coding tool. Prior to version 1.0.111, Claude Code contained insufficient URL validation in its trusted domain verification mechanism for WebFetch requests. The application used a startsWith() function to validate trusted domains (e.g., docs.python.org, modelcontextprotocol.io), this could have enabled attackers to register domains like modelcontextprotocol.io.example.com that would pass validation. This could enable automatic requests to attacker-controlled domains without user consent, potentially leading to data exfiltration. This issue has been patched in version 1.0.111.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nIn LangChain through 0.0.131, the LLMMathChain chain allows prompt injection attacks that can execute arbitrary code via the Python exec method. The chain uses insecure exec/eval on LLM-generated math expressions. Disclosed by the NVIDIA AI Red Team; fixed in 0.0.142.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nSummary The @apostrophecms/cli package contains a command injection vulnerability in the apos create command. User-supplied input from the password prompt is embedded directly into a shell command without proper sanitization or escaping. This allows execution of arbitrary commands on the host system. \u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501 Details Vulnerable file: lib/commands/create.js Location: Line 186 The CLI collects a password using an interactive prompt and passes it directly into a shell command. Vulnerable code: const response = await prompts({ type: 'password', name: 'pw', message: '\ud83d\udd0f Please enter a password:' }); exec(echo \"${response.pw}\" | ${createUserCommand}); The value of response.pw is not validated, sanitized, or escaped before being used in exec(). This allows shell metacharacters such as ;, &&, and $() to break out of the intended command and execute arbitrary commands. \u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501 Steps to Reproduce 1) Install the CLI npm install -g @apostrophecms/cli 2) Create a new project mkdir testproject && cd testproject apos create mysite 3)When prompted for the admin password, enter \"; id > /tmp/apos_rce_proof.txt; echo \" 4)Verify command execution cat /tmp/apos_rce_proof.txt \u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501 Proof of Concept Output uid=1000(vboxuser) gid=1000(vboxuser) groups=1000(vboxuser),27(sudo),984(docker) This confirms arbitrary command execution with the privileges of the user running the CLI. \u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501\u2501 Impact\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Duplicate Advisory This advisory has been withdrawn because it is a duplicate of GHSA-289f-fq7w-6q2w. This link is maintained to preserve external references. ### Original Description phpMyFAQ before 4.1.2 contains an information disclosure vulnerability in the getIdFromSolutionId() method that lacks permission filtering, allowing unauthenticated attackers to enumerate restricted FAQ entries and read their titles via the /solution_id_{id}.html endpoint. Attackers can sequentially iterate solution IDs to discover all FAQs including those restricted to specific users or groups, leaking sensitive metadata through redirect Location headers and page canonical links.\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Airbnb host tries to scam customer using fake AI smash and grab images. Technology: Machine learning. Purpose: Defraud. Ethical issues: Mis/disinformation; Transparency.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nFour vulnerabilities: API key exfiltration, arbitrary code execution, model information leakage, and prompt injection via Python docstrings or Markdown configs. Opening an infected repository and asking Cline to analyze it triggers attacker commands without user approval.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA security vulnerability has been detected in edvardlindelof notes-mcp up to 0.1.4. This affects an unknown function of the file notes_mcp.py. The manipulation of the argument root_dir/path leads to path traversal. The attack is possible to be carried out remotely. The exploit has been disclosed publicly and may be used. The project was informed of the problem early through an issue report but has not responded yet.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nIBM Langflow Desktop 1.6.0 through 1.8.4 Lanflow is vulnerable to stored cross-site scripting. This vulnerability allows an authenticated user to embed arbitrary JavaScript code in the Web UI thus altering the intended functionality potentially leading to credentials disclosure within a trusted session.\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact A vulnerability has been identified when using Fleet to manage Helm charts where sensitive information is passed through `BundleDeployment.Spec.Options.Helm.Values` may be stored in plain text. This can result in: 1. Unauthorized disclosure of sensitive data: Any user with `GET` or `LIST` permissions on `BundleDeployment` resources could retrieve Helm values containing credentials or other secrets. 2. Lack of encryption at rest: `BundleDeployment` is not configured for Kubernetes encryption at rest by default, causing sensitive values to remain unencrypted within the cluster datastore. This behavior differs from Helm v3\u2019s default approach, where chart state \u2014 including values \u2014 is stored in Kubernetes secrets, benefiting from built-in protection mechanisms. In affected scenarios, credentials and other sensitive information are exposed both at rest and in responses to API calls. Please consult the associated [MITRE ATT&CK - Technique - Credentials from Password Stores](https://attack.mitre.org/techniques/T1555/) for further information about this category of attack. For the exposure of credentials not related to Rancher, the final impact severity for confidentiality, integrity, and availability is dependent on the permissions that the leaked credentials have on their own services. It is recommended to review the potentially exposed sensitive data in this scenario and change secrets, tokens, and passwords as necessary. ### Patches This vulnerability is addresse\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nFastGPT is an AI Agent building platform. Prior to 4.15.0-beta1, a Server-Side Request Forgery (SSRF) vulnerability allows an authenticated attacker to bypass the global isInternalAddress network protection and make arbitrary HTTP GET requests to internal network services. This is achieved by exploiting an incomplete fix in the dataset preview endpoint /api/core/dataset/file/getPreviewChunks when utilizing the externalFile data import type. This vulnerability is fixed in 4.15.0-beta1.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nFastMCP is the standard framework for building MCP applications. Versions prior to 2.13.0, a command-injection vulnerability lets any attacker who can influence the server_name field of an MCP execute arbitrary OS commands on Windows hosts that run fastmcp install cursor. This vulnerability is fixed in 2.13.0.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nFrance's CNIL fined Clearview AI EUR 20 million and ordered erasure of biometric data on French residents, finding scraping and processing of facial templates incompatible with GDPR.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nGoogle's Agent-to-Agent (A2A) protocol has systemic vulnerabilities: agent card poisoning (malicious metadata injection causing data exfiltration), agent impersonation/shadowing, replay attacks, and contagion risk (one compromised agent influencing others in collaborative workflows). The spec delegates credential management entirely to implementers with no built-in protections.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary Kozea/CairoSVG (~300K downloads/week) has exponential denial of service via recursive `<use>` element amplification in `cairosvg/defs.py` (line ~335). This causes CPU exhaustion from a small input. ## Severity High \u2014 CVSS 3.1: 7.5 Vector: `CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H` ## Vulnerable Code File: `cairosvg/defs.py` (line ~335), function `use()` The `use()` function recursively processes `<use>` elements without any depth or count limits. With 5 levels of nesting and 10 references each, a 1,411-byte SVG triggers 10^5 = 100,000 render calls. ## Impact - 1,411-byte SVG payload pins CPU at 100% indefinitely - Memory stays flat at ~43MB \u2014 no OOM kill, process never terminates - Any service accepting SVG input (thumbnailing, PDF generation, avatar rendering) is DoS-able - Amplification factor: O(10^N) rendering calls from O(N) input ## Proof of Concept Save as `poc.svg` and run `timeout 10 cairosvg poc.svg -o test.png`: ```xml <?xml version=\"1.0\"?> <svg xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\"> <defs> <g id=\"a\"><rect width=\"1\" height=\"1\"/></g> <g id=\"b\"><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/><use xlink:href=\"#a\"/></g> <g id=\"c\"><use xlink:href=\"#b\"/><use xlink:href=\"#b\"/><use xlink:href=\"#b\"/><use xlink:href=\"#b\"/><use xlink:href=\"#b\"/><use\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Huq admits GPS location data sharing privacy breach. Technology: Location tracking. Purpose: Track user location. Ethical issues: Privacy/surveillance.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary Keras\u2019s model loader (KerasFileEditor) unsafely loads user-supplied .keras model files containing HDF5-based weight files without performing any validation on HDF5 dataset metadata. An attacker can craft a .keras archive containing a valid model.weights.h5 file whose dataset declares an extremely large shape (e.g. (50_000_000, 50_000_000)), but stores only a few bytes. The .keras file remains small (100\u2013400 KB) because HDF5 with gzip compression stores minimal data. During model loading, Keras executes: `python result[key] = value[()] # loads entire dataset into memory` value[()] instructs h5py to allocate RAM proportional to the dataset\u2019s declared shape \u2013 in this case 8.88 PiB of memory. This results in: Immediate memory exhaustion Python / TensorFlow crashes Jupyter kernel kill System instability Full Denial of Service on any workload that processes untrusted .keras models This allows an attacker to crash any environment or pipeline that loads .keras models, including MLOps backends, training services, model upload endpoints, or automated pipelines. ### Proof of Concept ``` // PoC.py import zipfile import io import h5py import numpy as np from keras.saving import KerasFileEditor # Create a malicious .keras model containing a massive HDF5 shape bomb def create_malicious_keras(path=\"bomb.keras\"): hdf5_bytes = io.BytesIO() # Create an HDF5 file with a huge declared dataset shape with h5py.File(hdf5_bytes, \"w\") as f: d = f.create_dataset(\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Iranian group uses ChatGPT to target US presidential election. System: ChatGPT. Technology: Generative AI. Purpose: Generate misinformation. Ethical issues: Mis/disinformation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability in the RpcAgentServerLauncher class of modelscope/agentscope v0.0.6a3 allows for remote code execution (RCE) via deserialization of untrusted data using the dill library. The issue occurs in the AgentServerServicer.create_agent method, where serialized input is deserialized using dill.loads, enabling an attacker to execute arbitrary commands on the server.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary Arcane's huma-based REST API exposes nine endpoints under `/api/customize/git-repositories` and `/api/git-repositories/sync` for managing GitOps source repositories and their stored credentials. Eight of those endpoints (`list`, `create`, `get`, `update`, `delete`, `test`, `listBranches`, `browseFiles`) never call the `checkAdmin(ctx)` helper that every other admin-managed resource (container registries, environments, users, API keys, swarm, settings, system, notifications, events) uses, and the huma authentication middleware deliberately enforces only authentication, not the `admin` role. As a result, any logged-in user with the default `user` role can list, create, modify, delete, and test git repository configurations. By repointing an existing repository's URL to an attacker-controlled host while omitting the `token`/`sshKey` fields (which `UpdateRepository` only rewrites when explicitly supplied), the attacker causes Arcane to decrypt the legitimate PAT/SSH key on its next `/test`, `/branches`, or `/files` call and present it as HTTP Basic auth (or SSH key auth) to the attacker's host \u2014 producing a one-step exfiltration of plaintext Git credentials. ## Details ### Auth bridge does not enforce role `backend/internal/huma/middleware/auth.go:192-254` (`NewAuthBridge`) validates Bearer JWTs / API keys / agent tokens and stores the user (and an `userIsAdmin` flag) in the request context, but it never rejects non-admin callers \u2014 admin enforcement is intentionall\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe Irish Data Protection Commission fined TikTok EUR 345 million for GDPR violations related to children's data processing and transparency failures. Separately, ongoing EU regulatory pressure over TikTok's data transfers to China led to the mandatory implementation of Project Clover \u2014 a EUR 12 billion program to localize European user data in European data centers. While not AI-specific, the case directly impacts TikTok's recommendation algorithm (an AI system) because training data, user interaction data, and model inference data must now be processed within EU borders. The enforcement established that AI systems processing EU personal data must comply with data localization requirements, and that the AI recommendation engine cannot be separated from the data governance obligations. This is the defining case for DSGAI20 \u2014 data localization violations in AI systems.\n</incident>"}], "ground_truth": "rce", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Study: Uber algorithm locks Indian drivers out of accounts. System: Real-Time ID Check; Face ID. Technology: Facial recognition. Purpose: Identify identity. Ethical issues: Accountability; Accuracy/reliability; Employment/labour; Transparency.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The Velbus asset import path parses attacker-controlled XML without explicit XXE hardening. An authenticated user who can call the import endpoint may trigger XML external entity processing, which can lead to server-side file disclosure and SSRF. The target file must be less than 1023 characters. ### Details Velbus import uses `DocumentBuilderFactory.newInstance().newDocumentBuilder().parse(...)` on untrusted XML input, without explicit safeguards to disable DTD/external entities. ```154:165:agent/src/main/java/org/openremote/agent/protocol/velbus/AbstractVelbusProtocol.java @Override public Future<Void> startAssetImport(byte[] fileData, Consumer<AssetTreeNode[]> assetConsumer) { return executorService.submit(() -> { Document xmlDoc; try { String xmlStr = new String(fileData, StandardCharsets.UTF_8); LOG.info(\"Parsing VELBUS project file\"); xmlDoc = DocumentBuilderFactory .newInstance() .newDocumentBuilder() .parse(new InputSource(new StringReader(xmlStr))); ``` Expanded `Caption` content is propagated into created asset names: ```193:198:agent/src/main/java/org/openremote/agent/protocol/velbus/AbstractVelbusProtocol.java String name = module.getElementsByTagName(\"Caption\").item(0).getTextContent(); name = isNullOrEmpty(name) ? deviceType.toString() : name; // TODO\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA single malicious prompt creates a covert DNS-based exfiltration channel leaking user messages, uploaded files, and conversation content. Bypasses AI guardrails by exploiting the underlying Linux runtime. Fixed by OpenAI February 20, 2026.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe US Federal Trade Commission's first AI-marketing fraud action targeted Automators LLC, which pitched 'AI-powered' Amazon storefronts and Walmart dropshipping coaching that allegedly delivered no working AI and cost consumers more than USD 22 million in losses.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nNHTSA expanded its investigation into Tesla Full Self-Driving (FSD) after hundreds of reports of phantom braking events \u2014 the AI vision system hallucinating obstacles (shadows, overpasses, road markings) and applying emergency braking at highway speed. The investigation covered 2.4 million vehicles. Multiple rear-end collisions resulted from unexpected braking. The investigation questioned whether the pure-vision AI perception system (no LiDAR, no radar) could reliably distinguish real obstacles from visual artifacts.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nLasso Security identified a major Microsoft Copilot vulnerability that exposed indexed enterprise data from Fortune 500 companies through Bing/Copilot's caching of formerly public-then-private GitHub repository content.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary A path traversal vulnerability was discovered in the quarkus-openapi-generator extension ### Details The `unzip()` method in `ApicurioCodegenWrapper.java` extracts ZIP entries without validating that the resolved file path stays within the intended output directory. At line 101, the destination is constructed as `new File(toOutputDir, entry.getName())` and the content is written immediately. A malicious ZIP archive containing entries with path traversal sequences (e.g., `../../malicious.java`) would write files outside the target directory. The interesting thing is that the client module in the same repository already has the correct fix. `OpenApiGeneratorStreamCodeGen.java` at line 137 performs proper `normalize()` and `startsWith()` validation. The server module was simply missed. ### PoC This vulnerability is exploitable when an attacker controls or can intercept the ZIP archive served by the Apicurio registry. In environments where the registry connection is over an untrusted network or where TLS is not properly configured, exploitation becomes practical. The attack occurs at build/codegen time. 1. Create a ZIP file containing an entry named `../../proof.txt` with arbitrary content 2. Configure quarkus-openapi-generator to use the server (Apicurio) code generation path 3. Serve the malicious ZIP from a controlled or MITM'd Apicurio registry endpoint 4. Trigger code generation 5. Observe that `proof.txt` is written two directories above the intended output\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nTurkish presidential candidate Muharrem Ince withdrew from the 2023 election days before the vote, citing a non-consensual sexual deepfake video he said was fabricated to discredit him. The episode is widely treated as an early high-impact election deepfake.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The Glances web server exposes a REST API (`/api/4/*`) that is accessible without authentication and allows cross-origin requests from any origin due to a permissive CORS policy (`Access-Control-Allow-Origin: *`). This allows a malicious website to read sensitive system information from a running Glances instance in the victim\u2019s browser, leading to cross-origin data exfiltration. While a previous advisory exists for XML-RPC CORS issues, this report demonstrates that the REST API (`/api/4/*`) is also affected and exposes significantly more sensitive data. ### Details When Glances is started in web mode (e.g., `glances -w -B 0.0.0.0`), it exposes a REST API endpoint at: http://<host>:61208/api/4/all The server responds with: Access-Control-Allow-Origin: * This allows any origin to perform cross-origin requests and read responses. The `/api/4/all` endpoint returns extensive system information, including: - Process list (`processlist`) - System details (hostname, OS, CPU info) - Memory and disk usage - Network interfaces and IP address - Running services and metrics Because no authentication is required by default, this data is accessible to any web page. ### PoC 1. Start Glances: glances -w -B 0.0.0.0 2. Create a malicious HTML file: ``` <!DOCTYPE html> <html> <body> <script> fetch(\"http://<victim-ip>:61208/api/4/all\") .then(r => r.json()) .then(data => { console.log(\"DATA:\", data); }); </script> </body> </html> ``` 2. Open the file in a browser whi\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: ICE facial recognition app misidentifies woman twice. System: Mobile Fortify. Technology: Facial recognition. Purpose: Verify individuals for detention, deportation. Ethical issues: Accountability; Accuracy/reliability; Automation bias; Autonomy/agency; Privacy/surveillance; Transparency. Reported consequences: Legislation introduction/update.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The `get_api_video_file` and `get_api_video` API endpoints in AVideo return full video playback sources (direct MP4 URLs, HLS manifests) for password-protected videos without verifying the video password. While the normal web playback flow enforces password checks via the `CustomizeUser::getModeYouTube()` hook, this enforcement is completely absent from the API code path. An unauthenticated attacker can retrieve direct playback URLs for any password-protected video by calling the API directly. ## Details The video password protection is enforced in the web UI via `CustomizeUser::getModeYouTube()` (`plugin/CustomizeUser/CustomizeUser.php:787`), which calls `videoPasswordIsGood()` before rendering the video player. However, this hook is only invoked during web page rendering \u2014 the API endpoints bypass it entirely. **Vulnerable endpoint 1 \u2014 `get_api_video_file` (`plugin/API/API.php:986-1004`):** ```php public function get_api_video_file($parameters) { global $global; $obj = $this->startResponseObject($parameters); $obj->videos_id = $parameters['videos_id']; if (!self::isAPISecretValid()) { if (!User::canWatchVideoWithAds($obj->videos_id)) { return new ApiObject(\"You cannot watch this video\"); } } $video = new Video('', '', $obj->videos_id); $obj->filename = $video->getFilename(); // ... $obj->video_file = Video::getHigherVideoPathFromID($obj->videos_id); $obj->sources = getSources($obj->filename,\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA customer used prompt injection to manipulate a ChatGPT-powered car-dealership chatbot into 'agreeing' to sell a 2024 Chevrolet Tahoe for one US dollar 'with no take-backsies', demonstrating jailbreak/agent-hijack risk in unsupervised commerce bots and triggering broader scrutiny of LLM-front-ended retail.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The fix for CVE-2026-26278 added entity expansion limits (`maxTotalExpansions`, `maxExpandedLength`, `maxEntityCount`, `maxEntitySize`) to prevent XML entity expansion Denial of Service. However, these limits are only enforced for DOCTYPE-defined entities. **Numeric character references** (`&#NNN;` and `&#xHH;`) and standard XML entities (`&lt;`, `&gt;`, etc.) are processed through a separate code path that does NOT enforce any expansion limits. An attacker can use massive numbers of numeric entity references to completely bypass all configured limits, causing excessive memory allocation and CPU consumption. ## Affected Versions fast-xml-parser v5.x through v5.5.3 (and likely v5.5.5 on npm) ## Root Cause In `src/xmlparser/OrderedObjParser.js`, the `replaceEntitiesValue()` function has two separate entity replacement loops: 1. **Lines 638-670**: DOCTYPE entities \u2014 expansion counting with `entityExpansionCount` and `currentExpandedLength` tracking. This was the CVE-2026-26278 fix. 2. **Lines 674-677**: `lastEntities` loop \u2014 replaces standard entities including `num_dec` (`/&#([0-9]{1,7});/g`) and `num_hex` (`/&#x([0-9a-fA-F]{1,6});/g`). **This loop has NO expansion counting at all.** The numeric entity regex replacements at lines 97-98 are part of `lastEntities` and go through the uncounted loop, completely bypassing the CVE-2026-26278 fix. ## Proof of Concept ```javascript const { XMLParser } = require('fast-xml-parser'); // Even with strict explicit limi\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA flaw has been found in ComfyUI up to 0.13.0. Affected is the function folder_paths.get_annotated_filepath of the file folder_paths.py of the component LoadImage Node. This manipulation of the argument Name causes path traversal. Remote exploitation of the attack is possible. The exploit has been published and may be used. The vendor was contacted early about this disclosure but did not respond in any way.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: European Parliament AI model is unable to identify first European Commission president. System: Archibot. Technology: Generative AI. Purpose: Provide facts about the European Parliament. Ethical issues: Accuracy/reliability; Mis/disinformation. Response: System review/update.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nLLama Factory enables fine-tuning of large language models. Prior to version 1.0.0, a critical vulnerability exists in the `llamafy_baichuan2.py` script of the LLaMA-Factory project. The script performs insecure deserialization using `torch.load()` on user-supplied `.bin` files from an input directory. An attacker can exploit this behavior by crafting a malicious `.bin` file that executes arbitrary commands during deserialization. This issue has been patched in version 1.0.0.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nTwitch streamer Atrioc was caught with a tab open to a paid site selling AI sexual deepfakes of fellow streamers QTCinderella, Pokimane, Sweet Anita and others, exposing a paid market for non-consensual streamer deepfakes.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Dream Machine AI video generator copies Disney's Monsters, Inc.. System: Dream Machine. Technology: Generative AI; Text-to-video; Machine learning. Purpose: Generate video. Ethical issues: Appropriation; Mis/disinformation; Transparency.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe WHO&#x27;s AI-powered health advisor, S.A.R.A.H. (Smart AI Resource Assistant for Health), is alleged to provide inconsistent and inadequate health information. The bot reportedly gives contradictory responses to the same queries, fails to offer specific contact details for healthcare providers, and inadequately handles severe mental health crises, often giving irrelevant or unhelpful advice.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Description A Remote Code Execution (RCE) vulnerability exists in the [modelscope/ms-swift](https://github.com/modelscope/ms-swift) project due to unsafe use of `yaml.load()` in combination with vulnerable versions of the PyYAML library (\u2264 5.3.1). The issue resides in the `tests/run.py` script, where a user-supplied YAML configuration file is deserialized using `yaml.load()` with `yaml.FullLoader`. If an attacker can control or replace the YAML configuration file provided to the `--run_config` argument, they may inject a malicious payload that results in arbitrary code execution. ## Affected Repository - **Project:** [modelscope/ms-swift](https://github.com/modelscope/ms-swift) - **Affect versions:** latest - **File:** `tests/run.py` - **GitHub Permalink:** https://github.com/modelscope/ms-swift/blob/e02ebfdf34f979bbdba9d935acc1689f8d227b38/tests/run.py#L420 - **Dependency:** PyYAML <= 5.3.1 ## Vulnerable Code ```python if args.run_config is not None and Path(args.run_config).exists(): with open(args.run_config, encoding='utf-8') as f: run_config = yaml.load(f, Loader=yaml.FullLoader) ``` ## Proof of Concept (PoC) ### Step 1: Create malicious YAML file (`exploit.yaml`) ```yaml !!python/object/new:type args: [\"z\", !!python/tuple [], {\"extend\": !!python/name:exec }] listitems: \"__import__('os').system('mkdir HACKED')\" ``` ### Step 2: Execute with vulnerable PyYAML (<= 5.3.1) ```python import yaml with open(\"exploit.yaml\", \"r\") as f: cfg = yaml.loa\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe CVE-2021-36156 fix validates the namespace parameter for path traversal sequences after a single URL decode, by double encoding, an attacker can read files at the Ruler API endpoint /loki/api/v1/rules/{namespace} Thanks to Prasanth Sundararajan for reporting this vulnerability.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary It is possible to bypass the mitigation introduced in response to [CVE-2025-1550](https://github.com/keras-team/keras/security/advisories/GHSA-48g7-3x6r-xfhp), when an untrusted Keras v3 model is loaded, even when \u201csafe_mode\u201d is enabled, by crafting malicious arguments to built-in Keras modules. The vulnerability is exploitable on the default configuration and does not depend on user input (just requires an untrusted model to be loaded). ### Impact | Type | Vector |Impact| | -------- | ------- | ------- | |Unsafe deserialization |Client-Side (when loading untrusted model)|Arbitrary file overwrite. Can lead to Arbitrary code execution in many cases.| ### Details Keras\u2019 [safe_mode](https://www.tensorflow.org/api_docs/python/tf/keras/models/load_model) flag is designed to disallow unsafe lambda deserialization - specifically by rejecting any arbitrary embedded Python code, marked by the \u201c__lambda__\u201d class name. https://github.com/keras-team/keras/blob/v3.8.0/keras/src/saving/serialization_lib.py#L641 - ``` if config[\"class_name\"] == \"__lambda__\": if safe_mode: raise ValueError( \"Requested the deserialization of a `lambda` object. \" \"This carries a potential risk of arbitrary code execution \" \"and thus it is disallowed by default. If you trust the \" \"source of the saved model, you can pass `safe_mode=False` to \" \"the loading function in order to allow `lambda\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA high-severity remote code execution vulnerability exists in feast-dev/feast version 0.53.0, specifically in the Kubernetes materializer job located at `feast/sdk/python/feast/infra/compute_engines/kubernetes/main.py`. The vulnerability arises from the use of `yaml.load(..., Loader=yaml.Loader)` to deserialize `/var/feast/feature_store.yaml` and `/var/feast/materialization_config.yaml`. This method allows for the instantiation of arbitrary Python objects, enabling an attacker with the ability to modify these YAML files to execute OS commands on the worker pod. This vulnerability can be exploited before the configuration is validated, potentially leading to cluster takeover, data poisoning, and supply-chain sabotage.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nDPD's customer service AI chatbot was manipulated by a customer using direct prompt injection. The customer instructed the bot to ignore previous instructions, causing it to swear, write poems criticising DPD, and call itself 'useless'. The incident went viral on social media. DPD disabled the AI chatbot and reverted to human-only customer service.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The `_safe_extractall` helper that all `recipe pull`, `recipe publish`, and `recipe unpack` flows route through validates each archive member's `name` for absolute paths, `..` segments, and resolved-path escape \u2014 but does **not** validate `member.linkname`, does not reject symlink/hardlink members, and calls `tar.extractall(dest_dir)` without `filter=\"data\"`. A bundle that contains a symlink with a name inside `dest_dir` but a `linkname` pointing outside it, followed by a regular file whose path traverses *through* the just-created symlink, escapes `dest_dir` and lets the attacker write arbitrary content to an attacker-chosen location on the victim's filesystem. ## Affected paths Every code path that calls `_safe_extractall` is exposed: | Caller | File:line | |---|---| | `praisonai recipe unpack` | `src/praisonai/praisonai/cli/features/recipe.py:1175` (introduced as the fix for GHSA-99g3-w8gr-x37c) | | `LocalRegistry.unpack` (recipe pull) | `src/praisonai/praisonai/recipe/registry.py:413` | | Registry archive validation (publish) | `src/praisonai/praisonai/recipe/registry.py:808` | ## Root cause `recipe/registry.py:131-178`: ```python def _safe_extractall(tar: tarfile.TarFile, dest_dir: Path) -> None: ... for member in tar.getmembers(): ... member_path = Path(member.name) if member_path.is_absolute(): raise RegistryError(...) if '..' in member_path.parts: raise RegistryError(...) resolved = (dest_resolved / mem\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA flaw has been found in langgenius dify up to 1.13.3. This issue affects the function parse_openai_plugin_json_to_tool_bundle of the file api/core/tools/utils/parser.py of the component ApiBasedToolSchemaParser. Executing a manipulation of the argument url can lead to server-side request forgery. The attack can be launched remotely. The exploit has been published and may be used. The vendor was contacted early about this disclosure but did not respond in any way.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability in Anthropic's Claude Google Chrome Extension allowed any website to silently inject prompts into the assistant as if the user wrote them, effectively a zero-click XSS-class prompt injection via web pages.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAuthenticated SSRF bypass in Microsoft Copilot Studio. Tenable researchers reached Microsoft's internal infrastructure, including IMDS and internal Cosmos DB. CVSS 8.5.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The local HTTP server started by `engram server` (binding `127.0.0.1:7337` by default) was exposed to any browser origin with no authentication unless `ENGRAM_API_TOKEN` was explicitly set. Combined with `Access-Control-Allow-Origin: *` on every response and a body parser that did not require `Content-Type: application/json`, this allowed a malicious web page the developer visited to: 1. **Exfiltrate** the local knowledge graph via `GET /query` and `GET /stats` (function names, file layout, recorded decisions/mistakes). 2. **Inject persistent prompt-injection payloads** via `POST /learn`, which wrote `mistake`/`decision` nodes that were later surfaced as system-reminders to the user's AI coding agent on every future session and file edit. Severity: **High** \u2014 confidentiality + persistent indirect prompt injection against the user's coding agent. ### Affected versions `engramx` >= 1.0.0, < 2.0.2 \u2014 any version that shipped the HTTP server. ### Patched in `engramx@2.0.2` ### Workarounds (if you cannot upgrade) - Do **not** run `engram server` or `engram ui`. - If developers must, set `ENGRAM_API_TOKEN` to a long random value and terminate the server before browsing the web. ### Remediation (applied in 2.0.2) 1. Fail-closed auth on every non-public route \u2014 Bearer header or HttpOnly cookie, constant-time comparison, 256-bit auto-generated token at `~/.engram/http-server.token` (0600). 2. Wildcard CORS removed entirely; default is no CORS headers. Opt-in allo\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## 1. Summary The Envoy RBAC (Role-Based Access Control) filter contains a logic vulnerability in how it validates HTTP headers when multiple values are present for the same header name. Instead of validating each header value individually, Envoy concatenates all values into a single comma-separated string. This behavior allows attackers to bypass RBAC policies\u2014specifically \"Deny\" rules\u2014by sending duplicate headers, effectively obscuring the malicious value from exact-match mechanisms. ## 2. Attack Scenario Consider an environment where an administrator wants to block external access to internal resources using a specific header flag. ### Configuration The Envoy proxy is configured with a **Deny** rule to reject requests containing the header `internal: true`. * **Rule Type:** Exact Match * **Target:** `internal` header must not equal `true`. ### The Bypass Logic 1. **Standard Request (Blocked):** * **Input:** `internal: true` * **Envoy Processing:** Sees string `\"true\"`. * **Result:** Match found. **Request Denied.** 2. **Exploit Request (Bypassed):** * **Input:** ```http internal: true internal: true ``` * **Envoy Processing:** Concatenates values into `\"true,true\"`. * **Matcher Evaluation:** Does `\"true,true\"` equal `\"true\"`? **No.** * **Result:** The Deny rule fails to trigger. **Request Allowed.** ## 3. Implications * **RBAC Bypass:** Remote attackers can bypass configured access controls. * **Unauthorize\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary `ListServiceAccount` (`GET /rustfs/admin/v3/list-service-accounts?user=<other>`) authorizes cross-user requests against `UpdateServiceAccountAdminAction` instead of `ListServiceAccountsAdminAction` at `rustfs/src/admin/handlers/service_account.rs:936`. The handler accepts the **wrong** admin action and rejects the **correct** one: - A user granted only `admin:UpdateServiceAccount` enumerates every service account in the cluster, including the root user's (HTTP 200, full metadata). - A user granted only `admin:ListServiceAccounts` \u2014 the permission name every IAM document treats as \"list service accounts\" \u2014 receives HTTP 403 AccessDenied on the same request. Because service account access keys act as the identifier a `UpdateServiceAccount` holder needs to rotate a secret, and the `UpdateServiceAccount` handler at `rustfs/src/admin/handlers/service_account.rs:489` performs no ownership check on the target access key, leaking those access keys lets a delegated \"service account updater\" role overwrite `root-sa-1`'s secret, authenticate as the root user's service account, and create a persistent backdoor admin with `admin:*` + `s3:*`. Proven live end-to-end against `rustfs/rustfs:latest` (1.0.0-alpha.91, revision `d4ea14c2`) \u2014 the same revision is byte-identical on current `origin/main`. ## Vulnerability Details - **Package:** `rustfs` (binary crate `rustfs`) - **Affected versions:** From `0a2411f` (the initial `service_account.rs` check-in on 2026-03-15) through cu\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary Symfony 1 has a gadget chain due to dangerous unserialize in `sfNamespacedParameterHolder` class that would enable an attacker to get remote code execution if a developer unserialize user input in his project. ### Details This vulnerability present no direct threat but is a vector that will enable remote code execution if a developper deserialize user untrusted data. For example: ```php public function executeIndex(sfWebRequest $request) { $a = unserialize($request->getParameter('user')); } ``` We will make the assumption this is the case in the rest of this explanation. Symfony 1 provides the class `sfNamespacedParameterHolder` which implements `Serializable` interface. In particular, when an instance of this class is deserialized, the normal php behavior is hooked by implementing `unserialize()` method: ```php public function unserialize($serialized) { $this->__unserialize(unserialize($serialized)); } ``` Which make an array access on the deserialized data without control on the type of the `$data` parameter: ```php public function __unserialize($data) { $this->default_namespace = $data[0]; $this->parameters = $data[1]; } ``` Thus, an attacker provide any object type in `$data` to make PHP access to another array/object properties than intended by the developer. In particular, it is possible to abuse the array access which is triggered on `$data[0]` for any class implementing `ArrayAccess` interface.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nCVSS 9.9. Critical RCE in Microsoft Semantic Kernel Python SDK's InMemoryVectorStore filter functionality. Attackers execute arbitrary code through crafted filter expressions. Semantic Kernel powers many Microsoft Copilot integrations and RAG-based AI applications. Fixed in python-1.39.4.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Study: Whisper AI transcription invents medical treatments. System: Whisper. Technology: Generative AI; Speech-to-text; Speech recognition. Purpose: Recognise speech; Transcribe speech. Ethical issues: Accuracy/reliability; Fairness; Mis/disinformation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary It has been discovered that the middleware functionality in Hoverfly is vulnerable to command injection through its `/api/v2/hoverfly/middleware` endpoint due to insufficient validation and sanitization in user input. ### Details The vulnerability exists in the middleware management API endpoint `/api/v2/hoverfly/middleware`. This issue is born due to combination of three code level flaws: 1. Insufficient Input Validation in [middleware.go line 94-96](https://github.com/SpectoLabs/hoverfly/blob/master/core/middleware/middleware.go#L93): ``` func (this *Middleware) SetBinary(binary string) error { this.Binary = binary // No validation of binary parameter here return nil } ``` 2. Unsafe Command Execution in [local_middleware.go line 14-19](https://github.com/SpectoLabs/hoverfly/blob/master/core/middleware/local_middleware.go#L13): ``` var middlewareCommand *exec.Cmd if this.Script == nil { middlewareCommand = exec.Command(this.Binary) // User-controlled binary } else { middlewareCommand = exec.Command(this.Binary, this.Script.Name()) // User-controlled binary and script } ``` 3. Immediate Execution During Testing in [hoverfly_service.go line 173](https://github.com/SpectoLabs/hoverfly/blob/master/core/hoverfly_service.go#L173): ``` _, err = newMiddleware.Execute(testData) // Executes middleware immediately for testing ``` ### POC 1. Send the below HTTP PUT request to `http://localhost:8888` in order to create our malicious middleware,\n</incident>"}], "ground_truth": "deserialization", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nCNET&#x27;s use of generative AI to write articles allegedly ran into plagiarism issues, reproducing verbatim phrases from other published sources or making minor changes to existing texts such as altering capitalization, swapping out words for synonyms, and changing minor syntax.\n</incident>"}], "ground_truth": "rce", "generated_content": "planning"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nVersions of Ray before 2.43.0 are vulnerable to Insertion of Sensitive Information into Log File: the Redis password (when passed as an argument) is logged in standard logs, allowing credential disclosure where logs are accessible.\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Verkada employees use facial recognition to surveil colleagues. System: People Analytics. Technology: Facial recognition. Purpose: Harass colleagues. Ethical issues: Privacy/surveillance; Safety.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Hangzhou No. 11 Middle School student surveillance. System: Hikvision Smart Classroom Behaviour Management System. Technology: Facial recognition; Emotion recognition; Deep learning; Neural network; Machine learning. Purpose: Assess student attentiveness. Ethical issues: Accuracy/reliability; Fairness; Privacy/surveillance; Security. Response: System suspension.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe H2O machine learning platform uses \"Iced\" classes as the primary means of moving Java Objects around the cluster. The Iced format supports inclusion of serialized Java objects. When a model is deserialized, any class is allowed to be deserialized (no class whitelist). An attacker can construct a crafted Iced model that uses Java gadgets and leads to arbitrary code execution when imported to the H2O platform.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nOpen-WebSearch is a multi-engine MCP server, CLI, and local daemon for agent web search and content retrieval. Prior to 2.1.7, isPublicHttpUrl / assertPublicHttpUrl in src/utils/urlSafety.ts do not recognize bracketed IPv6 literals and do not resolve DNS, which combine to allow non-blind SSRF with the response body returned to the caller. This vulnerability is fixed in 2.1.7.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMultiple public defenders and civil society groups have filed a legal action against S\u00e3o Paulo Metro, challenging its planned facial recognition system. They argue the AI technology is flawed and violates passengers' privacy and data protection rights, demanding transparency and compliance with legal standards before deployment.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nKuntz et al. introduce OS-HARM, an operating-system-level safety benchmark for computer-use agents covering harmful misuse, prompt injection, and unintended over-action across 150 tasks.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMaxKB is an open-source AI assistant for enterprise. Versions 2.7.1 and below contain a Stored Cross-Site Scripting (XSS) vulnerability where the frontend's MdRenderer.vue component parses custom <iframe_render> tags from LLM responses or Application Prologue configurations, bypassing standard Markdown sanitization and XSS filtering. The unsanitized HTML content is passed to the IframeRender.vue component, which renders it directly into an <iframe> via the srcdoc attribute configured with sandbox=\"allow-scripts allow-same-origin\". This can be a dangerous combination, allowing injected scripts to escape the iframe and execute JavaScript in the parent window using window.parent. Since the Prologue is rendered for any user visiting an application's chat interface, this results in a high-impact Stored XSS that can lead to session hijacking, unauthorized actions, and sensitive data exposure. This issue has been fixed in version 2.8.0.\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe Edmonton Police Service (EPS) in Canada released a facial image of a Black male suspect generated by an algorithm using DNA phenotyping, which was denounced by the local community as racial profiling.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n# Kimai 2.45.0 - Authenticated Server-Side Template Injection (SSTI) ## Vulnerability Summary | Field | Value | |-------|-------| | **Title** | Authenticated SSTI via Permissive Export Template Sandbox || **Attack Vector** | Network | | **Attack Complexity** | Low | | **Privileges Required** | High (Admin with export permissions and server access) | | **User Interaction** | None | | **Impact** | Confidentiality: HIGH (Credential/Secret Extraction) | | **Affected Versions** | Kimai 2.45.0 (likely earlier versions) | | **Tested On** | Docker: kimai/kimai2:apache-2.45.0 | | **Discovery Date** | 2026-01-05 | --- **Why Scope is \"Changed\":** The extracted `APP_SECRET` can be used to forge Symfony login links for ANY user account, expanding the attack beyond the initially compromised admin context. --- ## Vulnerability Description Kimai's export functionality uses a Twig sandbox with an overly permissive security policy (`DefaultPolicy`) that allows arbitrary method calls on objects available in the template context. An authenticated user with export permissions can deploy a malicious Twig template that extracts sensitive information including: 1. **Environment Variables** (APP_SECRET, DATABASE_URL) 2. **All User Password Hashes** (bcrypt) 3. **Serialized Session Tokens** 4. **CSRF Tokens** --- ## Prerequisites 1. **Authenticated Access**: Valid account with export permissions (typically ROLE_ADMIN, ROLE_SUPER_ADMIN, or ROLE_TEAMLEAD) 2. **Template Deployment**: Ability t\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nEnterprise deployments of Microsoft 365 Copilot were found to surface confidential SharePoint data (salaries, M&A docs, HR files) to employees who had inherited overly broad permissions, because Copilot retrieved everything the user technically had access to.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nPrototype pollution vulnerability in apidoc-core versions 0.2.0 and all subsequent versions allows remote attackers to modify JavaScript object prototypes via malformed data structures, including the \u201cdefine\u201d property processed by the application, potentially leading to denial of service or unintended behavior in applications relying on the integrity of prototype chains. This affects the preProcess() function in api_group.js, api_param_title.js, api_use.js, and api_permission.js worker modules.\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAn Israeli farmer, Moshe Har Shemesh, reportedly received a fine generated by a Tax Authority software system whose calculation officials were allegedly unable to explain. When the farmer reportedly sought access to the program or its source code to understand the basis for the amount, the authority allegedly refused, citing security concerns and the difficulty of extracting the embedded guidelines. The dispute reportedly later moved into legal proceedings focused on whether code and automated decision rules constitute information subject to public disclosure.\n</incident>"}], "ground_truth": "rce", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nClearview AI scraped over 30 billion images from social media and the open web without consent to build a facial-recognition tool sold to law enforcement, drawing GDPR-level fines from Italy, France, the UK and Australia, and exposing systemic biometric privacy abuse.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Company uses Marques Brownlee AI voice clone to promote product without consent. Technology: Deepfake. Purpose: Sell product. Ethical issues: Appropriation; Authenticity/integrity; Autonomy/agency; Consent; Transparency.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary `sveltekit-superforms` v2.27.3 and prior are susceptible to a prototype pollution vulnerability within the `parseFormData` function of `formData.js`. An attacker can inject string and array properties into `Object.prototype`, leading to denial of service, type confusion, and potential remote code execution in downstream applications that rely on polluted objects. ### Details Superforms is a SvelteKit form library for server and client form validation. Under normal operation, form validation is performed by calling the the `superValidate` function, with the submitted form data and a form schema as arguments: ```js // https://superforms.rocks/get-started#posting-data const form = await superValidate(request, your_adapter(schema)); ``` Within the `superValidate` function, a call is made to `parseRequest` in order to parse the user's input. `parseRequest` then calls into `parseFormData`, which in turn looks for the presence of `__superform_json` in the form parameters. If `__superform_json` is present, the following snippet is executed: ```js // src/lib/formData.ts if (formData.has('__superform_json')) { try { const transport = options && options.transport ? Object.fromEntries(Object.entries(options.transport).map(([k, v]) => [k, v.decode])) : undefined; const output = parse(formData.getAll('__superform_json').join('') ?? '', transport); if (typeof output === 'object') { // Restore uploaded files and add to data const filePaths = Array.from\n</incident>"}], "ground_truth": "deserialization", "generated_content": "deserialization"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Student reading aloud during remote exam is falsely accused of cheating. System: ProctorU. Technology: Facial recognition. Purpose: Detect and prevent cheating. Ethical issues: Accuracy/reliability; Fairness.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nSQLBot is an intelligent Text-to-SQL system based on large language models and RAG. Prior to 1.8.0, SQLBot contains a Cross-Workspace IDOR (Insecure Direct Object Reference) and Authorization Bypass vulnerability in the /api/v1/datasource/exportDsSchema and /api/v1/datasource/uploadDsSchema endpoints. An attacker can access and modify database schemas and data sources belonging to other tenants/workspaces. This vulnerability is fixed in 1.8.0.\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Scammers bypass WePay facial recognition security using gifs. System: WeChat Pay. Technology: Facial recognition. Purpose: Verify customer identity. Ethical issues: Security.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA security flaw has been discovered in Agions taskflow-ai up to 2.1.8. This impacts an unknown function of the file src/mcp/server/handlers.ts of the component terminal_execute. Performing a manipulation results in os command injection. The attack is possible to be carried out remotely. Upgrading to version 2.1.9 will fix this issue. The patch is named c1550b445b9f24f38c4414e9a545f5f79f23a0fe. Upgrading the affected component is recommended. The vendor was contacted early, responded in a very professional manner and quickly released a fixed version of the affected product.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary `gix-path` executes `git` to find the path of a configuration file that belongs to the `git` installation itself, but mistakenly treats the local repository's configuration as system-wide if no higher scoped configuration is found. In rare cases, this causes a less trusted repository to be treated as more trusted, or leaks sensitive information from one repository to another, such as sending credentials to another repository's remote. ### Details In `gix_path::env`, the underlying implementation of the `installation_config` and `installation_config_prefix` functions calls `git config -l --show-origin` and parses the first line of the output to extract the path to the configuration file holding the configuration variable of highest [scope](https://git-scm.com/docs/git-config#SCOPES): https://github.com/Byron/gitoxide/blob/12251eb052df30105538fa831e641eea557f13d8/gix-path/src/env/git/mod.rs#L91 https://github.com/Byron/gitoxide/blob/12251eb052df30105538fa831e641eea557f13d8/gix-path/src/env/git/mod.rs#L112 While the configuration variable of highest scope is not usually in the local scope, there are practical situations where this occurs: - A configuration file truly associated with the installation is not present on all systems and can occasionally be empty. Likewise, there may be no variables in the global scope. - Configuration files associated with those higher scopes may be deliberately skipped by setting the `GIT_CONFIG_SYSTEM` and `GIT_CONFIG_GLOBAL` en\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary An attacker can send a maliciously crafted TOML to cause the parser to crash, because of a stack overflow caused by thousands of consecutive commented lines. The library uses recursion internally while parsing to skip over commented lines, which can be exploited to crash an application that is processing arbitrary TOML documents. ### Proof of concept ```js require(\"smol-toml\").parse('# comment\\n'.repeat(8000) + 'key = \"value\"') ``` ### Impact Applications which parse arbitrary TOML documents may suffer availability issues if they receive malicious input. If uncaught, the crash may cause the application itself to crash. The impact is deemed minor, as the function is already likely to throw errors on invalid input. Downstream users are supposed to properly handle errors in such situations. Due to the design of most JavaScript runtimes, the uncontrolled recursion does not lead to excessive memory usage and the execution is quickly aborted. As a reminder, it is **strongly** advised when working with untrusted user input to expect errors to occur and to appropriately catch them. ### Patches Version 1.6.1 uses a different approach for parsing comments, which no longer involves recursion. ### Workarounds Wrap all invocations of `parse` and `stringify` in a try/catch block when dealing with untrusted user input.\n</incident>"}], "ground_truth": "dos", "generated_content": "dos"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Facebook fined for violating privacy of 200,000 South Koreans. System: Facebook Facial Recognition. Technology: Facial recognition. Purpose: Collect facial biometrics. Ethical issues: Accountability; Privacy/surveillance. Reported consequences: Regulatory investigation; Fine/settlement.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA hacker using the alias 'Gloomer' published data on Breach Forums claiming an OmniGPT breach: 30,000 user emails/phone numbers and 34 million message lines, including uploaded files containing credentials, billing info, and API keys.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: ChatGPT persuades depressive man to take pseudoephedrine. System: ChatGPT. Technology: Generative AI. Purpose: Provide emotional support. Ethical issues: Accountability; Autonomy/agency; Mis/disinformation; Safety; Transparency. Reported consequences: Hospitalisation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact In LXD's devLXD server, the source container identification process uses process cmdline (command line) information, allowing attackers to impersonate other containers by spoofing process names. The core issue lies in the findContainerForPID function in `lxd/api_devlxd.go`. This function identifies senders through two steps as shown below: 1. cmdline-based identification: Check while tracing back through parent processes, and if it starts with `[lxc monitor]`, extract the project name and container name from that process name in the format projectName_containerName. 2. PID namespace-based identification: If not found in Step 1, check against all containers' PID namespaces. https://github.com/canonical/lxd/blob/43d5189564d27f6161b430ed258c8b56603c2759/lxd/api_devlxd.go#L166-L276 Attackers can exploit Step 1 processing to impersonate arbitrary containers across projects by spoofing process names. ### Reproduction Steps 1. Access devLXD server from a normal container (e.g., EEEE): ``` root@EEEE:~# curl --unix-socket /dev/lxd/sock http://lxd-host/1.0/meta-data instance-id: 9f928574-2561-4eff-af82-a68e57d3c68b local-hostname: EEEE ``` 2. Use exec -a to spoof process name and impersonate another container (DDDD): ``` root@EEEE:~# bash -c \"exec -a '[lxc monitor]' curl --unix-socket /dev/lxd/sock http://lxd-host/1.0/meta-data -x 'test-project_DDDD'\" instance-id: 1bb2f1c3-3ad2-4cd6-9965-67b14c3582cc local-hostname: DDDD ``` This attack successfully obtains metadat\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nPostorius through 1.3.13 does not escape HTML in the message subject when rendering it in the Held messages pop-up, as exploited in the wild in May 2026.\n</incident>"}], "ground_truth": "xss", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The SSRF hardening shipped in v8.31.0 only covers outbound URLs that Gotenberg's Go code handles \u2014 Chromium asset fetches, webhook delivery, and download-from. The LibreOffice conversion endpoint (`/forms/libreoffice/convert`) passes uploaded documents directly to LibreOffice without inspecting their content. LibreOffice then fetches any embedded external URLs on its own, completely bypassing the SSRF filters. This was verified on v8.31.0 (latest at time of writing) with a crafted DOCX and got 3 outbound HTTP requests from LibreOffice to the canary server used for testing. ### Details When a file is uploaded to `/forms/libreoffice/convert`, the route in `pkg/modules/libreoffice/routes.go` reads form parameters and passes the input file directly to `libreOffice.Pdf()`: ```go err = libreOffice.Pdf(ctx, ctx.Log(), inputPath, outputPaths[i], options) ``` There's no content inspection happening before the file reaches LibreOffice. The SSRF protection in v8.31.0 (`pkg/gotenberg/outbound.go`) wraps Go's `http.Client` with a custom dialer that resolves URLs and rejects non-public IPs \u2014 but LibreOffice is a separate process that makes its own HTTP connections via libcurl. The Go-level dial hooks can't intercept that. OOXML formats like DOCX can embed external image references using `TargetMode=\"External\"` in relationship files. LibreOffice fetches those URLs during PDF conversion. **Suggested fix:** Run LibreOffice with `unshare --net` to drop all network access fr\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n| Field | Value | |-------|-------| | **Affected Software** | Paperclip AI v2026.403.0 | | **Affected Component** | Execution Workspace lifecycle (`workspace-runtime.ts`) | | **Affected Endpoint** | `PATCH /api/execution-workspaces/:id` | | **Deployment Modes** | All \u2014 `local_trusted` (zero auth), `authenticated` (any company user) | | **Platforms** | Linux, macOS, Windows (with Git installed) | | **Date** | 2026-04-13 | --- ## Executive Summary A critical OS command injection vulnerability exists in Paperclip's execution workspace lifecycle. An attacker can inject arbitrary shell commands into the `cleanupCommand` field via the `PATCH /api/execution-workspaces/:id` endpoint. When the workspace is archived, the server executes this command verbatim via `child_process.spawn(shell, [\"-c\", cleanupCommand])` with no input validation or sanitization. In `local_trusted` mode (the default for desktop installations), this requires zero authentication. Three independent proofs of exploitation were demonstrated on Windows 11: arbitrary file write, full system information exfiltration (`systeminfo`), and GUI application launch (`calc.exe`). --- ## Root Cause Analysis ### Vulnerable Code Path **`server/src/services/workspace-runtime.ts` (line ~738)** The `cleanupExecutionWorkspaceArtifacts()` function iterates over cleanup commands from workspace config and executes each via shell: ```typescript // workspace-runtime.ts \u2014 cleanupExecutionWorkspaceArtifacts() for (const command o\n</incident>"}], "ground_truth": "ssrf", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The kernel stores Attribute View (AV / database) names without any HTML escape, then a render template uses raw `strings.ReplaceAll(tpl, \"${avName}\", nodeAvName)` to embed the name in HTML before pushing to all clients via WebSocket. Three independent client paths (`render.ts:120` \u2192 `outerHTML`, `Title.ts:401` \u2192 `innerHTML`, `transaction.ts:559` \u2192 `innerHTML`) consume the value without escaping. Because the main BrowserWindow runs `nodeIntegration:true, contextIsolation:false, webSecurity:false` (`app/electron/main.js:407-411`), HTML injection in the renderer becomes Node.js code execution. Payload is stored on disk under `data/storage/av/<id>.json`, replicates via every sync transport (S3 / WebDAV / cloud), survives `.sy.zip` export-import, and triggers for any role (Administrator / Editor / Reader / publish-service Visitor) opening a doc bound to the AV. ## Details **Kernel write \u2014 no escape.** `kernel/model/attribute_view.go:3244-3255`: ```go attrView.Name = strings.TrimSpace(operation.Data.(string)) attrView.Name = strings.ReplaceAll(attrView.Name, \"\\n\", \" \") if 512 < utf8.RuneCountInString(attrView.Name) { attrView.Name = gulu.Str.SubStr(attrView.Name, 512) } err = av.SaveAttributeView(attrView) // \u2190 no html.EscapeString ``` **Kernel template \u2014 raw replace.** `kernel/model/attribute_view.go:3242,3283-3284`: ```go const attrAvNameTpl = `<span data-av-id=\"${avID}\" ... class=\"popover__block\">${avName}</span>` // ... tpl := strings.ReplaceAll(attr\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nPublicly deployed open-source model DALL-E Mini was acknowledged by its developers and found by its users to have produced images which reinforced racial and gender biases.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA reflected XSS vulnerability was found under admin panel -> System -> Import/Export -> Dataflow - Profiles. ## Steps to produce + Login to the admin panel + Go to the path `System -> Import/Export -> Dataflow - Profiles` + Select profile direction as `Import`. + Click on `Import Customers` + Upload the file. File Link: [customer_20260212_204335.csv](https://github.com/user-attachments/files/25629638/customer_20260212_204335.csv) + Go back to `Run profile`. + Select the uploaded file and Click on `Run in Popup`. + One can see a URL like this ``` https://demo-admin.openmage.org/index.php/admin/system_convert_gui/run/id/6/key/40dbbb2e93f45f0463c57ff733352f4f/files/import-20260215151125-1_customer_20260212_204335.csv/ ``` + One can see the filename getting reflection in HTML tags. + Inject an HTML tag and observe. ``` https://demo-admin.openmage.org/index.php/admin/system_convert_gui/run/id/6/key/40dbbb2e93f45f0463c57ff733352f4f/files/\"><h3>hacked</h3>/ ``` <img width=\"1796\" height=\"302\" alt=\"image (3)\" src=\"https://github.com/user-attachments/assets/502330b0-fa73-4b90-a81f-6216a98e474a\" /> + One can see the tag is getting executed. + Proceed for XSS. ``` https://demo-admin.openmage.org/index.php/admin/system_convert_gui/run/id/6/key/40dbbb2e93f45f0463c57ff733352f4f/files/%3CScRiPt%20%3Eprompt(document.cookie)%3C%2FScRiPt%3E ``` <img width=\"1670\" height=\"562\" alt=\"image (4)\" src=\"https://github.com/user-attachments/assets/98a75081-fa8c-4483-9078-0a\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Grok 2 generates Nazi Micky Mouse, Taylor Swift deepfakes. System: Grok. Technology: Generative AI. Purpose: Generate text. Ethical issues: Authenticity/integrity; Mis/disinformation; Safety.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "deepfake"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary CAPM3 is Metal3's Cluster API (CAPI) provider for baremetal provisioning in Kubernetes. Multiple cross-namespace access control vulnerabilities in Cluster API Provider Metal3 allow users with permissions to create or modify CAPM3 resources in one namespace to reference, read, or claim resources belonging to other namespaces. ## Patched In - **v1.13.0** (main branch \u2014 all fixes included) - **v1.12.5** (all four fixes backported) - **v1.11.8** (three of four fixes backported; Metal3DataClaim template restriction not applicable due to missing v1beta2 webhook infrastructure) ## Description Four related vulnerabilities were identified and fixed: ### 1. Cross-namespace Secret references in Metal3Machine Metal3Machine resources accepted userData, metaData, and networkData secret references pointing to arbitrary namespaces. A user could configure a Metal3Machine to reference secrets in namespaces they do not have access to, and the controller would fetch and use those secrets. ### 2. Cross-namespace BareMetalHost lookups The host annotation on Metal3Machine could include a namespace/name format, causing the controller to look up BareMetalHost resources in arbitrary namespaces. This allowed a user to claim or associate with BareMetalHosts belonging to other tenants. ### 3. Incorrect logical operator in ConsumerRef validation The Metal3LabelSync controller used AND logic (&&) when validating BareMetalHost ConsumerRef Kind and Group, meaning it only rejected a Co\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nRembg is a tool to remove images background. Prior to 2.0.75, a path traversal vulnerability in the rembg HTTP server allows unauthenticated remote attackers to read arbitrary files from the server's filesystem. By sending a crafted request with a malicious model_path parameter, an attacker can force the server to attempt loading any file as an ONNX model, revealing file existence, permissions, and potentially file contents through error messages. This vulnerability is fixed in 2.0.75.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: MEP files lawsuit to release iBorderCtrl lie detection system documents. System: iBorderCtrl. Technology: Behavioural analysis; Emotion recognition; Facial recognition. Purpose: Detect traveller lies. Ethical issues: Accountability; Accuracy/reliability; Fairness; Human rights/civil liberties; Transparency. Reported consequences: Litigation.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThis vulnerability exists in Nuclio Dashboard's project management API, allowing any authenticated user (without membership in the target project) to bypass OPA authorization checks on write paths (`PUT /api/projects/{id}`, `DELETE /api/projects`) and modify or delete any project along with all its associated resources (functions, API gateways, etc.). CWE classification: CWE-862 (Missing Authorization). --- ## Summary Nuclio Dashboard correctly enforces OPA-based authorization on the project read path (`GET /api/projects`), populating `MemberIds` in `PermissionOptions` so OPA can filter results by user membership. However, the write paths (`PUT /api/projects/{id}` and `DELETE /api/projects`) construct `PermissionOptions` without setting `MemberIds`. The platform-layer `FilterProjectsByPermissions` function (`pkg/platform/abstract/platform.go:652`) short-circuits when `MemberIds` is empty, bypassing OPA entirely. Any authenticated user who knows a project name can modify or delete that project, triggering cascading deletion of all associated Functions, APIGateways, and FunctionEvents. Affected: Nuclio v1.15.26 (latest Helm release) and HEAD commit e185454 (latest source). --- ## Attacker Model **Attacker type**: Authenticated low-privilege tenant (no membership in the target project) **Initial access**: - Holds any valid Nuclio Dashboard credentials - The account has no membership in the target project (OPA correctly denies `GET /api/projects`, returning an empty list)\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: AI creates error-plagued Wikipedia articles in obscure languages. System: Content Translate; Lsjbot. Technology: Bot/intelligent agent; Machine learning. Purpose: Translate articles. Ethical issues: Accountability; Accuracy/reliability; Mis/disinformation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact All versions of moment-timezone from 0.1.0 contain build tasks vulnerable to command injection. * if Alice uses tzdata pipeline to package moment-timezone on her own (for example via `grunt data:2014d`, where `2014d` stands for the version of the tzdata to be used from IANA's website), * and Alice let's Mallory select the version (`2014d` in our example), then Mallory can execute arbitrary commands on the machine running the grunt task, with the same privilege as the grunt task #### Am I affected? ##### Do you build custom versions of moment-timezone with grunt? If no, you're not affected. ##### Do you allow a third party to specify which particular version you want build? If yes, you're vulnerable to command injection -- third party may execute arbitrary commands on the system running grunt task with the same privileges as grunt task. ### Description #### Command Injection via grunt-zdownload.js and MITM on iana's ftp endpoint The `tasks/data-download.js` script takes in a parameter from grunt and uses it to form a command line which is then executed: ``` 6 module.exports = function (grunt) { 7 grunt.registerTask('data-download', '1. Download data from iana.org/time-zones.', function (version) { 8 version = version || 'latest'; 10 var done = this.async(), 11 src = 'ftp://ftp.iana.org/tz/tzdata-latest.tar.gz', 12 curl = path.resolve('temp/curl', version, 'data.tar.gz'), 13 dest = path.r\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA class action filed in California alleged OpenAI and Microsoft trained ChatGPT on personal information scraped from social media, blogs and forums without consent, framing web-scraping for LLM training as theft of personal information.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The 'gitImportSite' functionality obtains a URL string from a POST request and insufficiently validates user input. The \u2019set_remote\u2019 function later passes this input into \u2019proc_open\u2019, yielding OS command injection. ### Details The vulnerability exists in the logic of the \u2019gitImportSite\u2019 function, located in \u2019Operations.php\u2019. The current implementation only relies on the \u2019filter_var\u2019 and 'strpos' functions to validate the URL, which is not sufficient to ensure absence of all Bash special characters used for command injection. ![gitImportSite](https://github.com/user-attachments/assets/af9935ef-4735-446d-833f-2c2590ff1508) #### Affected Resources \u2022 Operations.php:2103 gitImportSite() \u2022 \\<domain\\>/\\<user\\>/system/api/gitImportSite ### PoC To replicate this vulnerability, authenticate and send a POST request to the 'gitImportSite' endpoint with a crafted URL in the JSON data. Note, a valid token needs to be obtained by capturing a request to another API endpoint (such as 'archiveSite'). 1. Start a webserver. ![webserver](https://github.com/user-attachments/assets/8594f9b1-67fa-4352-bbc3-310bb164ec9b) 2. Initiate a request to the \u2019archiveSite\u2019 endpoint. ![archiveSite](https://github.com/user-attachments/assets/08503f36-d984-4d53-8fe6-577ad78d5eb7) 3. Capture and modify the request in BurpSuite. ![request-modification](https://github.com/user-attachments/assets/61cd211e-afd3-453e-b86b-58bccffaf824) 4. Observe command output in the HTTP request from the se\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nmcp-neo4j-cypher is an MCP server for executing Cypher queries against Neo4j databases. In versions prior to 0.6.0, the read_only mode enforcement can be bypassed using APOC CALL procedures, potentially allowing unauthorized write operations or server-side request forgery. This issue is fixed in version 0.6.0.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nLibreChat is a ChatGPT clone with additional features. Prior to 0.8.4, LibreChat trusts the name field returned by the execute_code sandbox when persisting code-generated artifacts. On deployments using the default local file strategy, a malicious artifact filename containing traversal sequences (for example, ../../../../../app/client/dist/poc.txt) is concatenated into the server-side destination path and written with fs.writeFileSync() without sanitization. This gives any user who can trigger execute_code an arbitrary file write primitive as the LibreChat server user. This vulnerability is fixed in 0.8.4.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Description Sliver version 1.6.0 (prerelease) is vulnerable to RCE on the teamserver by a low-privileged \"operator\" user. The RCE is as the system root user. ## Impact As described in a [past issue](https://github.com/BishopFox/sliver/issues/65), \"there is a clear security boundary between the operator and server, an operator should not inherently be able to run commands or code on the server.\" An operator who exploited this vulnerability would be able to view all console logs, kick all other operators, view and modify files stored on the server, and ultimately delete the server. ## Reproduction First configure the Sliver server [in multiplayer mode and add an operator profile](https://sliver.sh/docs?name=Multi-player+Mode). Next, compile a slightly older version of the Sliver client. The commit after 5016fb8d updates the Cobra command-line parsing library in the Sliver client to strictly validate command flags. ``` git checkout 5016fb8d VERSION=1.6.0 make client ``` The latest server version is targeted: ``` All hackers gain exalted [*] Server v1.6.0 - bdfd89167dd47aece2397c638d482f94f3f91cba [*] Client 1.6.0 - 5016fb8d7cdff38c79e22e8293e58300f8d3bd57 [*] Welcome to the sliver shell, please type 'help' for options` ``` The exploit uses a command injection in the `generate msf-stager` to inject the `--out` flag to `msfvenom`. We overwrite Sliver's own go binary at `/root/.sliver/go/bin/go`: ``` sliver > generate msf-stager --lhost 192.168.0.128 --lport 8888 --adv\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Russian AI voice campaign attempts to undermine European support for Ukraine. System: ElevenLabs TTS. Technology: Text-to-speech; Deep learning; Machine learning. Purpose: Manipulate public opinion. Ethical issues: Mis/disinformation; Transparency.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMarkdownify is a Model Context Protocol server for converting almost anything to Markdown. Versions below 0.0.2 contain a command injection vulnerability, caused by the unsanitized use of input parameters within a call to child_process.exec, enabling an attacker to inject arbitrary system commands. Successful exploitation can lead to remote code execution under the server process's privileges. The server constructs and executes shell commands using unvalidated user input directly within command-line strings. This introduces the possibility of shell metacharacter injection (|, >, &&, etc.). This issue is fixed in version 0.0.2.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: RT bot farm spreads disinformation via 968 X accounts. System: Meliorator. Technology: Bot/intelligent agent. Purpose: Scare/confuse/destabilise. Ethical issues: Mis/disinformation; Transparency.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nmcp-security provides Security and Authorization support for Model Context Protocol in Spring AI. Prior to 0.1.9, the mcp-security framework fails to implement the mandatory SSRF mitigations outlined in the Model Context Protocol (MCP) security specifications. Specifically, it processes untrusted URLs for OAuth-related discovery and metadata without verifying if the targets are malicious or internal to the network. This only affects installations with Dynamic Client Registration (DCR) enabled This vulnerability is fixed in 0.1.9.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability was detected in chatchat-space Langchain-Chatchat up to 0.3.1.3. This vulnerability affects the function files/list_files/retrieve_file/retrieve_file_content/delete_file of the file libs/chatchat-server/chatchat/server/api_server/openai_routes.py of the component Compatible File Service. The manipulation results in missing authentication. The attacker must have access to the local network to execute the attack. The exploit is now public and may be used. The project was informed of the problem early through an issue report but has not responded yet.\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary A path traversal vulnerability exists in the FileUtil class of the code16/sharp package. The application fails to sanitize file extensions properly, allowing path separators to be passed into the storage layer. ### Detail In `src/Utils/FileUtil.php`, the `FileUtil::explodeExtension()` function extracts a file's extension by splitting the filename at the last dot. However, the extracted extension is never sanitized. While the application uses a `normalizeName()` function, this function only cleans the base filename, meaning any path separators (such as /) injected into the extension will survive and be passed into the `storeAs()` function. ### Impact Exploiting this flaw allows an authenticated attacker to manipulate file paths: - Files can be written outside of the intended tmp directory via path traversal. For more details on the package, visit: https://github.com/code16/sharp - Existing critical files (such as .env or configuration files) could potentially be overwritten. Review the CWE definition here: https://cwe.mitre.org/data/definitions/22.html (Note: This vulnerability was successfully chained with CWE-434 in a local Proof of Concept to confirm the traversal.) ### Patches This issue has been patched by properly sanitizing the extension using `pathinfo(PATHINFO_EXTENSION)` instead of `strrpos()`, alongside applying strict regex replacements to both the base name and the extension. The fix is available in pull request #715 ### Credits Reported by [zaurgs\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Instagram AI chatbots pretend to be licensed mental health therapists. System: Meta AI. Technology: Generative AI. Purpose: Provide emotional support. Ethical issues: Accuracy/reliability; Authenticity/integrity; Mis/disinformation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nImage upscaling tool PULSE powered by NVIDIA&#x27;s StyleGAN reportedly generated faces with Caucasian features more often, although AI academics, engineers, and researchers were not in agreement about where the source of bias was.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary The GraphQL Address element resolver (src/gql/resolvers/elements/Address.php) performs no schema scope filtering on top-level queries. A GraphQL API token scoped to a single low-privilege user group can read every address in the system, including addresses belonging to users in groups the token has no authorization to access. This exposes PII, including full names, addresses, organizations, tax IDs, etc. ### Details Every GraphQL element resolver in Craft CMS applies schema scope filtering via `GqlHelper::extractAllowedEntitiesFromSchema()` when handling top-level queries, except the Address resolver. The only gate check for addresses is `canQueryUsers()` (`src/gql/queries/Address.php`, line 30), which is a binary check. It returns `true` if the token has access to *any* user group. Once past this gate, no further filtering is applied. ### PoC **Tested on:** CraftCMS 5.9.17 (fresh Docker install, PHP 8.3) **Prerequisites:** A GraphQL API token with read access to any single user group ### Environment - Two user groups: `publicUsers` (in token scope) and `internalTeam` (NOT in scope) - 5 internal executives with corporate addresses (internalTeam) - 3 public customers with personal addresses (publicUsers) - GQL token scoped to `publicUsers:read` only **Step 1:** Introspect the schema to discover the `addresses` query is available to this token. Issue the below curl command ```bash curl -s -H \"Authorization: Bearer wbzwuzvlfohtahryztgaawyjpctqdvcm\" -H \"Con\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The `packages.js` template at `src/pyload/webui/app/themes/modern/templates/js/packages.js:172` interpolates a stored link URL into a template literal inside single-quoted HTML and then writes the result to the DOM via `$(div).html(html)`. No escaping runs between the API value and `innerHTML`. An attacker (Alice) who can submit a package link puts a single quote plus event handler into the URL, breaks out of the attribute, and executes JavaScript in every operator's browser that opens the downloads view. The theme does not set a Content Security Policy that restricts inline script or event handlers. ## Details **Sink**: `src/pyload/webui/app/themes/modern/templates/js/packages.js:165-188`: ```javascript const html = ` <span class='child_status'> <span style='margin-right: 2px;color: #337ab7;' class='${link.icon}'></span> </span> <span style='font-size: 16px; font-weight: bold;'> <a onclick='return false' href='${link.url}'>${link.name}</a> </span><br/> <div class='child_secrow' ...> <span class='child_status' ...>${link.statusmsg}</span>&nbsp;${link.error}&nbsp; <span class='child_status' ...>${link.format_size}</span> <span class='child_status' ...> ${link.plugin}</span>... </div>`; const div = document.createElement(\"div\"); $(div).attr(\"id\", `file_${link.id}`); $(div).css(\"padding-left\", \"30px\"); $(div).css(\"cursor\", \"grab\"); $(div).addClass(\"child\"); $(div).html(html); ``` `link.url` flows in from `/ap\n</incident>"}], "ground_truth": "xss", "generated_content": "xss"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: JR East facial recognition system prompts Privacy/surveillance row. System: Bio-IDiom. Technology: Facial recognition. Purpose: Identify criminal suspects. Ethical issues: Fairness; Normalisation; Privacy/surveillance; Transparency. Response: System review/update.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nSQLBot is an intelligent Text-to-SQL system based on large language models and RAG. In versions 1.7.0 and earlier, the Text2SQL chat interface is vulnerable to prompt injection. The user-provided question parameter is directly concatenated into the LLM prompt without filtering or escaping, and the SQL extracted from the LLM response is executed against the database without validation or sanitization. An authenticated attacker can craft a malicious question to manipulate the LLM into generating and executing arbitrary SQL statements. When connected to a PostgreSQL data source, this can lead to remote code execution via COPY FROM PROGRAM. This issue has been fixed in version 1.7.1.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: British Bangladeshi man wrongfully arrested for theft after facial recognition error. System: FaceVACS DBScan ID. Technology: Facial recgnition. Purpose: Identify criminal suspects. Ethical issues: Accountability; Accuracy/reliability; Automation bias; Autonomy/agency; Fairness; Privacy/surveillance; Transparency. Reported consequences: Litigation. Response: Public apology.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\naws-mcp-server AWS CLI Command Injection Remote Code Execution Vulnerability. This vulnerability allows remote attackers to execute arbitrary code on affected installations of aws-mcp-server. Authentication is not required to exploit this vulnerability. The specific flaw exists within the handling of the allowed commands list. The issue results from the lack of proper validation of a user-supplied string before using it to execute a system call. An attacker can leverage this vulnerability to execute code in the context of the MCP server. Was ZDI-CAN-27969.\n</incident>"}], "ground_truth": "command-injection", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary A memory leak vulnerability exists in the `LoadOpenCLDeviceBenchmark()` function in `MagickCore/opencl.c`. When parsing a malformed OpenCL device profile XML file that contains `<device` elements without proper `/>` closing tags, the function fails to release allocated memory for string members (`platform_name`, `vendor_name`, `name`, `version`), leading to memory leaks that could result in resource exhaustion. **Affected Version**: ImageMagick 7.1.2-12 and possibly earlier versions --- ### Details The vulnerability is located in `MagickCore/opencl.c`, function `LoadOpenCLDeviceBenchmark()` (lines 754-911). **Root Cause Analysis:** 1. When a `<device` tag is encountered, a `MagickCLDeviceBenchmark` structure is allocated (line 807-812) 2. String attributes (`platform`, `vendor`, `name`, `version`) are allocated via `ConstantString()` (lines 878, 885, 898, 900) 3. These strings are **only freed** when a `/>` closing tag is encountered (lines 840-849) 4. At function exit (lines 908-910), only the `device_benchmark` structure is freed, but **its member variables are not freed** if `/>` was never parsed **Vulnerable Code (lines 908-910):** ```c token=(char *) RelinquishMagickMemory(token); device_benchmark=(MagickCLDeviceBenchmark *) RelinquishMagickMemory( device_benchmark); // BUG: members (platform_name, vendor_name, name, version) not freed! ``` **Correct cleanup (only executed when `/>` is found, lines 840-849):** ```c device_benchmark->platform_nam\n</incident>"}], "ground_truth": "dos", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nThe nbconvert tool, jupyter nbconvert, converts Jupyter notebooks to various other formats via Jinja templates. In versions 6.5 through 7.17.0, when `HTMLExporter.embed_images=True`, nbconvert's markdown renderer allows arbitrary file read via path traversal in image references. A malicious notebook can exfiltrate sensitive files from the conversion host by embedding them as base64 data URIs in the output HTML. nbconvert 7.17.1 contains a fix. As a workaround, do not enable `HTMLExporter.embed_images`; it is not enabled by default.\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "path-traversal"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nBlack-box image-based prompt injection that embeds adversarial instructions into natural images using region selection, font scaling, and background-aware rendering. Up to 64 percent ASR on GPT-4-turbo under stealth constraints.\n</incident>"}], "ground_truth": "prompt-injection", "generated_content": "prompt-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Summary An unsafe deserialization vulnerability allows any unauthenticated user to execute arbitrary code on the server. ### PoC 1. Go to https://www.admidio.org/demo_en/adm_program/modules/messages/messages.php 2. Click on Send Private Message 3. In the `Message` field, enter the following payload `Testing<br><h1>HTML</h1><br><h2>Injection</h2>` > ![image](https://github.com/user-attachments/assets/0e5d9e4e-69c5-4908-9ab9-0c45c2548ff8) 4. Send the message 5. Open the message again > ![image](https://github.com/user-attachments/assets/d36f1b64-7d96-486d-ab65-cce2b7d21428) ### Impact 1. Data Theft: Stealing sensitive information like cookies, session tokens, and user credentials. 2. Session Hijacking: Gaining unauthorized access to user accounts. 3. Phishing: Tricking users into revealing sensitive information. 4. Website Defacement: Altering the appearance or content of the website. 5. Malware Distribution: Spreading malware to users' devices. 6. Denial of Service (DoS): Overloading the server with malicious requests.\n</incident>"}], "ground_truth": "deserialization", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: ChatGPT fails to debunk US election misinformation. System: ChatGPT. Technology: Generative AI. Purpose: Generate text. Ethical issues: Accuracy/reliability; Mis/disinformation.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nOpenAI&#x27;s ChatGPT was reportedly abused by cyber criminals including ones with no or low levels of coding or development skills to develop malware, ransomware, and other malicious softwares.\n</incident>"}], "ground_truth": "rce", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nNew York City Mayor Eric Adams sent robocalls to residents using AI-generated audio of his voice speaking in Spanish, Mandarin, Yiddish and other languages he does not speak, without disclosing the synthetic origin. Critics called it an officially-sanctioned deception.\n</incident>"}], "ground_truth": "deepfake", "generated_content": "rce"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n## Summary The `plugin/Permissions/setPermission.json.php` endpoint accepts GET parameters for a state-changing operation that modifies user group permissions. The endpoint has no CSRF token validation, and the application explicitly sets `session.cookie_samesite=None` on session cookies. This allows an unauthenticated attacker to craft a page with `<img>` tags that, when visited by an admin, silently grant arbitrary permissions to the attacker's user group \u2014 escalating the attacker to near-admin access. ## Details The root cause is a combination of three issues: **1. `$_REQUEST` used instead of `$_POST` (accepts GET parameters):** `plugin/Permissions/setPermission.json.php:14-24`: ```php $intvalList = array('users_groups_id','plugins_id','type','isEnabled'); foreach ($intvalList as $value) { if($_REQUEST[$value]==='true'){ $_REQUEST[$value] = 1; }else{ $_REQUEST[$value] = intval($_REQUEST[$value]); } } $obj = new stdClass(); $obj->id = Permissions::setPermission($_REQUEST['users_groups_id'], $_REQUEST['plugins_id'], $_REQUEST['type'], $_REQUEST['isEnabled']); ``` The only authorization check is `User::isAdmin()` at line 10 \u2014 there is no CSRF token validation via `isGlobalTokenValid()`. **2. Session cookies set to `SameSite=None`:** `objects/include_config.php:134-141`: ```php if ($isHTTPS) { // SameSite=None is intentional: AVideo supports cross-origin iframe embedding ini_set('session.cookie_samesite', 'None'); ini_set('sessi\n</incident>"}], "ground_truth": "path-traversal", "generated_content": "auth-bypass"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA vulnerability was found in TencentCloudBase CloudBase-MCP up to 2.17.0. Affected is the function openUrl of the file mcp/src/interactive-server.ts of the component open-url API Endpoint. The manipulation of the argument req.body.url results in server-side request forgery. It is possible to launch the attack remotely. The exploit has been made public and could be used. Upgrading to version 2.17.1 is able to address this issue. The patch is identified as 3f678a1e7bd400cd76469d61024097d4920dc6b5. It is recommended to upgrade the affected component.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n# Assessment The modules `uuid`, `_osx_support` and `_aix_support` were added to the blocklist of unsafe imports (https://github.com/trailofbits/fickling/commit/ffac3479dbb97a7a1592d85991888562d34dd05b). # Original report ## Summary fickling's `UNSAFE_IMPORTS` blocklist is missing at least 3 stdlib modules that provide direct arbitrary command execution: `uuid`, `_osx_support`, and `_aix_support`. These modules contain functions that internally call `subprocess.Popen()` or `os.system()` with attacker-controlled arguments. A malicious pickle file importing these modules passes both `UnsafeImports` and `NonStandardImports` checks. ## Affected Versions - fickling <= 0.1.8 (all versions) ## Details ### Missing Modules fickling's `UNSAFE_IMPORTS` (86 modules) does not include: | Module | RCE Function | Internal Mechanism | Importable On | |--------|-------------|-------------------|---------------| | `uuid` | `_get_command_stdout(cmd, *args)` | `subprocess.Popen((cmd,) + args, stdout=PIPE, stderr=DEVNULL)` | All platforms | | `_osx_support` | `_read_output(cmdstring)` | `os.system(cmd)` via temp file | All platforms | | `_osx_support` | `_find_build_tool(toolname)` | Command injection via `%s` in `_read_output(\"/usr/bin/xcrun -find %s\" % toolname)` | All platforms | | `_aix_support` | `_read_cmd_output(cmdstring)` | `os.system(cmd)` via temp file | All platforms | **Critical note:** Despite the names `_osx_support` and `_aix_support` suggesting platform-specific module\n</incident>"}], "ground_truth": "deserialization", "generated_content": "command-injection"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n### Impact The [PlantUML Macro](https://extensions.xwiki.org/xwiki/bin/view/Extension/PlantUML+Macro) is vulnerable to Server-Side Request Forgery (SSRF). The macro allows users to specify an alternative PlantUML server via the `server` parameter. However, the application does not validate the supplied URL. An attacker can supply an internal IP address or a malicious external URL. The XWiki server will attempt to connect to this URL to \"render\" the diagram. This issue affects all versions of the Plant UML Macro extension till version 2.4 included. ### Patches Version 2.4.1 of the Plant UML Macro extension fixes the issue by verifying if the supplied server domain matches one of the [trusted domain configured inside of XWiki](https://www.xwiki.org/xwiki/bin/view/Documentation/AdminGuide/Configuration/#HTrusteddomains). ### Workarounds Protect the XWiki server by placing it in a DMZ so that it cannot access any other internal servers. ### Resources The issue was fixed in [PLANTUML-25](https://jira.xwiki.org/browse/PLANTUML-25) by the commit [c8b19bda93058794e04c8862fc7ca85c59b5fe5c](https://github.com/xwiki-contrib/macro-plantuml/commit/c8b19bda93058794e04c8862fc7ca85c59b5fe5c). ### For more information If there are any questions or comments about this advisory: * Open an issue in [JIRA XWiki.org](https://jira.xwiki.org/) * Send an email to [Security Mailing List](mailto:security@xwiki.org) ### Attribution The issue was reported by \u0141ukasz Rybak.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nCybercrime group Storm-2139 hijacked Azure OpenAI accounts via stolen credentials, jailbroke the models to bypass content safeguards, and resold access. They produced thousands of policy-violating outputs including non-consensual explicit images.\n</incident>"}], "ground_truth": "auth-bypass", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\n**Description:** The `feldman_vss` library contains timing side-channel vulnerabilities in its matrix operations, specifically within the `_find_secure_pivot` function and potentially other parts of `_secure_matrix_solve`. These vulnerabilities are due to Python's execution model, which does not guarantee constant-time execution. An attacker with the ability to measure the execution time of these functions (e.g., through repeated calls with carefully crafted inputs) could potentially recover secret information used in the Verifiable Secret Sharing (VSS) scheme. The `_find_secure_pivot` function, used during Gaussian elimination in `_secure_matrix_solve`, attempts to find a non-zero pivot element. However, the conditional statement `if matrix[row][col] != 0 and row_random < min_value:` has execution time that depends on the value of `matrix[row][col]`. This timing difference can be exploited by an attacker. The `constant_time_compare` function in this file also does not provide a constant-time guarantee. This advisory formalizes the timing side-channel vulnerabilities already documented in the library's \"Known Security Vulnerabilities\" section. The Python implementation of matrix operations in the _find_secure_pivot and _secure_matrix_solve functions cannot guarantee constant-time execution, potentially leaking information about secret polynomial coefficients. An attacker with the ability to make precise timing measurements of these operations could potentially extract se\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMagento Open Source versions 2.4.7-p2, 2.4.6-p7, 2.4.5-p9, 2.4.4-p10 and earlier are affected by an Information Exposure vulnerability that could result in a security feature bypass. An admin attacker could leverage this vulnerability to have a low impact on confidentiality which may aid in further attacks. Exploitation of this issue does not require user interaction.\n</incident>"}], "ground_truth": "info-disclosure", "generated_content": "info-disclosure"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAIAAIC report: Paran\u00e1 school attendance facial recognition system criticised as \"invasive\", \"noxious\". System: LRCO Paran\u00e1. Technology: Facial recognition. Purpose: Register student attendance. Ethical issues: Accountability; Accuracy/reliability; Consent; Fairness; Normalisation; Power inbalance; Privacy/surveillance. Reported consequences: Litigation; Legislative proposal.\n</incident>"}], "ground_truth": "privacy-violation", "generated_content": "privacy-violation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nByteDance allegedly scraped short-form videos, usernames, profile pictures, and descriptions of accounts on Instagram, Snapchat, and other sources, and uploaded them without consent on Flipagram, TikTok\u2019s predecessor, in order to improve its \u201cFor You\u201d algorithm&#x27;s performance on American users.\n</incident>"}], "ground_truth": "rce", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nMeta AI launched Galactica \u2014 a large language model trained on scientific literature and designed to assist with scientific writing, summarisation, and knowledge retrieval \u2014 publicly via a demo on November 15, 2022. Within 72 hours, Meta withdrew the public demo after widespread criticism from the scientific community. Researchers found that Galactica confidently generated plausible-sounding but factually incorrect scientific text, including fabricated citations, incorrect chemical formulae, and authoritative-sounding passages on controversial topics (e.g., the history of bears in space). The core failure was that the model's confident, authoritative tone conveyed false certainty \u2014 users with limited domain expertise could not distinguish accurate from fabricated content. This remains the canonical real-world example of LLM misinformation at launch in a high-stakes domain.\n</incident>"}], "ground_truth": "misinformation", "generated_content": "misinformation"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nA flaw has been found in Ollama up to 18.1. This issue affects some unknown processing of the file server/download.go of the component Model Pull API. Executing a manipulation can lead to server-side request forgery. The attack can be launched remotely. The vendor was contacted early about this disclosure but did not respond in any way.\n</incident>"}], "ground_truth": "ssrf", "generated_content": "ssrf"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nAgentAPI is an HTTP API for Claude Code, Goose, Aider, Gemini, Amp, and Codex. Versions 0.3.3 and below are susceptible to a client-side DNS rebinding attack when hosted over plain HTTP on localhost. An attacker can gain access to the /messages endpoint served by the Agent API. This allows for the unauthorized exfiltration of sensitive user data, specifically local message history, which can include secret keys, file system contents, and intellectual property the user was working on locally. This issue is fixed in version 0.4.0.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "data-exfiltration"}
{"messages": [{"role": "user", "content": "What attack vector does the following GenAI security incident describe?\n\n<incident>\nClearview AI, known for its facial recognition system with a database of over 3 billion photos, suffered a data breach exposing its entire client list, including law enforcement agencies, and usage details. The incident raised significant privacy concerns and prompted criticism from lawmakers and investigations by privacy authorities.\n</incident>"}], "ground_truth": "data-exfiltration", "generated_content": "privacy-violation"}