LLMjacking: AI Model Hijacking Reaches Black Market Scale

Authors: Cloud Security Alliance AI Safety Initiative
Published: 2026-03-15

Categories: AI Security, Cloud Security, Threat Intelligence, Credential Security, Financial Risk
Download PDF

LLMjacking: AI Model Hijacking Reaches Black Market Scale

Cloud Security Alliance AI Safety Initiative | Research Note | March 15, 2026


Key Takeaways

  • LLMjacking — the unauthorized use of cloud-hosted LLM resources through compromised credentials, exposed APIs, or unauthenticated endpoints — has evolved from a novel proof-of-concept threat first documented by Sysdig in May 2024 into a commercially structured criminal enterprise with documented marketplace infrastructure by early 2026 [1][2].
  • Operation Bizarre Bazaar, jointly reported by Sysdig and Pillar Security Research in February 2026 [2][3], is the first publicly documented LLMjacking campaign to feature systematic attribution, full supply chain documentation, and confirmed marketplace monetization: researchers captured 35,000 attack sessions between December 2025 and January 2026 (averaging 972 attacks per day) against honeypot infrastructure [3].
  • The criminal infrastructure consists of a three-stage supply chain — automated reconnaissance using Shodan and Censys, quality-based endpoint validation, and commercial resale — culminating in the silver.inc underground marketplace, which resells unauthorized access to more than 30 LLM providers at 40–60% discounts via Telegram and Discord, accepting cryptocurrency and PayPal [3].
  • Victim organizations face financial losses that Sysdig researchers estimated could reach $46,000 per day for credential theft against AWS Bedrock at peak quota consumption [1], with one commercial assessment projecting potential exposure exceeding $100,000 per day when targeting the most capable frontier models [4].
  • A secondary MCP-focused campaign, distinct from the primary LLMjacking operation, accounted for 60% of total attack traffic against honeypot infrastructure by late January 2026, suggesting that attackers have begun targeting Model Context Protocol server endpoints as both an LLM access vector and a potential lateral movement pathway into broader enterprise infrastructure [3].

Background

When Sysdig’s Threat Research Team published the original LLMjacking research in May 2024, the threat category was genuinely novel. Researchers had discovered attackers exploiting CVE-2021-3129 [1], a known remote code execution vulnerability in outdated Laravel installations, to steal cloud credentials and use them specifically to access cloud-hosted large language model services [1]. The use case was economically motivated but organizationally narrow: the attacker script validated stolen credentials against ten distinct AI platforms — AI21 Labs, Anthropic, AWS Bedrock, Azure, ElevenLabs, MakerSuite, Mistral, OpenAI, OpenRouter, and GCP Vertex AI — and was designed to enable the attacker to consume inference capacity at the victim’s expense while reselling that access to paying customers through reverse proxy infrastructure [1]. The term “LLMjacking” was coined to describe this intersection of credential theft and AI resource abuse, distinguished from prior forms of cloud infrastructure hijacking (cryptojacking, for instance) by the specific targeting of language model services and the unique financial structure of AI inference billing.

The threat category gained attention in 2024 and early 2025 as AI adoption accelerated and organizations integrated LLM capabilities into cloud environments through services like Amazon Bedrock, Google Cloud Vertex AI, and Azure OpenAI Service. As these services became more widely deployed, they also became more frequently misconfigured: development and staging environments were left with unauthenticated API endpoints, long-lived API keys were committed to public repositories, and rate limiting was not applied consistently to AI-specific traffic. Each misconfiguration represented a potential LLMjacking entry point. The per-unit cost of AI inference — and the speed with which quota-saturating requests can accumulate charges — creates a financial exposure profile that can exceed equivalent misconfigurations in storage or compute environments. The financial structure of AI inference billing, with high per-unit costs and large available quotas, offers economic incentives that may have drawn actors previously focused on cryptocurrency mining or spam infrastructure toward AI-specific credential abuse, though the direct behavioral overlap between these populations cannot be confirmed from available evidence.

What was missing until late 2025 was evidence of organized, scaled, commercially structured criminal activity. LLMjacking was understood as a threat that attackers were exploring and, in some cases, operationalizing, but the full picture of systematic reconnaissance, quality-based victim triage, and commercial marketplace resale had not been documented and attributed. Operation Bizarre Bazaar, the campaign documented jointly by Sysdig and Pillar Security Research in February 2026, closed that gap.


Security Analysis

Operation Bizarre Bazaar: A Three-Stage Criminal Supply Chain

Operation Bizarre Bazaar represents the maturation of LLMjacking from opportunistic credential exploitation into a structured criminal enterprise with a documented supply chain, attributed threat actor infrastructure, and a functional commercial marketplace. The operation was characterized by Pillar Security Research as “the first public documentation of a systematic campaign targeting exposed LLM and Model Context Protocol endpoints at scale, featuring complete commercial monetization” [3].

The supply chain operates in three distinct stages. In the reconnaissance stage, distributed bot infrastructure systematically probes the internet for exposed AI endpoints, cataloging vulnerable Ollama instances (typically on port 11434 without authentication), unauthenticated vLLM servers (commonly on port 8000), accessible MCP endpoints, and OpenAI-compatible API surfaces exposed to the internet. Reconnaissance leverages public tools — Shodan and Censys foremost among them — to identify internet-facing AI infrastructure without requiring any prior knowledge of specific targets. The scale of this scanning is continuous; the honeypot data captured by Pillar Security shows that publicly exposed AI endpoints receive attack traffic within hours of becoming visible to internet scanners.

The validation stage refines the raw reconnaissance output through a process of endpoint testing that assesses both access and quality. Infrastructure tied to the silver.inc operation tests discovered endpoints through API calls, enumerates model capabilities, and evaluates response quality to determine which stolen access points are worth monetizing. Timing analysis from the honeypot data indicates that validation attempts follow public scans by an average of 2–8 hours, reflecting a structured workflow in which reconnaissance outputs are processed in near-real time [3]. The goal of validation is not merely to confirm that credentials work but to assess the economic value of each victim account: whether it has enabled high-capability models, what rate limits apply, and whether logging and monitoring configurations that might surface unauthorized usage are active. The Sysdig 2024 research documented that attackers specifically queried AWS Bedrock’s GetModelInvocationLoggingConfiguration API using invalid parameters — probing the environment without generating legitimate-looking traffic — to identify victim environments where logging was disabled before launching actual inference requests [1]. This logging evasion technique reflects attacker awareness of cloud monitoring capabilities and deliberate adaptation to avoid detection.

The monetization stage is the most consequential development in the threat’s evolution. silver.inc operates as what it describes to customers as “The Unified LLM API Gateway” — a commercial service offering discounted access to over 30 LLM providers through a single API interface [3]. In practice, the access being sold is entirely unauthorized: silver.inc resells stolen inference capacity at 40–60% below the retail pricing that victims are charged for the same capacity. The marketplace operates on Telegram and Discord, accepting both cryptocurrency and PayPal as payment methods. This payment flexibility suggests an intentional effort to lower the barrier for customers who may not be comfortable with cryptocurrency, broadening the potential customer base beyond technically sophisticated buyers.

Threat Actor Attribution

Pillar Security Research traced Operation Bizarre Bazaar to a threat actor operating under the alias Hecker, also known as Sakuya and LiveGamer101 [3]. Attribution was established through convergent evidence: the administrative panel for the operation displayed the message “Hiii I’m Hecker,” the silver.inc marketplace domain shared Cloudflare nameservers and DMARC DNS records with nexeonai.com (a service whose operators had previously been accused of conducting DDoS attacks against competitor platforms), and both domains were hosted on Netherlands-based bulletproof infrastructure with a documented history of thousands of abuse complaints. The attribution provides a degree of accountability relatively uncommon in financially motivated cybercrime at this scale, and the operational security failures that enabled it — shared infrastructure, self-identifying administrative interfaces, and domain reuse — suggest that the threat actor had not anticipated sustained investigative scrutiny.

The nexeonai.com connection is analytically significant beyond its role in attribution. The alleged DDoS attacks against competitors suggest a threat actor with financial motivations rooted in the AI access resale market itself: defending market share in underground AI services through coercive means against rival platforms. This profile is consistent with a criminal operation that views stolen LLM access as a revenue stream to be protected rather than a one-time exploitation opportunity, and it suggests organizational sophistication that extends beyond the technical attack capabilities alone.

Financial Impact and Victim Risk Profile

The financial exposure of LLMjacking victims has been documented through two complementary analyses. The original 2024 Sysdig research projected that an attacker consuming maximum quota against AWS Bedrock’s Claude 2 service across multiple regions could generate victim costs exceeding $46,000 per day, based on inference pricing of approximately $0.016 per 1,000 tokens at maximum rate limits [1]. One commercial assessment, attributing its estimate to frontier model rate card analysis, has projected exposure exceeding $100,000 per day for victims with access to the highest-tier models enabled in their cloud environments, though this figure has not been independently confirmed by primary research [4]. These projections assume continuous, rate-limit-saturating usage — the scenario an economically motivated attacker seeking maximum value extraction would pursue.

Real-world cases have been cited to illustrate that the financial impact can be severe even without rate-limit saturation. A cloud bill explosion from $12,000 to $143,000 attributed to credential compromise has been described in commercial security reporting; however, this case study originates from a single source without an independently verifiable primary reference and should be treated as an illustrative scenario rather than a documented incident [4]. What is well-established is the structural vulnerability: the detection latency present in cloud billing systems — in which charges may not surface in monitoring dashboards for several hours after they are incurred in default configurations, though real-time budget alert mechanisms from AWS, Google Cloud, and Azure can substantially reduce this window when properly configured — gives attackers a meaningful window to extract value before victim organizations can terminate compromised credentials.

The victim profile extends beyond the immediate billing impact. Because LLMjacking campaigns target cloud credentials rather than LLM-specific access tokens in isolation, compromised environments frequently contain credentials with permissions that extend well beyond AI service access. An attacker who validates stolen cloud credentials for LLMjacking purposes has, by definition, already demonstrated that those credentials work, and may subsequently use the same access for data exfiltration, lateral movement, or resource provisioning for other purposes. The validated-credential inventory assembled through Operation Bizarre Bazaar’s supply chain represents a persistent risk: credentials sold to multiple silver.inc customers for LLM access also exist in the attacker’s possession, where they could potentially be exploited for secondary purposes or sold through separate criminal channels — a risk inference consistent with standard criminal credential markets, though not specifically documented in reporting on the Bizarre Bazaar operation itself.

MCP Endpoints as a Secondary Attack Vector

A development distinct from but closely related to the primary LLMjacking campaign is the emergence of Model Context Protocol server endpoints as a targeted attack surface. By late January 2026, Pillar Security’s honeypot data showed that a campaign distinct in tooling and targeting patterns from the Hecker-attributed operation — assessed by Pillar as operated by a separate threat actor, based on those distinctions [3] — was responsible for 60% of total attack traffic against exposed AI infrastructure. This MCP-focused campaign prioritized endpoint enumeration and capability discovery rather than immediate inference consumption, a pattern that suggests a strategic orientation toward lateral movement preparation rather than cost-fraud monetization, though intent cannot be confirmed from behavioral data alone.

The threat model for MCP endpoint compromise is materially different from credential-based LLMjacking. Where credential theft enables unauthorized inference consumption, unauthorized access to an MCP server enables interaction with the resources that server exposes: file system operations, database queries, shell execution, API integrations, and whatever other tools the server has been configured to provide. MCP servers configured with access to internal services — such as development databases, code repositories, or internal APIs — represent a materially higher-value target than inference endpoints alone. Organizations should evaluate what tools their MCP servers expose as part of their threat model assessment, as the depth of internal connectivity varies considerably across deployments and determines the value of a given MCP server to a threat actor conducting espionage or ransomware staging.

The technical attack surface for MCP endpoint exploitation mirrors the patterns documented in the AI agent localhost vulnerability research published in early 2026: unauthenticated HTTP and WebSocket endpoints exposed to the internet without rate limiting, authentication, or origin validation. The same Shodan and Censys reconnaissance that identifies exposed Ollama instances and vLLM servers also surfaces unauthenticated MCP endpoints, making the attack surface continuous with the primary LLMjacking target set.

Newly Released Models as Preferential Targets

The targeting of DeepSeek-V3 within days of its public release reflects a tactical evolution in LLMjacking methodology [2]. Early LLMjacking attacks prioritized models that were broadly available and for which an established monetization demand existed among silver.inc-type customers. As the threat has matured, attackers have developed the capability to identify and pivot to newly released models rapidly, capturing early access before organizations have updated their security controls or billing anomaly baselines to account for the new models.

This behavior has significant implications for organizations that adopt new AI model offerings quickly. The same early-adopter dynamics that drive organizations to enable newly released models shortly after their availability — competitive advantage, capability evaluation, developer productivity — also create a detection gap: billing anomaly thresholds calibrated to existing model usage patterns will not flag zero-to-high-usage transitions on a model that was recently enabled as suspicious, and security teams may not have added new model names to monitoring queries. Attackers exploiting this gap can consume substantial resources during the detection window before organizations identify the unauthorized usage as anomalous.


Recommendations

Immediate Actions

Organizations with cloud accounts that have LLM services enabled should audit which models are active and which credentials have permissions to invoke them. API keys and IAM roles with LLM service permissions should be reviewed for necessity, scope, and exposure: any long-lived keys committed to source code repositories, stored in CI/CD pipeline environment variables without secrets management, or present in third-party integrations with more access than required represent a priority remediation target. Cloud provider access logging should be enabled for all AI service invocations, and billing anomaly alerts should be configured with thresholds appropriate for the cost-per-hour ceiling of the models available in each environment.

Environments running self-hosted LLM inference servers — Ollama, vLLM, or similar — should be audited for internet exposure. These servers commonly default to listening on all network interfaces without authentication, and the combination of internet exposure and default configuration is the primary target for the reconnaissance stage of Operation Bizarre Bazaar. Any inference server that does not require authentication should be treated as compromised if it has been internet-accessible, and access logs should be reviewed for evidence of unauthorized queries. Network segmentation that restricts inference server access to known internal client IP ranges represents the minimum required control.

Organizations that have deployed MCP servers in production environments should apply equivalent scrutiny to those endpoints. Any MCP server accessible from the internet without authentication should be treated as a high-priority remediation; given the documented shift of 60% of attack traffic toward MCP endpoint targeting, this exposure is actively being exploited.

Short-Term Mitigations

Detection coverage for LLMjacking should be implemented at two levels: billing-layer and API-layer. At the billing layer, cloud provider cost management tools should be configured with hourly budget alerts specific to AI service categories, with thresholds set to alert at a fraction of the maximum daily exposure for the models enabled in each environment. At the API layer, cloud-native logging services (AWS CloudTrail for Bedrock, Google Cloud Audit Logs for Vertex AI, Azure Monitor for OpenAI Service) should be queried for calls to logging configuration APIs (such as GetModelInvocationLoggingConfiguration in AWS Bedrock) using invalid parameters — the specific technique documented by Sysdig as a pre-exploitation reconnaissance pattern [1]. These calls should generate immediate security alerts rather than being treated as routine API noise.

Organizations should also monitor for traffic patterns consistent with the validation stage of the LLMjacking supply chain: rapid enumeration of available models, capability-testing calls with short prompts designed to assess model quality rather than accomplish a business task, and API calls that originate from IP ranges associated with datacenter hosting rather than expected user or application locations. Security teams should tune these signals against their specific application traffic baselines before treating them as high-fidelity alerts, as some patterns may overlap with legitimate automated API usage. In combination, however, they may be indicative of adversarial validation activity warranting further investigation.

For credential hygiene, the principle of short-lived credentials should be applied to AI service access wherever the cloud provider supports it. AWS IAM instance profiles and role assumption, Google Cloud Workload Identity Federation, and Azure Managed Identity all enable application code to access LLM services using temporary credentials that cannot be reused in the way long-lived API keys can, due to their short expiry windows. Organizations that have not migrated AI service authentication to short-lived credential mechanisms should treat this migration as a high-priority security initiative given the specific targeting of long-lived credentials by LLMjacking campaigns.

Specific network indicators from Operation Bizarre Bazaar that organizations can incorporate into threat detection include the silver.inc hosting subnet (204.76.203.0/24) and the autonomous system range AS135377 associated with the MCP-focused campaign [3]. These should be added to network block lists and monitored for connection attempts from internal systems that might indicate compromise.

Strategic Considerations

The commercialization of LLMjacking into a marketplace model suggests that the threat will continue to grow as AI adoption expands. silver.inc and analogous operations create an ecosystem in which the customer base for stolen AI access is far broader than the population of technically capable attackers who could conduct LLMjacking themselves: any individual willing to pay a 40–60% discount for AI inference capacity is a potential customer, regardless of their technical sophistication. This demand-side dynamic suggests durable economic pressure to maintain and expand LLMjacking supply chains, as long as organizations continue to misconfigure AI service credentials and the retail cost of frontier model inference remains high.

Organizations treating AI security as a subset of general cloud security should re-evaluate whether that framing captures the specific risks that AI services introduce. LLM inference billing is structurally different from other cloud resource billing in two important respects: the per-unit cost is high (frontier models charge materially more per request than storage or compute), and the abuse pattern is designed specifically to maximize cost without triggering conventional capacity-based anomaly detection. AI workloads should have dedicated security monitoring baselines, dedicated billing anomaly thresholds, and dedicated credential scope policies that reflect the financial risk profile of unauthorized access.

The MCP endpoint targeting observed in Operation Bizarre Bazaar’s adjacent campaign represents an emerging attack surface that most organizations have not yet incorporated into their threat models. As agentic AI deployments expand — bringing more MCP servers, more tool integrations, and more connections between AI agents and business-critical internal systems — the attack surface for LLM-adjacent credential and endpoint abuse will grow materially. Security architecture reviews for agentic AI deployments should explicitly model the scenario in which an MCP server endpoint is reached by an unauthorized external party, and should apply authentication, authorization, and network access controls commensurate with the sensitivity of the tools that server exposes.


CSA Resource Alignment

The MAESTRO agentic AI threat modeling framework provides the most directly applicable CSA analytical structure for understanding and defending against LLMjacking. MAESTRO’s seven-layer hierarchical model identifies Layer 3 (AI Model) as the inference access layer — covering the AI models themselves and their interfaces for receiving and responding to requests — and Layer 7 (Ecosystem) as the external integration and access control boundary governing how external systems, credentials, and APIs interact with the AI stack [5]. The credential theft and endpoint exploitation patterns in Operation Bizarre Bazaar both involve attackers achieving unauthorized access at the inference layer through failures in Ecosystem-layer access controls. Organizations applying MAESTRO-based threat models to their AI deployments should explicitly enumerate all surfaces through which their AI services can be reached, classify the authentication and authorization controls applied at each surface, and verify that long-lived credentials do not exist in contexts where they could be exfiltrated.

The CSA Cloud Controls Matrix (CCM) provides directly applicable control coverage across several domains relevant to LLMjacking. The Identity and Access Management (IAM) domain governs the provisioning, scoping, and lifecycle management of credentials used to access cloud AI services; CCM controls in this domain should be evaluated against the specific risk that long-lived, broadly scoped credentials represent in the LLMjacking threat model. The Threat and Vulnerability Management (TVM) domain addresses the detection and response requirements that LLMjacking demands: billing anomaly monitoring, API behavioral detection, and the patch and configuration management required to close the default-open authentication postures on self-hosted inference servers. The Logging and Monitoring (LOG) domain maps directly to the logging evasion technique documented in the original Sysdig research, in which attackers specifically probed for victim environments where inference logging was disabled before launching consumption attacks.

CSA’s AI Organizational Responsibilities: Core Security Responsibilities publication covers the organizational governance dimensions of credential security and AI service access control that LLMjacking exploits [6]. Its guidance on AI shared responsibility models — defining the security obligations of cloud AI service consumers versus providers — is particularly relevant here: the LLMjacking threat exploits the consumer-side responsibility for credential management and endpoint configuration, not vulnerabilities in the underlying AI services themselves. Organizations must understand that cloud AI providers secure the underlying infrastructure, but access control, credential hygiene, and endpoint exposure management are consumer responsibilities.

The CSA publication Securing LLM Backed Systems: Essential Authorization Practices provides implementation-level guidance on the authorization controls that can constrain LLMjacking exposure in applications built on cloud-hosted AI services [7]. Its recommendations for least-privilege credential scoping, external authorization checkpoints, and audit trail maintenance apply directly to the organizational posture improvements required to reduce LLMjacking attack surface. Its guidance on RAG system design and API integration security is also relevant to MCP server security architecture, as the authorization patterns that protect RAG pipelines from unauthorized access share structural characteristics with the controls required for MCP endpoint security.

CSA’s LLM Threats Taxonomy provides the definitional framework within which LLMjacking fits as a threat category [8]. The taxonomy’s Denial of Service and Model Theft categories are most directly applicable: LLMjacking constitutes both unauthorized consumption of model resources (a cost-impact form of service denial for the victim) and the commercial resale of model access (which, while not classical model theft, involves the unauthorized appropriation of model capability for third-party benefit). Organizations using the CSA LLM Threats Taxonomy as a risk assessment foundation should ensure that the credential theft pathway to model resource abuse is represented in their threat inventories alongside the direct attack vectors that the taxonomy’s primary threat categories address.

Zero Trust principles provide the architectural foundation for the most durable defenses against LLMjacking. The Zero Trust imperative — never trust, always verify — maps precisely to the credential and endpoint controls required: treating every API call to an AI service as requiring fresh authentication and authorization verification, regardless of whether the calling credential has been previously validated, and applying behavioral anomaly detection to all AI service usage rather than treating authenticated calls as inherently legitimate. Organizations that have applied Zero Trust architecture to their cloud infrastructure broadly should extend that architecture explicitly to AI service access, including AI-specific policies for credential verification, usage anomaly detection, and rate limiting enforcement.


References

[1] C. Morin, “LLMjacking: Stolen Cloud Credentials Used in New AI Attack,” Sysdig Blog, May 2024. https://sysdig.com/blog/llmjacking-stolen-cloud-credentials-used-in-new-ai-attack/

[2] C. Morin, “LLMjacking: From Emerging Threat to Black Market Reality,” Sysdig Blog, February 24, 2026. https://www.sysdig.com/blog/llmjacking-from-emerging-threat-to-black-market-reality

[3] Pillar Security Research, “Operation Bizarre Bazaar: First Attributed LLMjacking Campaign with Commercial Marketplace Monetization,” Pillar Security Blog, February 2026. https://www.pillar.security/blog/operation-bizarre-bazaar-first-attributed-llmjacking-campaign-with-commercial-marketplace-monetization

[4] NovaEdge Digital Labs, “LLMjacking: The $100K/Day AI Attack Draining Company Budgets,” NovaEdge Digital Labs Blog, 2026. https://www.novaedgedigitallabs.tech/Blog/llmjacking-100k-ai-attack-draining-budgets

[5] K. Huang et al., “MAESTRO: Agentic AI Threat Modeling Framework,” Cloud Security Alliance AI Safety Initiative, February 6, 2025. https://cloudsecurityalliance.org/blog/2025/02/06/agentic-ai-threat-modeling-framework-maestro

[6] J. Huang, K. Huang et al., “AI Organizational Responsibilities: Core Security Responsibilities,” Cloud Security Alliance, 2024. https://cloudsecurityalliance.org/artifacts/ai-organizational-responsibilities-core-security-responsibilities

[7] N. Lee, L. Voicu et al., “Securing LLM Backed Systems: Essential Authorization Practices,” Cloud Security Alliance, 2024. https://cloudsecurityalliance.org/artifacts/securing-llm-backed-systems-essential-authorization-practices

[8] S. Burke, M. Capotondi, D. Catteddu, K. Huang et al., “Large Language Model (LLM) Threats Taxonomy,” Cloud Security Alliance AI Controls Framework Working Group, 2024. https://cloudsecurityalliance.org/artifacts/csa-large-language-model-llm-threats-taxonomy

← Back to Research Index