ARTEX: Open-Source Agentic Pentest Tool Turned on Korean Banks

Authors: Cloud Security Alliance AI Safety Initiative
Published: 2026-10-10

Categories: Threat Intelligence
Download PDF

ARTEX: Open-Source Agentic Pentest Tool Turned on Korean Banks

Key Takeaways

CrowdStrike Intelligence reported on October 7, 2026 that a campaign against South Korean financial organizations, active from late September to early October 2026, combined ARTEX, a recently released open-source agentic penetration testing tool developed in China, with several large language models, and that the campaign resulted in data exfiltration [1][2]. The operator is unattributed. CrowdStrike assesses with moderate confidence that the operator is Chinese-speaking, and the evidence points to a single actor or small team motivated by financial gain rather than a state-directed program [1][2].

The campaign matters less for its technical novelty than for what it shows about the cost of entry. Earlier AI-orchestrated intrusions that CSA has documented involved a state-linked group using a frontier model or an LLM agent operating inside a compromised environment. In this case, a publicly available pentest framework apparently gave an individual operator machine-speed reconnaissance and attack-path planning against internet-facing financial applications [3][4]. Public reporting does not describe the initial access techniques in detail. CrowdStrike names a loan progress inquiry service and an employee mobile work-support system among the compromised systems [1][5], but the full range of techniques used is unknown, and defenders should not treat these examples as the limit of what the operator could do.

Three points deserve attention from security leaders. First, tools built for authorized testing and released openly can be used for intrusion with no modification, and although the developer later closed the source, a reported Korean-language fork suggests that the change may not have recalled copies already distributed [2][6]. Second, the operator’s infrastructure relied on LLM API resellers and several model providers, which creates both detection opportunities and attribution problems [1]. Third, officials in South Korea describe AI involvement as a possibility under investigation, not a completed attribution [7], even though some media coverage has described the attacks as AI-driven more definitively than the public evidence allows [9].

Background

ARTEX is described in public reporting as a multi-agent, LLM-driven autonomous penetration system developed by a developer using the handle Autumn-27. In the design reported by the press, a planner agent decomposes a target into tasks, and worker agents execute them with real tooling such as shell commands, HTTP requests, and port scans, while sharing logs so that one agent’s findings can inform another’s next step [8]. The Korea Times characterizes the tool as one that can scan systems for vulnerabilities and develop potential attack paths, and notes integration with models from several providers [9]. After the campaign became public, the developer reportedly took the project closed-source, stating that ARTEX “was originally designed for the purpose of learning and research” and that no further releases or maintenance would follow [2].

CrowdStrike identified the campaign through open directories on a Hong Kong-based IP address. The directories exposed Claude Code session histories, Claude memory files, and ARTEX configuration files [1][2]. A second server, at 38.244.50[.]120, hosted the ARTEX instance on port 18899 [1]. The configuration used DeepSeek v4.1-flash as its primary model backend and supplemented it with Z.ai’s GLM-5.3 and a Grok 4.6 model, with traffic routed through a domain, xcai[.]pro, that CrowdStrike identifies as a likely LLM API proxy or reseller [1][2]. CrowdStrike published indicators covering proxy IP addresses and the attack server, and mapped the activity to three ATT&CK techniques, all relating to infrastructure and tooling: T1583.003 (acquire VPS infrastructure), T1588.007 (obtain AI tooling), and T1090 (proxy) [1].

On the victim side, the reporting is less consistent than the infrastructure analysis. Press accounts name Shinhan Bank, KB Kookmin Bank, Hana Bank, BNK Busan Bank, and Yegaram Savings Bank among the affected institutions, and describe Woori Bank and NH NongHyup Bank as having detected attempts without confirmed data loss [2][9]. Shinhan reported roughly 25,000 affected customers, with exposed fields including names, phone numbers, annual income, and loan limits, while KB Kookmin and Hana reported 119 and 89 customers respectively [9]. Reporting that cites Yonhap and Seoul Economic Daily puts the number of affected firms at seven and the number of individuals in the secondary financial sector at 40,000, and notes that President Lee Jae Myung ordered a government probe on October 6 [7]. CrowdStrike itself states that the number of affected organizations is unconfirmed and gives no exfiltration volumes [1]. These figures come from different counting bases and should be read as press-reported, not verified totals.

The attribution picture is deliberately cautious. CrowdStrike links the activity to a suspected Chinese-speaking operator on the strength of the Chinese-developed tool and Chinese-language prompts, and notes that persona details found in the sessions, including a Telegram handle, cannot be definitively tied to the operator. The person using that handle has reportedly denied involvement [1][2]. The same sessions show the operator asking an AI assistant where threat actors typically sell Korean breach data, which supports a financial motive and may indicate an operator unfamiliar with that market [1][2]. South Korean officials have said that possible AI use in some of the incidents is under investigation and is distinct from an official conclusion [7].

Security Analysis

The analysis below separates what the public record establishes from what it merely suggests. CrowdStrike’s infrastructure findings are specific and largely corroborated, while the accounts of intrusion method, autonomy, and victim impact are thinner and come mostly from press reporting. Each subsection therefore states its evidentiary basis before drawing conclusions for defenders.

What the campaign does and does not show

It is tempting to read ARTEX as the next step after the GTG-1002 campaign, in which an AI-orchestrated operation reportedly executed 80 to 90 percent of tactical operations autonomously against roughly thirty organizations [3][14]. The two events differ in important ways. GTG-1002 involved a state-sponsored group using a frontier model’s agentic features as the orchestration layer. The ARTEX operator used a packaged, purpose-built offensive framework with commodity and low-cost model backends, and the operator’s own Claude Code sessions appear to have supplemented manual work [1][3]. The evidence in public reporting does not establish what share of the operation ARTEX performed autonomously, and CrowdStrike’s published analysis gives no step-by-step account of how access was obtained [1]. Claims that AI “planned and carried out” the attacks, including one academic comment that humans would need weeks or months for comparable work, are plausible but remain expert opinion, not forensic findings [9].

What the campaign does show is a lowered barrier. The CSA report on LLM agents as offensive post-exploitation tools describes the same kill-chain stages (credential harvesting, lateral movement, exfiltration) being compressed by LLM agents operating at machine speed [4]. ARTEX packages comparable planning and execution into a deployable framework that a single operator can run against many targets in parallel. For defenders, the practical consequence is that the time between the exposure of a weakly protected service and its exploitation may shrink, because automated tools reduce the marginal effort of testing each internet-facing asset.

Exposed internal tools

The most specific detail on the compromised systems comes from CrowdStrike’s own summary, which names a loan progress inquiry service at one bank and an employee mobile work-support system at another [1]. Trade reporting describes the same two systems [5]. Neither source establishes how access was obtained, and this note does not assume a particular authentication weakness. The examples nonetheless point to a class of exposure that defenders commonly encounter: partner- or employee-facing applications built for convenience, sitting outside the main authentication architecture, and receiving less testing than customer-facing channels. That pattern reflects the authors’ general experience and has not been shown for these specific applications. Agentic tools can automate enumeration and testing of such systems at scale, which makes them a natural fit for finding exposures of this kind [4].

Infrastructure patterns that defenders can use

The operator’s infrastructure choices create the most concrete detection opportunities. The campaign used multiple model providers through an API reseller, and the attack server was hosting a framework that continuously calls out to external LLM endpoints [1][2]. Egress traffic from a network segment that suddenly generates sustained, structured calls to unfamiliar LLM API endpoints or proxy services is an anomaly worth alerting on, particularly from servers that have no legitimate reason to call such services. The inverse also holds on the victim side: the traffic the tool generates against a target, such as rapid systematic enumeration, uniform request pacing, and probing sequences that adapt to responses, may differ from both conventional scanners and human testers. Published reporting does not describe signatures for these patterns, so organizations should treat them as hypotheses to validate against their own telemetry.

The reseller layer also complicates the response. CSA has documented how stolen or resold AI compute supports offensive tooling, and a reseller front end hides which accounts and models an operator actually uses [10]. Model providers can sometimes see abuse that victims cannot, which makes provider abuse-reporting channels and intelligence sharing with financial-sector ISACs more important than they were when offensive tooling ran entirely on the attacker’s own hardware.

Dual-use tooling and the limits of closing the source

The developer’s decision to stop releasing and closing the repository reflects a pattern that other offensive security projects have faced, but it has limited protective value once copies exist. A Korean-language fork of ARTEX under an AGPL-3.0 license reportedly appeared on GitHub on October 8, after the incidents became public [6]. This report comes from a secondary source and the fork’s provenance and functionality are not verified here. It illustrates the broader point regardless: a localized and maintained version of an agentic offensive tool aimed at a specific language market reduces friction for exactly the operators who are most likely to target that market. CSA’s earlier work on AI for offensive security describes legitimate uses of agents across reconnaissance, scanning, vulnerability analysis, exploitation, and reporting, and the same capabilities apply to authorized and unauthorized users [11]. Controls therefore need to assume that capable offensive agents are available to attackers and focus on the exposure those agents find.

Open questions

Several questions will determine how much of this note’s analysis holds up as the investigation proceeds. The Korean government’s investigation may reveal the actual intrusion paths, the degree of human involvement, and whether all seven firms were compromised by the same actor, since press reports suggest that more than one cause may be at play and officials have said that AI involvement is only suspected in some cases [7]. The affected-customer counts have not been reconciled. The operator’s identity is unknown, and the persona evidence is thin. Readers should treat any claim of state sponsorship, or of full autonomy, as unsupported by the public record as of the date of this note.

Recommendations

The recommendations follow the same sequence as the analysis: first closing the exposures that agentic tools are best at finding, then building detection around the operator’s infrastructure patterns, and finally adjusting longer-term planning assumptions. Because the public record on intrusion method is limited, the guidance favors broad hygiene over controls tuned to a single reported technique.

Immediate Actions

Organizations should inventory internet-facing applications that sit outside the primary customer authentication path, with particular attention to partner lookup tools, employee mobile applications, and internal portals that were exposed for operational convenience. For each, confirm that authentication is enforced server-side, that multi-factor authentication is required for anything that returns customer data, and that rate limiting and lockout controls cannot be bypassed by changing request parameters or headers. Teams should also add the indicators published by CrowdStrike to blocking and retrospective search, recognizing that the indicators are infrastructure addresses that the operator can replace easily [1]. Finally, security operations teams should review authentication and application logs for the period from late September to the present for systematic, evenly paced enumeration against the assets identified above.

Short-Term Mitigations

Over the next quarter, organizations should extend egress monitoring to flag outbound calls from servers and workstations to LLM API endpoints and known proxy or reseller domains that are not on an approved list, as a signal of both internal misuse and compromised hosts running agentic tooling. Penetration testing programs should add agentic tools to their own assessments, so that internal teams learn what an agent finds on their perimeter before an adversary does, and should do so under written scope and with the logging needed to separate test activity from attack activity. Data-minimization reviews are also warranted for the categories reported as exposed, including income and loan-limit fields, since reducing what a single lookup service returns limits what any compromise can expose. Where an organization has a relationship with an AI model provider, it should establish a reporting channel for suspected abuse.

Strategic Considerations

At a strategic level, the campaign supports planning on the assumption that the cost of discovering weaknesses will keep falling, and that remediation speed may become a more meaningful metric. Programs that measure exposure by the time between a service becoming reachable and its first adversarial test should expect that interval to shorten. Financial institutions should consider continuous automated testing of their external attack surface, tighter lifecycle governance for internal tools exposed to partners, and participation in sector information sharing so that infrastructure indicators and model-provider abuse patterns circulate quickly. Security leaders should also take part in any discussion of norms for the open release of dual-use offensive agents, with realistic expectations about what source-code withdrawal can achieve.

CSA Resource Alignment

CSA’s research note on LLM agents as active post-exploitation tools, along with the related white paper “LLM Agents as Offensive Post-Exploitation Tools,” most directly addresses the kill-chain compression this campaign illustrates, and its defensive guidance on credential handling, lateral movement, and anomaly detection applies to the exposure patterns described above [12][4]. CSA’s analysis of the first confirmed LLM-orchestrated attack chain places the incident in context with GTG-1002, which is the main reference point for earlier AI-orchestrated intrusions and shows how this campaign differs in actor profile and tooling [3][14]. CSA’s note on LLMjacking and stolen AI compute as attack infrastructure is relevant to the reseller and proxy layer that CrowdStrike observed, and to the question of how model providers and victims can detect abuse that neither can fully see on their own [10]. CSA’s earlier research report, Using AI for Offensive Security, covers the legitimate agentic testing capabilities that ARTEX-style tools share, and it frames the governance considerations organizations should apply when adopting such tools internally [11]. For the financial-sector context, The State of Cloud and AI for Financial Services 2026 examines how the sector is approaching cloud and AI adoption, including governance of autonomous agents [15]. As a standing framework, the AI Controls Matrix (AICM) v1.1 supports mapping the recommendations above to its threat and vulnerability management and identity and access management domains [13]. Organizations should treat these artifacts as complementary and should consult the full documents for details.

References

[1] CrowdStrike. “Unknown Threat Actor Uses AI-Driven ARTEX to Target South Korean Finance.” CrowdStrike Blog, October 7, 2026.

[2] The Hacker News. “ARTEX AI Pentesting Tool Used in Data Theft Attacks on South Korean Financial Firms.” The Hacker News, October 2026.

[3] Cloud Security Alliance. “LLM-Orchestrated Attack Chain Research Note.” CSA Labs, May 27, 2026.

[4] Cloud Security Alliance. “LLM Agents as Offensive Post-Exploitation Tools.” CSA Labs, 2026.

[5] The Paypers. “ARTEX agentic tool used in attacks on Korean banks.” The Paypers, October 2026.

[6] Hermes AI. “ARTEX Korean fork surfaces after AI pentesting tool hit seven banks.” Hermes AI News, October 8, 2026.

[7] Let’s Data Science. “South Korea Orders Probe Into Bank Data Breaches.” Let’s Data Science, October 2026 (citing Yonhap and Seoul Economic Daily).

[8] Security Point Break. “Beyond Shinhan: South Korea’s Bank Hacks Hit 66,000 With an AI Clue.” Security Point Break, October 7, 2026.

[9] The Korea Times. “EXPLAINER: How AI emerged as new threat in Korea’s bank hacking crisis.” The Korea Times, October 7, 2026.

[10] Cloud Security Alliance. “LLMjacking Evolves: Stolen AI Compute as Attack Infrastructure.” CSA Labs, June 18, 2026.

[11] Cloud Security Alliance. “Using AI for Offensive Security.” CSA, August 2024.

[12] Cloud Security Alliance. “LLM Agents as Active Post-Exploitation Tools.” CSA Labs, June 2, 2026.

[13] Cloud Security Alliance. “AI Controls Matrix (AICM) v1.1.” CSA, 2026.

[14] Cloud Security Alliance. “Agentic Post-Exploitation: First Documented Autonomous Attack Chain.” CSA, 2026.

[15] Cloud Security Alliance. “The State of Cloud and AI for Financial Services 2026.” CSA, 2026.

← Back to Research Index