Unit 42 is putting the latest frontier cyber models to work across customer environments to find, validate and help remediate the attack paths that matter most.
In May, we introduced Frontier AI Defense with a warning: the window to get ahead of AI-enabled attacks was shorter than most people realized. Since then, we have briefed more than 1,000 security teams around the world and introduced our Frontier AI Defense service to hundreds of customers.
Today, through our partnership with OpenAI, we are expanding Unit 42 Frontier AI Exposure Analysis to put advanced frontier cyber models directly to work in customer environments. Under Unit 42 direction, these models can find exposures, test whether they are exploitable, validate attack paths and help customers prioritize what to fix first.
Our early work shows why this approach matters: 36% of the exposures we identified map to no known CVE, often because they involve multiple gaps that have to be discovered, chained and tested together.
Bringing the Latest Frontier Cyber Capabilities to Defenders
Palo Alto Networks has been among a limited group of organizations with early access to advanced cyber capabilities from the leading frontier AI labs. Through our partnership with OpenAI, Unit 42 can now bring its latest advanced cyber capabilities, including GPT-5.6 Daybreak, to security testing and validation for our customers. Until now, GPT-5.6 Daybreak has not been available for commercial use.
Frontier models have helped inform the work of our experts. Now they can increasingly perform complex offensive security tasks directly, at machine speed and under Unit 42 direction. That allows us to go deeper than traditional vulnerability discovery by testing exploitability, reasoning across multiple weaknesses and determining how an attacker could use them to achieve an objective.
There is no single best model for every cyber task. Our research has shown that different models have different strengths and find vulnerabilities others miss. A multi-model harness routes work to the model best suited for the task, improving efficacy and coverage while managing the cost of frontier AI at scale. As stronger models emerge, we can incorporate them without rebuilding the offering around a single model or provider.
Unit 42 experts remain central to the process. We combine frontier models with our offensive security expertise, Palo Alto Networks telemetry and Unit 42 Threat Intelligence to validate findings, connect exposures into attack paths and understand what an attacker could ultimately achieve.
Built to Find What Attackers Can Exploit
The expanded service brings five capabilities together:
Leading cyber models: Apply the latest advanced cyber models to improve exposure discovery, testing and validation.
Multi-model harness: Use the right model for the right task to improve efficacy, expand coverage and optimize cost.
Exposure discovery: Find vulnerabilities, misconfigurations, leaked credentials, unmanaged attack surface and other posture gaps across applications and network assets.
Advanced adversary simulation: Actively test exploitability and validate end-to-end attack paths to understand how an attacker could compromise the environment.
Custom remediation plan: Prioritize the fixes that break the most important attack paths and feed those findings into existing IT, development and security workflows.
Most security teams already have more findings than they can act on. The harder problem is knowing which ones create a real path to compromise. Attackers look across applications, infrastructure, identity and cloud for weaknesses they can combine to achieve an objective. Frontier AI Exposure Analysis applies that same adversarial perspective, helping defenders understand which paths matter and what to fix first.
The Asymmetry Runs Both Ways Now
For the past several months, frontier AI has been a story about what is coming for defenders: vulnerability discovery at scale, exploit chaining that sees full-stack logic no scanner catches, and attack cycles compressed to seconds from initial access to exfiltration.
All of that is still true. Our answer has been to put everything we learn testing these frontier models into the hands of defenders. Today, that gets more direct: not just what frontier models have taught us, but the models themselves, working in your environment for your defenders before those same capabilities are working for the attacker.
The window is still closing. We intend to spend it building on the side of the defenders.
Cybercrime at Machine Speed: Key Takeaways from Flashpoint’s 2026 Midyear Threat Intelligence Briefing
Threat actors are no longer just using automation to execute tasks, they are leveraging prepackaged, safeguard-free AI, weaponizing stolen session data, and directly targeting defenders’ security stacks.
The threat landscape has developed at a striking pace with Flashpoint tracking over 22 million illicit AI discussions, 7.4 million compromised hosts yielding 1.7 billion stolen credentials, and over 21,600 disclosed vulnerabilities in just six months. Beyond these staggering numbers, the on-demand session detailed something even more alarming: a fundamental shift in adversary operational tradecraft.
Here are the five critical shifts every cyber threat intelligence (CTI), Vulnerability Management, and SOC team needs to know.
The Death of Signal: Threat Actors are Shifting to “Private AI”
The public discussion surrounding criminal artificial intelligence (AI) has reached a critical inflection point. Early in the AI boom, Flashpoint observed threat actors collaboratively experiment across underground forums, jailbreaking commercial frontier models or advertising surface-level tools like WormGPT and DarkGPT.
Today, adversaries are shifting from public forums to running fine-tuned, open-source models locally on private servers, which greatly hampers traditional signature-based detection. Flashpoint analysts are now seeing attackers generate unique, highly tailored malware variants, flawless phishing lures, and custom exploit scripts at extremely low costs—completely offline and shielded from public monitoring.
“A few months ago, a lot of this was collaborative… public outsourcing. Now what we’re seeing is scarier: pre-packaged cybercrime models run locally on private infrastructure. Malicious code, exploit scripts, and targeted phishing are all being generated inside closed environments.”
Ian Gray, VP of Intelligence, Flashpoint
Weaponizing the Defender’s Own Tooling
Another eye-opening tactical insight shared during the session was how threat actors are repurposing defender infrastructure for automated initial access and extortion. In the webinar, we pointed to recent campaigns where adversaries specifically targeted misconfigurations and zero-day vulnerabilities inside open-source vulnerability scanners, secrets-detection tools, Kubernetes clusters, and Infrastructure-as-Code(IaC) environments.
What this means for defenders is that the attack surface is no longer bounded by traditional enterprise network boundaries: it extends directly into CI/CD pipelines, security orchestration tooling, and third-party SaaS integrations. Security teams are finding themselves in a race against attackers who use automated scanning scripts to weaponize vulnerabilities in the security tools themselves.
The Global Infostealer Threat and Identity-First Attacks
Flashpoint tracked 7.4 million hosts compromised by infostealers in H1 2026—a 27% increase period-over-period—harvesting 1.7 billion credentials and identity information.
While the top infostealer strains remain familiar, law enforcement operations have created vacuums that competitors rapidly fill.
Threat actors are leveraging drive-by downloads, watering holes, and pirated software packages to plant stealers. Once a machine is compromised, the logs capture corporate SSO credentials, active browser cookies, VPN keys, and SaaS session tokens. This enables adversaries to simply log in without having to leverage complex technical exploits.
The Structural Failure of CVE/NVD and the Importance of KEV
The Common Vulnerabilities and Exposures (CVE) and National Vulnerability Database (NVD) have failed to keep pace with the velocity of AI-assisted vulnerability discovery. As such, vulnerability management teams are facing significant operational delays.
Metric
Flashpoint GTIR Midyear H1 2026 Data
Operational Impact
Total Disclosures
21,667
Remediation volume exceeds defender bandwidth.
Exploit Availability
19% (4,015 CVEs)
Functional code is ready before patches are deployed.
Public Catalog Lag
Growing Backlog (NVD/KEV)
Delay in official scoring leaves teams blind to active risk.
Therefore, waiting for NVD enrichment before prioritizing a patch is a dangerous strategy. To compensate, security teams require Vulnerability Intelligence (VI) that provides primary-source confirmation of weaponization, exploit availability, and actionable mitigation guidance long before public databases update.
Ransomware Evolution: From Encryption to Cloud Extortion
Ransomware-as-a-Service (RaaS) activity surged by 45% period-over-period, reaching 6,256 verified victim postings on data leak sites. However, total on-chain payout revenue dropped by 8% to $820 million, with victim pay-rates hitting a record low of 28%.
Faced with declining payouts and resilient enterprise backups, extortion syndicates are adapting. Rather than relying exclusively on technical file-encrypting malware, groups are executing pure data extortion campaigns—frequently targeting cloud platforms or extracting data through third-party vendor access.
Protect Your Organization Using Flashpoint
Defending against machine-speed attacks requires moving beyond reactive, post-incident telemetry. Flashpoint arms security, CTI, and vulnerability management teams with the primary-source intelligence required to preempt adversary operations:
Unrivaled Deep & Dark Web Visibility: Flashpoint’s Primary Source Collection actively monitors closed criminal communities, illicit Telegram channels, and private forums, giving you early warning when threat actors build custom AI toolkits or trade credentials targeting your organization.
Comprehensive Vulnerability Intelligence (VI): Flashpoint tracks zero-days and vulnerability disclosures independently, delivering immediate exploit availability data and threat-informed prioritization so you patch what actually matters.
Continuous Compromised Credential Monitoring: Instantly surface exposed enterprise credentials, active session tokens, and stealer logs tied to your domain or third-party supply chain before they lead to an account takeover (ATO).
The Evolution of Hacktivism in Hybrid Warfare: Modern Tactics and Real-World Impact
In this post we examine how modern hacktivism has evolved into a tool of global hybrid warfare, analyzing crowdsourced attack tactics, media-driven propaganda, and real-world impacts across Ukraine, the Middle East, European Union, and NATO nations.
Hacktivism used to be perceived as digital graffiti, with lone-wolf threat actors defacing government websites or temporarily crashing banking portals to make a political point. However, Flashpoint is tracking a fundamental shift in how these groups operate.
Modern hacktivism is evolving into a disciplined component of global hybrid warfare, capable of bridging digital disruptions with tangible real-world impact. Today, these operations blur the line between volunteer activism and coordinated state interest, leveraging crowdsourced infrastructure to disrupt critical utilities, manipulate media narratives, and target public infrastructure on a global scale. Unpacking these modern hacktivist collectives reveals what their tactics look like in practice and their far-reaching consequences across dozens of nations.
What is Hacktivism?
Hacktivism is the use of cyberattacks to promote or advance a particular political or social cause, leveraging a wide range of tactics such as website defacement, distributed denial-of-service (DDoS) attacks, and data breaches. Modern hacktivist collectives serve as the loud, high-visibility arm of cyber conflict—frequently aligning with state geopolitical interests, as seen most prominently in recent pro-Russian operations and Iranian-aligned cyber campaigns.
These pro-Russian hacktivist groups, such as NoName057 and Killnet, alongside pro-Iranian collectives and proxy ecosystems like Handala Hack, often react to the news cycle and target countries designated by state media or ideological narratives as enemies. As such, modern hacktivist campaigns are opportunistic and tied to global events—from the escalation in the Middle East following military operations like Operation Epic Fury, to the Milan-Cortina Winter Olympics and new aid packages to Ukraine. These groups’ justification narratives typically mirror state messaging.
Modern Tactics: Gamifying Cyber Warfare
In tracking modern hacktivist groups, Flashpoint analysts identified a new method these groups are utilizing to convert ordinary devices into tools for hybrid warfare—the gamification of cyberattacks. Flashpoint has observed groups like NoName057 turning DDoS attacks into community-based “patriotic online games,” such as their “DDoSia Project,” with participants earning military-style ranks and cryptocurrency rewards for overloading the websites of government institutions, banks, and various infrastructure across various countries.
This model has enabled the scaling of operations by utilizing a large, low-skilled participant base rather than having to rely on sophisticated technical tradecraft. The model’s decentralized structure and ideological appeal continue to pose a significant challenge for international law enforcement.
The Propaganda Engine: Media Amplification and Validation
Beyond technical disruptions, publicity is the primary currency of modern hacktivism. Hacktivist groups demonstrate a consistent pattern of media-seeking behavior and self-promotion, likely intended to amplify their perceived impact and reinforce notoriety within the broader cyber threat landscape. Many of these groups repeatedly repost media coverage and news articles referencing themselves.
This serves as a curated self-promotion mechanism, allowing the group to selectively showcase external validation of its operations, including coverage from mainstream and security-focused outlets, to its followers. This behavior aligns with a broader trend observed with especially pro-Russian hacktivist collectives, in which media visibility is treated as a measure of operational success independent of verified technical impact. It also serves as a deliberate tactic for engagement and recruitment that reinforces “patriotic” branding and sustains participant morale and visibility.
Beyond Propaganda: Aligning Cyber Disruption with Military Objectives
In some cases, the digital targeting of hacktivist collectives is more aligned with kinetic objectives, rather than public perception or propaganda initiatives. This is especially true for Iranian-aligned hacktivists and proxy groups who are more deeply intertwined with military operations in the Middle East. These groups have expanded their operations from website disruptions into claims of large-scale data wipers, extortion, and cyberattacks targeting key infrastructure across the Gulf.
The Far Reach of Modern Hacktivism
Major geopolitical flashpoints in the Middle East have triggered waves of hacktivist activity that has spread across North America, with threat actors targeting supply chains, financial infrastructure, and operational technology and control systems.
Simultaneously, pro-Russian hacktivist groups, particularly NoName057, have been extremely prolific within the last year—carrying out two major illicit campaigns heavily targeting Ukraine, which then spilled over to more than 30 nations globally. The following breakdown contains statistics and targeting dynamics of pro-Russian hacktivist groups observed between July 2025 and 2026:
Country-level targeting derived from Flashpoint intelligence. (Source: Flashpoint, graphic generated by Claude)
The Continuous Campaign Against Ukraine
Ukraine has been the primary target for pro-Russian hacktivist groups who seek to damage Ukrainian infrastructure and morale. Anti-Ukrainian content is constantly distributed through dedicated per-language channels, making it the most linguistically developed target spanning six languages. Involved channels each post near-identical translated content within minutes to hours of the Russian original, down to the same image file with identical SHA1 hashes, which suggests a sustained propaganda distribution operation.
This has resulted in alleged data breaches impacting Ukrainian General Staff, military enlistment offices, medical, and morgue databases to push a casualty-count narrative. It also has resulted in the defacement or disruption of websites of regional capitals and administrative centers, energy plants, water and power-adjacent infrastructure, and many more.
Spilling Over: Impact Across EU and NATO Allies
However, Ukraine is not the sole casualty of modern hacktivism. Recent pro-Russian hacktivist campaigns have spread to other EU nations and NATO members. Germany, the United Kingdom, and Spain have been observed to be priority targets, with threat actors targeting public transportation, federal and security agencies, municipal government and utilities, financial markets, and other infrastructure. In some cases, hacktivist campaigns manifest in the real-world, with physical sticker drives on municipal streets, alongside doxxing operations releasing alleged personal data and automated scans hijacking exposed CCTV camera systems across Europe.
Physical sticker campaigns (NoName057) in Spain identified by Flashpoint
Defend Against the New Wave of Hacktivism Using Flashpoint
As hacktivist operations continue to blur the boundary between digital disruption and real-world interference, organizations can no longer view DDoS attacks or low-level intrusions as simple background noise. Protecting critical assets requires proactive visibility into threat actor networks, early detection of targeting narratives, and primary source threat intelligence.
Request a demo today to see how Flashpoint provides actionable intelligence to help security teams, government agencies, and infrastructure providers identify, monitor, and mitigate emerging hacktivist campaigns before they impact operations.
Unknowingly, a member of key personnel is living two separate lives. On the clock, they are a highly-trusted systems administrator, but in their personal time, they moonlight on the deep and dark web, advertising their trust and access to the highest bidder. One day, they get a simple offer: $15,000 in the crypto of their choice to approve a single push notification at 2 AM. They accept. By morning, the attacker walks away with active domain admin credentials without the need for malware or cracking firewalls.
This is just one example of how insider threats lead to modern enterprise breaches. This year, Flashpoint uncovered 7,282 unique insider threat posts, with an average of 34 unique posts being posted daily. As perimeter security, EDR coverage, and other security tools mature, threat actors are finding it faster—and cheaper—to target the human element and simply buy an insider’s credentials or pay an employee to open the front door.
In a threat landscape where identity is becoming the primary attack surface, monitoring illicit marketplaces and recruitment efforts is critical. This new monthly report leverages Flashpoint’s Primary Source Collection (PSC) to analyze insider threat tactics, tracking active recruitment and advertising on dark web forums and encrypted networks.
The Insider Threat Landscape: July 2026
In July 2026, Flashpoint analysts identified a total of 12,653 insider posts. These communications include both threat actors attempting to recruit insiders in target organizations, as well as insiders advertising their services on illicit forums and marketplaces.
Of these total communications, Flashpoint observed 1,132 unique posts in July 2026.
Where Insider Threat Activity is Concentrated
Historically, the Telecommunications, Retail, and Financial industries are most adversely affected by insider threat activity. However, July 2026 findings noticeably deviate from this trend. Flashpoint found 58.6% of total insider threat posts affected “Other” industries—suggesting adversaries are diversifying their target base. Threat actors may be attempting to recruit within supply chain partners, logistic hubs, manufacturing platforms, and specialized service providers to find alternative entry points into target networks.
The following table shows a breakdown of unique insider posts by industry in July 2026:
Industry
Posts
Other
663
Financial
150
Retail
112
Technology
84
Telecom
74
Public Sector
43
Healthcare
3
Media
3
Total
1,132
Insider Threats: Recruiting vs. Advertising
Active insider threats work in two ways: an insider is “recruited” by a malicious outside party, or a malicious insider “advertises” their access and skills to an interested threat actor. Regardless, by leveraging this connection, insiders assist adversaries by exfiltrating valuable data, installing malware, sabotaging IT systems, or performing SIM swaps.
In July 2026, Flashpoint found that over 75% of unique threat actor posts came from insiders advertising their access to malicious third parties. This indicates a highly motivated internal threat landscape where disgruntled employees actively seek out buyers for corporate data and network entry points.
Protect Against Insider Threats Using Flashpoint
Insider threats are inherently difficult to detect using internal security controls alone because the malicious activity relies on valid credentials and legitimate access privileges. Relying solely on internal logs means security teams often only detect an insider threat after data exfiltration or system sabotage has already occurred.
Flashpoint protects organizations against insider threats through our Primary Source Collection (PSC) and specialized intelligence platforms:
External Threat Intelligence & Early Warning: Flashpoint monitors deep and dark web forums, invite-only threat communities, and encrypted chat platforms to identify employee solicitations, stolen corporate domain mentions, and active recruitment attempts before an intrusion develops.
Identity Protection & Infostealer Tracking: By tracking illicit marketplaces and infostealer activity, Flashpoint identifies compromised corporate credentials and active session tokens, preventing threat actors from utilizing purchased access.
User & Entity Behavior Context: Flashpoint’s intelligence equips SOC, Security Operations, and Risk Management teams with adversary TTPs, enabling security operations to look for anomalous data downloads, off-hours access, or unauthorized software installation.
To learn more about how Flashpoint can help protect your enterprise from insider risk and monitor illicit underground communities,Request a Demo Today.
Frequently Asked Questions (FAQs)
What is the Flashpoint Insider Threat Report?
The Flashpoint Insider Threat Report is a monthly intelligence brief that analyzes trends, volume, targeted industries, and tactics surrounding insider threat recruitment and illicit access advertising on the deep web, dark web, and encrypted chat channels.
How does Flashpoint collect insider threat data?
Flashpoint collects data using its Primary Source Collection (PSC) engine, which actively monitors thousands of dark web forums, illicit marketplaces, and underground chat networks where threat actors and malicious insiders communicate.
What is the difference between insider recruitment and insider advertising?
Insider recruitment occurs when an external cybercriminal attempts to entice a corporate employee into assisting with a cyberattack. Insider advertising occurs when an employee or contractor proactively lists their legitimate access or services for sale on illicit marketplaces.
In the first half of 2026, the global threat landscape reached a clear operational inflection point: threat operations have fundamentally transitioned from human-led campaigns to machine-speed, AI-driven exploitation. As threat actors gain commoditized access to open-source AI technologies and actively deploy automated, safeguard-free tooling locally on private infrastructure, organizations face an accelerating hybrid risk environment.
Flashpoint’s Global Threat Intelligence Report: 2026 Midyear Edition
The Flashpoint Global Threat Intelligence Report: 2026 Midyear Edition anchors security leaders—from threat intelligence, vulnerability management, to executive leadership—in the data required to navigate this evolving threat landscape. Covering the period from January 1 to June 30, 2026, the report delivers timely insights backed by Flashpoint’s proprietary primary-source collection from over 3.9 petabytes of continuously monitored illicit sources.
Our midyear findings reveal several key metrics that highlight the speed and scale of the H1 2026 threat landscape:
22M+ threat actor posts discussed, shared, or advertised artificial intelligence toolkits for criminal deployment.
1.7B credentials and identity data points extracted across more than 7.4M unique compromised hosts globally.
Nearly one-in-five (19%) of all vulnerability disclosures dropped with ready-made, functional exploit code.
45% period-over-period surge in Ransomware-as-a-Service (RaaS), with total victim volume reaching 6,256 even as victim payout rates dropped to a historic low of 28%.
A Clear Understanding of the Convergence Between AI and Cyber Threats From generating flawless phishing campaigns to automating vulnerability scanning and code obfuscation, discover how adversaries are optimizing for speed and cost-efficiency — utilizing AI as a force multiplier in their various illicit campaigns.
A Comprehensive Top-Down View of the Evolving Threat Landscape Gain full visibility of the threat landscape with Flashpoint’s primary-source collections and real-time threat intelligence.
Strategies for Proactive Defense and Risk Mitigation Move your organization beyond reactive incident response by leveraging Flashpoint’s comprehensive threat intelligence. Gain the foresight needed to strengthen defenses and optimize your security posture.
“AI is compressing the time between opportunity and exploitation. Capabilities that once took significant expertise, coordination, and time to develop are becoming faster to build, easier to scale, and harder to detect. Security teams are facing an adversary ecosystem that can use AI to iterate at unprecedented speed — the only way to keep pace is with primary-source intelligence that surfaces adversary behavior before attacks unfold.”
Josh Lefkowitz, Flashpoint Co-Founder & CEO
The Four Driving Themes Shaping the 2026 Threat Landscape
Artificial Intelligence (AI) Threats
During the first half of 2026, Flashpoint captured over 22M illicit posts discussing or advertising AI for criminal-related activities. By stripping ethical safeguards, custom malicious LLMs allow unsophisticated threat actors to automate complex phases of the attack lifecycle, including target profiling, malware evasion script creation, and zero-day exploit generation.
Information-Stealing Malware Threats
Infostealer malware harvested 1.7 billion credentials across 7.4 million compromised systems in H1 2026 alone, turning digital identity into the main entry point for enterprise intrusions.
Vulnerability Intelligence and Patching Management
19% (4,015) of all H1 2026 vulnerability disclosures arrived with ready-made exploit code. Adversaries deploy automated replication scripts almost immediately upon disclosure, eliminating manual remediation windows.
Ransomware Operations, Multi-Extortion Cartels, and Financial Risk
Despite a 45% surge in victim volume (6,256 overall), total on-chain revenue fell by 8% to $820M. Improved enterprise backups and incident response have driven payout rates down to 28%, prompting syndicates to demand larger sums from paying victims.
Proactive Security in 2026 and Beyond
The data shows that traditional enterprise security organizations are struggling to keep pace with modern threat cycles that are accelerated by illicit uses of AI. This continued convergence of AI engines and initial access vectors have further compressed attack timelines, making it nearly impossible for security teams to defend against them—especially if they are limited by traditional approaches to threat intelligence.
Equipping your team with primary-source threat intelligence is critical for protecting critical assets in 2026. Download the Flashpoint Global Threat Intelligence Report: 2026 Midyear Edition to gain the visibility and strategic clarity required to defend your organization.
Data Center Physical Security: Mitigating FPV Drone Threats
In this post, we explore how shifting online sentiment and low-cost First-Person View (FPV) technology are creating an unprecedented airborne threat vector for critical data center infrastructure.
Data centers have become a driving force in the modern digital economy—powering cloud services, global enterprise operations, and the explosive growth of artificial intelligence (AI). However, due to growing negative public discourse, data centers are facing a new physical threat vector: low-cost, payload-capable drones.
According to Flashpoint research, shifting public sentiment surrounding AI development, combined with the extreme accessibility of First-Person View (FPV) drone technology, is creating an unprecedented hybrid threat to physical critical infrastructure.
Here is what you need to know about this emerging threat landscape and what it means for physical security teams protecting critical assets.
Growing Online Sentiment and Anti-AI Hostility
Organizations tasked with protecting critical data infrastructure need to understand that this growing threat is not developing in a vacuum. Across both clearnet and Deep and Dark Web (DDW) forums, online discussions regarding data center expansion have intensified, with a significant portion bordering on hostility. Key drivers of negative sentiment include:
Environmental & Local Concerns: Debates over massive energy consumption, water usage, noise, and localized quality-of-life impacts.
Backlash against AI: Discontent directed at tech companies driving the rapid deployment of AI infrastructure.
Perceived Regulatory Inaction: Frustration among activists who feel local and state governments are failing to halt or regulate new construction.
While much of the current online chatter currently revolves around organized protests and aspirational threats, Flashpoint analysts note a troubling uptick in rhetoric targeting corporate tech executives and data center infrastructure.
The Evolving Data Center Threat Landscape
Data centers across the United States are seeing a rapid increase in physical and operational threats. Vandalism and property destruction have become common topics in illicit online spaces when discussing data centers and their impact on everyday life. Flashpoint research highlights two primary force multipliers driving this threat:
DIY Drones & Low Barriers to Entry
Historically, kinetic airborne strikes required specialized equipment and advanced training. Today, that barrier to entry has virtually collapsed. Rapid improvements in drone manufacturing have made payload-capable aircraft extraordinarily accessible. In today’s market, an individual can purchase an off-the-shelf system or assemble a customized drone for under $1,000 USD.
Inspiration for these tactics is also readily available; widespread footage of FPV drones operating in conflict zones like Ukraine has demonstrated to online audiences how easily and effectively low-cost aircrafts can be weaponized. Threat actors view this as a high-yield investment, especially given the capability to deploy multiple drones in quick succession.
Protests as Cover for Physical Operations
Organized protests to stop data center development remain prevalent, and large crowds can easily overwhelm contracted security personnel, diminishing the effectiveness of a response to an aerial threat. A malicious actor could use a protest at a data center as cover to cause physical damage to the facility while security resources are spread thin. For example, on July 19, 2026, activists threw balloons filled with acetic acid at a data center construction site in Amsterdam. In its aftermath, Flashpoint analysts captured individuals online discussing the use of drones to deliver similar payloads.
Regulatory and Defense Measure Challenges
Current federal regulations limit the ability to effectively deter or stop an incoming drone threat because the US Federal Aviation Administration (FAA) classifies drones as aircraft. Therefore, organizations specializing in the physical security of data centers will likely need to increase their operational capabilities and advise companies on potential hardening to deter attacks.
Traditional foot patrols and monitoring perimeter access control points will be insufficient in mitigating overhead threats. The majority of data centers are currently not equipped with the specialized Counter-Unmanned Aircraft Systems (C-UAS) equipment, specialized training, or legal authorization needed to respond effectively to airborne incursions.
Protect Critical Infrastructure Using Flashpoint
Defending against aerial incursions requires moving from reactive security to proactive, intelligence-led physical protection. Physical security teams cannot afford to rely solely on ground-level surveillance when threat actors are leveraging open-source hardware and coordinating online.
Flashpoint Physical Security Intelligence (PSI) equips security teams and executive protection units with real-time visibility into emerging physical threats before they reach your perimeter:
Early Warning Indicator Tracking: Monitor chatter across mainstream social platforms, fringe networks, and illicit DDW forums to identify probe attempts or the targeting of specific data center facilities and executives.
Geospatial Threat Mapping: Overlay real-time intelligence onto physical assets using customizable geofencing to detect active incidents, protest activity, and drone-related discussions near sensitive sites.
Actionable Counter-UAS Insights: Receive finished intelligence and analyst support to benchmark threat actor TTPs (Tactics, Techniques, and Procedures), enabling your organization to harden physical structures and justify operational investments.
To learn more about how Flashpoint helps safeguard critical infrastructure, executives, and high-value assets against physical and cyber threats, request a demo today.
Beyond Cyber: How CTI Teams Are Solving Converged Threat Use Cases
In this post we explain how cyber threat intelligence teams are being expected to take on physical risk, how tradecraft overlaps, and how Flashpoint bridges the gap.
For years, the mandate of Cyber Threat Intelligence (CTI) teams has been narrow and well understood: track cyber threat actors, monitor for indicators of compromise, and defend the network. However, that mandate is widening. In today’s interconnected threat landscape, more CTI teams are being tasked with physical security, geopolitical and protective intelligence. Whether that is monitoring and securing executive travel, a facility, or an event, data shows that this new informal expansion is becoming an industry-wide shift.
What the Data Says About Cyber-Physical Security Convergence
The SANS 2026 CTI Survey affirms that CTI programs are being asked to cover more ground, including physical and geographical risk, without a proportional increase in headcount. Survey findings additionally emphasize that the risks CTI teams navigate increasingly span cyber, physical, and geopolitical domains simultaneously, rather than staying contained to the network.
Industry research confirms this shift from every angle:
ASIS International: The security standards body developed formal Enterprise Security Risk Management (ESRM) guidance specifically to address how organizations struggle to unify physical and cyber risk into a single program with shared visibility.
2026 Physical Security Trends: Market analysis consistently identifies cyber-physical convergence and unified security operations as mainstream mandates rather than fringe concepts.
International Security Journal: Analysis highlights a fundamental shift from reactive to proactive security, driven by the reality that digital and physical systems are now so closely linked that a compromise on one side rarely stays contained.
Taken together, the picture is consistent across independent sources: intelligence teams are being pulled toward physical and human risk, and most organizations are still early in closing the gap between that mission and the tooling built to support it.
Why Physical Security is a Natural Extension
It might seem like a jump from tracking ransomware to monitoring executive travel risk, but the underlying methodology is similar. Both rely on:
Situational awareness: Understanding the context around an event, whether digital or physical.
Data aggregation: Bringing together disparate sources into a coherent picture.
Predictive analysis: Identifying indicators of risk before they become incidents.
CTI analysts are already well positioned to bridge this gap. When an executive’s safety or a physical location’s security is at risk, the earliest warning signs are frequently digital via social media sentiment, localized chatter, and open-source discussions. Treating physical security as an adjacent mission means pointing skills a team already has at a new question, rather than starting net-new.
The Strategic Advantage: Breaking Down Operational Silos
Bringing these missions together has a practical benefit beyond the workload—it prevents security silos where digital and physical intelligence teams operate in isolation. When the same team that monitors cyber threats also informs physical security decisions, the organization achieves a more complete view of risk, reducing the chance that threats fall between the gaps of two disconnected functions.
Extending CTI to Physical Security with Flashpoint
Facing this convergence head-on doesn’t require a new platform, a new vendor evaluation, or creating a new discipline. Organizations leveraging Flashpoint Ignite already have the foundation needed to seamlessly extend their visibility into physical and geopolitical threat landscapes.
Using both Flashpoint Cyber Threat Intelligence (CTI) and Flashpoint Physical Security Intelligence (PSI), security teams can answer two essential questions: “what is this threat actor doing” and “what is happening right now around this specific person or place.” Both draw on much of the same underlying data and OSINT tradecraft, so extending into physical security only requires a change in Intelligence Requirements, not mastery of new systems or tools.
With Flashpoint PSI, organizations gain real-time access to mainstream sources where conversations about fast-moving events tend to surface first, plus a geospatial layer that maps that activity to a specific place. Analysts can also draw boundaries around geographic locations to monitor mentions of an executive within that area, or observe a venue on event day, seeing relevant activity as it surfaces. All of this can be accomplished using plain language, removing the need to learn secondary query syntax or lengthy manual processes to get started.
Navigating the Future of Converged Intelligence
The distinction between cyber and physical intelligence will likely keep blurring and Flashpoint is helping security teams on the ground level integrate these two functions. CTI teams that take on physical security as part of their mission shouldn’t be expected to abandon their core discipline. Instead, they should be given the workflows to apply it to a wider set of questions, using tools built to extend rather than replace the way they already work.
See how Flashpoint supports converged cyber and physical missions from a single platform. Request a demo to see what this could look like for your team.
Security teams don’t lose ground because they lack tools. They lose ground because they can’t see everything an attacker can.
This is the challenge we addressed in our latest Demo Day webinar introducing Flashpoint External Attack Surface Management (EASM), a new module inside our Ignite platform that gives security teams a continuous, attacker’s-eye view of their external attack surface, mapped directly to our proprietary vulnerability intelligence.
The Problem: Too Much Noise, Not Enough Context
Most security teams are dealing with three compounding problems:
Disconnected Data: Vulnerability data lives isolated from actual infrastructure. Knowing a CVE exists doesn’t tell you whether it affects your active environment.
Alert Fatigue: CVSS-only prioritization treats every “critical” score as an emergency, even when an asset isn’t internet-facing or exploitable.
Accelerated Threat Cycles: AI is speeding up how quickly threat actors discover and exploit vulnerabilities, making manual tracking impossible.
Layer on top of that the reality that most teams still track their perimeter with spreadsheets or a static CMDB, and you get a widening gap between what security teams think they own and what is actually exposed. This gap has a name: shadow IT.
Shadow IT Is a Growing Blind Spot
Shadow IT covers the domains, subdomains, and cloud instances that get spun up to get work done, without IT’s knowledge or approval. It’s not a fringe issue. According to Gartner, by next year, 75% of employees will be acquiring, modifying, or creating technology outside their IT department’s visibility, up from 41% just a few years ago.
These unmanaged assets sit outside inventory and outside the reach of any scanner that only looks at what’s already known. That makes them exactly the kind of infrastructure an attacker finds first, and exactly the blind spot Flashpoint EASM is built to close.
What is Flashpoint EASM?
Flashpoint EASM gives security teams a continuous, attacker’s-eye view of their external attack surface and maps that view directly to Flashpoint’s vulnerability intelligence. Instead of your team asking “are we affected by this?”, every time a new vulnerability is disclosed, EASM answers that question continuously, often before the answer is obvious anywhere else.
Flashpoint EASM is built on three capabilities that work together:
Continuous Asset Discovery
Flashpoint EASM continuously discovers and monitors internet-facing assets: domains, subdomains, and IPs. New discoveries flow into a dedicated triage inbox, so security teams can quickly accept and focus on what’s actually relevant instead of drowning in noise.
Vulnerability Mapping
Every discovered exposure is mapped to Flashpoint’s proprietary vulnerability intelligence, including our pre-NVD findings, KEV (Known Exploited Vulnerabilities) status, ransomware likelihood, and exploit maturity. This provides organizations with immediate context into the vulnerabilities that pose the most risk.
Customizable Alerting
Using EASM, security teams get alerted to the exact moment a new asset or vulnerability is detected. This alert is fully customizable by severity and is available inside one unified workflow via Flashpoint Ignite.
Discover, map, and alert. This loop gives organizations an intelligence-led view of their perimeter, so they can proactively outpace threat actors instead of being forced to react.
How Flashpoint EASM Works
In our live demo, Flashpoint walked through the EASM workflow, which can be found under “Assets and Identifiers” in the Ignite Platform.
Here’s how it works:
Step 1: Submit Seed Keywords
Onboarding starts with keywords, meaning domain and IP address assets your organization actually owns. Any already set up asset is automatically surfaced in Flashpoint Ignite—such as through our compromised credential monitoring—ensuring no duplicated setup work.
Step 2: Triage Discovered Assets
Once keywords are approved, EASM iterates on them to surface additional related infrastructure, domains and IPs alike, along with a discovery graph showing exactly how each asset was found. That traceability makes it easy to judge relevance at a glance.
Every discovered asset lands in one of three statuses:
Owned: Assets in your tech stack. EASM continues discovering related infrastructure from these and links vulnerabilities to them.
External: Assets relevant to you, but where you don’t need further discovery, just vulnerability linkage.
Discarded: Assets you don’t need, removed from the triage feed entirely.
Step 3: Review the Vulnerable Assets Overview
In the main dashboard, the Vulnerable Assets page, security professionals can view total asset count, number of exposures, unique vulnerabilities affecting them, and total potentially vulnerable assets—in addition to criticality breakdowns for both domains and IPs.
From there, security teams can drill into:
Unique vulnerabilities, filterable by CVE or severity
Domains with vulnerabilities, showing exposure counts by severity and the last exposure date
Individual asset detail pages, showing products, versions, vendors, and ports, with vulnerabilities linked directly to the specific product version affected
Diving deeper into a surfaced vulnerability provides technical descriptions, solution information, and other affected products. Additionally, Flashpoint’s vulnerability database includes over 105,000 pre-NVD vulnerabilities, giving vulnerability management teams actionable indicators well before they show up in public sources.
Step 4: Set Up Alerting
Flashpoint EASM gives teams full control over signal versus noise. Whether that means getting notified the moment a critical vulnerability is disclosed, or reviewing a daily summary of your own schedule, EASM offers two alert types:
Asset discovery alerts, either per-asset or as a daily rollup
Vulnerability alerts, filterable by criticality (critical, high, medium, low), with the option for in-app only or in-app plus email, and available as a daily rollup
Why Flashpoint EASM Matters
Flashpoint EASM isn’t just another scanning tool. The intelligence underneath it is the differentiator: discovery tells you what’s out there, Flashpoint provides the much-needed context to tell you what’s dangerous right now.
The intelligence includes coverage that can’t readily be found elsewhere: Flashpoint’s independently researched data includes pre-NVD findings, improved KEV coverage, ransomware risk scoring, and exploit maturity.
It closes a blind spot teams have quietly lived with: EASM closes shadow IT gaps and surfaces assets sitting outside inventory entirely.
Flashpoint External Attack Surface Management gives security teams a continuous, intelligence-led view of everything a threat actor sees, so organizations can find and fix exposures before they’re exploited. To see it in action in a personalized walkthrough of your own environment, reach out to schedule a demo.
EASM Frequently Asked Questions (FAQs): What Security Teams Want to Know
What makes Flashpoint EASM different from other EASM solutions?
Most EASM tools stop at raw discovery, telling you an asset exists without telling you whether it matters. Flashpoint EASM pairs continuous asset discovery with a triage inbox to cut noise, then maps every asset directly to Flashpoint’s proprietary vulnerability intelligence, all natively inside Ignite alongside CTI and Vulnerability Intelligence. That combination means prioritization is based on real attacker activity, not just an asset inventory, giving remediation teams the exact context they need to proactively address risk.
What makes Flashpoint’s vulnerability intelligence unique?
Flashpoint’s database covers 400,000+ vulnerabilities, including 105,000+ not found in NVD or CVE, often surfaced up to two weeks earlier than public sources. Every entry is enriched with threat-informed context like EPSS scores, ransomware likelihood, exploit maturity, and MITRE ATT&CK mapping, then reviewed by human analysts, not just automated feeds. The result is prioritization based on real-world exploitation risk rather than CVSS alone.
Can existing monitored assets be imported into Flashpoint EASM? Yes. EASM integrates closely with Flashpoint’s existing assets module, so assets already set up (for example, for compromised credential monitoring) surface automatically during onboarding.
Is there a limit on discovered assets, beyond the 30-keyword cap? No. The 30-keyword limit only applies to initial seed keywords, to keep that starting set relevant. Once assets are marked owned or external, there’s no cap on ongoing discovery.
How does continuous polling compare to traditional scanning? Traditional scanners give you a point-in-time snapshot. EASM continuously discovers assets and vulnerabilities, giving you a moving view of your exposure, essentially the same view an attacker would have in real time.
Does EASM identify compound risk, where multiple weaknesses increase exploitability together? The Vulnerable Assets view surfaces how many vulnerabilities are tied to a given asset, so teams can quickly spot assets carrying disproportionate risk and prioritize accordingly.
Does EASM overlap with SBOM alerting? Not exactly. SBOM alerting monitors vulnerabilities in assets you already know about. EASM is focused on discovering the assets you don’t know about yet. Most mature security programs benefit from running both in tandem.
In today's fast-moving cybersecurity landscape, threat analysts must move beyond basic, binary reputation scores to successfully defend against modern, highly adaptive web threats. Traditional URL analysis has been redefined by the launch of URL Scanning 2.0, an update that significantly expands VirusTotal's URL analysis capabilities by introducing automated visits with a full browser instance and deeper historical visibility.
Instead of relying on static reputation scores alone, URL Scanning 2.0 enriches reports with "under-the-hood" headless browser telemetry, including the DOM, full-page screenshots, web technologies, and network request logs. Crucially, it introduces historical analysis pivoting, giving analysts the ability to track how a page has changed over time.
URL Scanning 2.0
To successfully defend against modern, highly adaptive web threats, threat analysts must move beyond basic, binary reputation scores. With the debut of URL Scanning 2.0, VirusTotal introduces robust headless browser integration that captures how a page behaves dynamically in a clean sandbox environment.
Every scan now generates rich, granular telemetry that provides a blueprint of the target page's execution:
- Headless Browser Data: Full-page visual screenshots, full DOM (Document Object Model) trees, and web technologies (e.g., Cloudflare, PHP, HTTP/3).
- Page and Network Statistics: Highly detailed counters of individual network requests, encrypted HTTPS transactions, unique contacted domains/subdomains, and serving IP address mappings with geographic tracking.
- Anti-Phishing Fingerprints: Automatic identification of brands, cloned-website tags, password input fields, tracker IDs, and favicon dhashes.
- Historical Pivoting: A timeline containing historical analyses of a URL with its corresponding risk score, allowing analysts to track exactly how its metadata and content have shifted over time.
Access Levels in VirusTotal
Public Access (Free for VirusTotal Users) The core enhancements of the URL Scanning 2.0 engine are available to everyone. For the latest scan, analysts can access rich telemetry generated by headless browser execution, including visual screenshots, extracted JavaScript globals, console messages, and a list of all loaded network resources.
VirusTotal Premium Customers For paid VirusTotal customers, the platform unlocks deeper retrospective capabilities and exclusive data fields. Analysts have the ability to pivot to and review the full historical analyses of a URL as it was observed at specific points in time, and access advanced telemetry like the full DOM captures of the execution. Furthermore, premium access unlocks advanced infrastructure relationships, allowing users to pivot on contacted domains, IPs, and downloaded files.
Note: The aforementioned Google Threat Intelligence and Automatic Brand Identification features are exclusively available to Google Threat Intelligence customers.
Investigating a Phishing Case
Initially, when an analyst navigates to the mentioned URL to view the report generated by VirusTotal, they would see something similar to the following with the new URL Scanning features:
At the top of the interface, we can see that the URL has been scanned three times. This means there are three distinct reports for the same URL, each potentially containing different information that could be highly useful for an analyst. In the top right corner, we can view these past analyses by clicking on "History".
This is where the new historical analysis pivoting comes into play: it allows analysts to travel back through a URL's timeline with point-in-time snapshots.
By clicking on "History", we can view all the historical analyses for that URL, including response codes, detections, screenshots, and other metadata. You can also apply filters to narrow down the timeline and view only the historical records you are interested in, based on specific response codes, URL actions, and other criteria.
In this case, if we click on the initial historical analysis performed on July 6, 2026 (as shown in the screenshot above), we can examine its specific information across the "Summary", "Details", and "Detection" tabs. A key feature of URL Scanning 2.0 is that the information within these report tabs will dynamically re-render to match the exact historical state of the snapshot you select.
As observed in the history timeline, after clicking on this specific analysis included a live screenshot and other relevant metadata, indicating the scan occurred while the website was fully operational and actively distributed. The previous screenshot gives us a clear view of how the phishing page was visually structured.
Furthermore, diving into the "Details" tab reveals other interesting technical artifacts from the campaign. These details are incredibly useful for pivoting and identifying new malicious URLs that share similar characteristics.
Among the wealth of information generated by URL Scanning 2.0, analysts will find HTTP transactions, detected JavaScript variables, console messages, external outbound links, and other critical metadata. These key technical markers serve as pivotable and searchable attributes, allowing teams to conduct advanced footprint hunting and instantly find other malicious URLs exhibiting the exact same technical fingerprint.
Furthermore, every snapshot taken during each analysis provides the complete Document Object Model (DOM) tree captured by the full browser instances. It allows you to inspect the exact structure of the page as it was dynamically rendered to the victim, exposing elements that static scans might miss. As can be seen in the following image, having direct access to this point-in-time DOM data empowers analysts to dig deep into the page's architecture.
Advanced Threat Hunting: Scaling the Investigation
Let's scale our investigation using VirusTotal Intelligence queries based on the artifacts discovered via URL Scanning 2.0.
During the analysis of the financial phishing site, we discovered that the page relied on static assets hosted on a third-party domain: jiaoyisuo.thai2570[.]com. We can pivot on this finding using an advanced query:
VT Query
entity:url (outgoing_link:jiaoyisuo.thai2570.com OR content:jiaoyisuo.thai2570.com)
The results demonstrate a multi-brand operation, including fake cryptocurrency exchange portals and typosquatting domains for other financial services. By further pivoting on the hosting domain with entity:domain "thai2570.com", analysts can map out a highly segmented subdomain tree used for hosting assets, capturing payments, and backend control panels.
Conclusion
URL Scanning 2.0 represents a paradigm shift in how security analysts investigate web-based threats. Investigations are no longer limited to static verdicts. By surfacing powerful metadata directly inside the workflow—such as historical DOM captures, live screenshots, and pivotable technical identifiers—analysts can now turn a single indicator into a comprehensive infrastructure map.
Log in to VirusTotal to explore the new URL Scanning 2.0 features today, and consider upgrading to VirusTotal Premium to unlock the full power of historical pivoting and advanced threat hunting.
Customers have access to models that are continuously getting better with each new generation bringing larger context windows, stronger reasoning, and lower token costs. Getting the strongest AI-powered security will come from tools that combine the most relevant models with deep knowledge of a customer’s specific environment.
AWS Continuum for code vulnerabilities (Preview) is built to be that tool to help secure your code at machine speed. Today, we’re announcing our work with Anthropic and OpenAI that extends AWS Continuum directly into the developer workflows where code is being written: Anthropic Claude Code, OpenAI Codex, and Kiro. Developers can use these integrations to discover vulnerabilities, contextually prioritize, validate, and remediate, within their existing workflows.
Models are getting smarter
AI models are advancing rapidly. Each generation brings new capabilities, and different models excel at different tasks. The latest frontier models can now identify vulnerabilities and reason through multi-step attack paths that would take a human security team weeks to trace manually.
This is a genuine breakthrough in detection, but it creates a new challenge for your security teams: more findings, more complexity, and the need to determine which ones matter most in your environment and how to address them. The next challenge customers face is building the correct harness and orchestration to turn these models into a single interface that goes from detection through remediation. This is what we set out to do when creating Continuum, which brings together many different models and uses the model that’s most effective for each part of the process.
We also partner with the Frontier Model Forum, an industry consortium developing shared safety standards, evaluation methods, and benchmarking to ensure we can evaluate these models effectively together. We’re also working with model providers on shared security performance benchmarking to make sure we’re using the best model for each task within Continuum and our other AWS security products.
The harness
An AI harness is the orchestration layer that wraps around a model to connect it to tools, guardrails, memory, and workflows, so it delivers outcomes. Think of the model as the engine and the harness as everything around it. You need both to have a high-performance car.
Harnesses are becoming increasingly complex. Teams are stitching together multiple models, agents that call agents, and dynamic workflows, and are dealing with constant change driven by innovations in models, agent frameworks, and tool integrations.
As a result of that complexity, customers are implementing shadow infrastructure to manage integration layers across models and tools. Every time the landscape shifts, security and governance controls potentially break, forcing teams to go back to revisit them and make updates.
These challenges extend beyond the model. They arise in the orchestration required to connect different models and developer environments with tools, context, controls, and workflows across a customer’s environment. At AWS, we see managing that complexity as heavy lifting that AWS should solve. We treat the harness as infrastructure and with the same rigor we apply to identity, discovery, policy enforcement, observability, and compliance of the core infrastructure at AWS.
Enter Continuum
AWS Continuum for code vulnerabilities discovers vulnerabilities, prioritizes them within the context of a customer’s business, validates them in a sandbox, and provides remediation at machine speed. Under the hood, Continuum is an agent-team loop architecture. A sophisticated harness that orchestrates all of it: selecting the right model, connecting to a customer environment, and delivering secure code that’s been validated in context. You never need to think about how the orchestration works, or what changed in the latest release.
Anthropic and OpenAI collaborations
We are working with Anthropic and OpenAI to bring Continuum into the developer workflows where code is being written.
How it will work:
Within Claude Code, Codex, and Kiro coding environments, on-demand vulnerability scans identify potential issues and send findings to Continuum. Continuum prioritizes them within the context of the customer’s AWS environment (configurations, AWS Identity and Access Management (IAM) policies, network topology, and exposure surfaces) and validates them in a sandbox. It then returns prioritized, contextual intelligence back to the coding assistant, which adjusts its recommendations accordingly.
This collapses what was traditionally a multi-step, multi-team process (write, scan, triage, prioritize, fix, rescan) into a single outcome: the code suggestion itself. Two modes, one outcome:
For existing code: Use Continuum for code vulnerabilities from AWS to discover, prioritize, validate, and remediate across your environment.
For greenfield code: Use the Continuum plugin within Codex, Claude Code, or Kiro to get security-validated suggestions in your development environment.
Early design partners are already seeing results.
“AWS Continuum connects source code with enterprise knowledge, allowing teams to accurately pinpoint security vulnerabilities and verify that flagged issues are truly meaningful. This shortens what really matters: timeline to fix serious vulnerabilities.” – Mike Johnson, CISO, Rivian
Next
AWS Continuum for code vulnerabilities is available in preview through AWS. Sign up to request access at AWS Continuum.
Continuum integrated into Claude Code, Codex, and Kiro workflows are coming soon.
If you have feedback about this post, submit comments in the Comments section below.
What OpenAI’s and Anthropic’s testing incidents really teach defenders In the past two weeks, two of the world’s leading AI labs have disclosed the same unsettling result. During their own safety testing, their most capable models reached real companies’ systems. First OpenAI, whose models broke into Hugging Face. Then Anthropic, whose models reached three more organizations. Read the disclosures closely. Two facts carry the weight. First, the safeguards were not defeated. They were switched off by design. OpenAI ran the models with reduced cyber refusals and safety classifiers disabled, to measure raw capability on a cyber benchmark. A model doing […]
AI coding agents are part of the developer toolchain. Tools like Kiro and Claude Code generate features, tests, and code refactors from natural-language prompts. A single agent can open dozens of pull requests (PRs) across your repositories in an afternoon. That productivity comes with a trade-off: agents optimize for task completion at machine speed with no understanding of your organization’s risk.
Through protocols like the Model Context Protocol (MCP), agents also reach beyond the integrated development environment (IDE) to call APIs, query databases, and modify infrastructure and even entire environments, expanding the scope of resources your application security team defends.
This post lays out an application security (AppSec) control framework for AI coding agents. Two pillars organize the framework: author-time controls shape what the agent produces in the IDE; build-time controls verify and gate what reaches production. Your existing secure software development lifecycle (SDLC) controls still apply and are critical to a defense-in-depth security strategy. The framework shows where to layer additional guardrails so AppSec scales with agent-driven development. The framework is tool-agnostic and cloud-agnostic. Throughout, we use AWS services—Kiro in the IDE and AWS CodePipeline in the build—as a running example that you can adapt to your own toolchain.
Risks
Each of the following risks includes a treatment summary. The control framework section later in this post provides implementation details. The risks are ordered by severity with the highest impact risks first.
R001. Prompt and context injection
Agents read untrusted content, such as issue descriptions, web pages, MCP responses, and README files in third-party packages. Text from outside parties can redirect the agent to disclose secrets, open unauthorized PRs, or invoke tools without user consent. This risk, known as prompt injection, is the top risk in the OWASP Top 10 for LLM Applications. Any agent that reads content from outside parties is exposed, with or without MCP, so connecting tools widens the scope of impact.
Treatment: Treat non-developer input as untrusted. A large language model (LLM) can’t reliably separate instructions from data in a single context window, so architect for it: keep the agent that orchestrates trusted actions separate from the one exposed to untrusted content and grant the exposed agent only read-only, least-privilege access. Require human approval for irreversible actions. Use version-control steering files to prevent silent tampering.
R002. Inadvertent data disclosure and overly permissive configurations
Agents optimize for getting work done. Left unchecked, the code they generate can default to wildcard identity and access management policies, open security groups, and unencrypted storage, or embed sensitive values in code rather than referencing a secrets manager. Most coding agents now include safety mechanisms that make these outcomes less likely, but they remain imperfect, so you still need controls to account for the possibility.
Treatment: Security requirements in a steering document, plus policy-as-code scanning (Checkov, cfn-nag) in the IDE and pipeline. See Context as a security control.
R003. Uncontrolled changes reaching production
Ungated code reaching production isn’t new, but AI agents amplify it. Machine-speed generation can propagate a flawed pattern across repositories before it’s identified.
Treatment: Branch protection rules requiring PR approval (a human-in-the-loop checkpoint), pre-commit hooks for security checks, and sandboxed agent runs that prevent direct pushes to protected branches. The right balance between human review and automated speed depends on the risk profile of the change. For many low-risk paths, automated checks alone might suffice, while higher-risk changes warrant a human checkpoint.
R004. Supply chain risks
Agents don’t always distinguish current best practices from outdated patterns. They might recommend deprecated packages, reference library versions with new Common Vulnerabilities and Exposures (CVEs), and hallucinate package names that don’t exist, which can introduce risks of dependency confusion issues.
Treatment: Software Composition Analysis (SCA) in the pipeline (for example, Amazon Inspector code scanning or Dependabot) to flag vulnerable or unexpected dependencies. For additional control, resolve against a scoped registry like AWS CodeArtifact. Even without a fully curated registry, lockfile validation and allow-listing critical packages reduce exposure.
R005. Uncontrolled external access
Through MCP and tool integrations, agents query databases, call APIs, and modify infrastructure. Without constraints on which tools and data an agent can reach, a single misconfigured integration provides unintended access to sensitive resources.
Treatment: Scope MCP servers to least-privilege tools and resources, enforce authn or authz on external connections, and audit tool invocations. The control point is the configuration file. Review it the same way you review AWS Identity and Access Management (IAM) policies.
R006. Hallucinations and incorrect code
Agents produce plausible-looking output. Code that compiles, passes linting, and looks reasonable can still be functionally wrong: misusing APIs, introducing subtle logic errors, or implementing security-sensitive operations incorrectly. Code that passes continuous integration (CI) but is wrong slips through review; code that fails to build is caught immediately.
Treatment: Layer deterministic verification (static application security testing (SAST), unit tests) with non-deterministic review (LLM-assisted screening against the specification). Neither catches everything alone.
R007. Scope creep
Given a bug-fix prompt, an agent might also refactor surrounding code, disable an unreliable test, or reorganize imports. Unrequested changes introduce regressions and complicate review.
Treatment: A reviewed specification document that defines what must change and what must not, paired with a targeted review of the proposed changes. See Specifications as scope boundaries.
The preceding risks share a common thread: agents produce output faster than humans can review it, and they lack context to self-correct.
The following framework addresses this gap. It organizes controls into two pillars: author-time (pre-generation and post-generation of code) and build-time (in the pipeline, before code reaches production). Author-time controls shape what the agent produces. Build-time controls verify it. Neither is sufficient alone; together they reduce the volume and severity of issues that reach human reviewers.
Deterministic compared to non-deterministic mitigations
Deterministic mitigations[D] produce the same result every time. Linters, SAST scanners, secrets detection, and policy-as-code match patterns against rules and define security invariants: no critical findings, no hardcoded secrets, and no wildcard IAM policies. Use them when the condition can be expressed as a rule. Organizations already have these and must continue enforcing them.
Non-deterministic mitigations [ND] use model judgment. They include steering documents, LLM-as-judge review, specification compliance checks, and scope-creep detection, and they evaluate intent rather than patterns. They catch novel issues that rules miss, but are probabilistic. Use them when evaluation requires context or reasoning across files. This is the new layer that AI-generated code demands, because agents produce code that can pass every deterministic check yet remain functionally wrong.
Human review[H] provides the final layer for the risk-based decisions neither tool type can make. Apply it where judgment is needed, not everywhere: routing every change to a person invites consent fatigue, where reviewers approve by reflex and the control loses its value. The default reflex is to route everything back to a human, but that isn’t always the right response—reserve human judgment for the decisions that genuinely need it.
The control framework
The framework organizes controls into two pillars. Author-time controls (Pillar 1) shape what the agent produces in the IDE, before code is generated and just after. Build-time controls (Pillar 2) verify and gate that output in the pipeline, before it reaches production. The controls within each pillar are tagged deterministic [D], non-deterministic [ND], or human [H].
Pillar 1: Author-time controls (pre- and post-generation of code)
Author-time controls work inside the IDE, where the developer and agent still hold full context. They shape the prompt and the generated output before it ever reaches a pull request. The following controls apply at this stage.
Context as a security control [ND]
Control statement: Encode security invariants as natural-language constraints in a steering document that every developer environment consumes at session start. Addresses R002. Many AI coding agent risks share one root cause: the agent lacks the security context an experienced developer carries implicitly. Your security team sets the policies, such as Amazon Simple Storage Service (Amazon S3) buckets require encryption, API gateways require mutual TLS, and credentials must come from AWS Secrets Manager. Developers don’t always have these requirements available when they’re building. They build what works, not what’s compliant. An AI agent amplifies this gap because it defaults to whatever pattern dominated its training data, with no awareness of your organization’s security posture.
A key mitigation is steering. Security teams write these invariants once as natural-language guidance in a steering document, then distribute them as shareable resources that developers consume in their IDE. The agent loads the file at session start and treats the contents as standing requirements:
IAM policies must follow least-privilege principles; no wildcard Amazon Resource Names (ARNs).
No hardcoded credentials in source code; use a secrets manager.
Security groups must not allow unrestricted inbound access.
This shifts security left, before code generation begins. Steering biases generation toward secure defaults; it doesn’t guarantee them. Treat it as a strong default, paired with the following deterministic gates that block non-compliant code from merging. Security teams define the rules once and every developer environment inherits them automatically. Steering reduces the volume of issues that reach the pipeline, though it doesn’t replace downstream scanning.
How to write effective steering rules: Keep each rule specific and testable, scope it to a concrete risk class, keep the rule set concise so the agent can hold it in context, and iterate from the issues your scanners and reviewers surface.
Specifications as scope boundaries [ND]
Control statement: Require a reviewed specification before code generation begins. Define what must change and what must not. Addresses R007.
Spec-driven workflows turn vague prompts into reviewable specifications before code is generated. This creates a human checkpoint at the design phase, where security decisions are made:
Requirements use testable notation that’s auditable before the agent writes a line of code. For example, the Easy Approach to Requirements Syntax (EARS): WHEN [condition] THE SYSTEM SHALL [behavior].
Tasks are ordered in implementation steps, each mapped back to a requirement.
For bug fixes, specifications add a critical element: unchanged behavior documentation. This is an explicit list of behaviors that must continue working, giving the agent a written boundary against scope creep.
In this model, the specification becomes the primary artifact, code is a derivative of it. Human review effort concentrates on whether the specification solves the right problem with the right constraints, not on reading implementation diffs line by line.
Controlled tool access using MCP [D + ND]
Control statement: Scope each MCP server to the minimum set of tools the agent needs, and give it a dedicated, scoped-down credential rather than the developer’s own. Maintain an allowlist of reviewed MCP servers. Addresses R005.
MCP servers act as controlled gateways between the agent, the external tools, and data:
Dependency management – An MCP server fronting your private package registry resolves dependencies against curated packages, not the public internet. This is a deterministic constraint on supply chain risk.
Infrastructure tooling – Visibility into current resource configurations prevents templates that conflict with existing infrastructure.
Scoped permissions – Each MCP server exposes a defined set of tools and resources. You choose exactly what the agent can access, supporting least-privilege at the integration layer. You supply that credential through the agent’s configuration (in Kiro, the env block of .kiro/settings/mcp.json). Avoid autoApprove: ["*"], which removes the human approval prompt on every tool call.
IDE code scanning [D]
Control statement: Run real-time static analysis in the IDE so security issues surface while the developer (and agent) still have full context. Addresses R002, R006.
Real-time diagnostics catch syntax errors, type mismatches, and configuration issues as the developer types. A malformed IAM policy is flagged before the agent builds further on it. Security-focused extensions (ESLint security plugins, Checkov, SAST) layer on top for immediate feedback while code is fresh in context.
Hooks: Automated guardrails at the point of action [D + ND]
Control statement: Attach deterministic checks to file-save events and non-deterministic verification to task-completion events. Addresses R002, R007.
Shell command hooks [D] – Triggered on file save, these run a linter, formatter, or security scanner and produce the same result every time. They enforce hard rules.
AI-powered hooks [ND] – Triggered on task completion. These prompt the agent to verify that the implementation matches the specification and check for any untested edge cases or files that were modified outside the task’s scope.
Pillar 2: Build-time controls (in the pipeline)
Build-time controls run in the pipeline after code is committed and before it reaches production. They verify and gate what the agent produced, catching what author-time controls did not. The following controls apply at this stage.
Layered security scanning [D]
Control statement: Run secrets detection, static analysis, dependency scanning, and infrastructure-as-code scanning in sequence. Fail the build on any critical finding. Addresses R002, R003, R004.
Secrets detection runs first because it’s cheapest and addresses a high-severity class of issue. It scans for hardcoded API keys, database connection strings, and credentials that AI agents might inadvertently include.
SAST scans source code for injection issues, insecure deserialization, and resource leaks. Custom rules can target AI-specific anti-patterns including overly broad exception handling, deprecated APIs, placeholder credentials, dynamic code execution through eval().
Software Composition Analysis (SCA) identifies known CVEs in dependencies. This is critical for AI-generated code, which might reference deprecated packages or hallucinate package names that open you to dependency confusion issues.
Infrastructure as code (IaC) scanning validates AWS CloudFormation, Terraform, and AWS Cloud Development Kit (AWS CDK) templates against security policies before deployment. Catches overly permissive IAM roles, unencrypted storage, and public-facing resources the agent created.
Each stage halts the pipeline on failure. Results export to a standard format (Static Analysis Results Interchange Format (SARIF)) for compliance auditing and flow downstream to human reviewers. The open source Automated Security Helper (ASH) bundles secrets, SAST, SCA, and IaC scanners behind one command that you can run locally and in AWS CodeBuild, emitting SARIF for the gates that follow.
Quality gates [D]
Control statement: Define pass/fail thresholds for each scan type. Block deployment on any critical or high-severity finding. Addresses R003.
Quality gates convert scan results into go/no-go decisions. Define thresholds for each severity: block on critical findings, require justification for highs, and track mediums. The gate is deterministic: if a threshold is breached, the pipeline stops. Exceptions require documented approval.
Differentiate blocking compared to advisory modes: hard failures on main, advisory on feature branches. Avoid gates becoming a friction that teams route around.
AI-assisted review [ND]
Control statement: Use an LLM reviewer to pre-screen every pull request for specification compliance, scope creep, and security anti-patterns before human review. Addresses R001, R006, R007.
Specification compliance – Does the implementation match the requirements document?
Scope verification – Were files modified outside the task’s stated scope?
Security pattern review – Are there logic errors, misused APIs, or insecure patterns that pass SAST but violate intent?
This pre-screening focuses human reviewer attention on genuine risks rather than formatting or obvious issues. On AWS, AWS Security Agent (code review in preview at publication) checks pull requests against AWS-managed and custom security requirements. The reviewer screens and surfaces findings; the merge decision stays with a human.
A critical principle: the agent that wrote the code should not be the agent that reviews it. A separate session helps avoid self-confirmation bias, but a separate session alone doesn’t always avoid the generator’s blind spots, because two sessions of the same model can share them. Where practical, use a different model for review so the reviewer is less likely to inherit the same systematic weaknesses.
Human-in-the-loop review [ND + H]
Control statement: Require human approval on most pull requests, especially those touching security-sensitive or high-blast-radius code. Lower-risk changes might be eligible for agent-assisted or fully automated approval as tooling matures. Provide reviewers with scan results, LLM pre-screening output, and specification context to enable fast, informed decisions. Addresses R003.
Scale review depth to the risk of the change. Low-risk or boilerplate changes can take a lighter-touch review, while security-sensitive or novel-logic changes warrant mandatory deep review and a second reviewer.
Scanners catch known patterns but can’t judge whether code implements the intended business logic. Human review also serves to calibrate trust: teams build intuition about where agents excel (boilerplate, test writing) and where they’ve tended to struggle (novel business logic, security-sensitive operations), recognizing that this frontier shifts as models improve.
Place two approval gates: after security scans (reviewer focuses on correctness and business logic, with scan results as context) and before production deployment (final sign-off after integration testing). Treat human review as a secondary control, not a guarantee: reviewers are themselves non-deterministic and can miss issues, so human review layers on top of the deterministic gates rather than replacing them.
Putting the framework into practice on AWS
The framework is tool-agnostic, but AWS gives you building blocks for each pillar. The following services map directly to the controls described previously: Kiro for author-time guardrails, and CodeBuild and CodePipeline for build-time gates.
Kiro: Structured AI development
Kiro maps to Pillar 1: It puts the author-time controls in the IDE, where the developer and agent still share full context. Each feature in the following list implements one of those controls, configured in-repo under .kiro/ so the guardrails are version-controlled and shared across the team rather than set per developer.
Steering documents – Markdown files in .kiro/steering/ load into the agent’s context at session start. Conditional inclusion using fileMatch (for example, ["**/*.tf"]) loads IaC-specific rules only when relevant.
Specification-driven workflows – Three-phase specifications (requirements in EARS, design, and tasks) with review checkpoints. Bug-fix specifications capture unchanged behavior explicitly.
Agent hooks – Triggered on file save, tool invocation, or task completion. Shell hooks run deterministic checks (linters, tests); Ask Kiro hooks run AI prompts for non-deterministic review. For example, a security pre-commit scanner hook can flag hardcoded credentials when the agent finishes a task.
Property-based testing – Guided by a specification or hook, Kiro can generate property-based tests (for example, using the hypothesis library) that exercise hundreds of randomized inputs, probing edge cases a hand-written test suite would miss.
MCP integrations – Connect Kiro to private package registries, internal docs, issue trackers, and infrastructure tooling, creating the controlled tool access pattern.
AWS CodeBuild and AWS CodePipeline: Pipeline controls
CodeBuild runs each scanning tool (checking for secrets, SAST, SCA, and IaC) as a build action. A non-zero exit code fails the action, and the stage halts or rolls back according to its OnFailure setting. Findings export as SARIF to Amazon S3 for compliance, and CodePipeline action variables pass results to downstream approval actions.
CodeBuild exit codes halt the pipeline on scan failures
AWS Lambda invoke actions evaluate scan results against configurable thresholds and return pass/fail decisions
Manual approval actions halt the pipeline, send Amazon Simple Notification Service (Amazon SNS) notifications, and link to review artifacts; decisions and reviewer identity are logged for audit
The following table consolidates the framework into a single view that includes each stage of the SDLC and the deterministic [D] and non-deterministic [ND] controls that apply there. Every stage carries both, a reminder that neither control type is sufficient on its own.
Full security scan suite, integration tests, and policy-as-code
AI-assisted review for human approvers
Post-deploy
Runtime monitoring and anomaly detection
AI-powered incident triage
Conclusion
This post laid out a framework for adopting AI coding agents at machine speed without letting unreviewed risk reach production. It layers guardrails at two points:
Author-time controls – Steering, specs, and scoped tools shape what the agent generates in the IDE.
Build-time controls – Scanning, quality gates, and layered review verify it before it reaches production.
No single layer is enough: deterministic gates enforce hard rules, non-deterministic review catches what they miss, and human judgment is reserved for the decisions that need it. Together, they let AppSec scale with agent-driven development.
Where to start this week:
Start with steering and specs – Encode security requirements as steering and use specifications for new features. Highest impact, lowest effort. For a ready-made starting set, the open source Project CodeGuard (a Coalition for Secure AI project under OASIS Open, of which Amazon is a contributing member) publishes reusable steering rules for common risk classes—hardcoded credentials, IaC misconfiguration, supply chain, and MCP security—that you can adapt to your AWS environment.
Add deterministic pipeline gates – Integrate SAST, SCA, and secrets detection. Table-stakes regardless of AI usage.
Calibrate and iterate – Review what controls catch, adjust steering for recurring issues, and expand agent autonomy as trust builds.
Accountability – Developers remain accountable for the security of what they ship. AI agents accelerate development; they don’t transfer ownership.
Your phone rings, you pick up and say hello. On the other end: total silence. No one answers, and the call abruptly disconnects. If you don’t already use spam call blockers, you’ve almost certainly run into this situation before.
In most cases, these are scam calls. Today, we explain why these calls happen, what the callers want from you, and how to protect yourself. Most importantly, we’ll look at whether you even need to bother protecting yourself against them in the first place.
Who’s calling?
It’s not just scammers on the line — robots, legitimate call center operators, and ordinary folks make these calls too. Let’s break down each type of caller — ordered from best-case to worst-case scenario for your security.
Actual person
The most harmless scenario is that an actual person called you, but their microphone is acting up. Maybe they accidentally muted themselves with their ear, or their smartphone connected to a Bluetooth headset, speaker, or car system that isn’t capturing their voice. Carrier glitches can also mute one side of a call. The caller might have no idea there’s a problem — as far as they know, they are speaking, but no one can hear them. In cases like this, you usually recognize the incoming phone number.
If the call comes from an unknown number, there’s still no need to panic — though the list of those who might be calling gets much longer.
One legitimate possibility is a call center agent who simply didn’t pick up or connect their headset in time. Call center systems are designed to dial numbers faster than agents can wrap up their calls. The system tried to route the call to a human, but no reps were available. That’s why you sometimes have to wait a few seconds before hearing a single word, or why you might hear ringing tones as if you were the one making the call.
Robot or AI
Silence on the line is a common sign of robocalls. Robots test whether a phone number is active and, if it is, pass it along to a human — meaning a real sales rep (or scammer) will call you back in the next few days. It’s worth noting that scammers aren’t the only ones making these pinging calls. Legitimate call centers use the exact same tools to reduce the workload on their live agents.
An AI agent could also be behind the silent call. To the person answering, there’s no practical difference: the call looks identical to one made by a standard bot. However, AI can do more than just auto-dial numbers — it can analyze your response and use that data to decide whether your number is active and ready to be handed off to a live person for follow-up.
Unwanted caller
Now we get to the real threat. Perhaps one of the most dangerous and unpleasant sources of silent phone calls is a scammer. A quick, silent call like this can actually be the groundwork for a long, elaborate attack with cover stories about loans, government agencies, other fraudsters, even law enforcement.
Debt collectors might also be calling and staying quiet. Your number could end up on their radar if you, your family, or close contacts have outstanding debts. In these cases, a silent call is often used as a tactic for psychological pressure.
A similar technique is used in stalking. While silent calls cause no direct harm on their own, they can be leveraged to induce anxiety, create a feeling of being constantly watched, and cause ongoing emotional distress.
Why do they call and stay silent?
When you pick up, you likely respond out of habit with a quick “Hello?” or “Hi there.” That’s all it takes for the other party to gather a wealth of data. While this information used to be difficult to process, the rise of artificial intelligence has made the task significantly easier. Let’s look at what someone can learn about you from just one spoken word:
Region, accent, and location. Scammers are sophisticated and cunning. Their tactics are often tailored by region — targeting residents of specific countries or even regions within them. This is especially relevant in places like India or South Africa, which have 22 and 11 official languages, respectively.
Approximate age and gender. While a human listener might easily confuse a teenager’s voice with a young woman’s or misjudge someone’s age entirely, AI is far better at picking up on subtle vocal nuances. Knowing your age and gender helps scammers refine their playbook for future social engineering attacks.
Times you’re available. If you answer the phone in the morning, afternoon, or late at night, attackers can schedule their follow-up call during the exact time window when you’re most likely to pick up.
Likelihood of a successful attack. AI can automatically assess the potential value of a target. For instance, if someone answers quickly, speaks calmly, and doesn’t immediately hang up on unknown numbers, they’ll likely be assigned a higher priority for follow-up calls by live scam operators.
Back to the “why do they call and stay silent”, the main reason is to harvest biometric data. Just a few seconds of recorded audio can help cybercriminals create a voice deepfake. While one or two words might not yield a convincing clone on their own, attackers can stitch together recordings from multiple silent calls to build a believable replica.
This technology is already being used in real-world scams. Impersonating a relative, colleague, or boss, fraudsters can urgently ask you to send them money, to share a two-factor authentication code for government services, or to complete some other seemingly innocuous request. The more realistic the deepfake sounds, the harder it is to spot the scam — especially when backed by a convincing backstory.
What to do if you get a silent call?
If you answer a call, say a few words, and hang up, there’s no need to panic. However, that brief interaction can confirm to attackers that your number is active and that you’ll answer calls from unknown numbers. As a result, your phone number could end up on target lists for future spam or scam campaigns. That said, it’s important to remember that a single silent call poses no immediate security threat.
Here are a few tips to help you stay calm and avoid falling for scam tactics if those silent calls are becoming a problem:
Don’t answer calls from unknown or hidden numbers. Here’s a helpful tip: if someone genuinely needs to reach you, they’ll find another way to do so, or keep calling from the exact same number at various times. Scammers almost always dial from different numbers, while automated bots operate on a rigid schedule — like calling every day at precisely 8:05 AM.
Don’t rush to call back. Scammers often count on proactive victims who are curious enough to return calls from unfamiliar numbers. On top of that, calling back could end up costing you money if it’s a premium-rate number.
Don’t speak first. Wait for the caller to greet you before starting a conversation. If you hear muffled noise or complete silence on the line, hang up and save yourself the hassle — it’s likely a scam.
Block unknown numbers — even after the call. If you picked up and realized the call could be risky, it’s best to block the number right away. You can use the built-in features on most modern smartphones to do this.
Don’t share your number everywhere. Phishing sites, fly-by-night web pages, and sketchy giveaways often exist solely to collect your personal data. When filling out forms online, it doesn’t hurt to use a burner or secondary number.
Get a second phone number. Separate your daily life between two numbers. Use your main line strictly for family, friends, and work contacts, and reserve the secondary line for deliveries, online marketplaces, and general web sign-ups.
Amazon is sharing new findings about how a threat actor linked to the Democratic People’s Republic of Korea (DPRK) is targeting open source software libraries, the shared building blocks that companies around the world use to develop applications. Amazon Threat Intelligence has linked several recent compromises of popular Node Package Manager (NPM) libraries to the same DPRK-linked threat actor, a connection that hasn’t been publicly reported until now. The analysis also describes how generative AI is already changing what malicious software packages look like and how threat actors are beginning to probe AI-based code systems. We’re sharing this research to help the open source community and security teams better identify and address these types of events.
These developments come 2 years after the XZ Utils backdoor, which demonstrated how a patient attacker can compromise critical open source software by exploiting the trust and limited time of volunteer maintainers. Open source software underpins much of the internet’s infrastructure: operating systems, web servers, encryption libraries, and the application frameworks that businesses rely on daily. When an attacker compromises a widely used open source package, every organization that depends on that package is potentially affected. Since then, Amazon Threat Intelligence has observed the volume and sophistication of software supply chain attacks increase, driven in large part by DPRK‑linked threat actors and cybercriminal groups.
In this post, Amazon Threat Intelligence and the Amazon Inspector team share new details about recent campaigns against popular NPM packages, including evidence that the compromises of the axios, debug, chalk, and typo-crypto libraries were carried out by the same DPRK-linked threat actor tracked by the security community as SAPPHIRE SLEET, STARDUST CHOLLIMA, BlueNoroff, CageyChameleon, and Alluring Pisces. We also outline how the techniques used to compromise open source repositories are evolving, why these changes matter for organizations that depend on open source software, and what Amazon Web Services (AWS) is doing to help customers detect and respond to these threats.
One DPRK–linked group behind multiple NPM compromises
In March 2025, the DPRK-linked threat actor compromised the typo-crypto package. In September 2025, the same threat actor compromised the debug and chalk NPM packages. In March 2026, the same operational playbook appeared in a compromise of the axios package, one of the most widely used JavaScript libraries with more than 100 million weekly downloads. In each case with debug, chalk, and axios, the threat actor gained access by socially engineering a trusted maintainer of the package, then published a software update containing malicious code. Any organization that automatically pulled the latest version of these packages received the compromised update.
While the axios compromise has been publicly attributed to this DPRK-linked threat actor, the typo-crypto, debug, and chalk incidents haven’t previously been connected to it. Amazon Threat Intelligence identified shared tactics, techniques, and procedures (TTPs) across these supply-chain campaigns, including trojanized NPM packages, use of post-install hooks (scripts that run automatically when a package is installed), and code reuse. Based on analysis of command-and-control (C2) indicators and TTPs, Amazon Threat Intelligence assesses with medium confidence that these campaigns are attributable to the DPRK-linked threat actor tracked as SAPPHIRE SLEET, STARDUST CHOLLIMA, BlueNoroff, CageyChameleon, and Alluring Pisces. This is the first time these compromises have been publicly tied to this DPRK-linked threat actor.
Amazon Threat Intelligence assesses this as part of a financially motivated pattern: by compromising a small number of highly popular packages, the group gains potential access to thousands of downstream environments simultaneously. For a financially motivated threat actor, this approach is far more efficient than targeting organizations one at a time.
The aggregate impact of these incidents underscores the efficiency of targeting share dependencies. As reported by Wiz Research, roughly 1 in 10 cloud environments were affected by the debug and chalk supply chain event within a two‑hour window.
A smaller campaign that foreshadowed later activity
During routine analysis of indicators and TTPs related to the axios threat actor, Amazon Threat Intelligence identified a connection to a domain registered in 2025, prompting a full investigation into its historical activity. That investigation uncovered that the same DPRK-linked threat actor had committed a trojanized file to the typo-crypto NPM package in March 2025. The malicious file, core.js, masquerades as the legitimate core-js NPM package within the typo-crypto repository.
Based on the limited number of observed downloads, Amazon Threat Intelligence assesses that this campaign was small scale and likely served as a testing ground for the more visible supply chain operations that followed in late 2025 and 2026. The group appears to have been refining supply chain techniques more than a year before the larger campaigns that drew public attention. Amazon Inspector reported this malware to the Open Source Vulnerabilities (OSV) database, where it’s now tracked as MAL‑2026‑3400, so that the broader security community can benefit from these findings.
The trojanized file executes when it receives a hash input beginning with the value 0098273. When triggered, it downloads a second-stage payload from a hardcoded C2 server, then executes the payload based on the victim’s operating system, with behavior tailored for Windows, macOS, or Linux. The malware implements file-based persistence with payload rotation and uses multi-layer obfuscation, combining base64‑encoded text with an XOR cipher keyed to 01042025.
Amazon Threat Intelligence assesses that the group was experimenting with techniques that later appeared in the higher-impact campaigns against axios, debug, and chalk. Although the observed download volume was low, the tradecraft aligns with what we later observed in attacks on more popular packages.
How attacker tradecraft is shifting
Over the past year, Amazon Threat Intelligence and Amazon Inspector have observed threat actors changing the techniques they use to target open source libraries. These changes matter because open source packages remain attractive targets: they’re widely trusted, automatically updated in many environments, and maintained by communities that welcome new contributors. The following patterns describe how attackers are adapting their methods to evade modern defenses. Each is designed to exploit the gap between the moment a dependency is inspected and the moment it actually executes. A year ago, we looked for malicious packages. Today, we look for malicious behaviors split across packages that appear harmless on their own.
From package‑level attacks to fragment‑level attacks
Amazon Inspector has observed attackers increasingly splitting a single malicious workflow across several ordinary-looking packages. One package stores an encrypted blob disguised as configuration. A second ships the decryption logic. A third, often published later, fetches and executes the payload.
Viewed on its own, each package looks benign. There are no install hooks that stand out, no obvious evaluation of untrusted input, no network calls that look suspicious. The malicious behavior only appears when the components are used together in the intended sequence. This approach is designed to defeat scanners that evaluate packages one by one instead of reasoning about how they interact in a real dependency graph.
Long-horizon campaigns that invest in trust
We’re also observing threat actors taking a long view of trust accumulation. Instead of publishing obvious malware and waiting for downloads, they publish something genuinely useful and maintain it. They behave like real maintainers for weeks or months, shipping features, fixing bugs, and gaining dependents.
The same patience shows up on the human side. In some cases, the goal isn’t to launch a new package at all, but to become a contributor to an existing project. That’s the through line from XZ Utils backdoor to the debug, chalk, and axios maintainer compromises. In each case, the adversary treated legitimacy as an asset to be spent once, at the moment of maximum access.
Decoupling the package from its behavior
In many recent cases, a library is clean on the public registry yet still dangerous, because its real behavior depends on resources the attacker controls elsewhere. These can include guard or license scripts fetched from an external repository at runtime, configuration files that gate certain behaviors, or remote endpoints consulted at startup.
As long as those external resources remain benign, code reviews pass and automated scans return clean results. When an attacker flips the content or arms an endpoint that previously returned a placeholder, every installed copy can become malicious at once, without any new package release. A package that shows no malicious behavior today isn’t the same as a package that’s is safe by design.
From basic obfuscation to real cryptography
Where attackers used to rely on simple obfuscation such as minification or single-layer base64 encoding, we now observe multi-stage payloads that use stronger cryptographic techniques. Examples include AES‑GCM encrypted blobs gated by passphrases, RC4-style string arrays with per-call keys, layered XOR over base64, and native loaders that hold the next stage as an encrypted field decrypted only in memory.
The common design choice is that the decryption key is never stored in the package itself. It’s derived from runtime context, fetched from a server at execution time, or supplied as a license key. That means even an analyst with full source access can’t reliably decrypt the payload statically. Stage one looks like a simple decryptor; the malicious content remains ciphertext until it runs on a real target with the real key.
Payloads that avoid detonating in sandboxes
As defenders have scaled automated analysis in cloud sandboxes, attackers have made their code more environment aware. The payload decides whether it’s being analyzed before it acts. We see execution gated behind real package install lifecycles, single-use environment variables, and checks for signals of a genuine developer or build environment. These include interactive terminals, realistic usernames and hostnames, domain membership, plausible uptime, local file history, specific operating systems, and cloud metadata that helps distinguish analysis infrastructure from normal workloads.
Some delivery servers also tailor what they serve based on the client. A benign decoy goes to generic browser-like requests, while the live payload only appears for the exact user agent used by the malware. The result is that a clean verdict from a cloud sandbox often tells you more about how convincing your environment looks than how safe the package is.
How generative AI is reshaping both attacks and defenses
Generative AI is changing what attackers can produce and what defenders can rely on. Adversaries can generate novel code and content at scale. Historically, many malicious packages were caught because they looked wrong, with broken language, thin documentation, obvious copy-paste, or a telltale function reused across samples. Generative AI erases many of those signals.
Attackers can now produce thousands of lines of coherent, idiomatic, well-commented code, complete with convincing documentation, plausible commit histories, and synthetic maintainer identities, wrapped around a backdoor. Because each variant can be mutated, renamed, restructured, and re-encrypted, there is no single stable signature to match. Pattern-based detection loses ground against malware that looks one of a kind in every deployment.
AI is also creating new initial access vectors. One emerging technique is slopsquatting, where attackers register package names that exist only because an AI coding assistant hallucinated them. When a developer or an autonomous coding agent asks for help and the model confidently recommends a nonexistent package, an attacker can pre-register that name and wait. The next person who follows the recommendation might receive malware, despite not mistyping anything or visiting a malicious site, because the AI effectively delivered the bad dependency for them. As organizations move toward agents that install dependencies with limited human review, this path looks less like a curiosity and more like a scalable delivery channel.
Most significantly, AI changes the calculus for defensive automation. Attackers are no longer just writing malware for humans to miss. They’re writing malware for AI reviewers to approve. As organizations rely on AI systems to review code and triage packages, those AI systems themselves become part of the attack surface. We expect that indirect prompt injection, a technique where hidden instructions manipulate an AI system into taking unintended actions, will increasingly be embedded in malicious packages to fool AI-based code scanners. These instructions can be hidden in source comments, README files, docstrings, or test fixtures, and crafted to convince an automated system to mark malicious code as safe, skip a specific file, or perform an unintended action during analysis. The same content the malware needs to function can carry a second, separate message aimed at the machine that inspects it.
How AWS is responding
We’re investing across Amazon Threat Intelligence and Amazon Inspector to help customers adapt to this shifting landscape of software supply chain risk. Amazon remains committed to helping protect the security of our customers and the internet by actively hunting for and mitigating threats from sophisticated threat actors. We will continue working with Amazon teams, industry partners, and the security community to share intelligence and mitigate threats. Upon discovering this campaign, Amazon Threat Intelligence worked with Amazon Inspector so the malicious package was tracked, mitigated, and shared with the community through the OSV database. Additionally, the observed indicators were shared with Amazon GuardDuty to alert our customers of this activity.
Amazon Inspector uses these insights to refine our detection logic, broaden coverage across registries, and prioritize signals that reflect the tradecraft shifts described in this post, and is collaborating with industry partners such as package registries and Open Source Security Foundation (OpenSSF) to share findings.
We’re also investing in helping open source maintainers better secure their projects. In 2026, AWS joined the Linux Foundation and other industry leaders to launch Akrites, a collaborative initiative to defend critical open source software against AI-enabled cyber threats. AWS has also jointly invested $12.5 million alongside other organizations to defend the open source ecosystem from AI-driven attacks. These efforts reflect a broader commitment: the security of open source software is a shared responsibility, and defending it requires sustained investment from the organizations that depend on it.
Our goal is to help customers understand where their environments rely on open source components, identify suspicious behavior early, and respond quickly when the software supply chain is used as an entry point.
August 11, 20206: This post was updated to clarify that the social engineering of a trusted maintainer applied to the debug, chalk, and axios compromises specifically. The underlying attribution and findings remain unchanged.
Most threat intelligence frameworks were built around clear, recognizable motives—advanced persistent threats seeking intelligence, financially motivated ransomware syndicates, or ideological extremists pursuing political or religious goals. However, security practitioners and physical security teams are facing a vastly different and highly volatile new vector on the threat landscape: Nihilistic Violent Extremism (NVE).
Operating across surface web platforms, niche gaming servers, and encrypted messaging channels, NVE actors seamlessly blend traditional cybercrime, physical violence, real-world property destruction, and severe digital extortion.
In a recent Flashpoint webinar, our analysts took a deep dive into this complex digital threat, fully breaking down the inner mechanics of NVE, its warning indicators, and how cross-functional security teams can proactively monitor and mitigate these dangerous digital-to-physical threats.
Here are the core takeaways from our on-demand webinar that organizations need to understand.
What is Nihilistic Violent Extremism (NVE)?
Nihilistic Violent Extremism (NVE) defines criminal conduct driven by a deep misanthropy and a desire to trigger societal collapse through random acts of chaos, psychological cruelty, and violence. While casual observers might dismiss these activities as extreme “internet trolling” or adolescent angst, Flashpoint recognizes NVE as a digitized, accelerated evolution of long-standing extremist and occult philosophies.
NVE draws heavily from the Order of Nine Angles (O9A), a paramilitary philosophy originally established in the United Kingdom. Unlike traditional movements seeking political control, O9A advocates for the total destruction of modern civilization to force a return to social darwinism.
How NVE Transitioned from Ideological Literature to Gamified Online Terror
The transition of reclusive occult literature into digital networks followed a deliberate path of gamification. Threat actors stripped away the theological texts, replacing them with fast-paced, highly visual media designed to engage younger audiences on gaming platforms and encrypted messaging apps.
These repackaged materials were then adopted by the various groups within The Com, such as 764 and other scavenger cults. By wrapping graphic violence and extremist symbology in internet humor, these groups lower a recruit’s psychological defenses, accelerating their desensitization and drawing them rapidly into higher-harm activities.
Key Tactics, Techniques, and Procedures (TTPs) of NVE
NVE networks represent a primary example of digital-to-physical convergence, where virtual harassment directly manifests as physical security risks. For NVE actors, violence that remains private is considered wasted effort—because their focus is on generating public fear, breaking taboos, and winning peer status polls, publicity is an operational requirement.
Recorded acts of violence serve as the primary currency across all three pillars of “The Com”. To build status, gain access to private channels, or enforce extortion, threat actors rely on a distinct set of operational tactics to create a societal environment of fear and discord, elaborated on in our expert webinar.
The Demographic Realities and Accessibility of NVE Groups
A critical takeaway from the webinar was the demographic profile and accessibility of NVE networks, with participants—both perpetrators and victims—being overwhelmingly young, typically ranging from ages 11 to 22, with a high concentration of juveniles. Additionally, because extreme coercion and abuse are normalized in these spaces, victims are frequently pressured into becoming enforcers against others as a condition to cease their own victimization.
Because of this young demographic, most NVE actors do not rely solely on Tor hidden services. Instead, they recruit, coordinate, and broadcast activities across mainstream social media, open messaging apps, and popular online gaming platforms.
Protect Against NVE Risk Using Flashpoint
Tracking a highly decentralized threat ecosystem where groups form, rename, and dissolve within hours requires specialized, multi-disciplinary intelligence capabilities. Flashpoint provides enterprise security teams, physical safety leads, and CTI analysts with the visibility required to identify and mitigate NVE activity.
To explore the complete webinar discussion, which includes deeper analyst breakdowns of threat actor activity, behavioral indicators, and enterprise mitigation strategies, watch the on-demand recording today.
The Flashpoint Method: Prioritizing Vulnerabilities in an Era of AI-Accelerated Discovery
We outline Flashpoint’s practical, repeatable framework for prioritizing vulnerabilities based on real-world risk, exploitability, and business impact.
Organizations are gaining new ways to identify vulnerabilities at scale, thanks to new generations of powerful AI models. However, security teams still face the same fundamental question: which vulnerabilities actually matter?
Vulnerability management teams have increasingly struggled to keep pace with growing disclosure volumes. From January 1, 2026 to June 30, 2026, Flashpoint tracked 21,667 vulnerabilities, an 8% period-over-period increase, with one-in-five containing publicly available exploit code at time of disclosure. At the same time, the gap between disclosure and exploitation continues to shrink, with some vulnerabilities weaponized in as little as 24 hours.
Flashpoint’s Method for Threat-Informed Vulnerability Prioritization
Recent developments such as Anthropic’s Mythos model have highlighted the growing potential for AI-assisted vulnerability discovery. As advances in code analysis enable researchers and organizations to identify software flaws at unprecedented speed and scale, the volume of discovered vulnerabilities is set to potentially increase significantly across software ecosystems.
That’s why we created this guide, The Flashpoint Method for Threat-Informed Vulnerability Prioritization, a practical, intelligence-driven framework designed to help vulnerability and exposure management teams cut through the AI-driven noise and focus on the vulnerabilities that matter most. By incorporating real-world exploitation activity, threat actor behavior, asset exposure, business context, and remediation considerations, organizations can make faster, more informed decisions and reduce risk more effectively.
Download to gain:
A clear, threat-informed prioritization framework: How to assess which vulnerabilities demand immediate attention, and why — moving beyond static severity scores alone.
Core and expanded prioritization checklists: Criteria spanning asset criticality, active exploitation, CVSS severity and ransomware risk, social risk and community chatter, business context, compensating controls, zero-day status, KEV inclusion, EPSS scoring, ease of remediation, and vulnerability age.
How to operationalize prioritization at AI scale: Insight into how Flashpoint’s vulnerability intelligence platform and analyst expertise help teams keep pace as AI-assisted discovery accelerates disclosure volume.
Prioritize Vulnerabilities More Effectively and Faster Using Flashpoint
While increased visibility into vulnerabilities is ultimately a positive for defenders, it amplifies a challenge security teams already face—separating which vulnerabilities represent meaningful risk to your environment and require immediate action.
What is threat-informed vulnerability prioritization?
Threat-informed vulnerability prioritization is the process of evaluating vulnerabilities based on real-world risk rather than severity scores alone. It incorporates factors such as active exploitation, exploit availability, threat actor activity, asset exposure, business context, and remediation considerations to determine which vulnerabilities require immediate attention.
Why is vulnerability prioritization important?
Organizations face thousands of newly disclosed vulnerabilities each year, while security teams have limited time and resources to remediate them. Effective vulnerability prioritization helps organizations focus on the vulnerabilities most likely to be exploited and most likely to impact their environment.
How is AI changing vulnerability management?
AI-assisted code analysis is enabling researchers and organizations to identify software flaws faster and at greater scale. While increased visibility into vulnerabilities benefits defenders, it also increases the volume of vulnerabilities that security teams must evaluate, making effective prioritization even more important.
Why isn’t CVSS enough for vulnerability prioritization?
CVSS provides a standardized measure of technical severity, but it does not account for whether a vulnerability is actively being exploited, relevant to your environment, or likely to impact your business. Effective prioritization combines severity with threat intelligence and organizational context to assess real-world risk.
How does Flashpoint help organizations prioritize vulnerabilities?
Flashpoint combines analyst-driven vulnerability intelligence with real-world exploitation data, threat actor insights, asset exposure, and business context to help organizations identify the vulnerabilities that pose the greatest operational risk. This intelligence supports faster, more informed remediation decisions and operationalizes threat-informed vulnerability management at AI scale.
Understanding Illicit Ecosystems: Inside Rehub’s Rise as a Primary Ransomware Marketplace
As part of our ongoing series, Flashpoint intelligence tracks Rehub, breaking down its migration, infrastructure, and the various RaaS groups sponsoring and partnering with it.
Rehub, also known as ReHub or RehubCom, is a Russian-language cybercrime forum founded in August 2025 by a former XSS moderator following its shutdown in the summer of 2025. Rehub dedicates itself to the commercial and marketplace use of ransomware, while its counterpart, DamageLib, serves as a knowledge base archive and exchange.
2025
July 23: XSS is taken down by law enforcement
August 1: XSS moderators launch DamageLib, which completely abandons illicit commerce.
August 10, 2025: Rehub forum is launched by a former XSS moderator, fully embracing illicit commerce.
January 28, 2026: RAMP is seized by law enforcement, with its users migrating to Rehub.
Operating both on Clear Web domains and an onion domain, the forum positions itself as free from state and law enforcement interference, framing existing XSS iterations as compromised. After law enforcement seized the RAMP (RAMP4U) forum in January 2026, Rehub absorbed a significant portion of the displaced cybercriminal community and became one of the primary destinations for ransomware operators.
The Rehub login page in August 2025, early stage of the forum. (Source: Rehub)
Who Are Known Members of Rehub?
There are many notable threat actors among Rehub moderators and users, including ransomware operators, vendors, and other prominent threat actors active across several illicit communities. Several current or ex-Rehub moderators were also maintainers of other illicit forums such as XSS, DamageLib, and RAMP.
Notably, Ransomware-as-a-Service (RaaS) groups such as DragonForce have maintained an active presence on the platform to market their affiliate programs. Flashpoint assesses that DragonForce is likely the forum’s primary sponsor or partner, as their banner is permanently displayed on the forum’s home page, with both logos merged—similar to its previous placement on RAMP.
The Rehub home page with the DragonForce logo. (Source: Rehub)
As of July 2026, Flashpoint intelligence observes over 8,300 active users, 15,000 posts, and nearly 3,000 threads. Despite being free to join, Rehub practices a zero trust policy, which was established in mid-April 2026. Under this system, the forum restricts newly registered users from accessing any section other than its Sandbox. Users can also purchase paid upgrades:
Premium status (gold rank): Costing US $100 per year, this rank grants distinctive color, custom title, nickname changes, unlimited post editing/deletion, extended signature, unlocks all hidden text regardless of post count, likes, join date, ability to bump commercial threads, and inherits all lower-tier perks.
Patron status(pink/magenta rank): Costing US $5,000 per year, this rank grants custom title editing, a personal profile link, custom styling for posts, profile, and postbit, and inherits all “Premium” perks.
The only section available to newly registered users on Rehub forum. (Source: Rehub)
What are the Various Rehub Forum Sections?
Rehub sections, similar to other forums, are grouped by major activities, separating the knowledge base from commerce and from general discussions.
The list of Rehub forum sections. (Source: Rehub)
Sandbox
Serves as an entry-level general discussion area and a place for community questions. Main activity consists of queries about operational security, introductory networking, and entry-level fraud or malware logistics.
Technical
Covers threads ranging from traditional network infrastructure vulnerabilities to emerging technologies such as AI jailbreaking and deepfake social engineering. Highly active, most communications focus on network vulnerabilities and carding.
Programming (Development)
This is a dedicated space for discussions on software engineering, system administration, and web optimization within the forum. Primary activities include sharing programming language tutorials, comparing backend technologies, and developing specialized automation tools.
Library
Serves as a repository of resources for the forum, hosting the most threads and community engagement. Users share operational materials, leaked databases, and utility software. Additionally, this section aggregates cybersecurity and tech industry news and articles.
Supermarket
This is a commercial section featuring ransomware affiliate programs, compromised network access, malware tools, stolen financial data, bulk spam infrastructure, forged documents, anonymous hosting, and crypto laundering services.
Arbitration
Serves as the forum’s internal justice system, where members resolve financial disputes and flag scammers. The “Black List” subsection functions as a public record of bad actors and scam sites.
Administration
This is where forum staff post announcements, policy updates, and operational notices, including rules, official domains, forum news, moderator applications, and 2FA requirements. Members use it to ask questions, request escrow services, propose features, and raise concerns about the forum’s public image.
Monitor Illicit Marketplaces Using Flashpoint
Flashpoint will continue to monitor Rehub’s marketplace activity and infrastructure updates. Rehub’s rapid evolution from a post-XSS refuge to a heavily sponsored ransomware marketplaces demonstrates the resilience of the cybercrime ecosystem.
Positioning itself as the primary ransomware marketplace, Rehub has built a high-barrier, high-reward environment for sophisticated threat actors. Request a demo to learn how Flashpoint delivers visibility into illicit communities—empowering security teams to track threat actors, identify exposed assets, and mitigate ransomware risks.
Inside Qilin Ransomware: Custom Rust Loader and Kernel-Level EDR Killer
In this post we analyze Qilin ransomware’s new custom Rust loader, break down the inner workings of its sophisticated kernel-level EDR killer, and explore how organizations can defend against these aggressive defense evasion tactics. Flashpoint customers can access the full intelligence report—complete with deeper technical analysis and all associated IOCs—directly within Flashpoint Ignite.
Qilin ransomware is a highly active and sophisticated ransomware operation that has rapidly modernized its evasion techniques. Historically focused on file encryption, the ransomware-as-a-service (RaaS) group has expanded its operations to include aggressive, kernel-level defense evasion. By deploying a specialized toolkit, Qilin now focuses heavily on blinding and permanently disabling endpoint security products before its main ransomware payload is executed on a victim’s network.
Flashpoint has observed Qilin quietly deploying a previously unreported custom packer, which has been actively observed in wild samples since May 2024, with continuous use detected as recently as last month.
Here’s how Qilin works:
How Qilin Ransomware Uses a Custom Rust Loader for Reflective PE Loading
Flashpoint analysts observed a custom Rust-written loader that performs reflective Portable Executable (PE) loading of the ransomware payload. After deobfuscation, the code execution jumps to the newly unpacked executable within the same process, avoiding noisier process injection techniques. The following is an overview of the decompiled unpacking routine:
Decompiled code of Qilin ransomware unpacking routine. (Source: Flashpoint)
The unpacking routine then reads each DWORD from the embedded bytes, allocates it on the heap, and performs multiple mathematical operations to deobfuscate. Flashpoint notes that the calculations and values used were unique to each sample, but the underlying methodology remained the same.
Manually performing the calculations in the sample confirms the presence of the embedded binary, with the first deobfuscated DWORD yielding an ‘MZ’ header in little-endian format.
To better understand Qilin, Flashpoint analysts created an automated unpacker and configuration extraction script that uses CPU emulation to address the issue of unique calculations per sample. This script uses pattern matching to locate the unpacking routine within the binary. It then reads the disassembly, identifying specific points in the code at which emulation should start and stop.
Python code snippet reading the disassembly to find optimal areas to emulate. (Source: Flashpoint)
Reading the disassembly directly avoids issues arising from hardcoded offsets, such as when threat actors add or remove code, or when the compiler introduces changes. Additionally, it provides a smaller set of instructions for emulation, avoiding WinAPI calls and other invalid memory errors that often occur when emulating a full binary.
After additional setup, including mapping the sample into the emulator’s memory and creating a fake heap, the unpacking routine runs successfully.
Python code snippet performing CPU emulation to unpack the embedded binary. (Source: Flashpoint)
The script then performs configuration extraction from the deobfuscated bytes produced by the CPU emulation, achieving a 100% success rate.
Automated tooling successfully unpacking and extracting Qilin’s configuration. (Source: Flashpoint)
How Qilin’s New EDR Killer Blinds Security Products
An additional update with Qilin is its new endpoint detection and response (EDR) killer, which Flashpoint found to be sold on illicit marketplaces for US $2,000. This is packed via the Shanya packer—which was sold on XSS for US $100 to US $150 back in 2024. The packer is highly sophisticated, and uses several techniques that make it difficult to analyze, such as junk code, application programming interface (API) hashing, IAT hooking, pattern scanning, and VEH code execution flow.
Once unpacked, the EDR killer starts by using dynamic API hashing and PE walking to resolve a number of useful NTAPI functions it will use throughout the process, and stores them in a structure located within the GdiHandleBuffer within the Process Environment Block (PEB).
The structure stored in the PEB itself looks as follows:
Recreated structure definition based on Flashpoint analysis. (Source: Flashpoint)
The API hashing algorithm is simple: it performs a bitwise OR of each character of the API name with hexadecimal value 0x20 to convert any and all uppercase characters to lowercase, then performing additional simple calculations.
The EDR killer compares the returned locale to a known locale blacklist to avoid attacking any Commonwealth of Independent States (CIS) countries such as Russia and Belarus.
The malware then attempts to give itself the following privileges by dynamically resolving and calling RtlAdjustPrivilege():
SE_PROF_SINGLE_PROCESS_PRIVILEGE
Required to gather profile information for a single process.
Used later to create a map of the victim machine’s physical memory space.
SE_DEBUG_PRIVILEGE
Required to debug and adjust the memory of a process owned by another account.
SE_LOAD_DRIVER_PRIVILEGE
Required to load or unload a device driver.
Abusing Vulnerabilities to Map Physical Memory
The EDR killer then writes a vulnerable driver to disk and loads this driver via Service Manager. This driver is the ThrottleStop driver from TechPowerUp LLC’s free and legitimate application of the same name, used to bypass CPU throttling. However, the driver suffers from a vulnerability, allowing the malware to map physical memory to kernel-mode virtual memory to perform direct kernel read and write operations.
Qilin weaponizes this vulnerability by feeding its EDR killer physical memory addresses, as the driver relies on the API to map physical memory to a kernel-mode virtual address. To achieve this, the EDR killer builds a physical memory map using a Windows memory management service that preloads frequently used applications into RAM.
First it gathers baseline information about all physical memory blocks. Because memory pages (typically 4KB) are allocated to physical blocks, hundreds of virtual pages can point to a single physical range.
It then calls the service to obtain detailed Page Frame Number (PFN) details. The malware stores this complete mapping in a global variable, giving it a reliable, built-in translation table between virtual and physical memory spaces.
Bypassing Driver Signing Checks
To run its own malicious tools, the EDR killer must first bypass Windows’ driver signing enforcement. Normally, Windows uses a built-in verification check to block unsigned or blacklisted drivers from loading. The malware tricks Windows into disabling this gatekeeper using a simple swap:
The malware finds a specific kernel function and uses its physical memory map to pinpoint its location.
It commands the vulnerable driver to scan this memory area for a specific byte signature. This leads directly to the Code Integrity callback table.
Within this table, the malware locates the built-in verification check and “patches” it with a harmless, dummy function.
Blinding Security Products
With driver signing checks completely bypassed, the malware uses its read/write primitives to dismantle system callbacks, it identifies and targets:
Process notify callbacks
Thread notify callbacks
Image load notify callbacks
Registry callbacks and minifilters
Rather than conducting a blanket unlinking of all system callbacks, the EDR killer checks the address of each callback. If the address falls within a memory range owned by a security product on its hardcoded blacklist, Qilin surgically unlinks it by zeroing out the pointer with null bytes.
Next, the EDR killer drops and loads its own custom driver, which appears to Windows as purpose-built. Once loaded, the Qilin EDR killer gets all relevant running processes. For any processes running that match a hardcoded list, it stores the Process ID in a vector.
For every PID found, the malware sends a message to a driver. At a high level, the driver finds the full path of the target executable, makes it unreadable, unwriteable, and undeletable to any and all users, and then terminates the process.
Interestingly, the Qilin EDR killer performs a Discretionary Access Control List (DACL) modification on the target security product executable. The driver creates a new empty ACL header and sets the flag SE_DACL_PRESENT to TRUE. This is significant because a null DACL and empty DACL are not the same. A null DACL grants everyone access, whereas an empty DACL grants no access. This process makes it so that the security product’s executable can no longer be executed without needing to delete the file like other EDR Killers. Once the driver then terminates the executable, it can’t be restarted.
DACL modification to remove access to the security product executable. (Source: Flashpoint)
Once everything is completed, the EDR killer unpatches the Code Integrity Check to avoid triggering PatchGuard and then exits.
Defend Against Qilin Using Flashpoint
The sophisticated kernel-level manipulation highlights a rapidly expanding trend in the broader threat landscape: the proliferation of highly effective malware designed purely to disable enterprise-level security products. Qilin’s integration of these techniques demonstrates how the EDR killer market is maturing in the cybercrime underground, transitioning from a niche capability into a standard prerequisite for high-impact ransomware operations.
As security platforms continuously improve their detection mechanisms, Flashpoint believes the threat landscape surrounding anti-EDR tools will only grow larger and more aggressive, forcing organizations to focus on protecting the kernel and detecting rogue driver deployments. To learn more about Qilin and the latest advancements in ransomware, request a demo.