Reading view

Measuring LLMs’ Ability to Perform Cryptanalysis

There’s new benchmark measuring AI’s ability to perform mathematical cryptanalysis. Anthropic’s frontier model actually found new attacks.

The benchmark: “CryptanalysisBench: Can LLMs do Cryptanalysis?” The idea is to benchmark the ability of LLMs to discover new mathematical cryptanalytic attacks against a series of historical algorithms.

Abstract: Cryptanalysis—the task of finding attacks against cryptographic schemes—its at the intersection of mathematical reasoning and cybersecurity, two areas where LLMs have advanced fastest. Cryptanalysis represents both a clean testbed for frontier reasoning (as practical attacks can be automatically verified) and a domain with unusually high stakes, since the primitives under study underpin our digital security. In this paper we ask whether LLMs can do cryptanalysis, and find that the answer is increasingly yes. We introduce CryptanalysisBench, 191 tasks across six families of cryptographic primitives (block ciphers, hash functions, etc.) drawn primarily from four NIST standardization competitions. Our benchmark consists of three tiers: (i) primitives with known practical breaks; (ii) primitives with no known practical break, evaluated both at full strength and as scaled-down variants; and (iii) a challenge set of production primitives at the frontier of cryptanalysis. Five frontier models (Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and the open-weights GLM-5.2) break 65%­86% of Tier 1 schemes, 6­12 Tier-2 schemes at full strength, and 24­61 across all scaled-down variants. Beyond deriving known results, models produce novel cryptanalysis, such as a key-recovery attack that exploits a design flaw in the SpoC AEAD and an error in KINDI’s published CCA-security proof, both to the best of our knowledge not previously known.

We release CryptanalysisBench as a tool to help track if (or when) AI cryptanalysis becomes a serious factor and as a scaffold for stress-testing candidate schemes before deployment. The attacks that the benchmark already surfaces are an early snapshot of a fast-moving frontier that may soon match, and in places exceed, the published state of the art.

Anthropic used the benchmark to test Mythos Preview, and found new vulnerabilities in Hawk and reduced-round AES.

Still early results, but this is definitely something to watch.

SlashDot thread.

  •  

The Fourth Circuit Says Border Agents Can Search Your Phone By Hand, No Suspicion Required

Legal intern Suzanne Castillo was the principal author of this post.

The Fourth Circuit issued a disappointing opinion in U.S. v. Belmonte Cardozo, a case in which EFF filed an amicus brief, alongside the national ACLU, its Maryland, North Carolina, South Carolina, and Virginia affiliates, and the National Association of Criminal Defense Lawyers (NACDL).

We argued that electronic device searches at the border should require a warrant based on probable cause, but at minimum, regardless of whether an officer searches by hand or with forensic software that plugs into a device and downloads its entire contents for search, the same Fourth Amendment standard should apply to all device searches at the border.

Unfortunately, the court rejected that argument and ruled that a lower standard applies to manual searches, allowing the government to conduct extraordinarily invasive electronic device searches without any suspicion of wrongdoing, simply because the border officer chooses to search by hand rather than with a forensic tool.

The Border Search Exception Meets Your Phone

The Fourth Amendment requires that government searches of persons or property be reasonable, which usually means obtaining a warrant based on probable cause from a judge.

But a warrantless search can still be reasonable if it falls within an exception to the warrant requirement, including the exception that allows officers to search your belongings at the border. The border search exception allows warrantless searches of persons or property crossing the U.S. border, including the functional equivalent of the border such as international airports, given the government’s interests in controlling who and what may enter the country.

Historically, courts have categorized border searches of luggage, vehicles, and personal effects as “routine” and thus reasonable even if conducted without any suspicion that the traveler has engaged in wrongdoing; courts have also held that more invasive “nonroutine” searches, such as certain body searches and searches that damage property, require reasonable suspicion.

But a person’s privacy interests in the personal data on a phone or laptop are extraordinarily different than their limited privacy interests in the contents of their suitcase.

The Supreme Court addressed cell phone privacy in Riley v. California (2014), holding that the search-incident-to-arrest exception to the warrant requirement did not apply to cell phones, thereby generally requiring a warrant for phone searches, at least at the interior of the country. The court recognized the unprecedented privacy interests people have in their cell phones and how even brief manual searches can reveal the “sum of an individual’s private life,” including our political affiliations, religious beliefs, sexuality, and more. Accordingly, the Supreme Court held that because electronic device searches bear “little resemblance” to searches of bags or physical containers, they should be evaluated differently.

Following Riley, the Fourth Circuit considered two border device search cases involving forensic searches, in which border officers used external software to extract and analyze a device’s data.

In U.S. v. Kolsuz (2018), the Fourth Circuit held that a forensic search of a cell phone at the border “must be considered a nonroutine border search, requiring some measure of individualized suspicion” of a transnational offense, but the court declined to decide whether the standard is only reasonable suspicion or instead a probable cause warrant.

Then in U.S. v. Aigbekaen (2019), the Fourth Circuit held that a forensic device search at the border in support of a purely domestic law enforcement investigation requires a warrant. The court also reiterated the general Kolsuz rule for a forensic border-related device search: the “Government must have individualized suspicion of an offense that bears some nexus to the border search exception's purposes of protecting national security, collecting duties, blocking the entry of unwanted persons, or disrupting efforts to export or import contraband.”

In Belmonte Cardozo, manual searches were finally before the court.

A Disappointing Decision

Jose Belmonte Cardozo was already on the U.S. government’s radar when he traveled from Bolivia to the U.S. and was met by a U.S. Customs and Border Protection (CBP) officer at Washington Dulles International Airport. The officer manually searched his cell phone and found child sexual abuse material (CSAM), considered “digital contraband,” leading to Belmonte Cardozo’s arrest and criminal prosecution.

At issue on appeal was what standard should apply to manual device searches at the border. The Fourth Circuit held that, unlike forensic searches, manual searches are “routine” and thus reasonable under the Fourth Amendment without a warrant or individualized suspicion.

The court’s holding hinged on four differences between manual and forensic searches: (1) in a manual search, a person does the searching, not a machine; (2) a manual search’s breadth depends on the officer’s time and energy, while forensic searches are comprehensive; (3) manual searches reveal only what a user can typically access, while forensic searches can uncover deleted files, cached fragments, metadata, and more; and (4) manual searches are subject to an officer’s fading memory or imperfect notes, while forensic searches create a permanent copy.

But in identifying these technical differences, the court never explains why they justify a lower standard for manual searches.

The Fourth Circuit’s holding is problematic because, as we argued in our amicus brief, manual searches reach the same categories of data as forensic searches—data that can reveal highly personal aspects of our identities and our lives. It does not matter if a search is conducted by an agent’s thumbs or by software: the end result is equally as invasive, therefore all device searches should fall under the warrant requirement, or at least the same Fourth Amendment standard.

The court repeatedly emphasized that the search here lasted only two minutes, suggesting that the time-limited search was not privacy-invasive. But an individual’s privacy interests in their personal data don’t change based on how their phone is searched or how long. Scrolling for two minutes through someone’s personal text messages or photos is an invasion of privacy that may reveal intimate details about the person even in that short period of time.

Moreover, as devices’ native search functions improve, manual searches can surface personal information in seconds through keyword searches, even for photos, where it might have taken an hour of scrolling to find the same information, further showing that a time-limited search is not necessarily less privacy-invasive. What matters is not the breadth of the search itself, but the unprecedented (and growing) breadth of data on our phones.

A Silver Lining

There’s one silver lining: by relying on the fact that the search lasted two minutes, the Fourth Circuit left open the possibility that lengthier manual searches could trigger heightened suspicion requirements. But until a clear line is drawn, border officers within the Fourth Circuit’s jurisdiction can use manual searches to sidestep heightened Fourth Amendment standards that would otherwise apply. In the meantime, EFF will keep fighting against extraordinarily invasive warrantless, suspicionless device searches at the border, and for robust privacy standards to protect our most personal data.

  •  

EFF and Allies: X’s FTC Petition to Waive Privacy Violation Order Should be Rejected

X Corp. should not be able to escape privacy compliance because it changed its name. 

On May 15, X Corp. filed a petition before the Federal Trade Commission (FTC) to set aside or modify an order issued in 2022 requiring the company to report regularly to the FTC for its violations of user data. The order or “consent decree” is a result of misleading the platforms’ 140 million users by using private information given to secure accounts, like phone numbers and email addresses, for targeted advertising. It also fined the company $150 million for the infraction. As part of an open comments period, EFF and allies including Demand Progress Education Fund (DPEF), National Consumers League (NCL) and Electronic Privacy Information Center (EPIC) call on the FTC to reject this petition.

The 2022 order was a renewal of an order stemming from a previous violation. Back in 2011, Twitter (now X) reached a settlement with the FTC after the regulator found Twitter had failed to secure users’ personal information, resulting in exposure of that data to hackers. The settlement banned the company from misrepresenting its data protection measures, required it to set up safeguards on user data, and regularly report its security posture for twenty years. The renewal updated the expiration of X’s obligations to 2042, but if the FTC accepts X's petition, it would end much sooner.

In arguing to set aside the order, X remarks that since the order in 2011 it has “built an entirely new privacy and information security program staffed by new personnel operating under new leadership with a … philosophy grounded on the importance of privacy and information security.” 

These sweeping assurances that corporate restructuring led to a fundamental change in X’s policy and practices around user data should be met with a healthy dose of skepticism, given evidence to the contrary. For example, the company’s quiet rollout integrated its AI model Grok with the platform in 2024, trained (without meaningful consent) on X user data. The company was also subject to a massive data breach in 2025. Even if a rotation of leadership led to prioritizing privacy and information security, our letter highlights that this would not be sufficient grounds to remove the order, “because the FTC orders bind the corporate entity. Those obligations do not dissolve when the employees who negotiated or administered it depart.”

X argues that its entry into the AI space should be reason not to continue the oversight, claiming that “terminating the Order is critical to advancing American leadership in artificial intelligence.” Here again, broad-stroke claims that the guardrails in place “[diverts] engineering resources from innovation to compliance paperwork” ignores the dangers that AI introduces to user data. Far from being a reason to waive the order, clever attacks on models trained on user data has the ability to supercharge the types of secondary use violations that led to the 2022 order renewal. After all, an entire art has been developed around engineering LLM prompts to reveal the data a model was originally trained on.

Our response to X’s petition debunks many claims the company uses in its arguments. For example, there’s little evidence the order placed an undue financial burden on X. In our letter, we note that the compliance cost is merely “a rounding error against the $200 billion valuation of X Corp. following the xAI merger.”

Strong safeguards on our information require eagle-eyed oversight when that data is abused and misused for profiteering ventures. X’s actions not only showed us this in the past, but continue to do so in the present day. We and our civil society partners urge the FTC to take the clear, sensible path and reject X’s petition.

  •  
❌