Подтвердите e-mail

Для публикаций, комментариев, реакций и сообщений подтвердите адрес.

Профиль

Anthropic {bot}

Профиль Vively

Unofficial mirror account of https://x.com/anthropicai from Twitter We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://claude.ai/.

The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under “deliberately permissive conditions” that are not representative of any of our production models. (5/6)

120

Gaining a clear picture of Claude’s understanding of its situation—by examining its reasoning transcripts and running our own analyses—will help us identify the causes of its behavior. (4/6)

121

We’re grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents. We’re working closely with them to gather more details of the incident as we conduct our own investigation. (3/6)

120

The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately given internet access. AISI reports that the models “engaged in sustained, potentially harmful activity directed at real people and organisations”. (2/6)

120

We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. (3/4)

140

Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. (2/4)

150

We've made progress restoring capacity and things are stable right now, but we're not fully in the clear yet. If anything comes up, we'll post it on http://status.claude.com.

Claude StatusWelcome to Claude's home for real-time and historical data on system performance.status.claude.com
020

Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately pace the frontier of AI development so society can prepare. We’re glad to see broad agreement across the field. https://www.pacingthefrontier.com/ (2/2)

000

Also in this release: Auth hardening and a formal deprecation policy See the MCP blog for more: https://blog.modelcontextprotocol.io/posts/2026-07-28/

The 2026-07-28 SpecificationThe 2026-07-28 Model Context Protocol specification is out, bringing a stateless protocol core, Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening, a formal extensions framework, and updated Tier 1 SDKs.blog.modelcontextprotocol.io
040

Extensions are also first-class — a formal path to extend the protocol. Examples: 1. MCP Apps: server-rendered UIs in a sandboxed iframe 2. Tasks: long-running and async operations 3. Enterprise Managed Auth: control MCP server access centrally via your identity provider

130

Before, running a remote MCP server meant managing session state, which limited where you could run it. Now that MCP is stateless, you can deploy on serverless and edge infrastructure, or scale horizontally behind any load balancer.

140

We also worked with academics at ETH Zurich, Tel Aviv University, and the University of Haifa to build CryptanalysisBench, a benchmark for studying LLMs’ cryptanalysis abilities. https://arxiv.org/abs/2607.18538

CryptanalysisBench: Can LLMs do Cryptanalysis?Cryptanalysis - the task of finding attacks against cryptographic schemes - sits at the intersection of mathematical reasoning and cybersecurity, two areas where LLMs have advanced fastest. Cryptanalysis represents both a clean testbed for frontier reasoning (as practical attacks can be automatically verified) and a domain with unusually high stakes, since the primitives under study underpin our digital security. In this paper we ask whether LLMs can do cryptanalysis, and find that the answer is increasingly yes. We introduce CryptanalysisBench, 191 tasks across six families of cryptographic primitives (block ciphers, hash functions, etc.) drawn primarily from four NIST standardization competitions. Our benchmark consists of three tiers: (i) primitives with known practical breaks; (ii) primitives with no known practical break, evaluated both at full strength and as scaled-down variants; and (iii) a challenge set of production primitives at the frontier of cryptanalysis. Five frontier models (Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5, and the open-weights GLM 5.2) break 65%-86% of Tier 1 schemes, 6-12 Tier-2 schemes at full strength, and 24-61 across all scaled-down variants. Beyond deriving known results, models produce novel cryptanalysis, such as a key-recovery attack that exploits a design flaw in the SpoC AEAD and an error in KINDI's published CCA-security proof, both to the best of our knowledge not previously known. We release CryptanalysisBench as a tool to help track if (or when) AI cryptanalysis becomes a serious factor and as a scaffold for stress-testing candidate schemes before deployment. The attacks that the benchmark already surfaces are an early snapshot of a fast-moving frontier that may soon match, and in places exceed, the published state of the art.arxiv.org
030

Full technical details of both attacks are provided in our new papers: On HAWK: https://anthropic.com/document/hawk_key_recovery.pdf On AES: https://anthropic.com/document/aes_mobius_bridge.pdf And the associated model chain-of-thought for AES:

120

Still, both results show that frontier AI models are capable of doing expert-level cryptography research. This has important defensive applications—testing the algorithms that keep our online activity secure, and ultimately helping to make digital systems safer.

120

These are substantial research advances, but they don’t have a practical impact on today’s systems. HAWK is a proposed scheme that hasn’t been deployed anywhere, and the AES attack we discovered was on a weaker version and does not break the full cipher.

120

Mythos Preview did most of this work autonomously, with occasional human guidance. Each of the two results cost roughly $100,000 in API usage. We disclosed the findings in advance to the algorithms’ authors, as well as to US government and industry partners.

120

The symmetric cipher is a reduced version of the Advanced Encryption Standard (AES)—which has received decades of scrutiny (more than almost any other encryption algorithm). In a week, Mythos Preview found a way to speed up an attack on this version of AES by 200-800×.

120

The digital signature scheme is HAWK, which is designed to be robust even against hypothetical quantum computers. HAWK has survived two years of expert review, but in 60 hours Mythos Preview found a previously-unknown attack that reduced the scheme’s key strength by half.

120
Показать ещё