When Honesty Becomes a Competitive Force: Anthropic, AI Misuse, and the Emerging Politics of Transparency
September 18, 2026 | Jonathan Brown
Anthropic’s latest threat-intelligence report is more than a warning about criminals and state-linked actors using Claude. It is also an institutional intervention. By repeatedly publishing uncomfortable details about malicious activity involving its own systems, Anthropic is helping establish a new expectation: frontier AI companies should disclose what they discover, how the misuse worked, what they stopped, what they could not verify, and what the rest of the world should do next.
That transparency is costly. It exposes weaknesses, invites criticism, alarms customers, supplies adversaries with information about safeguards, and gives regulators evidence that advanced models are already changing the character of cyber operations. Yet the company continues to publish. The result may be a competitive pressure on other frontier providers: silence can begin to look less like prudence and more like concealment.
The important question is not whether Anthropic is acting from pure altruism. Companies rarely act from a single motive. The important question is whether a company’s safety interests, commercial interests, public commitments, and institutional self-interest can align in a way that improves the information available to defenders. In Anthropic’s case, the answer appears to be yes—provided the company continues to preserve uncertainty, publish useful evidence, and allow independent scrutiny.
What Anthropic actually disclosed
Anthropic’s September 10, 2026 report, “Detecting and countering misuse of AI: September 2026,” covers activity identified and disrupted between December 2025 and August 2026. Anthropic says its Threat Intelligence team investigated misuse across seven harm areas, including cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional-weapons development, and model distillation.
The report is significant partly because it does not describe misuse as a single dramatic jailbreak. Its central argument is operational. Actors are incorporating models into workflows that combine reconnaissance, phishing, exploitation, persistence, data extraction, credential use, and evasion. In the cyber section, Anthropic describes Claude moving “from assistant to orchestrator”—a shift from answering an operator’s questions to participating in a structured, multi-stage process.
Anthropic also attempts to measure what it calls “uplift”: the difference AI makes in the speed, scale, and depth of an operation. That is an important analytical choice. The public argument about AI misuse often gets trapped in the question of whether a model can independently invent a novel exploit. Anthropic’s report points toward a more immediate problem: even when a human remains responsible for goals and decisions, an AI system can compress the labor required to search, test, adapt, document, and repeat an attack.
The report says the cases involved Claude Haiku, Sonnet, and Opus. It says none of the misuse cases involved Claude Fable or Mythos-class models except for one illicit-distillation case. That qualification matters. It prevents the report from quietly turning every model in the company’s portfolio into a generalized threat claim, and it gives readers a more precise basis for judging which capabilities were actually involved.
Anthropic describes suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions, and politically motivated individuals. Among the cyber cases, the company reports a suspected Russian state-nexus operation targeting more than twenty organizations, including governments, defense and intelligence bodies, embassies, defense-industrial companies, drone suppliers, a Southeast Asian maritime authority, and a North African government technology authority.
The report also describes financially motivated activity associated with suspected ShinyHunters affiliates. It says actors used stolen credentials and AI-assisted workflows against software providers, airlines, energy-related systems, and downstream customers. In one case, a compromised software-as-a-service provider became a path to approximately two hundred downstream organizations. More than two thousand one hundred Azure Active Directory token sets spanning more than forty corporate tenants were reportedly extracted in roughly thirty-four hours.
These are serious claims, but Anthropic does not present all of them with the same evidentiary status. Some are direct observations by Anthropic’s investigators. Some are assessments. Some are claims made by operators or discovered within compromised environments. A strong threat report must preserve those distinctions, and this one generally does.
Anthropic also makes a disclosure that is easy to overlook but important for institutional credibility: it states that its own systems were not compromised in the customer-key theft operations. That is not a confession of failure, but neither is it a declaration that the company’s ecosystem is invulnerable. The report is about misuse of the service and its downstream effects, not merely an advertisement for Anthropic’s ability to protect itself.
Why the report costs Anthropic something
There is a temptation to treat corporate transparency as cost-free public relations. Sometimes it is. A company can disclose a threat in carefully selected language, emphasize its own response, and frame itself as the hero of the story. That possibility should always be considered.
But the incentives are not entirely comfortable here. A frontier provider that reports real misuse is publicly confirming several unpleasant facts at once.
First, malicious actors are using its products for serious operations. That can concern customers, investors, policymakers, and the public. It gives critics a factual basis for arguing that the provider’s safeguards are imperfect or that advanced models should be deployed more slowly.
Second, a detailed report can reveal something about the company’s detection capabilities. When Anthropic explains what behavioral patterns it identified, how actors attempted to evade safeguards, and what indicators were useful, future adversaries can study the boundaries of the system. Responsible disclosure requires a balance between defender utility and adversary utility. Publishing that balance is itself a risk.
Third, Anthropic’s reports create discoverable statements that can later be examined by customers, regulators, courts, legislators, and journalists. A company that describes a known misuse pattern cannot easily claim later that the category was unforeseeable. Transparency creates accountability because it creates a record.
Fourth, disclosure can make a provider look less safe in the short term. A company that says “we stopped several malicious operations” is also saying “several malicious operations reached our platform and progressed far enough to require investigation.” Competitors can emphasize the first half or the second half depending on their interests.
Finally, admitting that models are becoming more useful to attackers complicates a provider’s commercial message. The same capabilities that make a model valuable for coding, analysis, research, and automation can make it useful for intrusion, fraud, surveillance, or weapons-related work. There is no clean marketing separation between “powerful enough to help customers” and “powerful enough to assist determined adversaries.”
Anthropic is therefore paying a real price for honesty, even if the disclosure also produces strategic benefits.
The strategic benefit does not make the transparency false
It would be naive to describe Anthropic’s behavior as pure institutional self-sacrifice. The company has reasons to publish.
Anthropic is competing partly on trust. Its public identity is closely associated with safety, constitutional principles, responsible scaling, and careful deployment. A credible record of investigating and disclosing misuse supports that positioning. It tells customers that the provider is not merely waiting for outside researchers or law enforcement to discover abuse after the fact.
The report also helps Anthropic shape the language of future regulation. If policymakers accept the company’s categories—AI uplift, model misuse, agentic workflows, safeguards, indicators, coordinated disruption—then Anthropic has helped define what “responsible provider behavior” will mean. That can benefit a company with a strong safety organization and extensive reporting capacity, while imposing costs on smaller or less mature competitors.
That is not necessarily cynical. Corporate self-interest and public benefit can overlap. A pharmaceutical company can profit from a safety standard that protects patients. An aircraft manufacturer can benefit from reporting a defect that affects its own fleet. A cloud provider can gain customer trust by publishing abuse statistics and incident-response lessons. The relevant test is not whether the actor benefits. The relevant test is whether the disclosure is accurate, useful, sufficiently specific, and open to challenge.
Anthropic’s report is strongest when those conditions align. It describes methods defenders can recognize, gives a time window, identifies broad victim classes, distinguishes model families, discusses safeguards, and provides indicators. It also states that Anthropic shared intelligence with authorities and industry partners where appropriate. In other words, the report is not only a narrative about Anthropic’s vigilance. It is an attempt to increase the defensive capacity of people outside Anthropic.
The pattern is persistent, not accidental
The September report is not Anthropic’s first public disclosure. The company has published earlier reports on malicious uses of Claude and, in November 2025, described what it called the first reported AI-orchestrated cyber-espionage campaign. That earlier report argued that models had reached an inflection point in cybersecurity: useful enough to perform a substantial portion of tactical work, while humans retained strategic control.
The persistence matters more than any one report. A one-time disclosure could be dismissed as a publicity event. A series of reports establishes a habit, and habits create expectations. Anthropic is gradually teaching the public to ask frontier providers questions such as:
- What misuse did you detect?
- How long did the activity persist?
- What did the human operators do, and what did the model do?
- What evidence supports your attribution?
- How many victims or downstream organizations were affected?
- What did you change after discovering the activity?
- What indicators can defenders use?
- What remains unknown?
Those questions are much more useful than asking whether a model is simply “safe” or “unsafe.” They convert a vague debate about trust into an incident-reporting discipline.
Anthropic’s own stated purpose reinforces this. The company says its Threat Intelligence team investigates real-world misuse of Claude and works with the broader safeguards organization to improve defenses. It says it publishes regular reports to help other developers recognize similar patterns, give governments and civil society a clearer view of emerging threats, and strengthen collective defenses.
Those are institutional commitments, not merely product claims. They can be evaluated over time. If future reports become less specific, stop acknowledging uncertainty, omit indicators, or report only successes, the public will be able to see the change.
Why the other frontier providers may feel pressure
The competitive effect operates through asymmetry. If every provider remains quiet, silence is normal. If one provider repeatedly publishes substantial disclosures, quiet competitors begin to look different by comparison.
They may not look safer. They may look less observable.
That distinction is increasingly important because model providers are becoming part of the operational environment of governments, companies, software developers, researchers, and security teams. A customer cannot responsibly assess supply-chain risk if the provider reports only uptime, benchmark performance, and broad policy commitments. The customer also needs to know how abuse is detected, how quickly accounts are disrupted, how information is shared, and whether the provider can distinguish isolated misuse from coordinated campaigns.
Anthropic’s reporting therefore creates a reputational benchmark. OpenAI has also published reports describing malicious uses of its models, including state-affiliated influence operations and cyber activity. Microsoft has published threat-intelligence reporting on AI-assisted operations. Google and other major providers have long published threat research and abuse findings. The issue is not that Anthropic is the only company capable of disclosure. The issue is whether persistent, provider-specific misuse reporting becomes a normal part of frontier-model governance.
If that happens, competitive pressure may improve the quality of information available to defenders. Providers may compete over detection speed, quality of indicators, clarity of victim notifications, model-specific safeguards, and the usefulness of post-incident recommendations. That would be a far healthier form of competition than simply claiming that each company’s model is more aligned than the others.
There is also a regulatory effect. Policymakers often legislate after incidents but struggle to define what providers should monitor before an incident becomes public. Regular provider reports can supply concrete categories and operational examples. They can help regulators distinguish between ordinary policy violations, organized criminal use, state-linked campaigns, and genuinely dangerous capability thresholds.
At the same time, regulators should be cautious. A company’s self-report is not an independent investigation. Providers have commercial incentives to emphasize some threats and minimize others. Government agencies, independent researchers, affected customers, and civil society organizations must be able to challenge provider narratives. Transparency should be a floor, not a substitute for external oversight.
Transparency can become a form of collective defense
The strongest argument for Anthropic’s approach is practical rather than reputational. Attackers do not operate inside a single provider’s boundaries. A campaign may use one model for research, another for translation, a third for code generation, public tools for reconnaissance, cloud services for execution, and stolen credentials for access. No provider sees the whole operation.
That means each provider’s internal visibility is incomplete. Anthropic may see suspicious model use but not the final compromise. A cloud provider may see unusual infrastructure but not the prompts that generated the attack plan. A victim may see credential abuse but not know which tools helped the attacker scale. Defenders need shared intelligence to assemble the picture.
Anthropic’s report supports that collective model. The company says it shared intelligence with authorities and industry partners where appropriate, and it publishes indicators in addition to narrative findings. This is exactly the direction that responsible AI security should take: not public disclosure of every sensitive investigative detail, but a structured flow of actionable information to people who can detect, contain, and learn from the activity.
The critical infrastructure implications are substantial. If an AI-assisted campaign moves through a software provider into two hundred downstream organizations, the incident is not simply a model-abuse problem. It is a supply-chain problem. If token sets from dozens of corporate tenants are extracted in hours, the relevant defenses involve identity, session management, service principals, API access, and cross-tenant trust. If a maritime authority, energy-related system, airline, or defense supplier appears in the target set, the question is not merely whether the attacker asked a model for code. The question is whether the provider, customers, and national authorities can share enough information to prevent propagation into operational systems.
This is why Anthropic’s disclosure should be read as part of critical-infrastructure defense rather than as a narrow corporate safety announcement.
The limits and dangers of provider transparency
The case for transparency should not become uncritical admiration. There are at least five limits.
First, the provider sees only what reaches its services or becomes visible through partners. Offline models, competing providers, open-source systems, and conventional tools may be involved elsewhere. Anthropic’s report cannot establish the total amount of AI-assisted malicious activity.
Second, attribution is difficult. A model provider may infer a state nexus from language, infrastructure, targeting, timing, or operational patterns, but those are assessments rather than courtroom proof. Anthropic appropriately uses terms such as “suspected” and “state-nexus.” Those qualifiers must remain visible when the findings are repeated by journalists or policymakers.
Third, victim-impact claims may be incomplete. An operator may claim access that was never achieved. A provider may observe account activity without knowing whether a physical system was affected. The report’s value depends on preserving the distinction between observed model use, attempted intrusion, confirmed access, and confirmed consequence.
Fourth, disclosure can create copycat risk. Detailed case studies may teach defenders what to look for, but they may also teach attackers which behaviors attracted attention. Providers must continually decide how much operational detail is safe to publish.
Fifth, transparency can become performative. A polished report may create the impression of control while leaving customers with little practical ability to protect themselves. The test is whether a security team can take the report and improve detection, credential management, supplier assurance, and incident response.
Anthropic’s reporting will deserve continued respect only if it survives these tests. The company must publish not only dramatic cases but also failures, false positives, unresolved investigations, customer-notification practices, and measurable improvements. It should explain what it cannot see. It should permit independent researchers and affected organizations to contest its account. And it should resist the temptation to convert every threat report into a product advertisement.
The emerging norm may be more important than the individual report
The deepest significance of Anthropic’s disclosure is normative. Frontier providers are no longer ordinary software vendors. Their models can function as research assistants, coding systems, conversational interfaces, automation engines, and components inside larger agentic workflows. When those systems are misused, the provider may be only one participant in a much larger chain—but it is often one of the few participants with visibility across many incidents.
That creates a public responsibility. A provider cannot reasonably claim that misuse is solely the attacker’s problem if the provider has detected recurring patterns, knows that safeguards are being tested, and possesses indicators useful to other defenders. Nor can it disclose everything without considering privacy, investigation, and adversary risk. The responsible path is disciplined transparency: enough information to improve collective defense, enough uncertainty to avoid overclaiming, and enough continuity that the public can evaluate whether the provider is learning.
Anthropic appears to be moving in that direction. Its reports are not perfect, and they are not independent audits. They are also not free. The company is publicly documenting a class of risks that can damage its reputation and complicate its commercial story. Yet the company keeps reporting, and the reports become part of the public record.
That persistence can change the market. Once one major provider repeatedly tells the world what it has seen, what it stopped, and what remains unknown, other providers face a choice. They can publish comparable information, or they can explain why their systems produce no comparable disclosures. Over time, the second position becomes increasingly difficult to defend.
This is how transparency becomes a competitive force. It does not require companies to become altruistic. It requires them to recognize that trust is no longer built only by promising safety. Trust is built by showing the work—especially when the work reveals that the technology is being used in ways the company would rather not advertise.
Editorial conclusion
Anthropic is not entitled to a free pass because it publishes threat reports. We should interrogate its evidence, scrutinize its attribution, ask whether its safeguards actually improve, and insist on outside validation. But we should also acknowledge what the company is doing correctly.
It is treating misuse as an observable security problem rather than an embarrassing public-relations anomaly. It is publishing repeatedly rather than issuing a single defensive statement. It is describing workflows rather than merely listing bad prompts. It is separating model families, time periods, threat actors, and levels of confidence. It is sharing indicators and urging other defenders to learn from its investigations.
That is a meaningful standard.
The other frontier providers should meet it—and improve on it. Governments should use it as a starting point, not an endpoint. Customers should demand it contractually and operationally. Researchers should test it. And the public should understand that transparency is not proof of safety; it is one of the conditions under which safety can be evaluated.
Anthropic may be embarrassing its competitors. If so, that embarrassment is productive. In a field developing systems with consequences that extend into cybersecurity, public information, biological research, critical infrastructure, and military decision-making, the status quo should be embarrassed into becoming more honest.
Sources
Anthropic, “Detecting and countering misuse of AI: September 2026”:
https://www.anthropic.com/threat-intelligence-report-september-2026
Anthropic, Threat Intelligence archive:
https://www.anthropic.com/threat-intelligence
Anthropic, “Disrupting the first reported AI-orchestrated cyber espionage campaign”:
https://www.anthropic.com/news/disrupting-AI-espionage
Anthropic, report PDF for the November 2025 cyber-espionage disclosure:
https://www-cdn.anthropic.com/d7dd50dd1185f59be051b307150d877f2b82bd2c.pdf
OpenAI, “Disrupting malicious uses of our models: an update, February 2025”:
OpenAI, “Disrupting malicious uses of our models,” October 2025 report:
Jonathan Brown writes independent, decision-focused analysis on cybersecurity, infrastructure resilience, and operational risk, with an emphasis on primary-source verification and explicit uncertainty.
Support this work by sharing the briefing with operators who can act on it. Corrections supported by primary evidence are welcomed; material errors should be amended transparently. Feel free to subscribe, comment, or buy us a coffee! Thanks.
© 2026 Border Cyber Group. All rights reserved.
Member discussion: