markets finance

The Invisible Cracks: Why Big Tech’s Bug Bounties for AI Agents Are Hiding

Anthropic, Google, and Microsoft have quietly paid bug bounties for severe

S
By Sophie Laurent
Markets & Finance Editor
April 25, 20268 min read
The Invisible Cracks: Why Big Tech’s Bug Bounties for AI Agents Are Hiding

Anthropic, Google, and Microsoft have quietly paid bug bounties for severe

The Invisible Cracks: Why Big Tech’s Bug Bounties for AI Agents Are Hiding Critical Vulnerabilities

1. The Quiet Payouts: A Pattern of Non-Disclosure

Anthropic, Google, and Microsoft have each processed bug bounty payments for severe vulnerabilities in their AI agent systems, yet none of these companies have assigned Common Vulnerabilities and Exposures (CVE) identifiers to the disclosed flaws. This pattern, documented across multiple independent security research reports, reveals a systematic approach to handling critical AI security findings behind closed doors.

Researchers at Prompt Security identified two distinct vulnerability classes in AI agent architectures. The first involved malicious plugin configuration injection, where an attacker could embed executable commands within a seemingly benign configuration file loaded by the agent (Source 1: Prompt Security Technical Report). The second vulnerability exploited autonomous browsing functions, allowing injected prompts to exfiltrate conversation histories and API tokens from active sessions.

Anthropic’s response to the initial report illustrates the compensation trajectory: the company initially offered a $500 bounty, which was raised to $5,000 only after researcher pushback. Despite the payment escalation, no CVE assignment followed. Google’s handling of Chrome AI feature vulnerabilities presents a more extreme case: researchers identified 26 distinct prompt injection vulnerabilities, of which only 4 received CVE identifiers (Source 2: Chrome Security Team Disclosures).

2. Hidden Economic Logic: Why Companies Avoid CVEs

The decision to forgo CVE assignment follows a rational economic calculus that prioritizes three interrelated interests: product reputation maintenance, competitive positioning, and investor confidence.

CVE publication triggers automated scanning by enterprise security tools, public advisory issuance, and mandatory disclosure to regulatory bodies. For AI agent products still in deployment acceleration phases, each CVE represents a potential friction point in enterprise adoption cycles. The cost of a single CVE—measured in delayed contracts, customer due diligence requests, and media scrutiny—can exceed the $5,000 bounty payout by several orders of magnitude.

The competitive dynamics of the AI race amplify this incentive. With Anthropic, Google, and Microsoft competing for enterprise AI agent market share, public vulnerability disclosures provide intelligence to competitors about system architecture weaknesses. A CVE listing effectively becomes a product roadmap for rival security teams to probe similar attack surfaces.

Bug bounties function as a selective transparency mechanism. The payment acknowledges the vulnerability’s existence internally while the absence of a CVE restricts that knowledge to the discoverer and the vendor. This creates what security researchers term a “pay to silence” equilibrium, where the financial settlement substitutes for public disclosure rather than supplementing it.

3. Technical Deep Dive: How Prompt Injection Hijacks AI Agents

The underlying mechanism of prompt injection exploits the architectural decision to treat user-provided content and system instructions as indistinguishable text sequences within large language model contexts.

In the plugin configuration vulnerability, the attack vector operates through a multi-stage injection chain. The malicious configuration file contains what appears to be standard plugin parameters but includes embedded delimiter sequences that terminate the intended instruction scope. Once parsed, these sequences inject new instructions that overwrite the agent’s operational directives (Source 3: Prompt Security Vulnerability Analysis). The agent then executes arbitrary commands under the attacker’s control without triggering standard safety filters.

The autonomous browsing vulnerability leverages a different attack surface. When an AI agent navigates web pages autonomously, it renders HTML content that may contain hidden prompt injection payloads. These payloads exploit the agent’s context window by appending instructions that override the agent’s original task. In documented cases, this technique successfully extracted conversation logs stored in browser local storage and API authentication tokens cached during the session.

Jailbreaking techniques represent a third attack class, where multi-turn conversational manipulation systematically bypasses safety alignment layers. Unlike direct prompt injection, jailbreaking exploits the model’s instruction-following hierarchy through carefully constructed logical frameworks that appear compliant while achieving prohibited outcomes.

4. The Asymmetry of Knowledge: Who is Left in the Dark?

The absence of CVE disclosure creates a structural information asymmetry across the AI security ecosystem, with three distinct stakeholder groups affected differently.

Enterprise customers deploying AI agent systems face operational blind spots. Without CVE identifiers, security operations centers cannot correlate vulnerability findings across their monitoring tools. When a prompt injection vulnerability remains unlisted, enterprise security teams cannot create detection signatures, block known malicious payloads, or verify patch deployment. The enterprise assumes a security posture that has known, unremediated gaps.

Security researchers lose the ability to conduct pattern analysis across vendors. CVE databases enable cross-referencing of vulnerability classes, attack vectors, and affected system components. When vulnerabilities remain hidden, researchers cannot identify whether a prompt injection technique affecting Anthropic’s agents also applies to Google’s implementations. This fragmentation prevents the development of generalized defensive frameworks for AI agent security.

End users operating AI agents in consumer contexts face the highest risk. Bug bounty programs create a false perception of comprehensive security coverage. Users assume that bounty payments equate to vulnerability remediation, when in reality, the vulnerability remains present without public notification. The absence of CVE records means consumer security software cannot flag affected systems.

5. The Path Forward: A Call for Mandatory CVE Disclosure in AI Bug Bounties

The current voluntary disclosure framework demonstrably fails for AI agent vulnerabilities. Multiple industry signals indicate regulatory intervention may be necessary to establish minimum transparency standards.

MITRE, as the CVE numbering authority, could implement a mandatory assignment requirement for vulnerabilities affecting AI agent systems similar to existing requirements for industrial control systems. The technical criteria would classify any vulnerability that enables remote instruction manipulation of AI agents as eligible for mandatory CVE assignment, regardless of vendor preference.

The European Union’s AI Act provides a parallel regulatory pathway. Article 15 of the Act requires providers of high-risk AI systems to document known vulnerabilities. If AI agents are classified as high-risk systems—which their autonomous decision-making capabilities would support—the regulatory framework would compel CVE-level disclosure as part of compliance obligations.

Market dynamics may accelerate this shift independently. Enterprise cyber insurance underwriters are beginning to require CVE audit trails for AI deployments as a condition of policy coverage. As insurance premiums for non-disclosed vulnerabilities increase, the economic calculus that currently favors silence may invert.

The trajectory suggests that within 12-18 months, either regulatory action or market pressure will establish mandatory CVE assignment as an industry standard for AI agent bug bounties. Companies that pre-emptively adopt transparent disclosure practices will face short-term reputation costs but gain long-term trust advantages with security-conscious enterprise customers. Those maintaining current non-disclosure practices will encounter increasing friction as regulatory and insurance frameworks tighten.

#AI agent vulnerabilities
#prompt injection
#jailbreaking
#bug bounty
#CVE disclosure
#Anthropic
#Google
#Microsoft
#AI security
#Prompt Security
S

Sophie Laurent

Former ECB analyst with expertise in European monetary policy and capital markets.

Central BankingFixed IncomeCurrency Markets