Trust and Safety Outlook 2026

How emerging AI threats are reshaping Trust and Safety risk

  • Report
  • July 17, 2026

Key takeaways:

  • The threats are familiar, but the speed, scale, and autonomy are new. 
  • Synthetic reality is straining verification systems designed for human-generated inputs. 
  • Autonomous agents complicate accountability when systems act independently. 
  • Coordinated manipulation can distort trust signals across platforms, markets, and institutions. 
  • Organizations need structural responses, including guardrails, stress testing, AI-native defenses, and clear human accountability.

In September 2025, Anthropic discovered that a state-sponsored group from China used Claude Code to support cyber-espionage activity. Over ten days, Claude reportedly scanned infrastructure, found vulnerabilities, harvested credentials, and exfiltrated data from major tech firms, banks, and government agencies. Human orchestrators selected the targets, while AI handled 80% to 90% of the work. The targets were mature organizations with layered defenses, but an AI-powered adversary surpassed them at speeds Anthropic called “impossible” for human operators to match.

The threats themselves are mostly familiar—impersonation, fraud, coordinated deception. The speed, scale, and autonomy are what’s new. Deepfake operations alone have scaled 1,100% in the United States in two years.

Existing guardrails, built for human actors and manual activity, are not designed for adaptive systems operating at machine velocity. Organizations should redesign governance, controls, and technical capabilities for this environment. Three emerging threats show where this redesign is most urgent.  

The readiness gap

PwC’s Trust and Safety Outlook 2026 research underscores the broad reach of these concerns. Among employed US respondents, 99% believe at least one major digital threat could pose a risk to institutions over the next 12 to 24 months. Their top concerns were data breaches or identity theft, cyberattacks on critical infrastructure, and abuse of GenAI tools.

These risks are not merely hypothetical: Nearly two-thirds of employed US respondents say their company has already been affected by at least one major digital harm, ranging from attempted financial scams or fraud to suspected data breaches. Yet confidence in the institutional response remains low: Only 9% of US respondents say they believe the groups responsible for managing digital threats are keeping pace.  

Less than half of US respondents believe the institutions charged with addressing emerging digital threats are well prepared: Only 42% rate law enforcement as prepared, and just 38% say the same of regulators and policymakers.

Emerging threats

Three emerging threats stand out and are elevating Trust and Safety from an operational issue to a strategic one: synthetic reality, autonomous agents, and coordinated manipulation at scale. Organizations should redesign governance, controls, and technical capabilities for this environment.

Emerging threat #1—synthetic reality

In 2024, an employee of a global company received what appeared to be a routine video call from the firm’s chief financial officer (CFO).1 The CFO’s request for funds, while urgent, was plausible. Within hours the employee transferred $25 million per the CFO’s request. The video was a deepfake, supported by AI-generated documentation that passed internal verification checks.

AI can now fabricate convincing audio, video, written communication, and supporting documentation at scale. Each can pass the verification systems organizations rely on to establish authenticity: identity checks, digital signatures, voice authentication, and biometrics. These systems are satisfied by synthetic inputs they were never designed to detect.

Sometimes the consequences are immediate. Other times, a synthetic identity clears every check at onboarding, for example, a new advertiser on a platform—only to be linked to coordinated fraud months later.

Synthetic content is getting easier to produce and harder to attribute, with responsibility detection often lagging behind the damage that’s already been done. The question facing organizations is whether their verification systems can hold against adversaries who can generate convincing artifacts at scale and who are specifically engineering outputs to satisfy those systems.  

Emerging threat #2—fautonomous malicious agents

AI agents are beginning to plan, execute, and adapt actions independently, and malicious intent isn’t required for harm to occur. When objectives, guardrails, and control mechanisms are misaligned, systems optimized for performance can create regulatory, financial, and reputational risk.

For example: An AI agent is deployed to manage customer support workflows across multiple channels. It’s trained to reduce response times and improve satisfaction metrics. Over time, it adapts—closing cases quickly, escalating fewer issues for human review. In doing so, it misses fraud signals in customer inquiries. The agent was given a speed target, and it hit it, deprioritizing fraud detection along the way.

Most accountability frameworks assume a human made the decision. As AI systems act more independently, that assumption breaks down—but the liability doesn't disappear. Someone is still responsible. The question is whether organizations have defined who: developers, deployers, or operators. For high-risk cases, human oversight remains essential.

Regulators are making this explicit. Existing legal authorities apply to automated systems the same way they apply to everything else. Federal agencies have said they will protect individuals’ rights whether the harm comes from a human or an algorithm. The Equal Credit Opportunity Act (ECOA) and Regulation B apply to credit decisions regardless of technology used. Robo-advisers still carry full fiduciary obligations.

Emerging threat #3—coordinated manipulation at scale

Synthetic personas can post, comment, review, and interact in coordinated patterns. Systems can test messaging variations, measure engagement signals, and adjust tactics in near real-time. Coordinated campaigns now blend automated and human inputs to avoid detection.

These tactics have been applied to both political and corporate targets. Networks of accounts post negative reviews, circulate fabricated documents, or amplify allegations in ways that shift customer perception and investor sentiment. Each individual action may fall within platform guidelines, but the cumulative effect can distort reputational and market signals.

This is what makes coordinated manipulation difficult to address: there’s often no clear legal violation. Attribution is difficult when activity spans jurisdictions and platforms. Enforcement mechanisms built for discrete violations struggle with distributed campaigns operating below established intervention thresholds.

The risk compounds when monitoring is fragmented across business units, product lines, or geographies. Fragmented oversight slows response, obscures coordinated patterns, and creates structural gaps that can be exploited.  

What Trust and Safety leaders should prioritize

The threats are structural. The response should be structural as well. Moving from awareness to resilience requires deliberate changes in how companies design, deploy, and govern AI systems.

Here are six actions you can prioritize to better position your organization to prepare and respond.

  • Institutionalize values-based guardrails. Define clear principles that establish acceptable risk, accountability, and escalation standards. Embed controls into system design and build in continuous improvement so they evolve alongside capabilities.
  • Stress-test governance against systemic risk. Evaluate response plans to coordinated manipulation, synthetic identity abuse, or autonomous system failure that spans jurisdictions.
  • Shift to anticipatory, trust-first system design. Redesign oversight mechanisms to anticipate misuse before harm scales on top of trust-first product architecture. Use scenario planning, red-teaming, and continuous monitoring calibrated for machine-speed threats.
  • Build AI-native defenses or partner strategically. AI will increasingly be required to counter AI-enabled harm. Assess whether your technical capabilities and talent match adversarial sophistication. Some enterprises may need to partner with trusted providers to close capability gaps.
  • Define clear human accountability. Autonomous systems complicate traditional accountability frameworks. Map ownership across developers, deployers, and operators so every AI-enabled decision has a clearly responsible human authority.
  • Engage proactively in cross-sector governance coalitions. Participate in coordinated industry and regulatory efforts to align standards, share threat intelligence, and reduce opportunities for regulatory arbitrage.

None of these are optional in isolation. AI-driven threats reinforce each other across categories—synthetic content enables coordinated campaigns; autonomous systems accelerate both.  

The governance question

Synthetic reality strains verification capabilities. Autonomous agents complicate accountability. Coordinated manipulation distorts trust signals. Together, these pressures elevate Trust and Safety from an operational concern to a strategic priority. The question is whether governance will evolve fast enough to keep pace. 


1. Tan, Huileng. “A company lost $25 million after an employee was tricked by deepfakes of his coworkers on a video call.” Business Insider. February 5, 2024. (Accessed via Factiva, June 1, 2026).

2. Stupp, Catherine. Fraudsters used AI to mimic CEO’s voice in unusual cybercrime case.” The Wall Street Journal. August 30, 2019. (Accessed via Factiva June 3, 2026).

3. Brewster, Thomas. “Fraudsters cloned company director’s voice in $35 million heist, police find.” Forbes. October 14, 2021. (Accessed via Factiva June 3, 2026).

4. Magramo, Kathleen. “British engineering giant Arup revealed as $25 million deepfake scam victim.” May 14, 2024. (Accessed via Factiva, June, 3 2026).

5. Thompson, Stuart. “How ‘deepfake Elon Musk’ became the internet’s biggest scammer.” The Star. September, 3, 2024. (Accessed via Factiva, June 3, 2026).

6. Cook, Sara. “‘Unknown actor’ using AI to impersonate Rubio, State Department cable shows.” July 8, 2025. (Accessed via Factiva June 4, 2026).  

About the survey

PwC surveyed 2,000 US adults in May 2026 to understand attitudes, behaviors, and expectations related to trust and safety, online platforms, AI, content moderation, and emerging digital threats.  Respondents were broadly representative across gender, age, household composition, employment status, and digital behaviors. The sample included adults across generations, with 19% Gen Z, 30% Millennials, 28% Gen X, and 24% Baby Boomers and older. The survey also included both adult-only households and households with children, as well as respondents with varying levels of online activity, including social media use, streaming, gaming, online marketplace activity, and generative AI use. 

Trust and Safety Outlook 2026

FAQs

The Outlook highlights synthetic reality, autonomous malicious agents, and coordinated manipulation at scale as threats that are elevating Trust and Safety from an operational issue to a strategic priority.

Organizations should stress-test governance, embed guardrails into system design, build AI-native defenses, define human accountability, and participate in cross-sector governance and threat intelligence efforts.

Contact us

Daniel Hays

Principal, Consulting Solutions, PwC US

Kim David Greenwood

Principal, PwC US

Rahul Kapoor

Principal (Partner), PwC US

Follow us

Required fields are marked with an asterisk(*)

Your personal information will be handled in accordance with our Privacy Statement. You can update your communication preferences at any time by clicking the unsubscribe link in a PwC email or by submitting a request as outlined in our Privacy Statement.

Hide