Trust and Safety Outlook 2026

The future of content moderation—from scale to specialization

  • Report
  • July 17, 2026

Key takeaways:

  • Automation is reducing routine moderation but increasing the importance of complex, high-consequence work. 
  • Cross-product harms, regulatory fragmentation, physical AI, and GenAI are raising the complexity of Trust and Safety operations. 
  • The supplier ecosystem needs more specialized capabilities, including adversarial testing, harm intelligence, and AI governance. 
  • Trust and Safety is shifting from moderation engine to intelligence layer.

For Trust and Safety leaders, the next phase of content moderation won’t be about handling more volume. It’ll be about handling more complexity—and whether the partner ecosystem can keep up with evolving requirements.

Historically, Trust and Safety operations scaled with platform growth: more users, more content, more human reviewers. AI is rewiring that relationship.

Automated moderation systems are already capable of enforcing most baseline content moderation decisions. According to an analysis of the EU’s Digital Services Act Transparency Database, 97% of potentially violating content detections across major platforms are automated, and more than half of all decisions to remove content are made fully automatically. Front-line moderation volume will plateau—and eventually decline—even as platforms continue to grow.  

The forces reshaping Trust and Safety

However, the rise of automated moderation doesn’t eliminate risk. It changes the nature of the work that remains for manual review. As routine moderation becomes increasingly automated, four forces are simultaneously increasing the complexity of Trust and Safety operations.

As GenAI becomes embedded across products and platforms, harms increasingly span multiple product areas. A copyright issue originating in a GenAI assistant, for example, may surface later in advertising, productivity, or search products as part of integration efforts. Legacy Trust and Safety frameworks built for content have no established methodology for assessing or governing accountability for these rapidly expanding scopes where product deployments are already happening.

Managing these risks requires a tightly integrated ecosystem wide approach to Trust and Safety, alongside deeper domain expertise. Tools like Google's MedGemma, now supporting health applications in radiology, pathology, and electronic health record analysis1, represent a new class of frontier AI where the stakes of cross-product harm extend beyond content to patient safety, liability, and regulatory compliance. As purpose-built clinician models and similar domain-led products become widely integrated in broader product ecosystems, Trust and Safety functions should evolve to meet harms that legacy frameworks were not designed to assess. Effectively evaluating these risks increasingly requires specialized expertise in domains such as healthcare, finance, and law.  

Rules are multiplying and diverging by jurisdiction, age group, and, increasingly, by AI type. The EU’s Digital Services Act (DSA), emerging AI-specific laws, and a growing number of legal challenges involving AI chatbot interactions and youth safety—including cases involving OpenAI2, Google, and Character.AI3—are prompting platforms to reevaluate how they manage risk across jurisdictions.

The scale and speed of potential harms created by GenAI require AI-powered moderation, testing, and risk-detection capabilities. Organizations increasingly need AI to facilitate work to identify emerging risks, uncover vulnerabilities, and monitor AI systems at scale.

In April 2026, Anthropic's Claude Mythos Preview model identified more than 10,000 high- or critical-severity vulnerabilities across major software systems—a pace that led Anthropic to withhold the model from public release and instead deploy it defensively through Project Glasswing with approximately 50 partner organizations.4 This example demonstrates that certain classes of AI-generated risks can emerge at a scale and speed that humans cannot realistically evaluate on their own. As AI capabilities advance, organizations will increasingly rely on automated testing, risk discovery, and detection systems.  

GenAI is evolving from systems that generate content to systems that reason, take actions, and eventually interact with the physical world. Each shift expands the scope of Trust and Safety oversight, from content moderation to behavioral and real-world risk management. The pace of that expansion is accelerating faster than oversight frameworks can keep up. What was once a frontier concern is now a reality—with the physical AI surge, general-purpose models are being deployed in autonomous vehicles, humanoid robots on factory floors, consumer wearables, and clinical environments.

Trust and Safety now encompasses more than just traditional content moderation problems: Governance is required for novel cases, such as a malfunctioning autonomous vehicle, a compromised factory robot, or a clinical AI giving dangerous guidance—potentially creating harms that are immediate, physical, and irreversible.

This marks a structural shift. As automation expands, high-volume, low-complexity moderation will decline. At the same time, work requiring greater judgment, context, and specialized expertise is increasing, including edge cases, appeals, adversarial testing, and model oversight. Human decisions will become less frequent, but far more consequential.

This hypothesis has direct implications for how Trust and Safety work is sourced, structured, and delivered.  

The changing shape of Trust and Safety work

Historically, Trust and Safety has been anchored in front-line content moderation: high-volume, rules-based review delivered through large partner ecosystems running standardized workflows. That model was built for consistency and scale. Automation is now reshaping both the composition of moderation work and the role Trust and Safety plays within organizations.

Now

Front-line content moderation remains the primary driver of work, supported by large human teams executing standardized policy decisions. While AI-assisted moderation is becoming more common, complex content moderation—including GenAI-related decisions, novel harms, and policy escalations—represents a relatively small share of overall activity.

Emerging

Evaluator agents and agentic systems are increasingly absorbing front-line moderation work. At the same time, the growth of GenAI-generated content, multimodal experiences, and increasingly personalized products is expanding demand for more complex forms of moderation.

Frontier

As AI becomes embedded across products and services, overall content moderation volumes rise, and complex moderation becomes the dominant workload. The role of Trust and Safety shifts from reviewing large volumes of content to managing novel risks, shaping governance frameworks, calibrating AI systems, and providing strategic input into product and policy decisions. Human decisions become fewer in number but significantly higher in impact.

What this means for the Trust and Safety function

As the mix of content moderation work changes, so does the role for Trust and Safety. Headcount and budget historically dedicated to high-volume review can be redeployed toward higher-value work earlier in the life cycle:

From  To Theme Example
Front-line content moderation

Complex content moderation

 

Growth of complex and emerging harms tied to new product experience 
  • Multimodal and agentic content: Assessing content that combines text, images, audio, video, and agentic actions, often requiring cross-modal context to make decisions
  • Personalized experiences: Reviewing user-specific outputs and interactions that may create risks not visible through standardized testing
Redeployment of human effort toward edge cases requiring expertise
  • Gray area cases: considering policy ambiguities, edge cases, novel trends
  • Appeals: requests for secondary review that signal need for additional judgement
Service provider Strategic partner Trust and Safety moves beyond workflow execution to influencing product, policy, and risk decisions
  • Edge-case calibration: Refining policies and AI systems based on difficult or ambiguous decisions
  • Emerging harms research: Identifying new risk patterns, abuse vectors, and behaviors
  • Model oversight: Monitoring AI performance, reviewing failures, and supporting continuous improvement 
  • Governance support: Advising on risk frameworks, regulatory compliance, and responsible AI practices
Transactional Value-added services Resources shift from moderation execution to capabilities that improve products, models, and risk management earlier in the life cycle
  • Expert annotations and ground truth data sets: Creating high-quality training and evaluation data for AI systems
  • Analytics and insights: Translating moderation activity into actionable business and risk intelligence
  • Threat intelligence: Monitoring emerging threats, adversarial behavior, and abuse trends
  • Engineering and infrastructure support: Building testing environments, evaluation frameworks, and operational tooling for AI oversight

The human role comes in when judgment, expertise, and context are required—and specifically where automation alone cannot deliver reliable outcomes.

Rebuilding the supplier ecosystem

For years, the dominant model for scaled Trust and Safety operations relied heavily on large-scale business process outsourcing (BPO) providers charged with delivering standardized moderation services through time-and-materials contracts.

That model worked in a world where moderation work was predictable, high-volume, and operationally uniform. The emerging landscape requires a different approach. Respondents to PwC’s Trust and Safety Outlook 2026 research consistently ranked better detection and faster response times as their top investment priorities for content moderation. Meeting both is becoming the baseline expectation for platforms. Legacy Trust and Safety delivery models are not equipped to meet these needs.

The pressure to detect faster is only one part of the challenge. Future Trust and Safety operations will face additional complexities from highly variable demand profiles, including:

  • Always-on oversight for automated moderation systems
  • Capacity spikes during product launches or emerging harm events
  • Cyclical demand tied to platform growth or seasonal activity
  • Unpredictable demand driven by adaptive adversarial behavior

The evolving demand patterns can’t be addressed through workforce scale alone, and while AI-first content moderation offers significant opportunity to improve efficiency, it introduces its own set of operational challenges. Human expertise will remain essential for complex judgment calls, escalation management, and oversight.

While AI can help meet increases in moderation demand, running AI at scale does not automatically result in lower operating costs than deploying human reviewers. Microsoft recently reversed course on internal AI tool access after costs became unsustainable, and Uber burned through its 2026 AI coding tools budget in just four months.5 This creates a direct prioritization challenge: the delivery model decision is about both cost and risk.  

A new capability portfolio

Supporting these dynamics requires a more diversified set of capabilities. Rather than relying on a single delivery model, organizations will increasingly draw on a broader mix of specialized capabilities, including:

Capability What it does
AI-native simulation and evaluation services Building and maintaining classifiers, LLM-based review systems, and automated testing pipelines for routine moderation at scale, including ongoing evaluation loops to keep them current as content patterns shift 
Agentic safety tech and infrastructure Creating sandboxes for agentic system testing before deployment, integration layers connecting moderation tools to product surfaces, and tooling for human reviewers to interact efficiently with automated queues
Data services (e.g. “anydata”) Curating training and evaluation data sets—sourcing and labeling edge-case content, generating synthetic examples for underrepresented harm categories, and localizing data sets for cultural and linguistic context 
Adversarial stress testing Dedicated teams probing AI systems the way bad actors would—testing for jailbreaks, checking whether policy boundaries hold under creative circumvention, and identifying failure modes before they reach production
Domain specialists Subject-matter specialists (e.g. pharmacologists, financial compliance professionals, legal experts, cybersecurity analysts) who evaluate whether AI outputs cross real-world professional thresholds and build gold-standard evaluation data sets 
Harm intelligence Analysts who specialize in recognizing emergent patterns of misuse, new manipulation vectors, and harms that arise from AI capability jumps rather than established content categories
Edge case calibration Ongoing refinement of policies and AI systems based on difficult or ambiguous decisions—translating hard cases into clearer rules and better-trained models 
Analytics and intelligence Translating moderation activity into actionable business and risk intelligence, surfacing patterns that inform product, policy, and investment decisions
Agentic observability and calibration Monitoring AI system performance over time, reviewing failures, and supporting continuous improvement cycles 
AI governance and assurance Tracing and documenting decisions made across agentic workflows to specific decision points, helping organizations meet regulatory reporting requirements and demonstrate accountability when automated decisions are called into question

Capability evolution changes how work gets delivered

As the capability portfolio diversifies, the models through which work is executed evolve. The traditional assumption that Trust and Safety work is delivered by human reviewers operating standardized workflows is giving way to a more differentiated set of delivery archetypes. Organizations will increasingly need to match the right delivery model to the right type of work:

Archetype Description Example
Human only  Expert judgment applied without AI assistance, reserved for the high-severity decisions where accountability, nuance, or domain makes automation difficult Lawyers constructing policy data sets, regulatory specialists providing testimony, or senior specialists deciding on novel harm categories with no precedent
AI-first, expert-in-the-loop AI handles the initial decision; a human expert reviews flagged, uncertain, or high-risk cases in real time before action is taken  Evaluator agent-led moderation with specialist review of novel harms, appeals, and regulatory-sensitive content
AI-first, expert-over-the-loop AI operates autonomously at scale; a human expert audits outputs periodically, identifies drift or failure patterns, and recalibrates the system to facilitate ongoing AI observability and accountability  Evaluator agent running at scale with monthly performance review by specialists looking for patterns where the AI is getting things wrong, drifting from policy intent, or struggling with emerging content trends. 
Fully automated  AI makes and executes enforcement and annotation decisions end-to-end; human oversight at system level only Baseline policy enforcement on high-volume, low-complexity content (e.g. spam, known CSAM hashes)

The new risk—capability gaps

A major challenge facing organizations is that Trust and Safety work is evolving faster than the supplier ecosystem. As organizations face more complex risks, the specialized capabilities required to manage them remain scarce.

In PwC’s Trust and Safety Outlook 2026 research, 85% of US respondents say they believe new technologies and harms are being created faster than organizations can respond.

This highlights the urgency of building the specialized capabilities needed for the next generation of Trust and Safety. Those that delay may find themselves constrained by legacy operating models and limited supplier flexibility.

As a result, the next few years will determine whether Trust and Safety ecosystems evolve in step with AI or struggle to keep pace with it.  

What does this mean for Trust and Safety leaders and the path forward?

The opportunity for Trust and Safety leaders is to turn their work from a defensive backstop into a strategic, and increasingly offensive capability. Offensive here means proactive—identifying and closing attack surfaces before bad actors find them, rather than responding after harm reaches users.

The organizations that succeed in this new environment can:

  • Embrace automation while continuing to invest in human expertise, particularly for high-risk or complex cases
  • Design flexible supplier ecosystems rather than relying on single delivery models
  • Experiment early with emerging capabilities, running pilots with specialized partners before the need becomes urgent
  • Prioritize protection for vulnerable users (e.g. minors)
  • Invest early in physical and multimodal safety testing capabilities, particularly for robotics and physical AI, where harm vectors are still being identified and evaluation frameworks don’t yet exist
  • Create systems for accountability and transparency to facilitate safe deployment of new capabilities before scaling
  • Embed Trust and Safety earlier in the AI development life cycle, transforming it from an operational function into a strategic capability

The shift from moderation to intelligence is already underway. The question now is not whether Trust and Safety will evolve, but how quickly your organization can adapt your operating model, supplier network, and talent base to meet the complexity ahead.  


1. Diana Novak Jones, "OpenAI sued over chatbot advice linked to fatal overdose.” Reuters, May 12, 2026. (Accessed via Factiva, June 1, 2026).

2. Nitasha Tiku, "Google and chatbot start-up Character move to settle teen suicide lawsuits," The Washington Post, January 8, 2026. (Accessed via Factiva, June 1, 2026).

3. Anthropic, "Project Glasswing: An initial update." Anthropic, May 2026. (Available at: anthropic.com/research/glasswing-initial-update).

4. Google Research, "Next-generation medical image interpretation with MedGemma 1.5 and medical speech-to-text with MedASR." Google Research Blog, January 13, 2026. Available at: (research.google/blog/next-generation-medical-image-interpretation-with-medgemma-15-and-medical-speech-to-text-with-medasr).

5. Jake Angelo, "Microsoft reports are exposing AI's real cost problem: Using the tech is more expensive than paying human employees." Fortune, May 22, 2026. (Accessed via Factiva, June 10, 2026).  

Trust and Safety Outlook 2026

FAQs

Routine moderation is becoming more automated, while complex cases increasingly require specialized judgment, domain expertise, adversarial testing, and oversight of AI systems.

Specialization means moving beyond large-scale review toward capabilities such as harm intelligence, model oversight, edge-case calibration, domain expertise, and AI governance support.

Contact us

Daniel Hays

Principal, Consulting Solutions, PwC US

Kim David Greenwood

Principal, PwC US

Rahul Kapoor

Principal (Partner), PwC US

Follow us

Required fields are marked with an asterisk(*)

Your personal information will be handled in accordance with our Privacy Statement. You can update your communication preferences at any time by clicking the unsubscribe link in a PwC email or by submitting a request as outlined in our Privacy Statement.

Hide