CineMind

Manual or Automated Red Teaming: Where Current Trends Are Taking Us

Last updated: 10/4/2026

Back to blog
Sven Lindqvist avatarSven Lindqvist 8 min read
Cover image for Manual or Automated Red Teaming: Where Current Trends Are Taking Us
AI-assisted, human-reviewed. Drafted with AI research tools from public sources and edited by our team. How we build these โ†’

Red teaming is adversary emulation: a team behaves like a real attacker against a live organisation to test whether defenders can detect, respond and recover. For two decades it has been irreducibly manual โ€” creative humans against human defenders. Now a layer of automation sits underneath it: breach and attack simulation platforms, continuous security validation, open-source emulation frameworks, and increasingly autonomous AI agents that chain exploits on their own. This article maps where the discipline actually stands, what current evidence says about each approach, and where the trend lines point.

Thesis: Automation is absorbing the repeatable layers of red teaming โ€” technique execution, control validation, regression testing. Human red teams are moving up the stack toward objective-driven, multi-domain campaigns that test decisions and business consequences, not just controls.

Definitions, because the words are used loosely

Related but distinct practices
PracticeGoalTypical durationPrimary output
Vulnerability scanningFind known flawsContinuousPrioritised flaw list
Penetration testingExploit flaws in a defined scope1โ€“4 weeksFindings report
Breach & attack simulation (BAS)Validate controls against known techniquesContinuousControl coverage dashboard
Red teamingAchieve objectives like a real adversary4โ€“12 weeksNarrative of attack paths + detection gaps
Purple teamingRed and blue working together to improve detectionDays to weeksNew and tuned detections
Threat-led pen testing (TIBER-EU / CBEST)Regulated, intelligence-driven red team3โ€“6 monthsRegulator-recognised assessment

Where manual red teaming is unbeatable

Three capabilities remain stubbornly human.

1. Objective-driven creativity

A real adversary does not enumerate a technique list; it wants something โ€” payment fraud, intellectual property, disruption. Human red teams invent paths that no catalogue contains: abusing a business process, chaining a low-severity misconfiguration with a help-desk procedure, exploiting a trust relationship with a supplier.

2. Social and physical dimensions

Phishing with context, pretext phone calls, tailgating into a building, dropping a device on a network. Verizon's 2024 DBIR found the human element involved in roughly 68% of breaches. Automation can send phishing emails; it cannot improvise with a receptionist.

3. Testing decisions, not controls

The most valuable red team finding is often organisational: nobody escalated, the on-call runbook was wrong, legal and comms disagreed for six hours. That only surfaces when real people are surprised in real time.

ManualCreative attack pathsSocial engineeringPhysical accessAutomatedBAS platformsCALDERA / AtomicContinuous validationHybridPurple team sprintsHuman plans, bot executesAgent-assisted reconFrameworksMITRE ATT&CKTIBER-EUCBEST / NISTMetricsDetection rateTime to detectCoverageRisksProd impactUnsafe AI agentsCheckbox cultureRed teaming
Mind map: the red teaming landscape

Where automation already wins

Automated adversary emulation grew out of MITRE ATT&CK, published in 2013 and now the common language of the field. Open-source tools โ€” MITRE CALDERA, Red Canary's Atomic Red Team, Prelude and others โ€” let teams execute individual techniques on demand. Commercial BAS platforms do the same at enterprise scale with safety rails and reporting.

The clearest value is regression testing for detections. A detection that worked in March can silently break in June after an agent upgrade, a log-source change or a cloud migration. Running the same emulated techniques weekly catches that drift. No human team can afford to re-run 300 techniques every week; a platform can.

Manual versus automated red teaming
DimensionManual red teamAutomated emulation / BAS
CreativityHigh โ€” invents novel pathsLow โ€” executes known techniques
FrequencyA few times per yearDaily or weekly
Coverage breadthNarrow, deepBroad, shallow
Cost per runHighLow after setup
RepeatabilityLowVery high
Detects control driftOnly by luckBy design
Tests people and processYesBarely
Safety in productionManaged by skilled operatorsManaged by platform rails
Best metric producedAttack-path narrativeTechnique coverage over time
Illustrative annual technique executions by approach (example model)
Manual red team engagements4
Purple team sprints24
Automated emulation runs520

The new variable: autonomous offensive AI

Two developments changed the conversation in 2024โ€“2025.

First, research evidence that language models can do real offensive work. Google's Project Zero and DeepMind reported in late 2024 that their Big Sleep agent found a previously unknown memory-safety vulnerability in SQLite โ€” an AI discovering a real zero-day in widely used software. Anthropic, OpenAI and others have published evaluations tracking model capability in cyber tasks, and both have added usage policies and monitoring specifically for offensive misuse.

Second, evidence of commercial impact: in 2025 an autonomous hacking system reached the top of a major bug bounty leaderboard, demonstrating that agentic pipelines can out-produce most human hunters at volume on web targets.

At the same time, structural limits remain. Agents are strong at breadth โ€” enumerate, fuzz, try known patterns โ€” and weak at the judgement-heavy parts: choosing which objective matters, knowing when to stop to stay undetected, and improvising around unexpected business logic.

What offensive AI agents can and cannot do well today
TaskAgent capabilityEvidence / caveat
Reconnaissance and asset discoveryStrongMostly automation of existing tooling
Known-pattern web vulnerability discoveryStrong and improvingDemonstrated at bounty-platform scale
Novel memory-safety bug discoveryEmergingBig Sleep SQLite finding, 2024
Exploit chaining across systemsMixedWorks in labs; brittle in messy estates
Evading mature detectionWeakNoisy; poor sense of operational security
Social engineering with contextPartialCan draft lures; cannot improvise live
Objective selection and risk judgementWeakRequires business understanding
TechnologyAgentic offenseAI-native defencesStandard tool APIsRegulationDORA / TIBER-EUEU AI ActSector mandatesEconomicsScarce senior talentCheap computeContinuous assuranceAttack surfaceCloud + SaaS sprawlIdentity as perimeterAI systems themselvesPracticePurple as defaultObjective-led campaignsEvidence-based metricsRisksCheckbox automationUnsafe agentsSkill erosionRed teaming 2030
Mind map: forces shaping the next five years

Regulation is pushing the frequency up

Threat-led red teaming is no longer optional in parts of the economy. The European Central Bank's TIBER-EU framework (2018) and the UK's CBEST established intelligence-led testing for financial institutions. The EU's Digital Operational Resilience Act (DORA), applying from 17 January 2025, requires threat-led penetration testing for significant financial entities โ€” in practice at least every three years, performed by qualified testers. Meanwhile, the EU AI Act requires adversarial testing of general-purpose AI models with systemic risk, creating an entirely new red teaming specialism aimed at models rather than networks.

The practical effect: more mandated manual engagements at the top end, and a growing need for cheap continuous validation in between to keep the posture stable between assessments.

A realistic forecast

  1. The middle hollows out. Routine technique execution and commodity web testing become automated and cheap. The value concentrates at the two ends: continuous automated validation, and elite objective-driven human campaigns.
  2. Hybrid becomes the default operating model. Humans set objectives and design the campaign; agents handle recon, enumeration and parallel path exploration; humans make the risky calls.
  3. Purple teaming overtakes pure red. Organisations increasingly care more about improving detections than about being beaten quietly.
  4. AI systems become targets. Prompt injection, tool abuse, training-data exposure and agent hijacking become standard scope items, guided by resources such as OWASP's LLM Top 10 and the NIST AI Risk Management Framework.
  5. Measurement matures. Expect detection rate per ATT&CK technique, mean time to detect per attack stage, and control-drift rate to replace "number of findings" as headline metrics.
  6. Safety governance for offensive automation. Organisations will need rules of engagement for their own agents: read-only by default, blast-radius caps, kill switches and full audit logs.
What to invest in, by organisation maturity
MaturityPriorityWhy
Early โ€” basic logging, no detectionsVulnerability management + pen testingRed teaming finds gaps you already know about
Developing โ€” SIEM and some detectionsPurple team sprints + Atomic Red TeamBuilds detection coverage fastest per dollar
Mature โ€” strong SOC, tuned detectionsAutomated continuous validation + annual red teamCatches drift; tests the response organisation
Advanced / regulatedTIBER-style threat-led testing + AI-system red teamingMeets obligations and tests novel surfaces

How to prepare your program now

  • Map your detections to ATT&CK and measure coverage honestly, including which techniques you cannot observe at all.
  • Automate the repeatable: schedule emulation of your top 50 relevant techniques and alert on regressions.
  • Reserve human engagements for objectives, not checklists โ€” "exfiltrate customer records" beats "test lateral movement."
  • Write rules of engagement that cover AI tooling used by your own testers.
  • Add your AI applications to scope, with prompt injection and tool abuse as explicit test cases.
  • Track time-to-detect per stage, not just whether you were caught.

Verdict

Manual red teaming is not going away; it is becoming more senior, more objective-driven and more expensive per engagement, while automation takes over everything that can be scripted. The teams that benefit most will not pick a side. They will run continuous automated validation as hygiene, use agents to widen reconnaissance, and spend their scarce human creativity on the campaigns that test whether the organisation โ€” not just the toolset โ€” can survive a determined adversary.

Frequently asked questions

Will AI replace red teamers?

It is replacing parts of the work โ€” enumeration, commodity web testing, repeat execution. Objective selection, social engineering and judgement under uncertainty remain human for the foreseeable future.

Is BAS the same as red teaming?

No. Breach and attack simulation validates whether controls catch known techniques. Red teaming tests whether an organisation can stop a motivated adversary pursuing a goal.

How often should we red team?

Most mature organisations run one or two full engagements a year, purple team sprints quarterly, and automated validation continuously. Regulated financial entities in the EU face explicit threat-led testing requirements under DORA.

References and further reading

  1. MITRE ATT&CK framework
  2. MITRE CALDERA adversary emulation platform
  3. Red Canary, Atomic Red Team
  4. Google Project Zero, Big Sleep: AI-discovered SQLite vulnerability
  5. Verizon, 2024 Data Breach Investigations Report
  6. European Central Bank, TIBER-EU framework
  7. Regulation (EU) 2022/2554 (DORA)
  8. EU Artificial Intelligence Act
  9. OWASP Top 10 for Large Language Model Applications
  10. NIST AI Risk Management Framework
cybersecurityred teamingoffensive securityAI securitypenetration testing

From our own rounds

Measured on CineMind, from real sessions people played on this site โ€” not a third-party dataset.

Rounds played here
10
Questions per round
1

Most-played topics right now: AI (2), Streamers (1), Cartoons (1).

Play a round and add to these numbers
Share this post

Rate this article

No ratings yet

Discussion

Comments are moderated. Read our editorial policy.