Overview
A red team engagement starts with an objective agreed with your leadership (read the board pack, move money between two accounts, obtain customer records, gain domain administrator) and then pursues it the way a funded adversary would, across whatever combination of people, process and technology gets there.
The output that matters is not the list of vulnerabilities used along the way. It is the detection timeline: at which point in the chain your monitoring produced a signal, whether anyone acted on it, and how long the gap was between our first move and your first response. That number is the actual measure of a security programme.
Engagement at a glance
- Typical duration
- 3 to 6 weeks, deliberately paced rather than time-boxed per asset
- Shape
- Objective-based: defined crown jewels, defined success criteria
- Knowledge
- Black box, or purple-team collaborative with your detection team
- Frameworks
- MITRE ATT&CK, TIBER-EU principles, CBEST-style threat-led scoping
- Prerequisite
- Mature vulnerability management, red teaming is wasted on unpatched perimeters
- Deliverable
- Attack narrative, detection timeline, and a gap analysis per ATT&CK technique
What is covered
Every item below is tested and recorded, so the report shows what held as clearly as what failed.
- Objective-based adversary simulation
- External reconnaissance and initial access
- Phishing and social engineering (where authorised)
- Payload development and defence evasion
- Command and control infrastructure
- Credential access and privilege escalation
- Lateral movement and persistence
- Data discovery and simulated exfiltration
- Detection and response measurement
- Purple-team replay and control tuning
- MITRE ATT&CK technique mapping
- Executive attack narrative
Test matrix
What is attempted in each class, and what it means when it works.
| Class | What is attempted | Typical impact |
|---|---|---|
| Initial access | Perimeter exposure, exposed credentials, supply-chain paths, phishing where authorised | A foothold obtained the way a real intrusion begins |
| Evasion | Whether endpoint and network controls detect tooling, or only known signatures | Attacker operating undetected inside the estate |
| Credential access | Harvesting, Kerberos abuse, delegation and certificate services paths | Escalation from foothold to privileged identity |
| Lateral movement | Movement toward the objective, testing segmentation at each hop | Reach from a compromised workstation to the crown jewels |
| Persistence | Whether a foothold survives reboot, password reset and incident response | Adversary who stays after you think you evicted them |
| Detection | Every action logged against your timeline; which produced alerts, which were investigated | The measured gap between intrusion and response |
What we commonly find
Alerts fired, nobody acted
The most common outcome is not that detection failed but that it worked and the signal died in a queue. That is a process finding, and it is fixable this quarter.
Segmentation that stops at the diagram
The route to the objective usually crosses a boundary the architecture says is closed.
Detection tuned for tools, not behaviour
Controls that recognise a named toolkit but not the technique it implements are defeated by recompiling it.
The service account nobody owns
Over-privileged, non-rotating, excluded from monitoring because it was noisy, and the fastest path to the objective.
How the engagement runs
Scoping and threat modelling
Map the asset, the attacker profile, and what “compromised” actually means for this business. Rules of engagement, testing windows, excluded techniques and an escalation contact are agreed in writing before anything is sent.
Reconnaissance and surface mapping
Enumerate everything reachable: subdomains from multiple passive sources, every endpoint referenced in JavaScript bundles, exposed services, third-party integrations, and the assets nobody remembers deploying. Coverage is recorded per host, so what was not tested is as visible as what was.
Manual exploitation
Authenticated testing from every role, with at least two accounts per role. Business logic, authorisation boundaries, injection, race conditions and state transitions, with each candidate reproduced live before it is written down. Automated tooling contributes coverage; it never contributes findings.
Verification and impact
Every finding is reproduced in a fresh session, isolated to the single parameter that causes it, and pushed to its maximum realistic impact. A finding that cannot survive a clean-room reproduction does not appear in the report.
Reporting and retest
The report is written twice over: once for the engineer who has to fix it, once for the auditor who has to file it. A walkthrough session follows, then a retest of every finding, closed only when re-exploitation fails.
What you receive
Executive summary
One page for the people who approve budget: what was tested, what was found, what it means in business terms.
Technical findings
Each finding with severity, CVSS, affected component, full request and response, reproduction steps and a working proof of concept.
Attack chains
Where findings combine, the chain is written out end to end, from first request to demonstrated impact.
Remediation guidance
A specific fix for your stack and framework, with the corrected pattern, not a link to a generic reference page.
Audit mapping
Findings mapped to SOC 2, ISO 27001, PCI DSS, HIPAA and OWASP ASVS as applicable, so the report drops straight into an audit pack.
Retest and attestation
Every finding retested in a clean session after remediation, with a signed attestation letter for customers and auditors.
When to run it
- After vulnerability management is mature, red teaming an unpatched perimeter just re-proves the perimeter is unpatched.
- When you have invested in detection and response and want to know what it actually catches.
- Ahead of a regulator-driven threat-led testing requirement, as preparation.
- Annually, with purple-team exercises between engagements.
Questions
How is this different from a penetration test?
Scope and success criteria. A penetration test enumerates and exploits vulnerabilities across a defined asset, and completeness of coverage is the goal. A red team pursues a single business objective by whatever authorised path works, and stealth plus detection measurement are the goal. Most organisations need the penetration test first, red teaming assumes you have already fixed the things a penetration test would find.
Will you phish our staff?
Only with written authorisation from someone empowered to give it, and within agreed limits. Social engineering is scoped explicitly: which techniques, which populations, what happens to captured credentials, and whether individuals are ever named in the report. Our default is that individuals are never named. The finding is about the control, not the person.
Who should know the test is happening?
As few people as possible, which is the point: usually the CISO, the CEO or board sponsor, and a single technical escalation contact. That contact exists so that if your team detects us and begins genuine incident response, we can stop the clock before it becomes expensive.
What if you do not reach the objective?
That is a valid result, and a good one, and the report says so plainly. It documents every path attempted and why each failed, which is a defensible evidence pack that your controls work. A red team that always succeeds is usually scoped too loosely.
How much does a red team engagement cost?
Red teaming costs more than penetration testing because it runs longer and produces fewer findings by design. Scope is driven by the objective, the duration, and which channels are permitted. A fixed quote follows scoping, and we will say plainly if a penetration test would serve you better for the money.
How long does a red team engagement take?
Typically three to six weeks, sometimes longer. A short engagement forces loud tradecraft, which tests your detection against an adversary who is not being careful. If the budget only supports two weeks, the honest advice is usually to run a penetration test instead and come back to red teaming later.
Should we tell our security team?
That is the decision that shapes the whole engagement. An unannounced test measures detection and response as they actually are. An announced one is a purple team exercise and teaches more per hour. Most organisations get more from running it unannounced once, then announced thereafter.
Do you use custom tooling or off-the-shelf frameworks?
Both, chosen against what your defences are likely to catch. Off-the-shelf tooling tests whether your detection recognises known tradecraft, which is worth knowing. Custom tooling tests whether it recognises anything else. An engagement that only uses one tells you half the answer.
What is purple teaming?
Running the same techniques with your defenders in the room, watching what fires and what does not. It produces less drama and more detection coverage per hour than a covert engagement, and every red team we deliver ends with a replay session so the findings reach the people who have to catch it next time.