The dual mandate at the core: wield AI to attack, and attack the AI itself, surrounded by the operator fundamentals, the AD and cloud attack chain, autonomous ops, defence, and reporting. Every track runs from foundational theory to exam-level practice.
11 tracks, foundational theory through to exam-level practice.
Before you point AI at a target, you need the fundamentals that make an engagement legal, safe, and effective. This on-ramp covers the penetration-test lifecycle, rules of engagement and scoping, the kill chain and MITRE ATT&CK, risk and CVSS, and the assumed-breach mindset, each framed with how AI shifts the work. Start here if offensive security is new to you, or skim it to align on language before the deeper tracks.
The discipline of attacking your own systems on purpose, under rules, to find what a real adversary would before they do. The opening map of the field.
The contract that turns hacking into a profession: what is in bounds, what is forbidden, who is authorised, and what happens when something goes wrong. Get this right before you touch anything.
Two shared languages for describing how an intrusion unfolds: the linear kill chain for the arc of an attack, and the ATT&CK matrix for the catalogue of techniques along it.
The arc every engagement follows: scoping, recon, enumeration, exploitation, post-exploitation, reporting and retest, and what done looks like in each phase.
External, internal, web, mobile, cloud, wireless, physical, social, phishing and red team: the common shapes of offensive-security work, what each tests, and when to choose it.
How to turn 'this is broken' into 'this matters, this much': the difference between a vulnerability and a risk, and how CVSS, likelihood and business context combine into a severity that means something.
Why modern offensive security starts from the premise that the attacker is already inside, and what that changes about how you test.
The defender's core strategy of layered, independent controls so that no single failure is fatal, and why it is also the lens an attacker and a good report use to prioritise.
Reconnaissance is where AI first earns its place: correlating scattered signals, drafting target profiles, and mapping an attack surface faster than any human could by hand. This track covers passive and active recon, network and service enumeration, agentic OSINT, and profiling a target with an AI copilot, alongside the discipline of verifying everything a model surfaces before you trust it.
The layered instruction set that tells an AI copilot who it works for, what it may touch, and how to work, assembled fresh for every engagement so an agent stays safe, on method, and in scope.
Mapping a target before you touch it: passive open-source intelligence and light active probing to find the attack surface, the people, and the soft spots, with an AI copilot to correlate and summarise, and a discipline for catching what it invents.
Turning a list of hosts into a map of attack surface: port scanning, service and version detection, and protocol-by-protocol enumeration, with an AI copilot to parse and prioritise output and a discipline for catching what it gets wrong.
Turning an AI agent loose on a target to build a dossier autonomously: correlating identities, infrastructure and exposure across sources, then verifying every claim before you trust it. The agent does the legwork, you own the truth.
Enumerating and prioritising a large external attack surface with AI help: discovering assets, pivoting through subdomains, certificates and ASNs, deduplicating the mess, and ranking what actually matters. Volume is easy; a verified, ordered list is the work.
Mapping the footprint an organisation left in the cloud and across SaaS: public buckets and blobs, exposed APIs, OAuth applications, tenant enumeration and third-party exposure, with an AI copilot summarising and prioritising the mess. Passive-first, in-scope, and verified before trusted.
People-focused OSINT for authorised social-engineering assessments: mapping the org, inferring roles, researching credible pretexts, and staying inside the ethical and legal boundaries. AI drafts org charts and pretexts fast, and invents people just as fast, so everything is verified and everything stays in scope.
This is one half of the dual mandate: wielding AI as a copilot across live exploitation. You will use models to craft payloads, reason about vulnerabilities, and chain findings into a working attack path, while learning exactly where copilots hallucinate, mislead, or fail dangerously, and how to keep a human owning every decision that matters.
How an AI copilot fits inside an offensive engagement: the model reasons and enumerates at machine speed, the human approves the dangerous moves, and guardrails keep the whole thing safe.
Where an LLM copilot genuinely accelerates an engagement, where it hallucinates dangerously, and the verification discipline that separates a professional from a script-runner.
How to get great results from an AI copilot during an engagement: set objectives not chores, review proposals like a lead consultant, keep scope tight, and let it do the tireless work while you keep the judgement.
Using a model to accelerate payload development and obfuscation reasoning in an authorised test, with the verification and encoding discipline that keeps it reliable.
The two hard stops that keep AI-assisted testing safe: credential actions and destructive web actions both require explicit human approval, every time, with no carry-over.
Moving from single findings to an attack path. Using a copilot to reason about how weaknesses combine, while keeping the copilot's optimism on a leash.
The application layer is the biggest attack surface most organisations have: the OWASP Top 10, the WSTG methodology, and the classes of flaw (injection, access control, auth) that matter most, plus where an AI copilot speeds the work and where it fabricates.
Online credential attacks: password spraying, credential stuffing, and default-credential hunting, the low-tech, high-yield techniques that turn weak human passwords into access, and how an AI copilot proposes them behind a hard approval gate.
A single authorised engagement walked from recon handoff to exploitation to chained impact, with a copilot accelerating each phase and a human approving every consequential action.
The other half of the mandate: turning your attention on the AI systems now embedded in every target. This track covers prompt injection, RAG and embedding attacks, MCP exploitation, and agent abuse: the OWASP LLM and agentic threat landscape, taught as hands-on offensive technique rather than theory.
A structured walkthrough of the OWASP LLM Top 10 for 2025, plus a nod to the Agentic Top 10. Each risk, a concrete example, and how it is tested. This is the map of the whole track.
How untrusted input becomes instructions, why the LLM can't tell data from commands, and the direct vs. indirect injection split that defines the attack surface.
Jailbreak techniques from role-play to obfuscation, encoding, many-shot, and multi-turn crescendo. Why safety filters are probabilistic rather than deterministic, and how to test guardrail robustness methodically.
Retrieval-augmented generation widens the attack surface, because every document in the store is untrusted input. Indirect injection, retrieval poisoning, and embedding-space attacks.
Injection through content the model ingests: retrieved documents, emails, web pages, and tool output. Cross-domain and cross-tool confusion, dormant payloads, and the exfiltration channels that turn a read into a leak.
When an LLM can call tools, prompt injection stops being a text problem and becomes a code-execution one. Tool abuse, confused-deputy chains, and exploiting Model Context Protocol servers.
Attacks on the model itself: extraction and stealing, model inversion, membership inference, training-data extraction, and embedding-space attacks. A grounded look at what is practical today versus what remains largely theoretical.
Attacks that arrive through images, audio, and files: hidden instructions and adversarial inputs. Then abusing tool-use and function-calling to reach real systems, chaining injection through to action and impact.
An end-to-end operator scenario: assessing a production-style AI application that combines a chatbot, a RAG pipeline, and tools. From recon through injection and tool abuse to impact and reporting, safely and in scope.
The frontier almost no one else certifies: operating and governing autonomous pentest agents. You will learn the autonomy spectrum, how to orchestrate an agent safely, guardrails and kill-switches, validating what an agent reports, and taking accountability for machine-speed engagements, the governance skills operators need as hackbots move from demo to production.
What an autonomous offensive agent (a hackbot) actually is, where it sits on the autonomy spectrum, and why operating one is a distinct skill rather than a button you press.
Human-in-the-loop, on-the-loop, and out-of-the-loop supervision, when each is defensible, and how to map the supervision model to the blast radius of what the agent can do.
How to set an autonomous pentest agent up to succeed without causing harm: give it objectives not chores, grant least-privilege tools, sandbox it, cap its time and budget, and enforce scope in code.
The load-bearing controls that keep an autonomous agent inside its lane: unoverwritable policy layers, kill switches, rate and spend and blast-radius limits, egress control, and hard gates on dangerous actions.
Why an agent's confident 'success' is a claim, not proof: triaging output, catching hallucinated and false-positive findings, holding evidence to a reproducibility standard, and requiring human sign-off.
One supervised autonomous engagement walked end to end: brief, scope, guardrails, launch, monitor, validate findings, and report, with the operator on the loop the whole way.
Lead-level governance for autonomous red teaming: accountability and audit at machine speed, scope enforcement when the agent outruns the human, the legal and ethical duty, and the methodology and policy that hold it all together.
The largest attack chain in the curriculum, AI-accelerated. This track walks the full Active Directory and cloud kill chain: enumeration and BloodHound, Kerberos and NTLM relay, AD CS abuse, credential access, lateral movement, and cloud attacks. AI reasons over the identity graph and attack paths, with a human owning every state change.
Why Active Directory is the battleground of internal engagements, the shape of a typical domain takeover, and a map to the techniques that get you there.
Active Directory describes itself, and BloodHound collects that description to compute the shortest path from where you are to Domain Admin, turning a tangle of permissions into a queryable graph.
How the design of Active Directory's authentication protocol becomes an attack surface: Kerberoasting, AS-REP roasting, delegation abuse, and ticket forgery.
How an attacker on the wire turns Windows' own authentication against it: poisoning name resolution with Responder, relaying NTLM with ntlmrelayx, and forcing authentication with coercion attacks.
Active Directory Certificate Services is a privilege-escalation goldmine. Learn the ESC1 to ESC11 family of misconfigurations, why certificates are such durable credentials, and how Certipy enumerates and exploits them.
Once you have a foothold, credentials are the currency of escalation. Learn where Windows stores secrets, how attackers harvest them from LSASS, DPAPI and NTDS, and how offline cracking turns hashes into passwords.
Turning one compromised host into many. Learn the protocols and techniques that move an attacker across a network, the tooling that does it at scale, and how to map the blast radius of a single credential.
Cloud compromise is usually a misconfiguration story, not an exploit story. Learn how identity, over-permissioned roles, exposed storage and logging gaps drive real incidents across M365, Azure, AWS and GCP, measured against benchmarks.
A full identity attack path walked from a single assumed-breach foothold to domain and tenant dominance, with AI reasoning over the graph while every credential and state-changing step stays behind a human approval gate.
Offense sharpens defense. This track flips the lens: detecting AI-driven attacks, hardening LLM applications and their retrieval layer, securing agents and MCP, and purple-teaming the result. Built for operators who need to defend the same AI systems they have learned to attack.
What AI-accelerated attacks look like from the blue side: the shift in tempo and volume, the detection signals that survive it, and how to frame your defence against the OWASP LLM and Agentic Top 10.
The observable signals of AI-generated recon, phishing, and exploitation, how to catch prompt injection in application logs, and an honest account of what is and is not reliably detectable.
Threat-modeling an LLM or agentic application from first principles: draw the trust boundaries and data flows, find the untrusted-input surfaces, run STRIDE adapted for AI, and turn the model into a prioritised control list.
Trust boundaries, input and output handling, system-prompt hygiene, least privilege for tools, and isolating untrusted content, so that an injection you cannot prevent still cannot do much.
The retrieval trust boundary in practice: document provenance and sanitisation, access control on the store, and defending against retrieval poisoning and embedding-space attacks.
Tool scoping and least privilege, approval gates, egress control, and sandboxing, plus defending MCP-server integrations against tool-description injection and confused-deputy chains.
Running incident response when the incident is a prompt injection or an abused agent: detect and triage the signal, contain by revoking tokens and cutting egress and disabling tools, do forensics from model and tool logs, and recover with the hole closed.
Building an offense/defense loop for an LLM application: test harnesses like garak, PyRIT, and promptfoo, mapping findings to controls, and continuous evaluation that measures resistance over time.
Standing up an AI security program at lead level: policy and governance that people follow, an honest risk-acceptance process, secure-by-design review gates, vendor and model supply-chain risk, and metrics that actually change decisions.
One narrated purple-team engagement against a production-style AI support agent: threat model it, attack it, derive controls from the findings, deploy them, re-test, and report resistance as a measured, defensible result.
An engagement is only as good as the report a client can act on. This track covers the anatomy of a strong report, AI-assisted drafting that never invents evidence, writing findings that land, evidence handling and reproducibility, and the anonymisation and data-protection discipline that keeps a client safe.
Clients do not buy testing, they buy the report. This module breaks down the sections a professional deliverable must carry, the two audiences it serves, and what separates a good finding from a weak one.
How to turn an engagement's accumulated findings, evidence and commands into a client-ready report: the AI drafts the prose, grounded strictly in real data, and a human reviews, verifies and signs.
A finding only matters if it changes what the client does next. This module covers impact statements, business-risk framing, honest severity, actionable remediation, and where AI helps you draft versus where it quietly fabricates.
A finding is only as strong as the evidence behind it. This module covers chain of evidence, capturing screenshots and command logs, writing reproducible steps, protecting artefact integrity, and what makes a finding defensible months later.
Engagement data is a map of a client's weaknesses, and reports often carry PII and secrets. This module covers de-identification discipline, handling sensitive data, redaction before sharing, and never leaking client data into third-party models.
The report is the artefact, but the debrief is where the work actually lands. This module covers running a technical readout and an executive readout, walking findings live, handling pushback and disagreement, and setting up remediation, with AI to prep and a human in the room.
The boundaries around AI in reporting work: keeping client data out of models unless authorised, a defined data boundary through your own key, metered and capped spend, and the practices that keep AI assistance safe, private and affordable.
Finding the bug is a fraction of the value; the client only gets safer when the fix ships and holds. This module covers prioritised remediation guidance, verifying fixes, scoping a retest, catching regressions, and the trade-off between point-in-time and continuous validation.
Leadership does not buy vulnerabilities; they buy decisions about risk. This module covers translating technical findings into business language, building a risk narrative without FUD or false precision, writing the one-page summary, and the metrics executives and boards actually use.
One good report is a craft; a hundred consistent reports across a team is a system. This lead-level module covers standardised severity and rubrics, trend metrics across engagements, quality assurance of AI-assisted reports, and building defensible consistency that holds up under scrutiny.
The product track for licensed StrikeOps firms and their operators. Go from zero to running a live engagement on the platform: proposals and scoping, brands and the client portal, deploying an agent, the engagement brief and playbooks, the live shell, findings and evidence, and building the report. Reserved for StrikeOps operators.
Your first hour operating StrikeOps: how you sign in, how to read the platform, and the setup path that takes a fresh tenant to a running engagement.
Configure the brand identity your reports and portal wear, and serve the client portal from your own domain with a single CNAME before any client is invited.
Build a proposal that bundles scope, Rules of Engagement, methodology and pricing, share it with the client over a portal link, and convert the agreed proposal into an engagement.
Register the client, open the engagement whose type and dates shape the work, and write the scope and Rules of Engagement that the AI operator respects exactly as recorded.
Stand up a StrikeOps agent host, the assumed-breach foothold your engagements run from. Pick one of three deployment paths, choose a transport, and assign the host to an engagement.
The layered CLAUDE.md that briefs your on-host AI operator: a guardrail base, a methodology floor, and the engagement layer you author. How to write a tight engagement layer, tune scope playbooks, and re-push to a live host.
The in-browser terminal into your agent host. How to open it in an engagement context, launch Claude, the passkey step-up and recording it enforces, the session limits, and the Tailscale SSH alternative.
The daily loop of an AI-assisted engagement: confirm the brief, give the operator an objective, approve the dangerous steps at the guardrail, watch findings and evidence stream in, and steer toward reporting while you own the outcome.
How to work the findings an engagement produces in StrikeOps: shaping a well-formed finding, driving the status workflow, linking artifacts and commands, and letting the AI draft from real evidence.
How the StrikeOps report engine assembles a branded deliverable from findings, evidence and commands: preparing inputs, triggering a versioned build, working the report workspace, interim advisories, and delivery.
How to build, send and track an authorised phishing engagement in StrikeOps Phishops: choosing static versus adversary-in-the-middle, picking the channel, crafting the lure, tracking interactions, and handling captured secrets responsibly.
How StrikeOps runs on your own accounts and keys: setting the Anthropic key first, connecting recon and messaging providers, understanding where secrets live, setting daily caps, and pointing encryption at your own AWS KMS key with a kill-switch.
How to run StrikeOps as an ongoing programme and a firm: recurring cadence and cycles, retest verdicts, flat versus enforced RBAC, the read-only client portal, and billing and usage meters.
Illustrative, fully fictional engagement walkthroughs that show the tradecraft end to end: how the pieces from every track come together on a realistic but invented target. Read these to watch method turn into outcome, from first foothold to final report.
A short guide to the Case Studies strand: what each walkthrough teaches, why every scenario is fictional and illustrative, and the anonymisation discipline that real reports demand.
A fictional internal assumed-breach test where a single low-privileged account was chained through Active Directory misconfiguration to domain admin, with AI reasoning over the path and a human approving every state-changing step.
A fictional web application test where one broken-access-control anti-pattern chained into tenant-wide data exposure, showing how a copilot helped enumerate the surface and how the finding was proven and reported responsibly.
A fictional assessment of a RAG-backed support chatbot with tool access, where an indirect prompt injection hidden in a knowledge-base document led to data exfiltration, plus how it was detected and the controls that fixed it.
A fictional cloud engagement where an over-permissioned role and a misconfigured trust policy chained to tenant-wide access, with AI triaging the IAM graph and a human confirming and containing each escalation.
The glossary, frameworks and standards, tool index, and scoring rubrics that underpin every track: the shared vocabulary of AI-augmented offensive security. Free to read, and designed to be linked to from anywhere in the curriculum when you need a definition fast.
Plain-language definitions of the offensive-security and AI-offsec terms you will meet across the Academy. A living reference.
The methodologies, scoring systems and control frameworks offensive-security work is measured against: what each one is, and when you will cite it.
A quick-reference catalogue of the tools referenced across the Academy: what each does, when to reach for it, and the job it belongs to.
The calibration reference for rating findings: what each severity level means, how to use CVSS without being ruled by it, and how to turn a score into a risk.
Quick answers to the questions people ask most about offensive security and AI-assisted testing.