Black Shard

Insights12 June 2026

Red Team or Penetration Test: Which Do You Need?

A penetration test measures the security of a system. Adversary simulation measures whether your organisation can detect and stop an intruder, and most buyers need the first before the second.

Chess board mid game on a dark table lit by cold side light

Two engagements, two different questions

A penetration test answers one question: how secure is this system? Adversary simulation, red teaming, answers a different one: can this organisation detect and stop a real intruder before the damage is done? Every other difference between the two, stealth, scope, duration, price, who gets told, follows from which question you are asking. Buyers who conflate them either pay red-team prices for a vulnerability list, or run a pen test and walk away believing it said something about their detection capability. It did not, and it was never designed to.

The market makes this harder than it should be. Plenty of providers sell red teaming that is a penetration test with a phishing campaign bolted on. The honest test is the success criterion. If the engagement succeeds by producing a ranked list of exploitable weaknesses, it is a penetration test, whatever the label on the proposal. If it succeeds by producing a timeline of what your defenders saw, missed and did while an operator pursued a specific objective, it is a red team. For most Australian businesses the sequence is settled: penetration testing first, and adversary simulation once there is a detection capability worth measuring.

What a penetration test actually buys you

A penetration test is an exercise in coverage. The scope is an asset list: this web application, this external perimeter, this Azure tenant, this API. The window is fixed, usually one to three weeks. Everyone relevant knows it is happening, the testers make no serious effort to stay quiet, and that is a feature: stealth costs time, and time spent hiding is time not spent finding the next flaw. The goal is to enumerate as many exploitable weaknesses as the window allows, chain them where chaining proves impact, rank them, and hand you a remediation plan you can actually execute.

It is the right instrument in most situations Australian businesses actually face: a new product before launch, an annual assurance cycle, a material infrastructure change, a client security questionnaire that asks for evidence of independent testing, or control-effectiveness testing for APRA-regulated entities under CPS 234. It is also the right instrument if you have never been tested at all, because the first test of an untested environment always produces a long list, and you want that list at pen-test prices, not red-team prices.

The blue team is normally informed, and detection is not the measure. A good tester will still tell you when nothing fired while they ran credential attacks for a week, but that observation is a bonus, not the deliverable.

What adversary simulation actually tests

A red team engagement is objective-driven. Instead of an asset list, you agree on flags that represent real business damage: access the payroll system, plant a marker file on the finance share, reach the environment that would make the front page. Domain admin is a means, never the end. The operators emulate a plausible adversary using tradecraft mapped to frameworks like MITRE ATT&CK, they move slowly, and they blend with legitimate administrative activity. Staying undetected is not vanity, it is the experimental condition: if the operators are loud, the exercise measures nothing.

The defenders are not told. A small white cell, typically two or three executives, holds the written authorisation and handles deconfliction. If the SOC detects something and escalates it as a real incident, the white cell decides whether to let the response run (watching that response is precisely the point) or to quietly stand the incident down. Engagements run weeks to months, not days. The deliverable is an attack narrative laid against your detection and response timeline: which actions generated telemetry, which alerts fired, which were triaged, where a human made the right call, and where the intruder would have completed the objective unchallenged.

Australia's financial sector has formalised this logic. The CORIE framework, coordinated under the Council of Financial Regulators, runs intelligence-led adversary simulations against financial institutions precisely because vulnerability-focused testing alone does not demonstrate operational resilience. The reasoning holds well below institutional scale.

The maturity gate most buyers skip

Here is the uncomfortable part: a red team tests your detection and response capability, so if you do not have one, there is nothing to test. No centralised logging, no alerting that reaches a human, no one rostered to look: the simulation succeeds trivially and teaches you what a moment of honest reflection would have taught you for free. If a pen tester making no effort to hide can walk through your environment today, an adversary simulation will prove the same thing at several times the cost and a fraction of the coverage.

If you fall between the stools, purple teaming is the honest middle step: attackers and defenders in the same room, running techniques deliberately, checking what telemetry appears, tuning detections on the spot. It trades realism for feedback density, and for a mid-maturity organisation it converts security spend into detection capability faster than a covert exercise ever will.

Before commissioning a red team, you should be able to answer yes to most of these:

  • EDR is deployed across the fleet and someone, an in-house SOC or an MDR provider, actually watches it
  • Logs are centralised and at least some alerts reach a human with authority to act
  • An incident response plan exists and has been exercised at least once
  • The findings from your last two penetration tests were substantially remediated
  • The Essential Eight basics (application control, patching, MFA, restricted admin privileges) are in place, not on a roadmap

Scope, rules of engagement and who gets told

The contracting looks different too. Pen test scoping is mechanical: assets, window, production or staging, test accounts, exclusions. Red team scoping is a negotiation about risk: which adversary you are simulating, which vectors are permitted (phishing, physical entry, third-party pivots) and which are hard exclusions: no destructive actions, marker files instead of real data exfiltration, no targeting of personal devices. Authorisation must come from someone with genuine authority over every entity in scope, in writing, because operators may be stopped and questioned during physical entry and staff may end up interviewed. You also cannot authorise what you do not own: cloud providers and other third parties have their own testing terms, and social engineering of your own staff deserves deliberate executive sign-off, not a line item in a statement of work.

Blue-team knowledge is a dial, not a binary. The most useful first red team for most organisations is assumed breach: the operators start from a granted foothold, a standard user workstation or a set of credentials, and the exercise tests everything after initial access. It is cheaper, it wastes no weeks proving that some phish eventually lands (one always does), and it aims the spend at the question that matters: once they are in, do you see them?

How to choose, in one pass

Never been tested, shipped a major change, need assurance evidence, or facing a compliance obligation: penetration test. Last two pen tests substantially remediated, monitoring in place, incident response plan exercised: assumed-breach red team. Mature SOC, prior simulations, a board asking harder questions: full covert exercise built on a real threat model. And it is a cycle, not a ladder you climb once. Penetration tests keep running against new and changed systems for coverage, while red teaming periodically tests whether the organisation, not the software, holds.

Black Shard runs both kinds of engagement, and the first honest conversation is about which one a client actually needs. A red-team request from a business still building its detection capability usually becomes a penetration test plus a monitoring uplift, because that sequence produces a red team worth running later. Expect the same discipline from whoever you engage: a provider happy to sell you stealth you cannot detect is selling theatre.

See what an attacker would find.

Australia-wide, from our Brisbane head office. Someone will contact you as soon as possible.

Open a briefinfo@blackshard.com.au