Skip to main content

Spear phishing versus bulk simulation: what changes in the test and in the numbers

A spear campaign and a bulk campaign measure different things over different populations. Why their rates cannot share a trend line, and how a six-person cohort is reported when a percentage would identify people.

By Yashodhan Sawant
August 31, 20267 min read

A spear phishing simulation is not a bulk campaign with better copy. It is a different instrument, run against a different population, and the number it returns answers a different question. The practical consequence is narrow, and getting it wrong costs a programme its series: a spear rate and a bulk rate must never sit on the same trend line.

Both send email, both count who acted on it, and the two percentages look identical in a spreadsheet. A programme that alternates between them and labels each round only with a date has built a series whose largest movements are artefacts of which instrument was used. Nobody reading it a year later can tell the artefact from the finding.

What changes in the test

A bulk simulation puts one pretext, or a small set of them, in front of a large population. The pretext is built from the organisation's general context: the payroll system everyone uses, the shared-drive invitation that is unremarkable in that company. Nothing in it is specific to any recipient, and the population is large enough to stand for the workforce, or for a defined part of it.

A spear campaign inverts each of those choices. Reconnaissance is conducted per target, or per small group holding the same role. The pretext is assembled from what is publicly discoverable about that person: the work they are visibly doing now, the counterparties they are doing it with, the panel they sat on last month, the vendor whose implementation their team is midway through. The population is small, and it is selected rather than sampled. It exists because somebody decided these particular roles were worth testing.

Bulk simulationSpear campaign
PopulationLarge, sampled to stand for a workforceSmall, selected by role
PretextOne, or a small set, from context common to everyoneOne per target or role group, from that target's discoverable material
What varies between recipientsNothingAlmost everything
What the rate is computed overA representative population facing one stimulusA chosen population, each facing a different stimulus
What it is evidence ofHow this workforce responds to a plausible messageWhether a determined attacker with a day of reconnaissance reaches this role
How it reaches the reportA cohort figure and a point on a trendFindings about the pretexts and the reconnaissance surface

What changes in the numbers, which is the part that matters

A bulk campaign's rate is a measurement over a representative population in which every recipient faced the same message. The constant stimulus makes the figure a property of the population rather than of the email: the variation that remains is behavioural. Provided the next round's pretext sits in the same difficulty band, the two figures can be compared.

A spear campaign's rate has neither property. The population was chosen rather than sampled, so the figure describes the selection at least as much as the people in it. And the stimulus was not constant: each target faced a message tuned to them, so the campaign ran as many one-person tests as it had targets and then averaged them. An average taken across different stimuli describes nobody.

The denominator finishes the argument. In a campaign against six people, one action moves the figure by roughly seventeen points. One different judgement from the same six people, and the quarter looks transformed.

This is why spear is not simply the hard end of one scale. Pretext difficulty banding keeps a series meaningful within a single instrument: record how demanding the pretext was, compare like with like, and an improvement cannot be manufactured by sending an easier email. But banding assumes a shared measurement and a shared population, and spear and bulk share neither. A programme that runs both keeps two series, and labels every round with the instrument, the population definition, the pretext count and the band.

What each one is evidence of

Both are legitimate. They answer different questions and belong in different rows of the report.

Bulk answers: how does this workforce respond to a plausible message? That is a statement about a population, which is the shape the standing obligations take. The RBI Directions, 2026 require that a bank "shall evaluate the awareness level of employees periodically" (Commercial Banks, ¶202, issued 31 July 2026; the same sentence is ¶201 for Payments Banks and Small Finance Banks, and ¶197 for Credit Information Companies). NBFCs carry it in a different construction, as a "formal mechanism to measure and track the effectiveness of such training through periodic assessments or testing" (NBFCs, ¶36). SEBI's CSCRF requires that regulated entities "shall periodically assess level of employee cybersecurity awareness, for e.g., through phishing test success rate, etc." (CSCRF v1.0, 20 August 2024, GV.RM Guidelines item 1(e), p. 87). Each describes a population assessed over time, which is what the population instrument measures.

Spear answers something narrower: can a determined attacker who spends a day on reconnaissance reach this specific role? The answer is not a rate. It is an account of what was built, out of what material, and what happened when it arrived – which control saw it, and where the process around the person held or gave way.

Who may be in a spear population, and why cohort size is the hard part

Two clauses of CERT-In's Comprehensive Cyber Security Audit Policy Guidelines (CISG-2025-02, Version 1.0, 25 July 2025) govern this, and they bind the empanelled auditing organisation conducting the test. §13.2.7(ii), p. 50, requires that "specific written permissions must be obtained" from the auditee before social engineering is conducted. §15.2.2(iii), pp. 55–56, governs who may be in the population and what may be said about them:

"When targeting general staff (e.g., untrained or non-security personnel), such testing must utilize anonymized or statistical techniques—ensuring no individual is personally identified or penalized. The purpose is to evaluate overall awareness and the effectiveness of security processes, not to single out individuals.

Social engineering and process testing must only target group of employees explicitly included within the agreed audit scope. These tests must not involve external entities such as customers, business partners, vendors, or other third parties, unless specific written consent is obtained from the target organization."

The scope rule is the one people see first, and spear campaigns walk straight into it. Reconnaissance surfaces the interesting counterparties: the vendor's project manager, the outsourced payroll desk, the customer whose name makes the pretext land. None is inside the agreed scope by virtue of being inside the story, and the clause admits them only on specific written consent from the target organisation.

The rule that bites hardest is the first one, and it bites arithmetically. Whether a cohort figure is anonymous is a function of cohort size. A rate published over six roles is not: anyone who knows which six were targeted holds a one-in-six list, and at nought or a hundred per cent it resolves completely to individuals. No wording placed around the number repairs that.

So a spear campaign against six people is not reported as a cohort rate. It is reported as findings, one per pretext:

  1. the pretext as constructed, and the material reconnaissance supplied to it;
  2. the route it took, including the infrastructure and sender reputation it had to survive;
  3. what the technical controls did with it, which is a finding about the mail path rather than about a person;
  4. whether an out-of-band check existed for the action requested, and whether it was invoked;
  5. the remediation, written against the process and the discoverable material rather than against the recipient.

Where the report needs a number about the workforce, it comes from the population instrument, which has the denominator to carry it. The cohort floor below which no rate is published belongs in the written authorisation, fixed before the campaign runs and before anyone has seen a result. Settled afterwards, it becomes a negotiation about one finding.

The reconnaissance surface is a finding on its own

The most durable output of a spear campaign is often produced before a message is sent. What was discoverable about the targets, assembled from public sources by someone with a day and no privileged access, is a finding whether or not anybody clicked. Report it as one.

Which roles and reporting relationships were published, and where; which live projects, vendors and counterparties were inferable from public material; which addresses were guessed from a naming convention; what automatic replies, event listings and document properties disclosed; and how much came from the organisation's own channels.

A campaign that produced no clicks and a rich reconnaissance surface has told the organisation something a clean rate would have hidden. That surface is the raw material of every pretext built against those roles next year, and it does not decay on the programme's cadence. It is also the part of the finding that survives being written down: unlike a six-person rate, an inventory of what is discoverable names nobody.

What to hold

  1. Two series, never one. Label every round with the instrument, the population, the pretext count and the difficulty band. A round that cannot be labelled cannot be plotted.
  2. Rates come from the population instrument. Trend reporting, cohort comparison and anything entering a compliance file as a figure comes from bulk rounds.
  3. Spear produces findings, under a cohort floor fixed in the authorisation. Below the floor the output is narrative and remediation, not a percentage.
  4. Report the reconnaissance surface in its own right, independent of what anyone did with the message.

Both belong in a programme that has to evidence itself over years. They do not belong in the same column.

About the author

Yashodhan Sawant, ISO/IEC 27001 Lead Auditor

Lead — ISMS & Certification Readiness

Leads ISO 27001 and SOC 2 readiness at Security Brigade — gap assessment, Statement of Applicability, clause 9.2 internal audit and management review, through to certification audit.