New: ESProfiler Services. Expert consultancy, augmented by AI.
ESPROFILER IconESPROFILER
Capability ExchangeCapability Exchange
Platform
How it worksHow you onboardHow you operate
Services
All ServicesSecurity Reality BaselineSecurity Consolidation Baseline
Use Cases
All
Resources
AllArticlesWebinarsEvents & ConferencesProduct Releases
AboutCareersStatus
Log InBook Demo
Back to all posts
2026-07-27
Articles

Reporting AI security tooling to the board: Why activity metrics fall short

Activity metrics show a security tool is busy, not effective. How structured staff interviews give boards real evidence on AI security tooling.

Security budgets increasingly carry AI claims on every line item. Detection tools, email security, endpoint platforms and SOC tooling now compete on machine learning performance, with vendors typically claiming higher detection rates and fewer false positives than the products they replace. Boards and audit committees have responded by asking a harder question: whether the spend produced a measurable improvement in the organisation's security position.

Standard security reporting does not answer that question. The metrics most security functions present to the board were designed for operational management, and they measure something different from effectiveness.

Activity metrics describe workload, not performance

The typical board pack contains finding counts, mean time to respond, and coverage percentages. Each of these measures throughput. Finding counts record how much a tool surfaces, without distinguishing findings that changed a decision from findings that consumed analyst time. Response-time metrics measure the team's handling of the queue rather than the quality of what enters it. Coverage figures confirm deployment, not detection performance.

The structural weakness is that two tools with very different detection quality can produce identical reporting. A poorly tuned scanner generating a high volume of low-value findings and a well-calibrated tool surfacing a small number of significant issues both appear, in an activity-based pack, as functioning systems. The reporting format cannot separate them.

Detection tooling can be evaluated, but the data rarely exists

There is an established method for assessing this class of technology. Any detection system can be evaluated as a classifier, on two measures: the proportion of its alerts that are correct, and the proportion of real events it identifies. Precision and recall are standard practice in evaluating machine learning systems in most other fields.

In security procurement, that evaluation rarely happens. Vendors do not generally publish performance figures against independent test sets, and buyers seldom have the means to construct one. Purchase decisions therefore rest on vendor documentation, and the level of scrutiny peaks before any real-world performance data exists. Post-deployment evaluation is rarer still. Once a tool is live, few organisations return to grade it against the claims that justified the purchase, and board reporting inherits the gap.

Operational teams hold the missing performance data

The practical performance of a security tool is observable to the people who operate it. Analysts accumulate evidence with every shift: which alert sources produce findings that get actioned, which have been suppressed or deprioritised, which capabilities were configured at deployment and which remain unused, and where the team maintains manual processes because the tooling falls short.

In most organisations this information is never captured in structured form. It is distributed across shifts and teams, held by individuals, and surfaces only informally. As evidence, it does not currently exist in any form a board could assess.

Structured interviews across the security function change that. When every operator answers the same set of questions about the tools they work with, individual observations aggregate into measurable data: the proportion of licensed capability in active use, the disposition of findings by source, and the compensating manual work that indicates where tooling underperforms. Interviewing at sufficient scale matters here. A handful of conversations produces anecdote; systematic coverage of the function produces a dataset.

A reporting structure boards can assess

With that data collected, board reporting can move from activity to evidence, structured in three layers.

Capability coverage establishes what the organisation has bought: the capabilities present across the stack, mapped against requirements, with overlaps and gaps identified. This frames the discussion at the level boards reason about, which is control and exposure rather than tool inventory.

Utilisation establishes what is operational. A capability that exists in the product documentation but was never configured represents cost without control. Operator-derived utilisation data separates deployed capability from licensed capability.

Performance in practice establishes how the tooling holds up in use. The proportion of findings actioned rather than dismissed, by source, is the closest measure of real-world precision available to most organisations, and it comes from observed behaviour rather than vendor claims.

Together, these layers address the question boards are actually asking: what the organisation can do, whether it is doing it, and whether the tooling performs as purchased.

Collecting the evidence at scale

The constraint is practical. Interviewing an entire security function manually requires weeks of senior time, which is why systematic collection of operational knowledge is rare and almost never repeated.

ESProfiler's Security Reality Baseline is a fixed-cost, fixed-time engagement that produces this evidence in four to six weeks. It catalogues every tool providing security capability across the estate, runs asynchronous agent-led interviews with the people who operate them, and has every finding audited and verified by cybersecurity experts before it reaches the report. The interviews are short, structured conversations that adapt to each person's role and can be paused and resumed, so covering the full function does not depend on workshops or calendar blocks.

The deliverable follows the structure boards need: capability coverage across the stack, utilisation grounded in operator evidence, and risks tied to how tools are actually deployed rather than how dashboards describe them. It is a baseline a security leader can present without relying on vendor documentation, because none of it comes from vendor documentation.

The standard boards will apply

Scrutiny of AI security spend is increasing, and activity metrics will not withstand it. The organisations able to demonstrate effectiveness will be those that can show what capabilities exist across their stack, which are in operational use, and how they perform in the hands of the teams running them. That evidence exists inside every security function. The difference lies in whether it is collected.

Ready to Optimize
Your Security Stack?

Talk to our team to see how ESPROFILER can help you gain full visibility and control over your security investments.

Book a Demo

Platform

  • Market Layer
  • Capability Layer
  • Commercial Layer
  • Tribal Layer
  • Architect Layer

Services

  • All Services
  • Security Reality Baseline
  • Security Consolidation Baseline

Company

  • About Us
  • Jobs
  • Resources
  • Changelog
  • Contact
ESPROFILER IconESPROFILERNCSC For Startups AlumniSupported By GoogletechUK Winner
© 2026 ESPROFILER. All rights reserved.
Policies & Terms