IT Brief Asia - Technology news for CIOs & IT decision-makers
Asia
Secure.com warns AI security tools can falter in production

Secure.com warns AI security tools can falter in production

Fri, 11th Sep 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Secure.com has published guidance on how security teams should validate AI-native security tools in live environments. The analysis was written by Cybersecurity Leader and Founding Member Yasir Zahid.

The guidance argues that AI security products often perform worse in production than in vendor testing, with live outputs potentially 45% to 50% less accurate than laboratory results suggest. It sets out a framework for assessing whether tools can be trusted before organisations allow them to influence alert handling, escalation and remediation.

The paper focuses on a problem many security teams face after deployment, not during procurement. Zahid says teams often discover flaws only when false alerts surge, analysts begin to ignore recommendations or a genuine threat is missed.

Production checks

A central recommendation is to treat the first weeks of deployment as a probation period. Instead of letting an AI tool act on its own, teams should run it in shadow mode alongside existing analyst workflows and compare the outcomes.

This is intended to show where human analysts and the system agree and, more importantly, where they diverge. Those disagreements can reveal whether a tool is missing important threats, generating unnecessary noise or reaching correct conclusions for weak reasons.

The guidance also advises teams to focus on false negatives as well as false positives. A product that reduces alert volumes but misses serious incidents may create a greater risk than the overload it was meant to solve.

Validation should extend beyond the final recommendation. Security teams should examine the enrichment and reasoning behind a decision, because a correct call reached for the wrong reason may fail when the environment changes.

Monthly review

Another recommendation is regular replay testing using known attack scenarios. Red team exercises should be replayed monthly so organisations can see whether a tool still detects threats after changes to infrastructure, user behaviour or data sources.

This is linked to what the paper describes as model drift. Once deployed, AI systems do not remain static in practice, and changes in the environment can alter their performance over time even if the software itself remains unchanged.

To monitor that, the guidance recommends setting a baseline at deployment and then tracking shifts in false positive rates, mean time to triage, alert closure rates and analyst override rates. If those metrics move materially over a 30-day period, teams should investigate the cause rather than assume the system is behaving as intended.

Analyst override rates are given particular weight. Frequent rejection of AI recommendations is presented as the clearest sign that the tool needs adjustment and may also indicate that confidence labels no longer match real-world accuracy.

Explainability issue

Zahid's guidance places strong emphasis on explainability. A security team cannot properly validate a recommendation if the tool cannot show which signals contributed to the decision, how those signals were connected and why a given level of confidence was assigned.

This reflects a broader concern in cybersecurity over whether AI systems can be audited after incidents. In security operations, the issue is not only whether a tool reaches the correct answer but whether analysts can reconstruct how it got there.

The paper argues that this is especially important for high-stakes actions such as containment, blocking or automatic escalation. It recommends keeping human approval gates in place until a tool has demonstrated accuracy over time in the specific environment where it is being used.

Board questions

Beyond day-to-day operations, the analysis links validation to governance. Boards and audit committees, it argues, should focus less on raw detection claims and more on whether the organisation can prove that an AI system stayed within approved boundaries, remained subject to override and produced a reliable audit trail.

For buyers, these checks should begin before a contract is signed. Requests for proposals should ask vendors about real-world accuracy in comparable environments, integration requirements, onboarding responsibilities, explainability standards and data portability.

Vendor lock-in is highlighted as another risk. If alert history, custom workflows and threat mappings sit inside one platform in proprietary formats, switching costs can become a security issue rather than just a commercial one.

Organisations should ask whether tools can operate with an existing security stack, whether data and logic can be exported in standard formats and whether migration would require significant extra work.

At an operational level, the paper presents validation as a continuing process rather than a one-off sign-off exercise. Changes in infrastructure, new integrations and fresh alert types can all alter model behaviour, meaning prior testing may no longer be enough.

"The goal isn't proving the tool is perfect. It's knowing, with evidence, where it can be trusted to act alone and where a human still needs to be in the loop," said Yasir Zahid, Cybersecurity Leader and Founding Member, Secure.com.