Contents
- opportunities for using AI in detecting network anomalies
- the role of AI in log analysis and reducing alert noise
- automating incident response with SOAR
- Threats arising from attackers’ use of AI
- Risks associated with AI models and LLM systems
- law and ethics in the context of AI and cyber security
- implementing AI in cyber security: strategies and best practices
- the future of AI in cyber security: trends and challenges
Share
opportunities for using AI in detecting network anomalies
AI increases the effectiveness of network anomaly detection (NDR), because it directly answers the question: “is this traffic normal for my network?”. Anomaly models can spot unusual patterns faster than manually maintained rules, e.g. sudden DNS connections to rare domains. The essence of NDR is building a baseline of host and user behaviour, so deviations are visible in the context of a specific environment. This approach works particularly well where classic signatures or simple rules cannot keep up with traffic dynamics.
Examples of this class of solutions include NDR platforms such as Darktrace, Vectra AI or ExtraHop. Their models observe behaviour over time and make it possible to detect events that “on their own” may look harmless, but when combined with the network profile start to raise suspicion. As a result, the analyst receives a signal based on deviation from the norm rather than solely on a match to a static rule. The biggest gain appears when an organisation can compare traffic against its own real behavioural baseline, rather than relying only on universal patterns.
- 01Establishing the baselineA profile of normal network behaviour.
- 02Fast pattern detectionAutomation of analysis.
- 03Contextual deviationsSpotting anomalies.
- 04Adapting to dynamicsEffectiveness beyond rules.
AI builds a contextual baseline to detect anomalies faster than static rules, which is crucial in dynamic environments. Examples include: Darktrace, Vectra AI, ExtraHop.
the role of AI in log analysis and reducing alert noise
AI in log analysis and reducing alert noise makes it easier to identify which signals actually require action when billions of events flow in every day. In SIEM with ML mechanisms (e.g. Microsoft Sentinel, Splunk Enterprise Security, Elastic Security), models combine events into incidents and reduce the number of notifications, e.g. by correlating many failed logins with geolocation and IP reputation. In practice, AI answers the question “which 0.1% of alerts is worth attention?”, instead of adding further layers of noise. When the data is well standardised, this approach shortens MTTD — often from hours to minutes.
AI will not work well without appropriate data quality, because the most common cause of poor results remains a lack of log normalisation (ECS/CEF), inconsistent timestamps and gaps in telemetry. The minimum set of sources that enables sensible analytics includes several critical log categories. Without them, correlation and building reliable incident context are limited, even with a large data volume.
- identity logs (Entra/AD)
- endpoint (EDR)
- DNS
- firewall/proxy
- critical applications
Reducing alert fatigue also relies on prioritisation. Risk scoring can combine asset criticality (CMDB), exposure (public IP), Threat Intelligence credibility and behavioural abnormality. The practical operational goal is to keep the volume at a manageable level (e.g. <50 alerts/day per analyst), rather than flooding the team with thousands of notifications. This approach makes it easier to distinguish a real breach from an accidental scan and organises SOC work around the highest-risk cases.
automating incident response with SOAR
Automating incident response with SOAR speeds up case handling because it moves repetitive tasks from the analyst to playbooks. Solutions such as Palo Alto Cortex XSOAR, Splunk SOAR or Microsoft Sentinel playbooks (Logic Apps) can automatically enrich IOC, isolate an endpoint, block a domain or create a ticket in the system. The biggest impact is visible in high-volume incidents (e.g. phishing, scans), where a person wastes time on routine work instead of proper risk analysis. SOAR delivers real value when it automates concrete operational steps, rather than just “beautifully visualising” alerts.
Automation must be safe, so for high-impact actions a human-in-the-loop approach or risk thresholds are used. In practice, this means that auto-quarantine of email and URL blocking can run without approval, but host isolation or password reset should require confirmation after correlation with EDR and Threat Intelligence. In well-designed playbooks, “safe failures” are also important, meaning a clear answer to what the system should do when the API does not respond or context is missing. Such boundaries reduce the risk that automation itself causes downtime.
In phishing handling, SOAR can collect headers, check the URL in VirusTotal, search for similar messages in the O365 tenant and remove them (Search & Purge), and then block the domain in the proxy. In practice, this shortens the handling time of a single report from 10–15 minutes to 1–3 minutes, provided the policies are set up correctly. For ransomware, the “first 5 minutes” matter. Automation can disconnect the host from the network (EDR), block accounts (Entra/AD), stop suspicious processes and force a snapshot/backup where the environment supports it. Testing playbooks in exercises is critical, because overly aggressive isolation can cause a downtime as severe as the incident itself.
- 01Routine automationMoves repetitive tasks into playbooks
- 02Operational realityAutomatic enrichment, endpoint isolation, ticket
- 03Impact at scaleEffective for phishing and scans
- 04Human-in-the-loopControl over key decisions and risk
Conclusion: SOAR speeds up incident handling, eliminating routine work in risk analysis, but leaves key decisions in human hands.
Threats arising from attackers’ use of AI
Threats arising from attackers’ use of AI come down mainly to greater scale of activity, stronger personalisation and faster preparation of attacks. LLMs can create grammatically correct phishing and BEC messages in a company’s style and tailor them to a specific role (e.g. HR, finance), which reduces the number of typical “red flags”. Deepfake audio can imitate a voice and speech rate in CEO fraud scams, and additional time pressure increases the effectiveness of the fraud. In such scenarios, procedures are key (call-back to a known number, payment limits, four-eyes principle), because deepfake detection technology alone can be unreliable.
AI also improves reconnaissance (OSINT) and target selection, because models can analyse employee profiles, technologies visible on a website and job adverts, building a map of potential weaknesses. In practice, campaign preparation time can be shortened from days to hours, and social engineering gets much closer to the point. At the same time, LLMs make it easier to write malware components (e.g. parsing, HTTP communication, encryption), which lowers the barrier to entry and encourages the emergence of mass “commodity” attacks. This does not remove the need for testing or domain knowledge, but it increases the number of working prototypes in circulation.
Attackers can also use AI to evade detection through mutations and polymorphism, generating payload variants or changing the “wrapper” to make signature-based detection harder. For this reason, behavioural detection (EDR) and application control (allowlisting) are becoming more important than relying solely on hashes. AI also supports password and MFA attacks thanks to better conversations and pretexts tailored to the victim, and techniques such as MFA fatigue can still work where processes are weak. The response is phishing-resistant MFA (FIDO2/WebAuthn), restricting logins from unusual locations and conditional access policies.
Risks associated with AI models and LLM systems
Risks associated with AI models and LLM systems stem from the fact that the model can be manipulated, can disclose data or generate incorrect recommendations that lead to poor operational decisions. Prompt injection involves “injecting” instructions into content (e.g. an email or document) that the model treats as guidance, for example to ignore rules and reveal API keys. The risk is particularly high in applications based on RAG, where the answer may inadvertently expose information from internal databases or chat logs. The basis of defence is role separation (system/developer/user), input filtering and restricting tools (tool calling) to a whitelist.
Data leaks can also result from the way the system maintains context and conversation logs, as well as whether information is used for training. That is why a precise policy is needed to define what content may be sent to external models. In practice, DLP mechanisms are used (e.g. Microsoft Purview) and PII masking (e.g. Presidio) to limit the spread of sensitive data in prompts and responses. Another threat is poisoning, i.e. “contaminating” the knowledge base or training data. It is enough to inject a fake document into a wiki or repository for the model to cite incorrect procedures or surface malicious links. Protection is based on access control, document signatures and versioning, as well as source validation before indexing.
Model errors in cybersecurity are also dangerous because LLMs can hallucinate the cause of an incident or commands that do not exist, and the “confident tone” of the answer can lull the team into a false sense of security. In addition, model inversion or membership inference attacks can attempt to determine whether a specific sample was in training, which becomes particularly risky with personal or medical data. When an assistant has access to tools (cloud API, EDR, IAM), incorrect interpretation or prompt injection can trigger high-impact actions, so granular permissions (least privilege), OPA policies and a “dry-run” mode for administrative commands are needed. The risk is reduced by a layered approach: input/output moderation, policy classifiers, RAG with source citation and a requirement for human confirmation for destructive actions.
The security of LLM systems also requires proper management of secrets and keys, because tokens can reappear in responses or end up in monitoring if they are present in logs and prompts. Secrets should be stored in dedicated vaults (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault), and the application should fetch them at runtime. The effectiveness of safeguards should be checked periodically through AI red teaming, covering tests for prompt injection, jailbreaks, data leakage from RAG and tool abuse. In practice, the results of such tests are translated into guardrail adjustments and automated regression tests before deployment.
- 01Model manipulationInjection of malicious instructions.
- 02Unwanted data disclosureExposure of confidential information from databases.
- 03Incorrect recommendationsMisguided operational decisions.
- 04Defence strategiesRole separation, filtering, whitelist.
It is crucial to understand attack vectors and implement multi-layered protection, including role separation and strict filtering of input data and tools.
law and ethics in the context of AI and cyber security
Law and ethics in the context of AI and cyber security come down to data minimisation, decision auditability and maintaining processes that AI cannot replace. GDPR requires minimisation and a clear purpose for processing, so if data is not necessary, it should be anonymised or pseudonymised, and access appropriately restricted. In practice, this means, among other things, masking PII and separating environments so that security logs do not contain unnecessary personal data. Implementing AI does not relieve you of organisational obligations (NIS2/KSC): you still need roles, escalation, a reporting method and consistent evidence and event timelines.
The AI Act may impose additional requirements depending on the use case, especially where the system genuinely affects decisions concerning users (e.g. automatic account blocks). In such a setup, risk management, data quality and robust documentation of how the system works and where its limitations lie come to the fore. Traceability is also important: input/output logs, model versions, parameters and data sources (e.g. citations in RAG), so that the basis of a decision can be reconstructed. A good practice is to store metadata in tools such as MLflow and maintain a history of SOAR playbooks.
Ethics includes, among other things, the risk of discrimination in scoring and behavioural analytics, because UEBA models may “punish” unusual work patterns (e.g. night shifts) if the training data is biased. Safeguards include bias testing, clear alerting criteria and the ability to appeal and explain decisions in the SOC process. In relationships with suppliers, contractual terms remain key: where data is processed (EU/EEA), how long it is retained, whether it can be used for training, and what the deletion conditions are. For sensitive data, “no-train” variants and private instances are often used (e.g. Azure OpenAI with region control). At the same time, it is worth limiting “Shadow AI” through approved tools, browser DLP and training based on concrete examples of prohibited data.
Third-Party Risk assessment should include, among other things, certifications (ISO 27001, SOC 2), security policies, encryption mechanisms, incident history and a bug bounty programme, as well as the ability to control prompt retention. A checklist covering access logging, tenant segregation, key management and the option to disable conversation retention also works well. AI-related incidents should be handled like standard IR incidents if there has been a breach of confidentiality, integrity or availability of data/systems, and then the notification obligations should be assessed (e.g. a personal data breach). In practice, a separate category in the incident register helps: “AI security”, to measure trends and the effectiveness of safeguards.
implementing AI in cyber security: strategies and best practices
AI implementation in cyber security should be carried out in stages, from pilot to production, so as not to get bogged down in multiplying POCs without real impact. The recommended path includes: (1) a pilot on one use case (e.g. phishing), (2) integration with the SOC process and metrics, (3) scaling to additional data sources and automations. Each stage should have measurable KPIs, such as reduced false positives, lower MTTR and the number of cases handled automatically. The best results come from an approach where AI is “embedded” in data, processes and metrics, rather than functioning as a separate experiment.
Selecting use cases with high ROI usually starts with areas dominated by repetitive work. In practice, the most effective are phishing triage, IOC enrichment, incident summarisation and vulnerability prioritisation, because the impact on handling time and task queue sizes quickly becomes visible. At the same time, deployments that automatically make high-impact decisions carry greater risk (e.g. blocking critical systems without approvals), so the automation scope should be precisely described in runbooks and risk thresholds. If there is no measurable improvement after 30 days, it is usually the data quality or the poorly chosen use case that is at fault, not the “model being too weak”.
Successful implementation also requires clearly defined roles and competencies, because alongside SOC analysts you need people in data engineering (log pipelines), application security (API integrations) and a product owner (business priorities). To reduce LLM risks, it makes sense to build an assistant based on RAG with access control and source citation, and to apply guardrails (input/output validation and redaction of PII/secrets), rather than relying solely on “trusting the answer”. In integrations with SIEM/EDR/SOAR, a proven pattern is an intermediary layer with clearly defined actions, which makes auditing, permission management and control of actions easier. At the same time, you should plan monitoring of model quality and drift (precision/recall, FP rate) and regression tests on historical incidents before deploying changes.
- Define a data policy: what must not be sent to external models.
- Choose one high-volume process (e.g. phishing) and prepare a playbook.
- Implement guardrails and decision logging.
- Set KPIs: FP rate, MTTR and cost per incident.
the future of AI in cyber security: trends and challenges
The future of AI in cybersecurity is moving towards more autonomous SOCs and agents that combine detection with action, which increases response speed but also raises the risk of mistakes. In practice, this means more automated actions in environments (e.g. auto-response in the cloud), while at the same time a greater susceptibility to abuse, including prompt injection targeting systems with access to tools. Growing autonomy will require stronger safeguards: least privilege, auditing, security testing and controlled execution environments. Where actions have a high impact, the need for decision approval and rigorous traceability will remain.
The most serious operational challenge will be maintaining model quality over time, because the environment does not stand still. New applications appear, there is migration to the cloud, and the way people work changes too, which translates into drift and a drop in detection effectiveness. For this reason, the importance of continuous monitoring of metrics (e.g. precision/recall and FP rate), drift alerts and cyclical retraining carried out with data quality control is growing. At the same time, the MLOps pipeline must be protected like a production system, meaning signing artefacts, isolating CI runners, scanning dependencies (SCA) and controlling access to the model registry. Completing the picture are regular AI red teaming tests (prompt injection, jailbreak, RAG leaks and tool abuse) and automated regression tests run before deployment.
Challenges will also cover architecture and costs, because an LLM in a SOC easily “bloats” through long prompts, large contexts and retention. To keep cost-effectiveness under control, summarisation and caching are used, as well as limiting RAG to top-k fragments (e.g. 3–5) instead of attaching the full documentation. It is worth calculating the pricing model as the cost per incident and comparing it with the time saved by the analyst, which makes decisions about scaling easier. In the longer term, organisations that combine autonomy with control will win: from least privilege and “dry-run” to auditable decisions and security testing on their own scenarios.
FAQ
Frequently asked questions
How does AI help detect anomalies in network traffic?
AI builds a baseline of host and user behaviour, which makes it quicker to spot deviations from the norm. It works well where classic signatures and rules cannot keep up with traffic dynamics.
Can AI in SIEM reduce the number of alerts for an analyst?
Yes, ML models can combine events into incidents and reduce notification noise. As a result, the analyst receives fewer alerts, but more relevant ones.
Why does AI in cybersecurity need good-quality logs?
Because the effectiveness of analysis depends on log normalisation, consistent timestamps and complete telemetry. Without this, event correlation and incident context building are limited.
When does SOAR automation deliver the biggest impact in incident handling?
Most strongly with high-volume incidents, such as phishing or scans. In those cases, playbooks take over routine steps and shorten response times.
How do attackers use AI for phishing and fraud?
AI helps create grammatically correct phishing and BEC messages tailored to a specific role, and also supports audio deepfakes in CEO fraud scams. It also cuts campaign preparation time from days to hours.
What can go wrong when using LLMs in cybersecurity?
The model can be manipulated through prompt injection, disclose data, or generate incorrect recommendations. The risk increases especially when it has access to tools and internal databases.





