Submitted:
25 November 2024
Posted:
17 December 2024
You are already at the latest version
Abstract
Cyber threats are growing increasingly sophisticated, Security Operations Centers (SOCs) require more efficient and intelligent response tools to manage and mitigate incidents effectively. In response, organizations have adopted comprehensive frameworks to enhance their cybersecurity postures, notably the National Institute of Standards and Technology (NIST) Cybersecurity and MITRE’s ATT&CK framework. However, the process of effectively mitigating and responding by mapping the frameworks derived from real-time security events remains challenging. To address this, we developed a Security Event Response Copilot (SERC) which guides security analysts in order to respond and mitigate security breaches more effectively. The SERC consists of two components that are Security Event Data Extraction using Retrieval-Augmented Generation (RAG) methods, and LLM-based Incident Response Guidance. Those systems were integrated with Wazuh as a Security Information and Event Management (SIEM) platform to capture the security events on targeted endpoints. By combining Wazuh’s monitoring capabilities with the structured intelligence of MITRE ATT&CK, this system identifies adversarial tactics, techniques, and procedures (TTPs) relevant to security events, while NIST standards ensure adherence to best practices in incident handling. Utilizing the RAG approach, the Copilot retrieves context-specific information from historical and real-time data sources, enhancing the generation of actionable insights and response recommendations. The RAG data sources were clustered into 3 database vector collections that are (1) incident response general knowledge, providing various information across cybersecurity platform; (2) NIST cybersecurity-related framework (CSF) 2.0 which offers a lifecycle approach to managing cybersecurity issues; and (3) MITRE ATT&CK framework to identify the tactics and techniques associated with the incident. The complementary use of NIST's strategic breadth and MITRE’s tactical detail allows organizations to build a multi-layered defense that integrates high-level risk management with actionable intelligence. To this end, we conducted end-to-end simulation and evaluated RAG performance output using the semantic coherence method. Alongside this, we also measure the end-to-end task completion rate. Overall, this research amplifies the potential of combining structured threat intelligence frameworks and powerful AI models to meet the dynamic needs of SOCs in a rapidly changing cybersecurity environment.
Keywords:
1. Introduction
2. Literature Review
2.1. Wazuh as a Versatile SIEM Solution
2.2. Role of Large Language Model Copilots in Enhancing Efficiency and Personalization
2.3. Emerging Role of Large Language Models in SOC Operations
2.4. RAG Components
2.4.1. Qdrant Vector Database
2.4.2. Embedding Model: BAAI/bge-Large-en-v1.5
- Efficiency and Scale: The model demonstrates high efficiency in processing large volumes of text data while maintaining high accuracy in similarity detection. It uses a technique known as Matryoshka Representation Learning [35], which enables the model to generate embeddings in multiple dimensionalities (e.g., 1024, 768, 512, down to 64 dimensions) without significant loss of accuracy.
- Performance Comparison: Compared to similar models such as OpenAI’s Ada [36] and traditional BERT-based models [37], BAAI/bge-large-en-v1.5 is optimized for faster responses and adaptability with large datasets. Its open-source nature also makes it more cost-effective than proprietary models, which is an important consideration for scalable cybersecurity applications.
2.4.3. Similarity Metric: Cosine Distance
- and are the vectors being compared,
- is their dot product,
- and are the magnitudes of the vectors.
2.5. Atomic Red Team: A Modular Framework for Adversary Simulation and Detection Validation
3. Methodology
3.1. Event Security Data Extraction and Refinement

| Prompt Format |
|---|
| Extract the primary issue or problem from the following Wazuh JSON log. |
| Focus on details like the alert description, severity level, associated tactics, compliance tags, and any specific event data that |
| clarifies the problem: |
| {input_event_json} |
| Present the extracted issue in a concise format, describing the main problem indicated in the log. |
| Expected Output Example: |
| ”’ |
| Given the JSON log provided, here’s how the response might look: |
| Extracted Problem: |
| Description: Wazuh agent ’suricata-nids’ has stopped, indicating a potential disruption in NIDS monitoring. |
| Alert Level: 3 (Medium severity) |
| Associated Tactic: Defense Evasion (MITRE ID: T1562.001 - Disable or Modify Tools) |
| Compliance Concerns: PCI DSS (10.6.1, 10.2.6), HIPAA (164.312.b), TSC (CC7.2, CC7.3, CC6.8), NIST 800-53 (AU.6, AU.14, AU.5), GDPR (IV_35.7.d) |
| Log Details: Full log message reads "Agent stopped: ’suricata-nids->any’," suggesting possible interruption in security monitoring. |
| ”’ |
| Expected output above focuses on the core issue, making it easily readable and actionable for SOC and RAG systems. |
3.2. RAG Workflow
3.2.1. Data Collections

| Category | Content |
|---|---|
| General Knowledge | Incident Response Guide [47,48,49,50,51]IRP [52,53] |
| NIST Knowledge | CSF 2.0 [54] |
| MITRE Knowledge | MITRE ATT&CK [55] |
- General Knowledge: This collection provides foundational knowledge for computer security incident handling, derived from sources such as security operations and automation response (SOAR) playbooks and general cybersecurity incident handling guides. It serves as a primary resource for addressing general security incident management needs.
- NIST Knowledge: Specifically structured to align with the NIST Cybersecurity Framework, it includes guidelines, policies, and standards that ensure relevance to regulatory frameworks and assist in the compliance verification process.
- MITRE Knowledge: Designed around the MITRE ATT&CK framework, this collection supports the retrieval of documents directly relevant to threat mitigation and response strategies. By mapping queries to specific MITRE tactics and techniques, this collection enhances the precision of actionable insights for mitigating cyber threats. To ensure that the collection remains aligned with the state-of-the-art MITRE tactics and techniques, a dedicated service periodically checks the MITRE servers for updates and synchronizes the local collection in the vector database, enabling real-time adaptability to emerging threats and maintaining the relevance of the system.
3.2.2. Chunking Techniques
3.2.3. Large Language Model Config
3.3. Copilot: Generative AI Incident Response Workflow System

| Parameter | Value |
|---|---|
| Temperature | Set to 0.1, ensuring high determinism by minimizing randomness in the model’s predictions, favoring the most probable outputs. |
| Top-k | Configured at 50, restricting token sampling to the top 50 most likely candidates, reducing the likelihood of low-probability tokens. |
| Top-p (Nucleus Sampling) | Set to 0.9, allowing dynamic token selection by considering tokens with a cumulative probability of 90%, ensuring a balance between determinism and contextual diversity. |
| Max Tokens | Defined as 4096, specifying the upper limit for the total number of tokens in the generated output, suitable for applications requiring concise yet comprehensive responses. |
| Copilot Prompt Format |
|---|
| This is the context information for general knowledge purposes: |
| {context_general} |
| This is the context information knowledge of planning for generating the incident response playbook based on The NIST Cybersecurity Framework (CSF) 2.0: |
| {context_nist} |
| This is the context information knowledge from MITRE ATT&CK for security incident response mitigation purposes: |
| {context_mitre} |
| Based on the above context information, hope you can use and elaborate on the knowledge you have to analyze this incident and tell me what action to take: |
| json |
| Based on that incident, what should be done to mitigate the risk? Make sure to use knowledge of the NIST CSF 2.0 and the MITRE ATT&CK |
| framework to identify the tactics and techniques associated with the incident. |
| Do not mention the source of JSON or text input, just tell what action to take with Markdown format. |
3.4. Simulation Scenario

4. Experiment Results and Analysis
4.1. Copilot Integration for Security Event Analysis in Wazuh
4.2. Performance Evaluation
5. Discussion and Future Work
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Vielberth, M.; Bohm, F.; Fichtinger, I.; Pernul, G. Security Operations Center: A Systematic Study and Open Challenges. IEEE Access 2020, 8, 227756–227779. [Google Scholar] [CrossRef]
- Perera, A.; Rathnayaka, S.; Perera, N.D.; Madushanka, W.W.; Senarathne, A.N. The Next Gen Security Operation Center. 2021 6th International Conference for Convergence in Technology, I2CT 2021. Institute of Electrical and Electronics Engineers Inc., 2021. [CrossRef]
- Arora, A. MITRE ATT&CK vs. NIST CSF: A Comprehensive Guide to Cybersecurity Frameworks.
- Farouk, M. The Application of MITRE ATT&CK Framework in Mitigating Cybersecurity Threats in the Public Sector. Issues In Information Systems 2024. [Google Scholar] [CrossRef]
- Stine, K.; Quinn, S.; Witte, G.; Gardner, R.K. Integrating Cybersecurity and Enterprise Risk Management (ERM), 2020. [CrossRef]
- Wainwright, T. Aligning MITRE ATT&CK for Security Resilience - Security Risk Advisors.
- Freitas, S.; Kalajdjieski, J.; Gharib, A.; McCann, R. AI-Driven Guided Response for Security Operation Centers with Microsoft Copilot for Security 2024.
- Fysarakis, K.; Lekidis, A.; Mavroeidis, V.; Lampropoulos, K.; Lyberopoulos, G.; Vidal, I.G.M.; i Casals, J.C.T.; Luna, E.R.; Sancho, A.A.M.; Mavrelos, A.; Tsantekidis, M.; Pape, S.; Chatzopoulou, A.; Nanou, C.; Drivas, G.; Photiou, V.; Spanoudakis, G.; Koufopavlou, O. PHOENI2X – A European Cyber Resilience Framework With Artificial-Intelligence-Assisted Orchestration, Automation and Response Capabilities for Business Continuity and Recovery, Incident Response, and Information Exchange 2023.
- Wazuh - Open Source XDR. Open Source SIEM.
- Companies Using Wazuh, Market Share, Customers and Competitors.
- Younus, Z.S.; Alanezi, M. Detect and Mitigate Cyberattacks Using SIEM. Proceedings - International Conference on Developments in eSystems Engineering, DeSE. Institute of Electrical and Electronics Engineers Inc., 2023, pp. 510–515. [CrossRef]
- Hello GPT-4o | OpenAI.
- Pixtral Large | Mistral AI | Frontier AI in your hands.
- wazuh. Wazuh documentation.
- Šuškalo, D.; Morić, Z.; Redžepagić, J.; Regvart, D. 34th DAAAM International Symposium on Intelligent Manufacturing and Automation: Comparative Analysis of IBM QRadar and Wazuh for Security Information and Event Management. [CrossRef]
- IBM Security QRadar vs Wazuh Comparison 2024 | PeerSpot.
- Dunsin, D.; Ghanem, M.C.; Ouazzane, K.; Vassilev, V. A Comprehensive Analysis of the Role of Artificial Intelligence and Machine Learning in Modern Digital Forensics and Incident Response Article info, 2023.
- Hays, S.; White, J. Employing LLMs for Incident Response Planning and Review 2024.
- Lin, G.; Feng, T.; Han, P.; Liu, G.; You, J. Paper Copilot: A Self-Evolving and Efficient LLM System for Personalized Academic Assistance 2024.
- Li, R.; Patel, T.; Wang, Q.; Du, X. MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents 2024.
- Haque, S.; Eberhart, Z.; Bansal, A.; McMillan, C. Semantic Similarity Metrics for Evaluating Source Code Summarization. IEEE International Conference on Program Comprehension. IEEE Computer Society, 2022, Vol. 2022-March, pp. 36–47. [CrossRef]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models 2022.
- Han, Z.; Gao, C.; Liu, J.; Zhang, J.; Zhang, S.Q. Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey 2024.
- Outshift | Fine-tuning methods for LLMs: A comparative guide.
- Challenges & limitations of LLM fine-tuning | OpsMatters.
- Wang, X.; Wang, Z.; Gao, X.; Zhang, F.; Wu, Y.; Xu, Z.; Shi, T.; Wang, Z.; Li, S.; Qian, Q.; Yin, R.; Lv, C.; Zheng, X.; Huang, X. Searching for Best Practices in Retrieval-Augmented Generation 2024.
- Explainer: What Is Retrieval-Augmented Generation? | NVIDIA Technical Blog.
- Xu, H.; Wang, S.; Li, N.; Wang, K.; Zhao, Y.; Chen, K.; Yu, T.; Liu, Y.; Wang, H. Large Language Models for Cyber Security: A Systematic Literature Review 2024.
- Tseng, P.; Yeh, Z.; Dai, X.; Liu, P. Using LLMs to Automate Threat Intelligence Analysis Workflows in Security Operation Centers 2024.
- Enhancing Cybersecurity: The Role of AI & ML in SOC and Deploying Advanced Strategies.
- Ferrag, M.A.; Alwahedi, F.; Battah, A.; Cherif, B.; Mechri, A.; Tihanyi, N. Generative AI and Large Language Models for Cyber Security: All Insights You Need 2024.
- What is Qdrant? - Qdrant.
- GitHub - qdrant/qdrant-rag-eval: This repo is the central repo for all the RAG Evaluation reference material and partner workshop.
- BAAI/bge-large-en · Hugging Face.
- Matryoshka Representation Learning 2022.
- openai. New and improved embedding model | OpenAI.
- bert. [1810.04805] BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.
- Steck, H.; Ekanadham, C.; Kallus, N. Is Cosine-Similarity of Embeddings Really About Similarity? 2024. [CrossRef]
- Guo, K.H. Testing and Validating the Cosine Similarity Measure for Textual Analysis.
- GitHub - redcanaryco/atomic-red-team: Small and highly portable detection tests based on MITRE’s ATT&CK.
- Test your defenses with Red Canary’s Atomic Red Team.
- Landauer, M.; Mayer, K.; Skopik, F.; Wurzenberger, M.; Kern, M. Red Team Redemption: A Structured Comparison of Open-Source Tools for Adversary Emulation 2024.
- Rules - Data analysis · Wazuh documentation.
- Event logging - Wazuh server · Wazuh documentation.
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; tau Yih, W.; Rocktäschel, T.; Riedel, S.; Kiela, D. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks 2020.
- Vulnerability detection - Use cases · Wazuh documentation.
- Cichonski, P.; Millar, T.; Grance, T.; Scarfone, K. Computer Security Incident Handling Guide : Recommendations of the National Institute of Standards and Technology, 2012. [CrossRef]
- IRM/EN at main · certsocietegenerale/IRM.
- socfortress/Playbooks: Playbooks for SOC Analysts.
- Diogenes, Y.; Ozkaya, E. Cybersecurity, attack and defense strategies : infrastructure security with Red Team and Blue Team tactics; Packt Publishing, 2018.
- Nccic.; Ics-cert. Recommended Practice: Improving Industrial Control System Cybersecurity with Defense-in-Depth Strategies Industrial Control Systems Cyber Emergency Response Team, 2016.
- Cybersecurity_incident_response_1731275000.
- Cybersecurity Incident & Vulnerability Response Playbooks Operational Procedures for Planning and Conducting Cybersecurity Incident and Vulnerability Response Activities in FCEB Information Systems.
- The NIST Cybersecurity Framework (CSF) 2.0, 2024. [CrossRef]
- MITRE ATT&CK®.
- How to split text by tokens | LangChain.






Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).