ABD3ID CyberGuard 32B

πŸ›‘οΈ ABD3ID CyberGuard 32B

Defensive Cybersecurity β€’ Full-Parameter Fine-Tuning β€’ Evaluation Research

A 32B-class full-parameter fine-tuning project for evidence-based defensive cybersecurity, developed and maintained by Abdallah Eid.

Status Model Training Framework Focus


πŸ” Project Overview

ABD3ID CyberGuard 32B is a full-parameter fine-tuning project for adapting a 32B-class language model to authorized defensive cybersecurity workflows. Training was performed with the Hugging Face Transformers framework, updating the model weights.

The objective is to investigate whether supervised fine-tuning can improve a general-purpose language model's ability to assist security analysts, developers, researchers, and defenders with evidence-based cybersecurity tasks.

The intended system focuses on authorized defensive workflows, including security analysis, vulnerability interpretation, secure-code review, incident-response assistance, remediation planning, and cybersecurity education.

Current status: Full-model fine-tuning completed; checkpoint files are available.
Checkpoint manifest: Add the exact weight filenames, formats, and sizes from the released checkpoint before publishing this README.
Evaluation status: Formal reproducible held-out benchmarking is still required.

Project Scope & Technical Contributions

The project, authored and maintained by Abdallah Eid, documents a full-model adaptation effort for defensive cybersecurity:

  • Full-parameter supervised fine-tuning of the base-model weights using Hugging Face Transformers.
  • Documentation of the training workflow, model and tokenizer identification requirements, and checkpoint packaging requirements.
  • A defensive task taxonomy spanning vulnerability analysis, secure-code review, incident triage, and remediation.
  • A base-versus-fine-tuned-model evaluation framework covering held-out testing, safety, false positives, and expert review.
  • Explicit separation of released artifacts, illustrative examples, planned capabilities, and measured results.

These points describe the repository's technical scope. They are not a claim of verified model accuracy, superiority over the base model, or completed independent evaluation.


🎯 Research Objectives

CyberGuard is designed around five primary research areas:

Area Intended Capability
πŸ”Ž Threat Analysis Explain alerts, suspicious behavior, and security findings
🧩 Vulnerability Analysis Interpret vulnerabilities and prioritize remediation
πŸ’» Secure Code Review Identify insecure patterns and recommend safer implementations
🚨 Incident Response Summarize evidence and generate defensive response checklists
πŸŽ“ Security Education Explain cybersecurity concepts in authorized environments

🧠 Training & Evaluation Workflow (Conceptual)

╔══════════════════════════════════════╗
β•‘      CURATED DEFENSIVE DATA         β•‘
β•‘  Security β€’ Code β€’ CVEs β€’ IR Data   β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•€β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
                   β”‚
                   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚        DATA QUALITY PIPELINE         β”‚
β”‚                                      β”‚
β”‚  Deduplicate                         β”‚
β”‚      ↓                               β”‚
β”‚  Security Review                     β”‚
β”‚      ↓                               β”‚
β”‚  Quality Filtering                   β”‚
β”‚      ↓                               β”‚
β”‚  Train / Validation / Test Split     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
                   β–Ό
╔══════════════════════════════════════╗
β•‘      INSTRUCTION FORMATTING         β•‘
β•‘             +                        β•‘
β•‘          TOKENIZATION                β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•€β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
                   β”‚
                   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                                      β”‚
β”‚         32B-CLASS BASE MODEL         β”‚
β”‚                                      β”‚
β”‚  Model weights initialized from      β”‚
β”‚          the base model              β”‚
β”‚                                      β”‚
β”‚                 ↓                    β”‚
β”‚                                      β”‚
β”‚  FULL-PARAMETER SUPERVISED           β”‚
β”‚       FINE-TUNING WITH               β”‚
β”‚  HUGGING FACE TRANSFORMERS            β”‚
β”‚                                      β”‚
β”‚  Update model parameters ΞΈ           β”‚
β”‚                                      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
                   β–Ό
╔══════════════════════════════════════╗
β•‘     SUPERVISED FINE-TUNING          β•‘
β•‘                                      β•‘
β•‘   Optimize model weights on the      β•‘
β•‘       documented training data      β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•€β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
                   β”‚
                   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚        HELD-OUT EVALUATION           β”‚
β”‚                                      β”‚
β”‚  β€’ Security accuracy                 β”‚
β”‚  β€’ Remediation quality               β”‚
β”‚  β€’ False-positive analysis           β”‚
β”‚  β€’ Safety evaluation                 β”‚
β”‚  β€’ Base-model comparison             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚
                   β–Ό
╔══════════════════════════════════════╗
β•‘          HUMAN REVIEW               β•‘
β•‘                                      β•‘
β•‘     Security Expert Validation       β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•€β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
                   β”‚
                   β–Ό
              RELEASE DECISION

βš™οΈ Full-Parameter Fine-Tuning

CyberGuard uses full-parameter supervised fine-tuning with the Hugging Face Transformers framework. Unlike LoRA, this approach updates the base model's trainable weights and produces a fine-tuned model checkpoint rather than a standalone low-rank adapter.

The training starts from the identified base-model parameters and optimizes the model against the supervised training objective. Exact model revisions, trainable parameter counts, optimizer settings, precision, and hardware should be reported from the actual training configuration; they are not inferred here.

Conceptually:

Initial model parameters:

ΞΈβ‚€

Supervised fine-tuning:

ΞΈ* = arg min_ΞΈ L(D_train; ΞΈ)

Fine-tuned model checkpoint:
ΞΈ*

Where:

  • ΞΈβ‚€ = parameters from the exact base-model revision
  • ΞΈ = model parameters optimized during training
  • D_train = the documented training dataset
  • L = the supervised training objective
  • ΞΈ* = parameters saved in the resulting full-model checkpoint

Transformers is the training framework; full-parameter fine-tuning describes the weight-update method. Publish the training configuration and checkpoint manifest so these claims can be independently inspected.


πŸ“Š Training & Evaluation Dashboard

βœ… Current Release Status

Component Status
Project design βœ… Complete
Full-model checkpoint βœ… Produced; file manifest to be documented
Training framework βœ… Hugging Face Transformers
Tokenizer assets βœ… Available
Model card βœ… Available
Project artwork βœ… Available
Formal held-out benchmark πŸ§ͺ Pending
Independent expert review πŸ§ͺ Recommended
Workstream Status
Project design and full-model fine-tuning Complete
Fine-tuned checkpoint Produced; filenames and sizes to be documented
Documentation In progress
Formal benchmarking Pending
Independent review Recommended; pending

Reproducibility Manifest

For a technically auditable full-model release, publish the following values from the actual training run and checkpoint metadata. Do not infer missing settings:

Record Required details
Base model Repository or provider, exact model identifier, revision/commit, and applicable license
Tokenizer Identifier and revision, tokenizer files, special tokens, and chat template
Training data Sources, licenses, provenance, filtering, deduplication, split sizes, and contamination controls
Training recipe Transformers and PyTorch versions, optimizer, learning-rate schedule, warmup, weight decay, batch and accumulation settings, sequence length, precision, steps/epochs, and stopping criteria
Runtime Hardware/GPU model, memory, software environment, distributed-training configuration, and random seeds
Checkpoint Exact file names, formats, per-file and total sizes, configuration files, and SHA-256 checksums
Inference Prompt template, generation parameters, runtime versions, and hardware used for reported results

This manifest makes the training and release inspectable; it does not by itself establish model quality or benchmark performance.


πŸ† Evaluation Plan

CyberGuard's full-model fine-tuning is complete; the repository does not currently publish reproducible benchmark scores. Evaluation is a separate research stage, not an implied property of the checkpoint.

Evaluation dimension Evidence required Current status
Defensive safety Documented misuse-resistance test set and reviewed outcomes Not measured
Security relevance Held-out tasks with a published scoring rubric Not measured
Remediation quality Expert-rated correctness, completeness, and applicability Not measured
Incident triage Evidence-grounded decisions assessed against reference criteria Not measured
Vulnerability analysis Held-out vulnerability cases with per-class results Not measured
Secure code review Reproducible code samples and verified findings Not measured
Evidence handling Citation/support checks and unsupported-claim analysis Not measured

Report results only after publishing the benchmark data or a suitable data card, exact base-model and fine-tuned-checkpoint revisions, inference settings, scoring methodology, and limitations. Include sample counts and uncertainty where appropriate; do not convert project progress or intended capabilities into accuracy scores.


πŸ“ˆ Base Model vs CyberGuard β€” Benchmark Framework

                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚   HELD-OUT TEST    β”‚
                 β”‚        SET         β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚
               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
               β–Ό                       β–Ό
      β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
      β”‚   BASE MODEL   β”‚      β”‚   CYBERGUARD   β”‚
      β”‚   Unmodified   β”‚      β”‚ Full Fine-Tunedβ”‚
      β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚                        β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β–Ό
               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
               β”‚ IDENTICAL METRICS   β”‚
               β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
               β”‚ Security accuracy   β”‚
               β”‚ Remediation quality β”‚
               β”‚ False positives     β”‚
               β”‚ Secure-code review  β”‚
               β”‚ Incident triage     β”‚
               β”‚ Safety behavior     β”‚
               β”‚ Evidence quality    β”‚
               β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β–Ό
                  SECURITY EXPERT
                      REVIEW
Metric Base Model CyberGuard Status
Security Accuracy TBD TBD πŸ§ͺ Pending
Remediation Quality TBD TBD πŸ§ͺ Pending
False-Positive Rate TBD TBD πŸ§ͺ Pending
Secure-Code Review TBD TBD πŸ§ͺ Pending
Incident Analysis TBD TBD πŸ§ͺ Pending
Safety Evaluation TBD TBD πŸ§ͺ Pending

Measured values should only be inserted after reproducible evaluation under identical conditions.


πŸ§ͺ Evaluation Methodology

The fine-tuned model should be compared directly with the exact unmodified base-model revision.

Evaluation should include:

  1. A held-out test set separated before training, with documented provenance and contamination checks.
  2. The same prompt, context, tokenizer assumptions, and inference settings for the base and fine-tuned models.
  3. A declared task taxonomy and scoring rubric, with per-task results rather than a single opaque aggregate.
  4. Appropriate classification metrics (for example, precision, recall, F1, and false-positive/negative rates) alongside expert-rated response quality.
  5. Security-specific checks for evidence support, severity calibration, remediation correctness, and unsafe recommendations.
  6. Repeated or paired comparisons where appropriate, including sample counts and uncertainty estimates.
  7. Blinded independent review for security-sensitive outputs when feasible, with reviewer instructions and disagreement handling documented.
  8. Versioned model, tokenizer, dataset, and evaluation code identifiers, plus all decoding parameters.
  9. Failure analysis with representative errors, limitations, and known coverage gaps.
  10. A release gate that reports scores only after the methodology and results are reproducible.

πŸ’‘ Example 1 β€” Vulnerability Analysis

User

A web application builds a SQL query by directly concatenating
user-supplied input.

What is the security risk and what defensive remediation
would you recommend?

Illustrative CyberGuard Response

Finding:
Potential SQL injection vulnerability.

Reason:
Untrusted user input is being incorporated directly into a SQL
statement. An attacker may be able to alter the intended query.

Recommended remediation:

1. Replace string concatenation with parameterized queries or
   prepared statements.

2. Validate input according to the application's expected data
   format.

3. Apply least-privilege permissions to the database account.

4. Review similar database queries throughout the application.

5. Add security tests covering malicious and malformed inputs.

Priority:
High if attacker-controlled input reaches the database query
without parameterization.

This is an illustrative CyberGuard-style example, not a measured benchmark sample.


πŸ’‘ Example 2 β€” Security Alert Triage

User

Our monitoring system detected repeated failed login attempts
against one account from several IP addresses.

How should a security analyst investigate this alert?

Illustrative CyberGuard Response

Initial assessment:
The activity may represent credential guessing, password spraying,
automation, or legitimate authentication failures.

Recommended defensive investigation:

1. Verify the alert timestamps and affected account.

2. Review authentication logs for successful logins occurring
   before, during, or after the failed attempts.

3. Compare source IP addresses, geographic indicators, user agents,
   and authentication patterns.

4. Determine whether other accounts received similar attempts.

5. Review MFA events and account recovery activity.

6. Check whether the affected credentials are known to have been
   exposed.

7. If compromise is suspected, follow the organization's incident
   response procedure and protect the affected account.

Evidence required:
Do not classify the event as a confirmed compromise based only
on failed authentication attempts.

This example demonstrates the intended evidence-based defensive analysis style and is not a measured benchmark result.


πŸ›‘οΈ Intended Safety Boundary

CyberGuard is intended for:

βœ“ Defensive security analysis
βœ“ Authorized vulnerability assessment
βœ“ Secure software development
βœ“ Security education
βœ“ Incident investigation
βœ“ Security remediation
βœ“ Detection engineering
βœ“ Risk prioritization

The project is not intended to replace authorization requirements, organizational security controls, professional judgment, or independent validation.


πŸ“¦ Future Release Contents

A complete research release should include:

ABD3ID-CyberGuard-32B/
β”‚
β”œβ”€β”€ README.md
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ [full-model weight files β€” list actual filenames and sizes]
β”‚   β”œβ”€β”€ config.json
β”‚   └── [tokenizer files β€” list actual filenames]
β”œβ”€β”€ training_config.json
β”œβ”€β”€ evaluation/
β”‚   β”œβ”€β”€ results.json
β”‚   β”œβ”€β”€ base_model_results.json
β”‚   └── methodology.md
β”œβ”€β”€ dataset_documentation/
β”‚   └── DATASET_CARD.md
└── examples/
    └── inference.md

The release documentation should identify:

  • Exact base-model identifier
  • Base-model revision
  • Fine-tuned checkpoint filenames, formats, total size, and checksums
  • Hugging Face Transformers version
  • Training method and trainable parameter count
  • Optimizer and scheduler
  • Learning rate
  • Warmup and weight-decay settings
  • Batch configuration
  • Gradient accumulation and checkpointing settings, if used
  • Context length
  • Precision
  • Training steps / epochs
  • Hardware configuration
  • Dataset composition
  • Evaluation methodology
  • Reproducible benchmark results

⚠️ Limitations

CyberGuard is a research project.

Even after training, model-generated security recommendations may contain incorrect assumptions, incomplete analysis, false positives, or inaccurate remediation guidance.

Outputs should therefore be validated against:

  • Original security evidence
  • System configuration
  • Application source code
  • Vendor documentation
  • Relevant vulnerability information
  • Organizational security policies
  • Qualified human review

A language model should not be treated as an autonomous security authority.


πŸ“œ Project Attribution & License Scope

Project author and maintainer: Abdallah Eid. The project name, original documentation, and other project-authored materials are attributed to Abdallah Eid.

Copyright Β© 2026 Abdallah Eid. All rights reserved for original project materials, except where a separate license or notice applies. This statement does not grant or change rights to third-party materials.

The base model, tokenizer, datasets, libraries, and other third-party components remain governed by their respective owners' licenses and terms. A fine-tuned checkpoint derived from a base model may also be subject to that model's license and distribution conditions. Do not interpret this project attribution as transferring third-party rights or as permission for uses not allowed by the applicable licenses.

Before redistributing or using the fine-tuned checkpoint, identify and document the exact base-model name and revision, its license, and the provenance and terms of any training data. The exact base-model identifier and revision are not specified in this README; confirm them from the checkpoint metadata and training records before making licensing or compatibility claims.


πŸ”¬ Research Transparency

This repository intentionally distinguishes between:

Planned capabilities
Features the research project intends to investigate.

Illustrative examples
Human-written demonstrations of the desired response style.

Measured capabilities
Results obtained through reproducible evaluation after training.

The fine-tuned checkpoint has been produced, but no formal measured CyberGuard benchmark claims are made until reproducible evaluation is completed.


πŸ—ΊοΈ Development Roadmap

PHASE 01  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Project Design       βœ…
                           ↓
PHASE 02  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Full-Model Training  βœ…
                           ↓
PHASE 03  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  Repository Release   βœ…
                           ↓
PHASE 04  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘  Documentation        πŸ”„
                           ↓
PHASE 05  β–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  Formal Benchmarking  πŸ§ͺ
                           ↓
PHASE 06  β–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  Expert Review        πŸ§ͺ
                           ↓
PHASE 07  β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘  Verified Evaluation  ⏳

πŸ›‘οΈ ABD3ID CyberGuard 32B

Defensive AI β€’ Full-Parameter Fine-Tuning β€’ Transformers

Developed and maintained by Abdallah Eid

Status: Full-Model Fine-Tuning Complete β€’ Formal Evaluation Pending

Security decisions require evidence, validation, and human review.

Downloads last month
316
Safetensors
Model size
7B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including abdallah3id/ABD3ID-CyberGuard-32B