Accessing ASGTransformer During Training and Loading Intermediate Checkpoints

#1
by markdived - opened

Accessing ASGTransformer During Training and Loading Intermediate Checkpoints

Hello ASGTransformer community,

I would like to document and discuss the recommended method for accessing and testing wasmdashai/asg-v1 while the model is being trained.

The main objectives are:

  1. Save intermediate training checkpoints safely.
  2. Upload selected checkpoints to the Hugging Face Hub.
  3. Load a specific checkpoint for evaluation or scenario generation.
  4. Resume interrupted training from the latest checkpoint.
  5. Verify that the custom generate_scenario() method remains available after saving and reloading.
  6. Avoid disrupting the active training process or unnecessarily duplicating the model in GPU memory.

The expected training workflow is:

Training Dataset
      ↓
Tokenizer
      ↓
ASGTransformer
      ↓
Periodic Local Checkpoints
      ↓
Evaluation and Scenario Tests
      ↓
Hugging Face Hub
      ↓
Final Production Model

A checkpoint should contain the complete ASGTransformer configuration, custom architecture files, tokenizer files, model weights, generation configuration, knowledge catalog, and the training state required to resume training.

The model should remain loadable through:

from transformers import AutoModelForCausalLM, AutoTokenizer

checkpoint_id = "wasmdashai/asg-v1"

tokenizer = AutoTokenizer.from_pretrained(
    checkpoint_id,
    trust_remote_code=True,
)

model = AutoModelForCausalLM.from_pretrained(
    checkpoint_id,
    trust_remote_code=True,
    torch_dtype="auto",
    device_map="auto",
)

result = model.generate_scenario(
    tokenizer,
    "Create an authorized defensive cybersecurity training scenario.",
    language="en",
    max_new_tokens=384,
)

print(result["text"])

Questions for discussion:

  • Should development checkpoints be stored in the main branch, a dedicated branch, or separate repositories?
  • What is the recommended checkpoint interval for ASGTransformer?
  • Should optimizer and scheduler states be uploaded, or only the deployable model files?
  • What automated validation tests should run before promoting a checkpoint to the production revision?
  • Should stable releases use semantic tags such as v1.1.0?

The intended outcome is a reliable training and deployment lifecycle in which ASGTransformer can be evaluated during training, resumed after interruption, and published without losing its custom architecture or scenario-generation capabilities.

Sign up or log in to comment