mindXtrain / docs /HANDOFF.md
Gregory-L's picture
fork mindXtrain from GitHub (Professor-Codephreak/mindXtrain@661bd41) as the mindX-specific line
dfb775d verified
|
Raw
History Blame Contribute Delete
10.9 kB

HANDOFF β€” what you need to do next

This is the ordered checklist for taking the mindxtrain repo from "code is done" to "demo is live." Each step is concrete; check it off when finished.

The repo state at handoff:

  • Single canonical package at mindxtrain/ (12 subpackages, ~100 modules).
  • All stub NotImplementedError paths replaced with real Python (lazy imports for heavyweight deps).
  • 112/112 tests pass on a CPU-only laptop (uv sync + uv run pytest -q).
  • Optional dep groups in pyproject.toml: ml, eval, data, serve, chain, obs. Install only what you need.
  • 12 YAML training recipes wired through the CLI.
  • Coach UI (/coach/) serves all 12 recipes without GPU.

1. Local setup (no GPU; 10 minutes)

cd /home/hacker/Desktop/mindXtrain
cp .env.example .env       # then edit .env to fill in HF_TOKEN, etc.
uv sync                    # base install
uv run pytest -q           # β†’ 112 passed
uv run mindxtrain --help   # all 9 verbs listed

What goes in .env (rest of the file is sane defaults):

Var Where to get it
HF_TOKEN https://huggingface.co/settings/tokens (write scope)
HF_HUB_USERNAME your HF handle
LIGHTHOUSE_API_KEY https://files.lighthouse.storage/dashboard/apikey
MINDXTRAIN_OPENAI_API_KEY optional; only if you want to use openai_compat backend

Optional (on-chain anchors): MINDXTRAIN_REGISTRY_ADDR (ERC-8004 contract), MINDXTRAIN_FACILITATOR_URL (x402 facilitator). The publish path skips these gracefully if unset.

2. Provision the MI300X droplet (sign-up + 30 min)

Fast path (Coach UI): if you've populated GITHUB_TOKEN, AMD_DEV_CLOUD_TOKEN, and AMD_DEV_CLOUD_SSH_KEY_ID in .env, you can skip the manual SSH dance entirely:

  1. uv run uvicorn mindxtrain.operator.app:app --port 8080
  2. Open http://localhost:8080/coach/, scroll to step 6 ("Deploy").
  3. Click β‘  Push to GitHub β†’ β‘‘ Provision MI300X droplet. The droplet boots, cloud-init clones the repo from the SHA you just pushed, pulls the container, and runs mindxtrain bench automatically. All output streams live in the browser via SSE.

Equivalent CLI: mindxtrain github push && mindxtrain droplet provision.

The manual sequence below is preserved for scripted / CI use and as a fallback when the Coach UI isn't available.

# Sign up at https://devcloud.amd.com β€” request a single MI300X.
# Wait for the droplet (typically same-day).
# SSH in:
ssh ubuntu@<droplet-ip>

# Install podman if missing:
sudo apt-get update && sudo apt-get install -y podman podman-compose

# Pull the canonical training container:
podman pull docker.io/rocm/primus:v26.2

# Snapshot the digest into the repo so others can reproduce:
podman inspect --format '{{index .RepoDigests 0}}' rocm/primus:v26.2 \
  | tee -a ops/containerfiles/digest.lock

# Verify the GPU is visible:
podman run --rm --device=/dev/kfd --device=/dev/dri rocm/primus:v26.2 \
  rocminfo | head -50
# β†’ should show gfx942, 192 GB HBM3

Cost watch: $1.99/hr Γ— planned hours. Budget $30 for the full demo pipeline (15 GPU-hours). Leave the droplet stopped when not actively training.

3. Install heavyweight deps inside the container

# On the MI300X:
git clone <your-repo-url> /workspace/mindxtrain
cd /workspace/mindxtrain
podman run -it --rm \
  --device=/dev/kfd --device=/dev/dri \
  -v /workspace/mindxtrain:/workspace/mindxtrain \
  -w /workspace/mindxtrain \
  rocm/primus:v26.2 bash

# Inside the container:
pip install -e ".[ml,eval,data,obs]"
# (skip `serve` and `chain` until you need them β€” they pull large wheels)

4. Run the autotune probe (real, ~60 s)

mindxtrain bench --gpu 0 --out plan.json
cat plan.json | jq '.attention_backend, .gemm_heuristic, .rccl_config'
# β†’ "ck", "hipblaslt_default", "1gpu_noop"

Snapshot plan.json into the repo so the run is reproducible:

cp plan.json ops/k8s/plan-mi300x.json
git add ops/k8s/plan-mi300x.json
git commit -m "snapshot autotune plan from mi300x"

5. Train + eval + quantize (~ 2 hours total for the demo recipe)

# Pick a recipe: instella_3b_lora is the AMD-on-AMD demo path (~30 min).
# Or qwen3_8b_sft_lora for the Qwen side prize (~75 min).
mindxtrain init --template instella_3b_lora --out run.yaml

# Optional: edit run.yaml for your project name, dataset, output path.
$EDITOR run.yaml

# Dataset prep (pulls + dedupes + tokenizes + packs):
mindxtrain dataset prep run.yaml --out ./out/dataset

# Training:
mindxtrain train run.yaml --plan plan.json
# β†’ ./out/runs/<run_name>/checkpoint/

# Evaluation (MMLU subset):
mindxtrain eval run.yaml
# β†’ ./out/runs/<run_name>/eval/lm_eval.json

# Quantize to FP8:
mindxtrain quantize run.yaml
# β†’ ./out/runs/<run_name>/quantized/

If mindxtrain train fails with accelerate not found: you forgot pip install -e ".[ml]" inside the container (step 3).

6. Build the manifest + verify

# Generate the provenance manifest by hashing every artifact:
uv run python -c "
from pathlib import Path
from mindxtrain.config.loader import load_config
from mindxtrain.provenance.manifest import emit_receipt, ProvenanceHashes
cfg = load_config('run.yaml')
run = Path('./out/runs') / cfg.meta.run_name
m = emit_receipt(
    cfg,
    cfg.meta.run_name,
    config_yaml_path=Path('run.yaml'),
    dataset_manifest_path=run / 'dataset_manifest.json',
    checkpoint_dir=run / 'checkpoint',
    eval_json_path=run / 'eval/lm_eval.json',
)
out = run / 'manifest.json'
out.write_text(m.model_dump_json(indent=2))
print(out)
"

# Verify it round-trips:
mindxtrain receipt ./out/runs/<run_name>/manifest.json --config run.yaml
# β†’ all BLAKE3 fields = true (config, checkpoint, autotune_plan; dataset/eval if present)

Auto-emitted receipts (operator + CPU lane). Runs launched through the operator β€” Coach UI or POST /v1/training/jobs β€” now write manifest.json automatically at completion via provenance.manifest.emit_receipt_for_run, alongside config.snapshot.yaml and autotune_plan.json in the run dir. The receipt binds the frozen AutotunePlan hash to the checkpoint hash β€” this is the AOT artifact that makes a run bitwise-verifiable (cf. Verde/RepOps). The Coach "Verifiable receipt" card re-checks it live; mindxtrain receipt does the same from a shell. The manual emit_receipt above remains the full GPU path (dataset + eval JSON included). On MI300X, also snapshot the AOTriton / hipBLASLt tuning caches next to autotune_plan.json so the compiled artifact β€” not just the plan β€” is reproducible across machines.

7. Publish (HF Hub + Lighthouse + mindX register)

# Push to HF (uses HF_TOKEN; private=False for the demo):
mindxtrain publish run.yaml --manifest ./out/runs/<run_name>/manifest.json
# β†’ updates manifest.json in-place with hf_repo_id + lighthouse_cid

If LIGHTHOUSE_API_KEY is unset, the pin step skips gracefully and the manifest gets a cid://stub-… placeholder.

8. Deploy contracts (optional)

The demo can ship without on-chain anchors. Do these once, when ready:

cd contracts
forge install
forge test                       # local Foundry tests pass
forge script script/Deploy.s.sol \
  --rpc-url $MINDXTRAIN_BASE_RPC_URL \
  --private-key $DEPLOYER_KEY \
  --broadcast
# β†’ records contract address; paste into .env as MINDXTRAIN_REGISTRY_ADDR

Once MINDXTRAIN_REGISTRY_ADDR is set, mindxtrain.provenance.erc8004.broadcast_attestation can anchor the manifest BLAKE3 on-chain.

9. Serve the model + wire the production URL

The production URL is https://mindx.pythai.net β€” the Coach UI is at /coach/ and the public training-jobs API is at /v1/training/jobs.

# Inside the rocm/vllm-dev container:
podman-compose -f ops/compose/compose_dev.yaml up -d
# β†’ vLLM-ROCm at :8000, mindxtrain operator FastAPI at :8080

# Verify locally:
curl http://localhost:8080/coach/api/health
# β†’ {"coach_version":"0.1.0", "recipes_available":>=14, ...}

# Public training-jobs API smoke (bearer auth via MINDXTRAIN_API_KEY):
curl -X POST http://localhost:8080/v1/training/jobs \
  -H "Authorization: Bearer $MINDXTRAIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"recipe":"mindx_fallback_qwen3_1_5b_cpu_smoke"}'
# β†’ {"job_id":"...", "status":"running", "backend":"trl_cpu", ...}

# Reverse-proxy mindx.pythai.net β†’ MI300X:8080 (Caddy/Cloudflare).

Once the proxy is live, curl https://mindx.pythai.net/coach/api/health returns 200 from the public internet.

10. Publish & demo

# Push code:
git push origin main

# End-to-end demo walk-through:
#   1. mindxtrain init  β†’ show CLI verbs
#   2. mindxtrain bench β†’ 60-second autotune (the differentiator)
#   3. mindxtrain train β†’ timelapse of training
#   4. mindxtrain quantize β†’ FP8 weights
#   5. curl /v1/chat/completions β†’ live inference
#   6. mindxtrain receipt β†’ BLAKE3 reverify
#   7. Open /coach/ β†’ click through the UI
#   8. Open /coach/dcoach β†’ Imprint & Prove (CPU recall proof)

11. Quality gates (run before every push)

uv run ruff check .
uv run mypy mindxtrain/config mindxtrain/provenance
uv run pytest -q       # β†’ 112 passed

All three must pass before pushing to main. CI runs the same gates on the main branch.


What's still TODO

These paths are wired but require runtime/contracts/services to actually flow end-to-end:

  • x402 metering (mindxtrain.provenance.x402) β€” wired to httpx, needs a deployed facilitator URL.
  • ERC-8004 broadcast (mindxtrain.provenance.erc8004.broadcast_attestation) β€” needs deployed attestation registry + signer key.
  • BANKON ENS allocation (mindxtrain.provenance.algorand.allocate_ens_subname) β€” needs the BANKON allocation service deployed.
  • AgenticPlace listing (mindxtrain.deploy.api_client.list_on_agenticplace) β€” needs agenticplace.pythai.net live.
  • mindX agent register (mindxtrain.deploy.api_client.register_with_mindx) β€” needs mindx.pythai.net/v1/agents live.

The framework itself ships as production-ready Apache-2.0; the integrations above are paid/external services you stand up at your own pace.


Quick reference

What Where
All CLI verbs mindxtrain --help
All recipes mindxtrain init --list
Coach UI http://localhost:8080/coach/
Per-module status docs/actualization_status.md
Architecture docs/architecture.md
Autotune detail docs/autotune.md
Coach detail docs/coach.md
CLI reference docs/cli.md
YAML schema docs/yaml_schema.md
dcoach proof loop docs/dcoach.md
Frozen blueprints docs/blueprints/{mindXtrain,mindXtrain2}.md