Instructions to use CK0607/Qwen3-1.7B-Swarm-Arena-RL-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use CK0607/Qwen3-1.7B-Swarm-Arena-RL-v1 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Qwen3-1.7B-Swarm-Arena-RL-v1
Four distinct LoRA policies trained over one frozen Qwen3 1.7B backbone in the
Swarm Arena 4v4 partially observed graph-control simulator. Policy directories
policy_blue_0 through policy_blue_3 retain separate optimizer identities and
must be assigned to their corresponding BLUE roles.
Selected trainer step: 3. Release status:
not-admitted. Mechanical reproducibility artifact only: development and communication gates did not pass; see results/REPORT.md.
Exact provenance, policy hashes, public input revisions, and compact evaluation
reports are in PROVENANCE.json, SHA256SUMS, and results/.
The reward is the zero-sum terminal control-margin delta. There is no speaking, silence, capture, or learned-judge bonus. Higher return is evidence of task learning; a communication claim additionally requires normal messages to beat dropped, shuffled, and delayed-message interventions on held-out cases.
This is research software for a discrete simulator. It is not evidence of broad swarm intelligence or real-world cybersecurity capability.
- Downloads last month
- -