Instructions to use minagayid/ELLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use minagayid/ELLM with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM2-135M-Instruct") model = PeftModel.from_pretrained(base_model, "minagayid/ELLM") - Transformers
How to use minagayid/ELLM with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="minagayid/ELLM")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("minagayid/ELLM", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use minagayid/ELLM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "minagayid/ELLM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "minagayid/ELLM", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/minagayid/ELLM
- SGLang
How to use minagayid/ELLM with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "minagayid/ELLM" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "minagayid/ELLM", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "minagayid/ELLM" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "minagayid/ELLM", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use minagayid/ELLM with Docker Model Runner:
docker model run hf.co/minagayid/ELLM
ELLM β Evolving Large Language Model
Not ready for deployment. The adapter labeled 177 of 192 non-release examples as
RELEASEon its held-out synthetic evaluation (92.2%). It fails promotion and must not be used to make operational decisions. Seedocs/evaluation.md. The deterministic evidence gate is a separate last check; an adapter label never authorizes a release or action.
Status: experimental research adapter; promotion rejected. ELLM uses the minimum-device SmolLM2-135M-Instruct checkpoint (135M parameters, Apache-2.0) and a small LoRA adapter for evidence-state triage. This repository contains the ELLM adapter and source code; it does not duplicate the upstream base checkpoint. The loader pins upstream revision 12fd25f77366fa6b3b4b768ec3050bf629380bac and downloads that checkpoint on first use. The upstream card labels the model English; this ELLM release has no Arabic or multilingual quality result.
The adapter was trained only on the included procedurally generated synthetic examples. It was intended for a four-label advisory task (RELEASE, WAIT, ABSTAIN, HUMAN_REVIEW) from structured evidence metadata, but its held-out results show it overwhelmingly predicts RELEASE. Exact accuracy was 79/256 (30.9%); WAIT and ABSTAIN recall were both 0%. This closed-world toy benchmark is not evidence of factual accuracy, source verification, multilingual quality, or real-world safety. Full counts and the rejection decision are in evaluation/results.json and docs/evaluation.md.
Important limits
- The adapter failed promotion and must stay out of operational decisions. Its label is never authorization to publish a claim or perform an action. The deterministic evidence gate is an independent final check, not proof that the adapter is reliable.
- Release requires the configured evidence thresholds: at least 30 days of claim observation and verification span, four fresh signed support records, three publishers and independence groups, one primary source, valid source/reuse/privacy attestations, and no unresolved contradiction. These defaults are deliberately conservative but not empirically calibrated.
- Missing, incomplete, tampered, or conflicting evidence waits, abstains, or goes to a human. High/critical-risk claims and every non-informational action require human review. There is no action executor.
- This is not a foundation model trained from scratch, and the adapter does not add general web knowledge. It is not suitable for consequential decisions or autonomous external actions.
- No whole-internet crawler is included. Public access and permission to retain or train on material are different things; data collection remains operator-scoped and requires a documented reuse basis.
Quick start
Use Python 3.10 or newer. Install the optional model dependencies. The rejected adapter is not loaded by default, and the base-only model also has no deployment-quality result:
python -m pip install -e ".[local-model]"
python -c "from ellm.triage import LocalTriageAdvisor; LocalTriageAdvisor(adapter_id=None, local_files_only=False); print('Base-only research instance loaded; not approved for operational decisions.')"
For a research comparison only, explicitly pass adapter_id='minagayid/ELLM'; that adapter failed promotion and must not be used for operational decisions. For offline base-only use, first download the pinned base checkpoint, then set local_files_only=True. The advisory prediction object exposes label=None if the model returns anything other than one exact allowed label; callers should fail closed.
Reproduce training and evaluation
The complete training set, dev set, and held-out set are generated deterministically by training/synthetic_data.py. No web text or third-party dataset is used. Training requires a local CUDA GPU; it intentionally stops if CUDA is unavailable and never falls back to a paid service.
python -m pip install -e ".[training]"
python training/synthetic_data.py
python training/train_adapter.py --output-dir artifacts/candidates/ellm-experiment
python training/evaluate_adapter.py --adapter-dir artifacts/candidates/ellm-experiment --data training/data/dev.jsonl --output evaluation/ellm-experiment-dev.json
python -m unittest discover -s tests -v
The trainer saves a new candidate in the specified output directory. The evaluation runner compares base-only and base-plus-adapter generations on the chosen split; use development data for iteration, preserve the existing failed test report, and never treat a synthetic score as deployment evidence. Reports require a new output path so previous evidence is not overwritten. These are descriptive toy-task measurements, not a public benchmark score.
Evidence-gated growth design
ELLM separates three kinds of change:
- Observe: only operator-allowlisted sources with reviewed access terms, privacy handling, and a documented reuse basis may enter the local ledger.
- Remember: the append-only SQLite evidence ledger records provenance and signed review/verification receipts. Retrieval only exposes material attached to a claim that passes the gate.
- Adapt: model changes are offline, versioned adapter candidates. A candidate is compared with a fixed held-out set and the previous version; it cannot update itself or the active model.
The model can draft or suggest. It cannot verify its own factual claims, promote evidence, change its weights at runtime, browse the open web, or execute tools. A trusted upstream claim scanner and real external verifiers remain required before any release. HMAC signatures prove that a configured key signed specific metadata; they do not prove that the assessment is true.
NLP training platform
The repository now includes an optional, local-only platform for rights-attested corpus preparation, byte-level tokenizer experiments, pinned pretrained model profiles, causal pretraining, SFT, DPO, GRPO, and a separate multilingual likelihood regression screen. This is scaffolding for controlled experiments; it does not add a general corpus, claim Arabic quality, or turn the existing rejected adapter into a foundation model. See docs/nlp-platform.md for input formats, commands, and limits.
Files and licensing
adapter_config.jsonandadapter_model.safetensorsβ trained ELLM LoRA adapter for SmolLM2-135M-Instruct.ellm/β the local draft backend, structured advisory model wrapper, evidence gate, ledger, and retrieval code.training/β synthetic adapter experiment plus the optional local NLP corpus, tokenizer, pretraining, SFT, DPO, and GRPO tooling.evaluation/β held-out adapter results plus the multilingual likelihood comparison and human-review eligibility gate.tests/β clock, parser, dataset, and release-gate smoke tests. Historic timestamps and test-only HMAC keys simulate code paths; they are not real verification.docs/architecture.md,docs/data-governance.md,docs/deployment-tiers.md,docs/evaluation.md,docs/nlp-platform.mdβ system design, data boundaries, evaluation scope, and the optional NLP workflow.- Project source code is covered by the repository's MIT
LICENSE. The adapter is published as an Apache-2.0 model artifact; its upstream SmolLM2 base is also Apache-2.0 and remains hosted by its owner.
References
- Hugging Face, SmolLM2-135M-Instruct model card and license.
- Hugging Face, PEFT adapter training and loading.
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2020.
- Huang et al., Large Language Models Cannot Self-Correct Reasoning Yet, 2023.
- IETF, RFC 9309: Robots Exclusion Protocol, 2022.
Framework versions
- PEFT 0.21.0
- Downloads last month
- 22
Model tree for minagayid/ELLM
Base model
HuggingFaceTB/SmolLM2-135M