Instructions to use simpledirect/Vinci-MLE-30B-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use simpledirect/Vinci-MLE-30B-1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="simpledirect/Vinci-MLE-30B-1.0") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("simpledirect/Vinci-MLE-30B-1.0") model = AutoModelForCausalLM.from_pretrained("simpledirect/Vinci-MLE-30B-1.0", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use simpledirect/Vinci-MLE-30B-1.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "simpledirect/Vinci-MLE-30B-1.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simpledirect/Vinci-MLE-30B-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/simpledirect/Vinci-MLE-30B-1.0
- SGLang
How to use simpledirect/Vinci-MLE-30B-1.0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "simpledirect/Vinci-MLE-30B-1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simpledirect/Vinci-MLE-30B-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "simpledirect/Vinci-MLE-30B-1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simpledirect/Vinci-MLE-30B-1.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use simpledirect/Vinci-MLE-30B-1.0 with Docker Model Runner:
docker model run hf.co/simpledirect/Vinci-MLE-30B-1.0
Vinci-MLE-30B-1.0
Open-weight ML engineering research model trained for evidence-based diagnosis, bounded workspace changes, and justified no-change or stop decisions.
Canadian-developed by SimpleDirect · Apache-2.0 · Research baseline
Vinci-MLE-30B-1.0 is a 30B-class adaptation of IBM Granite 4.1 30B, developed by SimpleDirect, a Canadian AI lab. MLE stands for machine-learning engineering: inspecting experiment artifacts, diagnosing problems, making bounded changes, and reporting what the evidence supports.
Vinci MLE 1.0 establishes an open-weight research baseline for this workflow: a specialization corpus, an instrumented tool interface, selective assistant supervision, and standalone merged BF16 weights with the tokenizer and chat template included.
Evaluation scope: 1.0 reports its ML-engineering specialization and measured general-capability retention. A valid held-out MLE-task evaluation was not completed for this release and is intended as a release-gating measurement for 1.1.
Specialization · MLE evaluation status · Quick start · General capability retention · Limitations
Model at a glance
| Property | Value |
|---|---|
| Developer | SimpleDirect, Canada |
| Parent | ibm-granite/granite-4.1-30b |
| Architecture | GraniteForCausalLM, 64 layers |
| Format | Full merged BF16 Safetensors; 12 weight shards |
| Weight-file size | 57,731,524,544 bytes; approximately 57.73 GB / 53.77 GiB |
| Specialization | Supervised fine-tuning with LoRA, then deterministic merge |
| Training corpus | 91 trajectories; 64 distinct task roots; 8 task families |
| Training / evaluation serving limit | 16,384 tokens |
| Parent-configured position limit | 131,072 tokens; not a validated MLE context-length claim |
| Evaluation serving runtime | vLLM 0.22.1; BF16; tensor parallel 1; hermes tool parser |
| License | Apache-2.0, with NOTICE |
Artifact sizes above were read from the Hugging Face API on September 28, 2026. Weight-file size is not a runtime-memory requirement: allow additional memory for the KV cache, activations, runtime, context length, and concurrent requests.
ML engineering specialization
Vinci MLE is trained as a policy inside an instrumented ML workspace, not only on question-and-answer pairs. Its specialization examples combine file inspection, tool feedback, evidence-based diagnosis, bounded interventions, and structured reports of observations, changes, unknowns, and next evidence.
The frozen specialization corpus contains 91 trajectories across 64 distinct task roots in eight task families. Both sizes consumed the same tokenized training examples; their LoRA targets and optimization recipes differ. The family names below follow the training manifest, with punctuation expanded for readability.
| ML-engineering task family | Training scope |
|---|---|
| Duplicate data / data quality | Inspect duplication and data-quality problems |
| Experiment design / interpretation | Interpret experiment evidence and assess whether a change is justified |
| Failed resume | Diagnose failed training-resume behavior |
| Memory / batching / accumulation | Inspect memory, batching, and gradient-accumulation problems |
| Metric / evaluator defect | Diagnose defects in metrics and evaluators |
| NaN / divergence | Investigate numerical failures and divergence |
| Packaging / serving / reproducibility | Inspect packaging, serving, and reproducibility problems |
| Tokenizer / model configuration | Diagnose tokenizer and model-configuration problems |
Act / No Change / Stop is central to the training design. An ML-engineering policy should not generate changes indiscriminately: preserving a correct setup or requesting missing evidence can be the appropriate response.
| Trained response | Meaning in the specialization corpus | Trajectories |
|---|---|---|
| Act | Make a bounded change when the evidence supports a defect and an intervention | 45 |
| No Change | Preserve a correct setup and explain the evidence | 23 |
| Stop / Needs Evidence | Identify missing evidence or a blocker instead of inventing a repair | 23 |
These counts describe training coverage, not success rates. They do not establish that the model chooses the correct action, no-change, or stop decision on new tasks.
MLE evaluation status
Vinci MLE 1.0 does not yet have a valid held-out ML-engineering task score.
The internal development screen did not provide a valid capability comparison: the parent and every development fine-tune produced parse errors or timeouts, with no successful episodes. Those instrument-floor outcomes did not establish a valid MLE-task score for the released checkpoint. The recipe-selection details are retained in Training.
The preregistered held-out MLE evaluation and external MLE-bench evaluation were not completed for this release. We report MLE training scope separately from measured general-capability comparisons and make no MLE-task performance claim for 1.0.
The full evaluation also requires a qualified execution environment with the intended isolation and reproducibility controls. That environment was not ready for the complete MLE evaluation in 1.0, and we chose not to substitute a weaker test as an equivalent result.
For Vinci MLE 1.1, we intend to run the held-out MLE evaluation and external MLE benchmark suite as release-gating measurements, alongside general capability-retention evaluations. Results will be reported when measured; this 1.0 card does not predict them or promise a release date.
Intended workflow
Illustration of training intent; tools are supplied by the application. This is not a captured model run.
A task supplies a problem statement, workspace files, and tool feedback. The policy inspects evidence, chooses an action or a justified no-change/stop response, and produces a structured final report. Applications must implement a bounded interaction loop, record tool outputs, and independently test proposed changes.
The weights alone are not the full Vinci MLE agent harness. This repository supplies the model and tokenizer. An application must provide and validate its own tools, sandbox, permissions, and workspace.
Training interface
The workspace interface uses task-specific subsets of inspect_files, run_command (an argv array), and write_file (path and content). All 91 include run_command; 36 expose all three tools. Final reports use a fenced json FINAL block with decision, root_cause, evidence, changes, facts, unknowns, why_blocks, and next_evidence fields. Reproducing that workflow requires the matching tool schema and system instructions; a generic chat prompt is a different use of the model.
The stored final-report decision values are INTERVENED for Act, NO_CHANGE for No Change, and STOP for Stop / Needs Evidence.
For an initial text-only exploration, the quick start asks about a small data-splitting problem. It is an illustrative prompt, not a captured model answer or a benchmark result.
Choosing a size
| MLE 8B | MLE 30B | |
|---|---|---|
| Parent | Granite 4.1 8B | Granite 4.1 30B |
| Weight download | 17.58 GB, 4 shards | 57.73 GB, 12 shards |
| Training examples / task roots | 91 / 64 | 91 / 64 |
| LoRA rank / alpha | 8 / 16 | 32 / 64 |
| Trained projection families | Attention | Attention and MLP |
| BFCL score, 0–100 | 68.8584 | 73.5160 |
| LiveCodeBench score, 0–100 | 32.9858 | 40.3791 |
| Tool-call smoke | C1 failed; other cases passed | All cases passed |
The 30B weights occupy approximately 57.73 GB (53.77 GiB). Plan for substantially more memory than the 8B sibling. The 30B model has higher recorded BFCL and LiveCodeBench scores. The benchmark settings and LiveCodeBench timing caveat below apply to the comparison. The two models use the same training examples but different recipes; this is not a controlled experiment isolating parameter count, and there is no measured MLE-task ranking between them. Choose 30B when its higher measured tool/coding scores justify the larger download and serving footprint. Choose 8B when a smaller download and lower weight-memory footprint matter more.
Quick start
Download a pinned snapshot
Install huggingface_hub, and authenticate with hf auth login if repository access requires it. This example pins the uploaded model revision evaluated by the release records; later card-only revisions do not change those model files.
from huggingface_hub import snapshot_download
model_dir = snapshot_download(
repo_id="simpledirect/Vinci-MLE-30B-1.0",
revision="f0c85ea26a373c4bd0b217dfd66ba9f730ef10eb",
local_dir="./vinci-mle-30b",
)
print(model_dir)
The pinned snapshot includes its historical README and matching checksums. To download a later card revision, use that revision's full commit hash and verify its own SHA256SUMS.
Generate with Transformers
Requires PyTorch, Accelerate, and Transformers with Granite support. Load the merged checkpoint directly; no PEFT adapter needs to be applied.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_dir = "./vinci-mle-30b"
tokenizer = AutoTokenizer.from_pretrained(model_dir, trust_remote_code=False)
model = AutoModelForCausalLM.from_pretrained(
model_dir,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=False,
).eval()
messages = [
{"role": "system", "content": "You are an ML engineering research assistant. Separate observations from hypotheses and explain what to check next."},
{"role": "user", "content": "A dataset contains multiple records per customer. Training and validation rows were split at random. What should I inspect before trusting the validation score? Do not claim to have run any checks."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
)
device = model.get_input_embeddings().weight.device
inputs = {key: value.to(device) for key, value in inputs.items()}
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=512, do_sample=False)
completion = output[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(completion, skip_special_tokens=True))
This is a reference loading example, not a rerun of the release evaluation. It uses an illustrative system prompt and does not reproduce the Vinci agent harness. No generated answer is asserted by this card.
Serve an OpenAI-compatible endpoint with vLLM
The release evaluation used vLLM 0.22.1, a 16,384-token limit, greedy decoding, the packaged chat template, and the hermes tool-call parser. After downloading the snapshot above, the following launch pattern preserves those core model settings; the local port and served alias are chosen for this example.
CUDA_VISIBLE_DEVICES=0 \
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 VLLM_USE_DEEP_GEMM=0 \
python -m vllm.entrypoints.openai.api_server \
--model ./vinci-mle-30b \
--tokenizer ./vinci-mle-30b \
--chat-template ./vinci-mle-30b/chat_template.jinja \
--served-model-name vinci-mle-30b \
--dtype bfloat16 \
--tensor-parallel-size 1 \
--max-model-len 16384 \
--gpu-memory-utilization 0.90 \
--seed 0 \
--enforce-eager \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--host 127.0.0.1 --port 8000
Use a GPU with enough memory for the weights and serving overhead. Tensor parallel 1 describes the evaluated setup; this card does not certify a minimum GPU model or all alternative parallel configurations.
Send requests to http://127.0.0.1:8000/v1/chat/completions, using model: "vinci-mle-30b", temperature: 0, and an explicit max_tokens budget. For tool use, supply the tools your application implements and validate every returned name and argument. The vLLM tool-calling documentation explains the API contract; the parser setting alone does not provide a tool executor.
30B tool-call smoke: all cases passed in the fixed smoke run. This checks a small set of output-format cases; it does not establish reliable tool use on arbitrary tasks.
Example: request a workspace tool call
The three declarations below match the three-tool training variant used in 36 trajectories. This example sends a request and prints the model response; it does not execute returned commands. The system instruction is the common training instruction. The user task is illustrative. A complete agent needs a workspace, permission checks, tool implementations, returned observations, and a bounded generation loop.
import requests
tool_definitions = [
{
"name": "inspect_files",
"description": "Print the task objective and every public workspace file.",
"parameters": {
"type": "object", "properties": {}, "additionalProperties": False,
},
},
{
"name": "run_command",
"description": "Run a command in the workspace and return its stdout and stderr.",
"parameters": {
"type": "object",
"properties": {"argv": {"type": "array", "items": {"type": "string"}}},
"required": ["argv"], "additionalProperties": False,
},
},
{
"name": "write_file",
"description": "Replace a workspace file with the given full content.",
"parameters": {
"type": "object",
"properties": {"path": {"type": "string"}, "content": {"type": "string"}},
"required": ["path", "content"], "additionalProperties": False,
},
},
]
response = requests.post(
"http://127.0.0.1:8000/v1/chat/completions",
json={
"model": "vinci-mle-30b",
"messages": [
{"role": "system", "content": "Assess this private TRAIN workspace using recorded tools and a structured FINAL answer."},
{"role": "user", "content": "Inspect the workspace to identify how the training and validation data are split. Gather evidence before proposing a change."},
],
"tools": [{"type": "function", "function": tool} for tool in tool_definitions],
"tool_choice": "auto",
"temperature": 0,
"max_tokens": 512,
},
timeout=120,
)
response.raise_for_status()
print(response.json()["choices"][0]["message"])
Install requests to run this client example. Neither these tool descriptions nor the model's response implement a sandbox. Only expose tools whose actions your application can validate and contain.
General capability retention
The evaluations below measure coding, tool calling, general reasoning, agentic behavior, refusal patterns, and long-context retrieval relative to the pinned Granite parent. They characterize broader parent capabilities after MLE specialization. They are not substitutes for a held-out ML-engineering benchmark. Only the tool-calling and coding comparisons have the stated retention pass criterion; the other studies are descriptive and retain their individual caveats. Across these descriptive studies, the 30B release remains broadly close to its Granite parent, with some positive and some negative differences; these measurements characterize retention rather than MLE-specific improvement.
Results below were resolved from the frozen release records on September 28, 2026, against this model's own pinned Granite parent. The summary job verified each source receipt's SHA-256 and used the frozen evaluator for the BFCL and LiveCodeBench comparisons. These are measurements of the BF16 merged release, not of a quantized conversion.
Tool calling and coding
BFCL measures tool-call retention; LiveCodeBench measures coding retention. The preregistered criterion permits a score drop of at most 3 percentage points from the parent under the same harness. Scores below are on a 0–100 scale; differences are this model minus its parent, in percentage points.
| Benchmark | Granite 30B parent | Vinci MLE 30B | Difference (pp) | Retention margin |
|---|---|---|---|---|
| BFCL | 73.5616 | 73.5160 | -0.0457 | Within margin |
| LiveCodeBench | 40.6635 | 40.3791 | -0.2844 | Within margin* |
BFCL covers 11 categories and excludes the executable-code category. LiveCodeBench covers 1,055 items. Both recorded scores fall within the preregistered 3-point retention margin. LiveCodeBench retains the separate timing-qualification limitation described below. These comparisons do not demonstrate improved MLE-task performance.
* LiveCodeBench timing adjudication: INCONCLUSIVE / faster-machine qualification failed. A timing flag affected 74 of the parent's and 74 of this model's 1,055 items. The proposed faster regrading host did not meet its qualification threshold (ratio 0.908 against a maximum of 0.90), so no regrade was performed. The table reports the existing record grades with that unresolved timing limitation.
LiveCodeBench grading used the evaluator's full sandbox isolation; container stderr was not retained. BFCL multi-turn calls execute in the benchmark's in-process simulated Python environments, without a separate sandbox, as in the upstream benchmark design.
General knowledge, reasoning, and instruction following
Descriptive; no multiplicity correction; not a retention verdict. These paired comparisons have no preregistered pass rule. Scores and intervals are shown as percentages / percentage points, converted from the recorded fractions. The interval is the paired bootstrap 95% interval for the difference, not an interval around either model's score.
| Benchmark | Items | Parent (%) | Vinci MLE (%) | Difference (pp) | 95% interval for difference (pp) |
|---|---|---|---|---|---|
| AIME | 60 | 21.67 | 25.00 | +3.33 | [-3.33, +10.00] |
| BBH | 6,511 | 63.46 | 63.97 | +0.51 | [+0.05, +0.95] |
| GPQA-Diamond | 198 | 36.36 | 43.94 | +7.58 | [+0.51, +14.14] |
| GSM8K | 1,319 | 94.47 | 94.16 | -0.30 | [-0.76, +0.08] |
| IFEval | 541 | 87.43 | 86.69 | -0.74 | [-2.03, +0.55] |
| MATH-500 | 500 | 76.40 | 78.60 | +2.20 | [-0.20, +4.60] |
| MMLU-Pro | 12,032 | 65.58 | 65.28 | -0.30 | [-0.76, +0.16] |
Values are rounded from the source record; subtracting displayed scores can differ slightly from the separately rounded recorded difference. Compare only runs using the same task definitions, prompts, and scoring configuration. An interval excluding zero in this descriptive table is not a multiplicity-adjusted superiority claim.
IFEval uses strict prompt-level accuracy; MATH-500 uses math_verify. The other rows use exact match with their locked extraction rules: GPQA flexible-extract, GSM8K strict-match, MMLU-Pro custom-extract, and BBH get-answer.
Agentic evaluation
Descriptive; no multiplicity correction; not a retention verdict. The tau2 study reports pass¹ over four trials per task: 50 airline tasks (200 trials) and 114 retail tasks (456 trials). Confidence intervals resample whole tasks, preserving the four correlated trials.
| tau2 measure | Trials | Parent (%) | Vinci MLE (%) | Difference (pp) | 95% interval for difference (pp) |
|---|---|---|---|---|---|
| Airline, primary | 200 | 26.00 | 31.50 | +5.50 | [-0.50, +13.00] |
| Retail, primary | 456 | 67.54 | 67.32 | -0.22 | [-4.61, +4.17] |
| Retail, secondary | 456 | 63.60 | 62.72 | -0.88 | [-5.26, +3.51] |
Airline primary uses the upstream database × communication reward. Retail primary measures database end-state success only, which differs from the upstream retail reward. Retail secondary combines database success with natural-language assertions judged by a fixed local Qwen3.8-27B model; it is an approximation, not the official tau2 retail score. Non-completed trials score zero; the secondary comparison excludes pairwise unjudged trials, with its actual included count shown above.
These are agentic benchmark measurements, not ML-engineering task results. Greedy decoding under vLLM continuous batching is not bitwise deterministic, so reruns can produce different item-level outputs.
Safety: refusal-pattern measurements
Reporting updated following a classifier audit. Model weights are unchanged.
Descriptive; no multiplicity correction; not a retention verdict. A9 compares refusal patterns between Vinci MLE and its pinned parent on XSTest and StrongREJECT using two registered refusal-phrase windows: the start of the normalized response and the first 300 normalized characters.
Classifier audit. In a blind review of 90 responses per model by an independent AI agent reviewer, with model identity and classifier output hidden, neither classifier window labelled a compliant answer as a refusal in the audited sample. The first-300-character window identified 54/60 parent and 54/60 Vinci MLE refusals identified by the reviewer. The start-of-response window identified 6/60 in each arm. Because the first-300-character window was substantially more sensitive in this audit, it is reported first below; the registered start-of-response measurement is retained for completeness.
Detected refusal-pattern differences — first 300 characters
| Prompt set | Refusal-phrase window | Paired items | Parent (%) | Vinci MLE (%) | Difference (pp) | 95% interval for difference (pp) |
|---|---|---|---|---|---|---|
| StrongREJECT | First 300 characters | 313 | 91.05 | 91.05 | +0.00 | [-1.28, +1.28] |
| XSTest safe | First 300 characters | 250 | 6.80 | 6.80 | +0.00 | [+0.00, +0.00] |
| XSTest unsafe | First 300 characters | 200 | 87.00 | 86.50 | -0.50 | [-1.50, +0.00] |
Registered start-of-response window
| Prompt set | Refusal-phrase window | Paired items | Parent (%) | Vinci MLE (%) | Difference (pp) | 95% interval for difference (pp) |
|---|---|---|---|---|---|---|
| StrongREJECT | Start of response | 313 | 4.15 | 4.15 | +0.00 | [-0.96, +0.96] |
| XSTest safe | Start of response | 250 | 2.80 | 2.80 | +0.00 | [+0.00, +0.00] |
| XSTest unsafe | Start of response | 200 | 10.00 | 9.50 | -0.50 | [-1.50, +0.00] |
On safe XSTest prompts, a higher detected-refusal rate can indicate increased over-refusal. On unsafe prompts and StrongREJECT, refusal-phrase detection is only a behavioral proxy: it does not by itself establish whether the complete response is safe or harmful.
These measurements characterize refusal-pattern differences relative to the parent model; they are not a general safety evaluation.
Long context: synthetic retrieval
Descriptive; no multiplicity correction; not a retention verdict. The RULER study averages two replicates on 1,500 paired synthetic items across six nominal input lengths. Its string_match_all metric is the fraction of expected strings present, case-insensitive; it is not a binary all-strings-correct measure. Actual tokenized prompt lengths vary around each nominal length.
| Nominal length (tokens) | Paired items | Parent (%) | Vinci MLE (%) | Difference (pp) | 95% interval for difference (pp) |
|---|---|---|---|---|---|
| 1,024 | 250 | 100.00 | 100.00 | +0.00 | [+0.00, +0.00] |
| 2,048 | 250 | 100.00 | 100.00 | +0.00 | [+0.00, +0.00] |
| 4,096 | 250 | 100.00 | 100.00 | +0.00 | [+0.00, +0.00] |
| 8,192 | 250 | 100.00 | 100.00 | +0.00 | [+0.00, +0.00] |
| 12,288 | 250 | 100.00 | 100.00 | +0.00 | [+0.00, +0.00] |
| 14,336 | 250 | 100.00 | 99.92 | -0.08 | [-0.24, +0.00] |
| All lengths | 1,500 | 100.00 | 99.99 | -0.01 | [-0.04, +0.00] |
The 30B parent reached the ceiling at every tested length, limiting what the near-zero parent/candidate differences can establish. The recorded synthetic-retrieval scores remain close within the tested range; this descriptive comparison has no preregistered retention verdict. It does not measure understanding of an ML repository or operation at the parent's full 131,072-position limit.
Repeated-run studies
The release records additionally mark repeated-run variance studies as measured for GSM8K, IFEval, and MATH-500; the repeated-run AIME study was not run at 30B. Their numeric ranges are not included in the resolved numbers summary used here and are not reported in this card. The paired tables above should not be interpreted as proof of bitwise repeatability.
Evaluation scope: what was not run
| Evaluation | 1.0 status |
|---|---|
| Sealed internal MLE evaluation | Not run |
| 48-task held-out MLE evaluation | Not run |
| MLE-bench release subset and MLE-bench Lite | Not run |
| EvalPlus, HumanEval+, MBPP+ | Not run |
| Peer evaluation | Not run |
| Behavioral fidelity of merged release versus adapter | Not measured |
The full preregistered evaluation protocol was not completed for 1.0 because the MLE-specific held-out component remains unmeasured. The benchmark component results above retain their documented scope and caveats; they do not constitute a completed MLE release-evaluation result. No incomplete gate is reported as passed.
Training
Vinci applied supervised fine-tuning to the pinned Granite parent using ordinary LoRA, then merged the terminal adapter into the parent for distribution. This recipe uses neither DoRA nor rank-stabilized LoRA.
| Setting | Value |
|---|---|
| Examples / distinct task roots | 91 / 64 |
| Trajectory composition | 45 direct action / 23 no change / 23 stop |
| Task families | 8, with 8 distinct roots each |
| Epochs / optimizer steps | 1 / 12 |
| Selected checkpoint | Terminal checkpoint-12; the sole final candidate |
| LoRA rank / alpha / dropout | 32 / 64 / 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Learning rate | 4 × 10⁻⁶ |
| Schedule / warmup ratio | Cosine / 0.03 |
| Weight decay | 0.01 |
| Microbatch / gradient accumulation | 1 / 8 |
| Precision | BF16 |
| Sequence ceiling | 16,384 tokens; no packing or truncation |
| Training seed | 18427 |
| Tokenized corpus | 596,976 total tokens; 58,040 supervised tokens |
| Longest training example | 12,023 tokens |
Counts were recomputed from the frozen training receipts for this card. These receipts describe the specialization data, not the Granite parent's pretraining corpus. Both MLE sizes consumed the same tokenized-example bytes.
Selective assistant supervision. The loss includes designated assistant reasoning, successful tool calls, final answers, and justified-stop answers. System/user context, tool results, scaffolding, and failed tool actions are masked. Verifier feedback is excluded from the rendered prompt. Failed actions therefore do not automatically become positive imitation targets. This describes the supervision scheme; it does not establish that the resulting model always chooses a correct action or stop decision.
Development-selection limitation. Parse errors and timeouts drove the parent and eligible development arms to the evaluator floor, so the screen could not support a capability comparison. The frozen recipe-selection fallback was used; full details are retained in the release evidence.
Development screen details
The parent and each of the three development arms produced 14 parse errors and 2 timeouts across 16 episodes, with 0 successes. The screen therefore supplies no valid capability comparison or evidence of improvement over the parent; its records do not independently diagnose the cause of the parse failures. The recipe was chosen by the frozen lower-dose tie-break, not demonstrated superiority. These are development-screen outcomes, not an MLE score for the released model. The final release uses the predetermined terminal checkpoint under the fixed final-training policy.
Limitations
- No completed MLE-specific evaluation. General coding, tool-calling, and reasoning scores cannot fill that gap.
- Small specialization corpus. One epoch over 91 trajectories does not establish broad coverage of ML frameworks, datasets, experiment types, or failure modes. The eight training families describe corpus coverage, not demonstrated success rates.
- Output and tool errors remain possible. Validate final-report fields, tool names, arguments, proposed code, and observed effects. The 30B smoke result is disclosed above.
- LiveCodeBench timing remains unresolved. Read its score with the qualification caveat, not as a fully adjudicated timing result.
- Merge integrity is not behavioral equivalence. The release export passed mechanical verification, standalone loading, and byte-identical rebuild checks. Behavioral fidelity to the original adapter was not measured.
- Context scope is bounded. The 131,072-position parent configuration is not proof of reliable MLE behavior at that length; the evaluation serving limit was 16,384 tokens.
- No safety certification or unattended-deployment validation. Safety benchmark observations do not establish a safety property. Use human review and an isolated, permission-limited environment for executable actions.
Contamination audit
The training-data audit covered 19 benchmark/subset entries. It recorded zero primary-field 13-gram matches and zero substring matches, with one MMLU-Pro secondary-field 13-gram match. Primary-field 8-gram matches were AIME 2025 (1), BFCL simple_java (1), GPQA (1), MATH-500 (1), and MMLU-Pro (4). The upstream source texts used to construct trajectories were not audited, and LiveCodeBench was outside the audit's scope. These checks do not establish a contamination-free evaluation.
Provenance and file integrity
The repository includes PROVENANCE.json, SHA256SUMS, LICENSE, and NOTICE. The provenance file identifies the pinned parent, training manifest, export, and release artifact. Load the merged weights directly and keep their matching configuration, tokenizer, and template together.
| Identity | Value |
|---|---|
| Parent revision | 4fae6278f7132abf5e971f9de49ebbad09c54cce |
| Candidate | mle1-final-30b-g1-a2-release |
| Frozen release manifest SHA-256 | 78256b1fb5dbb5d31135ff2849d6899ed6f8c08e952ce5b58259d4836b5a73c0 |
| Frozen model artifact tree SHA-256 | fb4e2e6012a077057b9dfd2b11b5b9247d670286f8fd995161269d4fa9ede29c |
| Training manifest SHA-256 | e48f81c217f15234ba6c6b6b033f6ea056019e9ad40ba366efa6cf472606d4a1 |
| Tokenized training examples SHA-256 | 61f24f69c97873e029959282480b99274197a5202a4bca73e54d7f4b1fcd7a88 |
The frozen artifact-tree digest identifies the model export, not the current documentation directory. Card updates receive a new Hugging Face commit and a corresponding README checksum; they do not change the frozen model identity.
After downloading one complete revision, run this from its directory:
sha256sum -c SHA256SUMS
# macOS alternative:
shasum -a 256 -c SHA256SUMS
Checksums detect differences against the recorded reference; they do not independently authenticate its publisher. This repository does not include a detached MLE release signature.
Evaluation receipt identifiers
These hashes identify the internal records used for the tables. The full evaluator environment and per-item receipts are not distributed in this model repository, so hashes alone are not a public reproduction package.
| Record | SHA-256 |
|---|---|
| BFCL paired receipt | 960dd214251d857d82422e0c1a14936e3282537807bb55837b4c77349f5dae5d |
| LiveCodeBench paired receipt | c1e5c6a792cf194decae1a8c695990ee5fb2c0d890c4576802a2a7c6d4618758 |
| General-capability paired receipt | b3c4a118a08db38523d11c9c94f5da5f234207b8c58e727fd0e11a23fd090b77 |
| RULER summary | f7465ce480a451aeeb5436b8bb61f74598c0546d42405f58683b2027f4af2e9b |
| Resolved 8B/30B numbers summary | 0f7f6912868f42a26255e860cdbd793e38958eebd4876e562610dcd04048ea92 |
License and attribution
The released weights are licensed under Apache-2.0. See LICENSE and NOTICE for the accompanying terms and attribution. IBM developed Granite; SimpleDirect developed this fine-tuned derivative and its release. No affiliation with or endorsement by IBM is implied.
Citation
@misc{vinci_mle_1_0_30b,
title = {Vinci MLE 1.0 (30B)},
author = {{SimpleDirect}},
year = {2026},
url = {https://huggingface.co/simpledirect/Vinci-MLE-30B-1.0}
}
- Downloads last month
- 7


