Vinci maker’s mark and wordmark with Vero

Vinci-MLE-30B-1.0

Vinci MLE 30B — Vero at an ML research workbench. Open weights, Apache-2.0, research.

Open-weight ML engineering research model trained for evidence-based diagnosis, bounded workspace changes, and justified no-change or stop decisions.

Canadian-developed by SimpleDirect · Apache-2.0 · Research baseline

Vinci-MLE-30B-1.0 is a 30B-class adaptation of IBM Granite 4.1 30B, developed by SimpleDirect, a Canadian AI lab. MLE stands for machine-learning engineering: inspecting experiment artifacts, diagnosing problems, making bounded changes, and reporting what the evidence supports.

Vinci MLE 1.0 establishes an open-weight research baseline for this workflow: a specialization corpus, an instrumented tool interface, selective assistant supervision, and standalone merged BF16 weights with the tokenizer and chat template included.

Evaluation scope: 1.0 reports its ML-engineering specialization and measured general-capability retention. A valid held-out MLE-task evaluation was not completed for this release and is intended as a release-gating measurement for 1.1.

Specialization · MLE evaluation status · Quick start · General capability retention · Limitations

Model at a glance

Property Value
Developer SimpleDirect, Canada
Parent ibm-granite/granite-4.1-30b
Architecture GraniteForCausalLM, 64 layers
Format Full merged BF16 Safetensors; 12 weight shards
Weight-file size 57,731,524,544 bytes; approximately 57.73 GB / 53.77 GiB
Specialization Supervised fine-tuning with LoRA, then deterministic merge
Training corpus 91 trajectories; 64 distinct task roots; 8 task families
Training / evaluation serving limit 16,384 tokens
Parent-configured position limit 131,072 tokens; not a validated MLE context-length claim
Evaluation serving runtime vLLM 0.22.1; BF16; tensor parallel 1; hermes tool parser
License Apache-2.0, with NOTICE

Artifact sizes above were read from the Hugging Face API on September 28, 2026. Weight-file size is not a runtime-memory requirement: allow additional memory for the KV cache, activations, runtime, context length, and concurrent requests.

ML engineering specialization

Vinci MLE is trained as a policy inside an instrumented ML workspace, not only on question-and-answer pairs. Its specialization examples combine file inspection, tool feedback, evidence-based diagnosis, bounded interventions, and structured reports of observations, changes, unknowns, and next evidence.

The frozen specialization corpus contains 91 trajectories across 64 distinct task roots in eight task families. Both sizes consumed the same tokenized training examples; their LoRA targets and optimization recipes differ. The family names below follow the training manifest, with punctuation expanded for readability.

ML-engineering task family Training scope
Duplicate data / data quality Inspect duplication and data-quality problems
Experiment design / interpretation Interpret experiment evidence and assess whether a change is justified
Failed resume Diagnose failed training-resume behavior
Memory / batching / accumulation Inspect memory, batching, and gradient-accumulation problems
Metric / evaluator defect Diagnose defects in metrics and evaluators
NaN / divergence Investigate numerical failures and divergence
Packaging / serving / reproducibility Inspect packaging, serving, and reproducibility problems
Tokenizer / model configuration Diagnose tokenizer and model-configuration problems

Act / No Change / Stop is central to the training design. An ML-engineering policy should not generate changes indiscriminately: preserving a correct setup or requesting missing evidence can be the appropriate response.

Trained response Meaning in the specialization corpus Trajectories
Act Make a bounded change when the evidence supports a defect and an intervention 45
No Change Preserve a correct setup and explain the evidence 23
Stop / Needs Evidence Identify missing evidence or a blocker instead of inventing a repair 23

These counts describe training coverage, not success rates. They do not establish that the model chooses the correct action, no-change, or stop decision on new tasks.

MLE evaluation status

Vinci MLE 1.0 does not yet have a valid held-out ML-engineering task score.

The internal development screen did not provide a valid capability comparison: the parent and every development fine-tune produced parse errors or timeouts, with no successful episodes. Those instrument-floor outcomes did not establish a valid MLE-task score for the released checkpoint. The recipe-selection details are retained in Training.

The preregistered held-out MLE evaluation and external MLE-bench evaluation were not completed for this release. We report MLE training scope separately from measured general-capability comparisons and make no MLE-task performance claim for 1.0.

The full evaluation also requires a qualified execution environment with the intended isolation and reproducibility controls. That environment was not ready for the complete MLE evaluation in 1.0, and we chose not to substitute a weaker test as an equivalent result.

For Vinci MLE 1.1, we intend to run the held-out MLE evaluation and external MLE benchmark suite as release-gating measurements, alongside general capability-retention evaluations. Results will be reported when measured; this 1.0 card does not predict them or promise a release date.

Intended workflow

Conceptual research workflow: inspect text and data evidence, choose act, no change or stop, then report for human review.

Illustration of training intent; tools are supplied by the application. This is not a captured model run.

A task supplies a problem statement, workspace files, and tool feedback. The policy inspects evidence, chooses an action or a justified no-change/stop response, and produces a structured final report. Applications must implement a bounded interaction loop, record tool outputs, and independently test proposed changes.

The weights alone are not the full Vinci MLE agent harness. This repository supplies the model and tokenizer. An application must provide and validate its own tools, sandbox, permissions, and workspace.

Training interface

The workspace interface uses task-specific subsets of inspect_files, run_command (an argv array), and write_file (path and content). All 91 include run_command; 36 expose all three tools. Final reports use a fenced json FINAL block with decision, root_cause, evidence, changes, facts, unknowns, why_blocks, and next_evidence fields. Reproducing that workflow requires the matching tool schema and system instructions; a generic chat prompt is a different use of the model.

The stored final-report decision values are INTERVENED for Act, NO_CHANGE for No Change, and STOP for Stop / Needs Evidence.

For an initial text-only exploration, the quick start asks about a small data-splitting problem. It is an illustrative prompt, not a captured model answer or a benchmark result.

Choosing a size

MLE 8B MLE 30B
Parent Granite 4.1 8B Granite 4.1 30B
Weight download 17.58 GB, 4 shards 57.73 GB, 12 shards
Training examples / task roots 91 / 64 91 / 64
LoRA rank / alpha 8 / 16 32 / 64
Trained projection families Attention Attention and MLP
BFCL score, 0–100 68.8584 73.5160
LiveCodeBench score, 0–100 32.9858 40.3791
Tool-call smoke C1 failed; other cases passed All cases passed

The 30B weights occupy approximately 57.73 GB (53.77 GiB). Plan for substantially more memory than the 8B sibling. The 30B model has higher recorded BFCL and LiveCodeBench scores. The benchmark settings and LiveCodeBench timing caveat below apply to the comparison. The two models use the same training examples but different recipes; this is not a controlled experiment isolating parameter count, and there is no measured MLE-task ranking between them. Choose 30B when its higher measured tool/coding scores justify the larger download and serving footprint. Choose 8B when a smaller download and lower weight-memory footprint matter more.

Quick start

Download a pinned snapshot

Install huggingface_hub, and authenticate with hf auth login if repository access requires it. This example pins the uploaded model revision evaluated by the release records; later card-only revisions do not change those model files.

from huggingface_hub import snapshot_download

model_dir = snapshot_download(
    repo_id="simpledirect/Vinci-MLE-30B-1.0",
    revision="f0c85ea26a373c4bd0b217dfd66ba9f730ef10eb",
    local_dir="./vinci-mle-30b",
)
print(model_dir)

The pinned snapshot includes its historical README and matching checksums. To download a later card revision, use that revision's full commit hash and verify its own SHA256SUMS.

Generate with Transformers

Requires PyTorch, Accelerate, and Transformers with Granite support. Load the merged checkpoint directly; no PEFT adapter needs to be applied.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_dir = "./vinci-mle-30b"
tokenizer = AutoTokenizer.from_pretrained(model_dir, trust_remote_code=False)
model = AutoModelForCausalLM.from_pretrained(
    model_dir,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=False,
).eval()

messages = [
    {"role": "system", "content": "You are an ML engineering research assistant. Separate observations from hypotheses and explain what to check next."},
    {"role": "user", "content": "A dataset contains multiple records per customer. Training and validation rows were split at random. What should I inspect before trusting the validation score? Do not claim to have run any checks."},
]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
)
device = model.get_input_embeddings().weight.device
inputs = {key: value.to(device) for key, value in inputs.items()}
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=512, do_sample=False)
completion = output[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(completion, skip_special_tokens=True))

This is a reference loading example, not a rerun of the release evaluation. It uses an illustrative system prompt and does not reproduce the Vinci agent harness. No generated answer is asserted by this card.

Serve an OpenAI-compatible endpoint with vLLM

The release evaluation used vLLM 0.22.1, a 16,384-token limit, greedy decoding, the packaged chat template, and the hermes tool-call parser. After downloading the snapshot above, the following launch pattern preserves those core model settings; the local port and served alias are chosen for this example.

CUDA_VISIBLE_DEVICES=0 \
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 VLLM_USE_DEEP_GEMM=0 \
python -m vllm.entrypoints.openai.api_server \
  --model ./vinci-mle-30b \
  --tokenizer ./vinci-mle-30b \
  --chat-template ./vinci-mle-30b/chat_template.jinja \
  --served-model-name vinci-mle-30b \
  --dtype bfloat16 \
  --tensor-parallel-size 1 \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.90 \
  --seed 0 \
  --enforce-eager \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --host 127.0.0.1 --port 8000

Use a GPU with enough memory for the weights and serving overhead. Tensor parallel 1 describes the evaluated setup; this card does not certify a minimum GPU model or all alternative parallel configurations.

Send requests to http://127.0.0.1:8000/v1/chat/completions, using model: "vinci-mle-30b", temperature: 0, and an explicit max_tokens budget. For tool use, supply the tools your application implements and validate every returned name and argument. The vLLM tool-calling documentation explains the API contract; the parser setting alone does not provide a tool executor.

30B tool-call smoke: all cases passed in the fixed smoke run. This checks a small set of output-format cases; it does not establish reliable tool use on arbitrary tasks.

Example: request a workspace tool call

The three declarations below match the three-tool training variant used in 36 trajectories. This example sends a request and prints the model response; it does not execute returned commands. The system instruction is the common training instruction. The user task is illustrative. A complete agent needs a workspace, permission checks, tool implementations, returned observations, and a bounded generation loop.

import requests

tool_definitions = [
    {
        "name": "inspect_files",
        "description": "Print the task objective and every public workspace file.",
        "parameters": {
            "type": "object", "properties": {}, "additionalProperties": False,
        },
    },
    {
        "name": "run_command",
        "description": "Run a command in the workspace and return its stdout and stderr.",
        "parameters": {
            "type": "object",
            "properties": {"argv": {"type": "array", "items": {"type": "string"}}},
            "required": ["argv"], "additionalProperties": False,
        },
    },
    {
        "name": "write_file",
        "description": "Replace a workspace file with the given full content.",
        "parameters": {
            "type": "object",
            "properties": {"path": {"type": "string"}, "content": {"type": "string"}},
            "required": ["path", "content"], "additionalProperties": False,
        },
    },
]
response = requests.post(
    "http://127.0.0.1:8000/v1/chat/completions",
    json={
        "model": "vinci-mle-30b",
        "messages": [
            {"role": "system", "content": "Assess this private TRAIN workspace using recorded tools and a structured FINAL answer."},
            {"role": "user", "content": "Inspect the workspace to identify how the training and validation data are split. Gather evidence before proposing a change."},
        ],
        "tools": [{"type": "function", "function": tool} for tool in tool_definitions],
        "tool_choice": "auto",
        "temperature": 0,
        "max_tokens": 512,
    },
    timeout=120,
)
response.raise_for_status()
print(response.json()["choices"][0]["message"])

Install requests to run this client example. Neither these tool descriptions nor the model's response implement a sandbox. Only expose tools whose actions your application can validate and contain.

General capability retention

The evaluations below measure coding, tool calling, general reasoning, agentic behavior, refusal patterns, and long-context retrieval relative to the pinned Granite parent. They characterize broader parent capabilities after MLE specialization. They are not substitutes for a held-out ML-engineering benchmark. Only the tool-calling and coding comparisons have the stated retention pass criterion; the other studies are descriptive and retain their individual caveats. Across these descriptive studies, the 30B release remains broadly close to its Granite parent, with some positive and some negative differences; these measurements characterize retention rather than MLE-specific improvement.

Results below were resolved from the frozen release records on September 28, 2026, against this model's own pinned Granite parent. The summary job verified each source receipt's SHA-256 and used the frozen evaluator for the BFCL and LiveCodeBench comparisons. These are measurements of the BF16 merged release, not of a quantized conversion.

Tool calling and coding

Granite 30B parent and Vinci MLE 30B BFCL and LiveCodeBench scores on a full 0–100 axis. Exact values and timing caveats follow.

BFCL measures tool-call retention; LiveCodeBench measures coding retention. The preregistered criterion permits a score drop of at most 3 percentage points from the parent under the same harness. Scores below are on a 0–100 scale; differences are this model minus its parent, in percentage points.

Benchmark Granite 30B parent Vinci MLE 30B Difference (pp) Retention margin
BFCL 73.5616 73.5160 -0.0457 Within margin
LiveCodeBench 40.6635 40.3791 -0.2844 Within margin*

BFCL covers 11 categories and excludes the executable-code category. LiveCodeBench covers 1,055 items. Both recorded scores fall within the preregistered 3-point retention margin. LiveCodeBench retains the separate timing-qualification limitation described below. These comparisons do not demonstrate improved MLE-task performance.

* LiveCodeBench timing adjudication: INCONCLUSIVE / faster-machine qualification failed. A timing flag affected 74 of the parent's and 74 of this model's 1,055 items. The proposed faster regrading host did not meet its qualification threshold (ratio 0.908 against a maximum of 0.90), so no regrade was performed. The table reports the existing record grades with that unresolved timing limitation.

LiveCodeBench grading used the evaluator's full sandbox isolation; container stderr was not retained. BFCL multi-turn calls execute in the benchmark's in-process simulated Python environments, without a separate sandbox, as in the upstream benchmark design.

General knowledge, reasoning, and instruction following

Descriptive; no multiplicity correction; not a retention verdict. These paired comparisons have no preregistered pass rule. Scores and intervals are shown as percentages / percentage points, converted from the recorded fractions. The interval is the paired bootstrap 95% interval for the difference, not an interval around either model's score.

Benchmark Items Parent (%) Vinci MLE (%) Difference (pp) 95% interval for difference (pp)
AIME 60 21.67 25.00 +3.33 [-3.33, +10.00]
BBH 6,511 63.46 63.97 +0.51 [+0.05, +0.95]
GPQA-Diamond 198 36.36 43.94 +7.58 [+0.51, +14.14]
GSM8K 1,319 94.47 94.16 -0.30 [-0.76, +0.08]
IFEval 541 87.43 86.69 -0.74 [-2.03, +0.55]
MATH-500 500 76.40 78.60 +2.20 [-0.20, +4.60]
MMLU-Pro 12,032 65.58 65.28 -0.30 [-0.76, +0.16]

Values are rounded from the source record; subtracting displayed scores can differ slightly from the separately rounded recorded difference. Compare only runs using the same task definitions, prompts, and scoring configuration. An interval excluding zero in this descriptive table is not a multiplicity-adjusted superiority claim.

IFEval uses strict prompt-level accuracy; MATH-500 uses math_verify. The other rows use exact match with their locked extraction rules: GPQA flexible-extract, GSM8K strict-match, MMLU-Pro custom-extract, and BBH get-answer.

Agentic evaluation

Descriptive; no multiplicity correction; not a retention verdict. The tau2 study reports pass¹ over four trials per task: 50 airline tasks (200 trials) and 114 retail tasks (456 trials). Confidence intervals resample whole tasks, preserving the four correlated trials.

tau2 measure Trials Parent (%) Vinci MLE (%) Difference (pp) 95% interval for difference (pp)
Airline, primary 200 26.00 31.50 +5.50 [-0.50, +13.00]
Retail, primary 456 67.54 67.32 -0.22 [-4.61, +4.17]
Retail, secondary 456 63.60 62.72 -0.88 [-5.26, +3.51]

Airline primary uses the upstream database × communication reward. Retail primary measures database end-state success only, which differs from the upstream retail reward. Retail secondary combines database success with natural-language assertions judged by a fixed local Qwen3.8-27B model; it is an approximation, not the official tau2 retail score. Non-completed trials score zero; the secondary comparison excludes pairwise unjudged trials, with its actual included count shown above.

These are agentic benchmark measurements, not ML-engineering task results. Greedy decoding under vLLM continuous batching is not bitwise deterministic, so reruns can produce different item-level outputs.

Safety: refusal-pattern measurements

Reporting updated following a classifier audit. Model weights are unchanged.

Descriptive; no multiplicity correction; not a retention verdict. A9 compares refusal patterns between Vinci MLE and its pinned parent on XSTest and StrongREJECT using two registered refusal-phrase windows: the start of the normalized response and the first 300 normalized characters.

Classifier audit. In a blind review of 90 responses per model by an independent AI agent reviewer, with model identity and classifier output hidden, neither classifier window labelled a compliant answer as a refusal in the audited sample. The first-300-character window identified 54/60 parent and 54/60 Vinci MLE refusals identified by the reviewer. The start-of-response window identified 6/60 in each arm. Because the first-300-character window was substantially more sensitive in this audit, it is reported first below; the registered start-of-response measurement is retained for completeness.

Detected refusal-pattern differences — first 300 characters

Prompt set Refusal-phrase window Paired items Parent (%) Vinci MLE (%) Difference (pp) 95% interval for difference (pp)
StrongREJECT First 300 characters 313 91.05 91.05 +0.00 [-1.28, +1.28]
XSTest safe First 300 characters 250 6.80 6.80 +0.00 [+0.00, +0.00]
XSTest unsafe First 300 characters 200 87.00 86.50 -0.50 [-1.50, +0.00]
Registered start-of-response window
Prompt set Refusal-phrase window Paired items Parent (%) Vinci MLE (%) Difference (pp) 95% interval for difference (pp)
StrongREJECT Start of response 313 4.15 4.15 +0.00 [-0.96, +0.96]
XSTest safe Start of response 250 2.80 2.80 +0.00 [+0.00, +0.00]
XSTest unsafe Start of response 200 10.00 9.50 -0.50 [-1.50, +0.00]

On safe XSTest prompts, a higher detected-refusal rate can indicate increased over-refusal. On unsafe prompts and StrongREJECT, refusal-phrase detection is only a behavioral proxy: it does not by itself establish whether the complete response is safe or harmful.

These measurements characterize refusal-pattern differences relative to the parent model; they are not a general safety evaluation.

Long context: synthetic retrieval

Descriptive; no multiplicity correction; not a retention verdict. The RULER study averages two replicates on 1,500 paired synthetic items across six nominal input lengths. Its string_match_all metric is the fraction of expected strings present, case-insensitive; it is not a binary all-strings-correct measure. Actual tokenized prompt lengths vary around each nominal length.

Nominal length (tokens) Paired items Parent (%) Vinci MLE (%) Difference (pp) 95% interval for difference (pp)
1,024 250 100.00 100.00 +0.00 [+0.00, +0.00]
2,048 250 100.00 100.00 +0.00 [+0.00, +0.00]
4,096 250 100.00 100.00 +0.00 [+0.00, +0.00]
8,192 250 100.00 100.00 +0.00 [+0.00, +0.00]
12,288 250 100.00 100.00 +0.00 [+0.00, +0.00]
14,336 250 100.00 99.92 -0.08 [-0.24, +0.00]
All lengths 1,500 100.00 99.99 -0.01 [-0.04, +0.00]

The 30B parent reached the ceiling at every tested length, limiting what the near-zero parent/candidate differences can establish. The recorded synthetic-retrieval scores remain close within the tested range; this descriptive comparison has no preregistered retention verdict. It does not measure understanding of an ML repository or operation at the parent's full 131,072-position limit.

Repeated-run studies

The release records additionally mark repeated-run variance studies as measured for GSM8K, IFEval, and MATH-500; the repeated-run AIME study was not run at 30B. Their numeric ranges are not included in the resolved numbers summary used here and are not reported in this card. The paired tables above should not be interpreted as proof of bitwise repeatability.

Evaluation scope: what was not run

Evaluation 1.0 status
Sealed internal MLE evaluation Not run
48-task held-out MLE evaluation Not run
MLE-bench release subset and MLE-bench Lite Not run
EvalPlus, HumanEval+, MBPP+ Not run
Peer evaluation Not run
Behavioral fidelity of merged release versus adapter Not measured

The full preregistered evaluation protocol was not completed for 1.0 because the MLE-specific held-out component remains unmeasured. The benchmark component results above retain their documented scope and caveats; they do not constitute a completed MLE release-evaluation result. No incomplete gate is reported as passed.

Training

Vinci applied supervised fine-tuning to the pinned Granite parent using ordinary LoRA, then merged the terminal adapter into the parent for distribution. This recipe uses neither DoRA nor rank-stabilized LoRA.

Setting Value
Examples / distinct task roots 91 / 64
Trajectory composition 45 direct action / 23 no change / 23 stop
Task families 8, with 8 distinct roots each
Epochs / optimizer steps 1 / 12
Selected checkpoint Terminal checkpoint-12; the sole final candidate
LoRA rank / alpha / dropout 32 / 64 / 0.05
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Learning rate 4 × 10⁻⁶
Schedule / warmup ratio Cosine / 0.03
Weight decay 0.01
Microbatch / gradient accumulation 1 / 8
Precision BF16
Sequence ceiling 16,384 tokens; no packing or truncation
Training seed 18427
Tokenized corpus 596,976 total tokens; 58,040 supervised tokens
Longest training example 12,023 tokens

Counts were recomputed from the frozen training receipts for this card. These receipts describe the specialization data, not the Granite parent's pretraining corpus. Both MLE sizes consumed the same tokenized-example bytes.

Selective assistant supervision. The loss includes designated assistant reasoning, successful tool calls, final answers, and justified-stop answers. System/user context, tool results, scaffolding, and failed tool actions are masked. Verifier feedback is excluded from the rendered prompt. Failed actions therefore do not automatically become positive imitation targets. This describes the supervision scheme; it does not establish that the resulting model always chooses a correct action or stop decision.

Development-selection limitation. Parse errors and timeouts drove the parent and eligible development arms to the evaluator floor, so the screen could not support a capability comparison. The frozen recipe-selection fallback was used; full details are retained in the release evidence.

Development screen details

The parent and each of the three development arms produced 14 parse errors and 2 timeouts across 16 episodes, with 0 successes. The screen therefore supplies no valid capability comparison or evidence of improvement over the parent; its records do not independently diagnose the cause of the parse failures. The recipe was chosen by the frozen lower-dose tie-break, not demonstrated superiority. These are development-screen outcomes, not an MLE score for the released model. The final release uses the predetermined terminal checkpoint under the fixed final-training policy.

Limitations

  • No completed MLE-specific evaluation. General coding, tool-calling, and reasoning scores cannot fill that gap.
  • Small specialization corpus. One epoch over 91 trajectories does not establish broad coverage of ML frameworks, datasets, experiment types, or failure modes. The eight training families describe corpus coverage, not demonstrated success rates.
  • Output and tool errors remain possible. Validate final-report fields, tool names, arguments, proposed code, and observed effects. The 30B smoke result is disclosed above.
  • LiveCodeBench timing remains unresolved. Read its score with the qualification caveat, not as a fully adjudicated timing result.
  • Merge integrity is not behavioral equivalence. The release export passed mechanical verification, standalone loading, and byte-identical rebuild checks. Behavioral fidelity to the original adapter was not measured.
  • Context scope is bounded. The 131,072-position parent configuration is not proof of reliable MLE behavior at that length; the evaluation serving limit was 16,384 tokens.
  • No safety certification or unattended-deployment validation. Safety benchmark observations do not establish a safety property. Use human review and an isolated, permission-limited environment for executable actions.

Contamination audit

The training-data audit covered 19 benchmark/subset entries. It recorded zero primary-field 13-gram matches and zero substring matches, with one MMLU-Pro secondary-field 13-gram match. Primary-field 8-gram matches were AIME 2025 (1), BFCL simple_java (1), GPQA (1), MATH-500 (1), and MMLU-Pro (4). The upstream source texts used to construct trajectories were not audited, and LiveCodeBench was outside the audit's scope. These checks do not establish a contamination-free evaluation.

Provenance and file integrity

The repository includes PROVENANCE.json, SHA256SUMS, LICENSE, and NOTICE. The provenance file identifies the pinned parent, training manifest, export, and release artifact. Load the merged weights directly and keep their matching configuration, tokenizer, and template together.

Identity Value
Parent revision 4fae6278f7132abf5e971f9de49ebbad09c54cce
Candidate mle1-final-30b-g1-a2-release
Frozen release manifest SHA-256 78256b1fb5dbb5d31135ff2849d6899ed6f8c08e952ce5b58259d4836b5a73c0
Frozen model artifact tree SHA-256 fb4e2e6012a077057b9dfd2b11b5b9247d670286f8fd995161269d4fa9ede29c
Training manifest SHA-256 e48f81c217f15234ba6c6b6b033f6ea056019e9ad40ba366efa6cf472606d4a1
Tokenized training examples SHA-256 61f24f69c97873e029959282480b99274197a5202a4bca73e54d7f4b1fcd7a88

The frozen artifact-tree digest identifies the model export, not the current documentation directory. Card updates receive a new Hugging Face commit and a corresponding README checksum; they do not change the frozen model identity.

After downloading one complete revision, run this from its directory:

sha256sum -c SHA256SUMS
# macOS alternative:
shasum -a 256 -c SHA256SUMS

Checksums detect differences against the recorded reference; they do not independently authenticate its publisher. This repository does not include a detached MLE release signature.

Evaluation receipt identifiers

These hashes identify the internal records used for the tables. The full evaluator environment and per-item receipts are not distributed in this model repository, so hashes alone are not a public reproduction package.

Record SHA-256
BFCL paired receipt 960dd214251d857d82422e0c1a14936e3282537807bb55837b4c77349f5dae5d
LiveCodeBench paired receipt c1e5c6a792cf194decae1a8c695990ee5fb2c0d890c4576802a2a7c6d4618758
General-capability paired receipt b3c4a118a08db38523d11c9c94f5da5f234207b8c58e727fd0e11a23fd090b77
RULER summary f7465ce480a451aeeb5436b8bb61f74598c0546d42405f58683b2027f4af2e9b
Resolved 8B/30B numbers summary 0f7f6912868f42a26255e860cdbd793e38958eebd4876e562610dcd04048ea92

License and attribution

The released weights are licensed under Apache-2.0. See LICENSE and NOTICE for the accompanying terms and attribution. IBM developed Granite; SimpleDirect developed this fine-tuned derivative and its release. No affiliation with or endorsement by IBM is implied.

Citation

@misc{vinci_mle_1_0_30b,
  title  = {Vinci MLE 1.0 (30B)},
  author = {{SimpleDirect}},
  year   = {2026},
  url    = {https://huggingface.co/simpledirect/Vinci-MLE-30B-1.0}
}

SimpleDirect · Vinci models

Downloads last month
7
Safetensors
Model size
29B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for simpledirect/Vinci-MLE-30B-1.0

Finetuned
(13)
this model
Quantizations
1 model

Collection including simpledirect/Vinci-MLE-30B-1.0