YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Nocturne is not a dense 12B model in practical runtime terms.

It combines a standard Transformer stack (multi-layer self-attention + feedforward) with sparse MoE layers: the model has many experts but only a small number of experts are routed-to (active) per token, so the active parameter footprint per forward pass is ~2.02B rather than 12B.

That sparse activation is what lets us actually claim “12B parameters” on disk but low memory at inference. 

It uses a top-k gating function that selects 2–4 experts per token and routes token activations to those experts; 32 total experts and 4 active experts per token for the 12B model.

Practical engineering choices are:

  • Grouped multi-query attention (GMQA)
  • Mixed sparse/dense attention patterns to trade compute for latency while retaining long-range context.
 Those choices reduce memory bandwidth and attention state size during long contexts.

The model supports extremely long contexts (up to 64k tokens), achieved thought careful attention memory and positional strategy.

Weights come natively quantized in a MXFP4 format which is designed to preserve accuracy while drastically lowering memory use. 

Trained it on a mostly-English, text-only dataset with an emphasis on STEM, coding, and general knowledge; tokenization uses a new tokenizer by OpenAI called o200k_harmony (a superset of the o4/o4-style tokenizers). after base pretraining, the model was post-trained on a structured format OpenAI calls.


license: mit datasets: - google/IFEval - Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b - HuggingFaceFW/finetranslations - nvidia/Llama-Nemotron-Post-Training-Dataset - lluisgomez/SISPI - 123olp/binance-futures-ohlcv-2018-2026 - meta-math/MetaMathQA language: - en - ru metrics: - f1 - accuracy - bleu base_model: - danielhanchen/new_ollama new_version: nvidia/personaplex-7b-v1 pipeline_tag: token-classification library_name: nemo

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support