📜 OPTIMIZATION HISTORY — LEGACY EDITION

This repository hosts a previous iteration of our handcrafted MiniPlus architecture. While not our current specification, it remains an outstanding, high-fidelity quantization that significantly outperforms any flat 3-bit community quants (Q3_K_S / IQ3_S) and generic 2-bit APEX Mini community releases.

We preserve this repository publicly with 100% transparency as a verified engineering record of continuous optimization within the strict 13–14 GB envelope.

👉 Current Definitive Specification (V2.1): Access the upgraded V2.1 release featuring zero AVX2 CPU stalls and maximal long-context stability at: IsValorum Hugging Face Catalog


⚡ Quick Navigation Index


📦 Bundled Model Files & Specifications

File Name File Size Memory Footprint Format / Precision Purpose
Thomson-1.0-Small.APEX-I-MiniPlus-V2.gguf 14.63 GB (13.63 GiB) 13.63 GiB Custom APEX-I (3.38 BPW) Main legal, auditing, tax & regulatory logic core
mmproj-thomsonreuters_Thomson-1.0-Small-Q8_0.gguf 610 MB (582 MiB) 582 MiB High-Precision Q8_0 Projector Required for scanned PDF analysis, tables & OCR vision
  • Base Architecture: Qwen3_5MoeForConditionalGeneration (40 layers, 256 fine-grained micro-experts with intermediate dimension 512, 8 active per token) + Vision Projector.
  • Active Parameters: approx. 3.2B active parameters per token (blazing generation speed combined with 35B domain intelligence).
  • Importance Matrix: 510 calibrated tensor entries derived from 550 dense chunks of statutory law, financial audits, and regulatory filings.
  • Memory Footprint: Lean 13.63 GiB weight size leaving ample VRAM for deep document context windows.

🔬 Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)

Also, don't confuse APEX-I-MiniPlus-V2 with a generic baseline APEX-I-Mini. Traditional APEX-I-Mini drops core experts aggressively to 2-bit IQ2_S and leaves output.weight at 3-bit Q3_K_M, which creates a noticeable perplexity hit on complex reasoning tasks. V2 was specifically re-engineered to avoid that quality floor (keeping core experts at calibrated IQ3_XXS, output in Q6_K, shared expert in non-linear IQ4_NL, and routers in F32).

To put the numbers in perspective: this cuts nearly 2 GB off a flat 3-bit quant (approx. 15.6 GB), and weighs only about approx. 1 GB more than a generic APEX-I-Mini (approx. 12.5 GB). For that single extra gigabyte of VRAM, you get a massive jump in reasoning and syntactic stability.

Take a look at the tensor-by-tensor comparison table below to inspect the exact architectural differences and see why this specific allocation is optimal. That's specifically what this was built for:

Architectural Component Generic Automated Quants (Flat Q3_K_S / IQ3_S) Generic APEX-I-Mini (Baseline Recipe) Our Handcrafted APEX-I-MiniPlus-V2 (IsValorum) Perceived Quality & Real-World Impact
Output Head (output.weight) Flat IQ3_S / Q3_K_S (approx. 3.44 BPW) Inherits base type Q3_K_M (approx. 3.44 BPW unarmored) Q6_K (approx. 6.56 BPW uncompromised) Eliminates Syntax & Vocabulary Hallucinations: Low-bit output heads cause tokenizer classification noise, breaking code indentation, brackets ({}, []), math symbols, and domain terms. Q6_K preserves near-FP16 output classification.
Expert Routers (ffn_gate_inp.weight) Blindly quantized to 3-bit / unoptimized Inherits base type Q3_K_M (approx. 3.44 BPW compressed) F32 uncompressed (32.0 BPW, 2 MB/layer) Zero Router Drift: In micro-expert models, even minuscule quantization errors in router logits misdirect tokens to wrong experts. Retaining uncompressed F32 guarantees 100% routing fidelity with virtually zero memory overhead (approx. 80 MB total).
Attention & Language (attn_output, attn_qkv) Flat IQ3_S / Q3_K_S Q3_K on 34 middle layers (L3–36), Q4_K on 6 edge layers Q6_K for attn_output, IQ3_S for attn_qkv Contextual Retrieval Precision: Generic APEX reduces attention and language projections to Q3_K across 85% of layers. Our V2 build protects attention output in high-precision Q6_K and uses calibrated non-linear IQ3_S, ensuring flawless needle-in-a-haystack retrieval across deep 128k–256k context windows.
Attention Gates (attn_gate.weight) Blindly compressed to 3-bit Compressed to Q3_K (middle) / Q4_K (edges) Q8_0 (8.50 BPW) Attention Head Stability: Attention gates modulate query-key routing across hybrid attention layers. Keeping them in 8-bit prevents attention crosstalk and hallucination over long contexts.
Shared Foundation Expert (ffn_*_shexp) Flat IQ3_S / Q3_K_S (3.44 BPW) Linear Q4_K (middle) / Q5_K (edges) IQ4_NL (4.50 BPW non-linear codebook) Foundational Knowledge Armor: The shared expert executes for 100% of tokens. In 256 micro-expert models, IQ4_NL non-linear codebooks preserve heavy-tailed outlier representations far better than standard linear quantization.
Core MoE Layers (Middle: 10–29) Flat IQ3_S / Q3_K_S (uniform bit-rate across all layers) Aggressive IQ2_S (2.50 BPW) IQ3_XXS (3.06 BPW) + calibrated imatrix Above the Quality Threshold: Generic 2-bit IQ2_S baselines drop below the critical quality floor for 35B MoEs, resulting in perplexity spikes on reasoning tasks. Our IQ3_XXS with imatrix achieves deep compression (272 MiB → 98 MiB per block) without sacrificing logic.
Edge MoE Layers (Layers 0–9 & 30–39) Flat IQ3_S / Q3_K_S (no layer-wise gradient) Q3_K (limited to first/last 5 layers only: L0–4, L35–39) IQ3_S (expanded to 10 input & 10 output layers) Protected Ingestion & Synthesis: Half of the model's layers (10 at input, 10 at output) form a non-linear armored envelope, preventing prompt misunderstanding and token degeneration across 256 micro-experts.
Multimodal Vision (mmproj) Often omitted, or left as uncompressed FP16 (approx. 900 MB) Often omitted or separate uncompressed FP16 Bundled Q8_0 (582 MB) with 27 critical F32/F16 fallbacks Saves approx. 320 MB VRAM with Zero Loss: Handcrafted quantization preserves normalization and bias tensors in F32/F16, ensuring razor-sharp OCR, DOM viewport reading, and coordinate detection without visual noise.
Normalization & Biases Often degraded Standard F32 uncompressed Numerical Stability: Prevents cumulative floating-point underflow/overflow across deep 40-layer computation.

👁️ Bundled Q8_0 High-Precision Multimodal Vision Projector

Standard community uploads often omit the multimodal projector or supply uncompressed FP16 files (approx. 857 MB), bloating memory.

  • Bundled Q8_0 Projector: Pre-quantized to Q8_0 (582 MiB / 610 MB), saving approx. 300 MB of VRAM.
  • Audited Layer Fallbacks: llama.cpp automatically preserved 27 critical normalization and embedding tensors in F32/F16, ensuring razor-sharp OCR of tiny contract footnotes, financial balance sheets, and scanned legal filings without artifacts.

💻 Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)

Estimated Projections on Consumer Hardware

You do not need enterprise infrastructure to perform automated legal analysis. Estimated throughput projections on an everyday consumer laptop (Intel Core i5 / AMD Ryzen, 4GB/6GB Laptop GPU, 32GB DDR4/DDR5 RAM):

  • GPU VRAM Allocation: Uses only approx. 3.8 GB VRAM (fits easily on budget laptop GPUs like RTX 3050, 4050, or 2060).
  • System Memory Offload: Standard 32GB system RAM accommodates the remaining layers.
  • Estimated Document Ingestion (Prefill): 300 to 420+ tokens/second sustained across dense briefs.
  • Estimated Streaming Generation: 20 to 24+ tokens/second sustained output across system RAM!

🔥 The 24GB Miracle: Full 256K Context Runs In VRAM!

Legal and auditing tasks require ingesting hundreds of pages of case law, depositions, and regulatory exhibits. Automated 3-bit/4-bit community quants exceed 16–19 GiB in weights alone, crashing with Out-Of-Memory (OOM) errors when context scales up.

Thomson APEX-I-MiniPlus-V2 fits the entire 256K context window within 24GB VRAM:

Context Length Model Weights (Est.) KV Cache (q8_0, 4 slots) Compute Buffers Total GPU VRAM (Est.) Hardware Feasibility
32,768 (32k) 13.63 GiB 0.58 GiB 1.80 GiB 16.01 GiB Full offload on 24GB; partial on 16GB
65,536 (64k) 13.63 GiB 0.92 GiB 1.95 GiB 16.50 GiB Effortless fit on 24GB GPUs
131,072 (128k) 13.63 GiB 1.58 GiB 2.22 GiB 17.43 GiB Effortless fit on 24GB GPUs
262,144 (256k) 13.63 GiB 2.92 GiB 2.80 GiB 19.35 GiB 🔥 FULL 256K BRIEF IN VRAM!

Note: Projections leave approx. 4.65 GiB of headroom on 24GB cards for display buffers and the Q8 vision projector.


🏎️ Hardware Throughput Projections (RTX 30 / 40 / 50)

| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Engineering Highlights | | :--- | :--- | :---: | :---: | : | | NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) + mmproj | 105 – 130+ tok/s | 2,400 – 3,500+ tok/s | Instantaneous contract audit on GDDR7 bandwidth | | NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) + mmproj | 75 – 100+ tok/s | 1,700 – 2,500+ tok/s | Real-time multi-page document review & synthesis | | NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) + mmproj | 62 – 78+ tok/s | 1,350 – 1,950+ tok/s | Full 256k legal brief ingestion in dedicated VRAM | | NVIDIA RTX 4080 / 5070 (16GB) | Partial offload (approx. 30 layers) | 32 – 42+ tok/s | 750 – 1,150+ tok/s | High-efficiency local legal workstation | | Consumer Laptop (4GB GPU + 32GB RAM)| Hybrid Offload | 20 – 24+ tok/s | 300 – 420+ tok/s | Smooth streaming from system DDR4/DDR5 RAM |

Projections represent theoretical estimates derived from hardware memory bandwidth and the approx. 3.2B active parameter MoE design.


⚖️ The Speed vs. Precision Trade-off

  1. Uncompressed F32 Router Selectors: In a 256 micro-expert model, even minor router rounding errors route clauses to the wrong domain expert (e.g. confusing tax compliance with intellectual property). ffn_gate_inp.weight is maintained in uncompressed F32 (2 MB/layer) to eliminate routing drift.
  2. Q6_K Output Head: Protects Latin legal terminology, statutory code citations, and quantitative tax calculation symbols from compression distortion.
  3. Non-Linear IQ Codebooks: IQ3_XXS, IQ3_S, and IQ4_NL preserve nuanced contractual semantics across edge and core layers.

🛠️ Surgical Tensor Quantization Map

Tensor Pattern Layer Scope Quant Type BPW Engineering Rationale
output.weight Vocabulary Head Q6_K 6.56 Uncompromised 6-bit precision for legal, fiscal & statutory vocabulary
token_embd.weight Embedding High-Prec High Preserves subtle token definitions and prompt grounding
ffn_gate_inp.weight Expert Routers F32 32.0 Uncompressed full-precision routers preventing domain expert drift
attn_gate.weight Attention Gates Q8_0 8.50 High-precision 8-bit gating for attention routing dynamics
ffn_*_shexp Shared Experts IQ4_NL 4.50 4-bit non-linear codebook for the 100% active shared foundational expert
ffn_down/up/gate Edges (0–9, 30–39) IQ3_S 3.44 Armored boundary layers protecting prompt ingest and final opinion synthesis
ffn_down/up/gate Core (10–29) IQ3_XXS 3.06 Deep compression (272 MiB → 98 MiB per block) calibrated via legal imatrix
mmproj (Vision) Visual Projector Q8_0 8.00 High-fidelity OCR and table rendering with 27 critical F32/F16 fallbacks
Norms & Biases All Layers F32 32.0 Absolute numerical stability across deep 40-layer computation

📖 Recommended Configuration & Setup

Unsloth Studio:

  1. Load Thomson-1.0-Small.APEX-I-MiniPlus-V2.gguf.
  2. Select mmproj-thomsonreuters_Thomson-1.0-Small-Q8_0.gguf as the vision projector.
  3. Configure KV Cache Dtype to q8_0 and Context Checkpoints to 1.
  4. Set GPU Offload to 100% (-ngl 99) on 24GB GPUs.

llama.cpp CLI:

llama-cli -m Thomson-1.0-Small.APEX-I-MiniPlus-V2.gguf \
  --mmproj mmproj-thomsonreuters_Thomson-1.0-Small-Q8_0.gguf \
  -ngl 99 \
  -c 32768

LM Studio / Ollama:

  1. Load the main model and attach the bundled mmproj adapter.
  2. Maximize GPU offload and define context buffer.
Downloads last month
63
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IsValorum/Thomson-1.0-Small-APEX-I-MiniPlus-V2-GGUF

Quantized
(9)
this model

Collections including IsValorum/Thomson-1.0-Small-APEX-I-MiniPlus-V2-GGUF