Himeros V2 27B · Long-context roleplay

Character-driven dialogue, romance and long-form creative writing

Himeros V2 is a 27B LoRA fine-tune of Qwen3.8-27B-Uncensored, trained with sequences up to 49,152 tokens. Its focus is English roleplay: sustained scenes, character interaction and narrative continuation.

This repository distributes GGUF precision variants of the same merged model. Each GGUF is standalone: download one file, with no separate adapter or base-model download required.

Model overview

Detail Configuration
Base model orcarouter/Qwen3.8-27B-Uncensored
Base revision 404ea47aaa5d8a8b00049c9e9750089aca011ab2
Model size 27B parameter class
Primary language English
Fine-tuning method LoRA, merged into the base before export
LoRA rank / alpha 16 / 32
Learning rate / schedule 1e-5 / cosine
Training sequence limit 49,152 tokens (48 × 1,024)
Distribution Single-file GGUF, K and IQ quantizations plus BF16

The training mixture combines roleplay conversations, synthetic dialogue and long-form prose. Source training texts are not included in this repository. The training context limit describes the fine-tuning configuration; it does not guarantee uniform quality across an entire 48K conversation.

GGUF collection

The export set comprises the following 23 formats. Consult Files and versions for files currently available; UPLOAD_COMPLETE.json records completion of the full transfer.

Family Formats
K quantizations Q2_K, Q2_K_S, Q3_K_S, Q3_K_M, Q3_K_L, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K
IQ quantizations IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ2_M, IQ3_XXS, IQ3_XS, IQ3_S, IQ3_M, IQ4_NL, IQ4_XS
Unquantized export BF16

Files follow this naming pattern:

Himeros-V2-27B-<FORMAT>.gguf

These are alternatives, not pieces of one model. Choose a precision that fits your available memory, leaving room for the context cache and runtime overhead. Quantization can affect output quality; this release does not provide a controlled comparison of every format.

Getting started

  1. Download a single .gguf from Files and versions.
  2. Import it into a compatible LM Studio or llama.cpp runtime.
  3. Use the embedded chat template.
  4. Set a context length and GPU offload configuration that fit your hardware.
  5. Supply your character description, scene setup and writing preferences in the system prompt.

Thinking mode

Thinking is not permanently disabled in the GGUF. Whether it is enabled depends on the client and inference backend.

A short local LM Studio test confirmed zero reasoning tokens with Off and reasoning output with Low. This verifies the tested integration, not identical behavior in every client. Exact Low/Medium/High budgets remain backend-dependent.

Export integrity

The original quantization shards were joined losslessly and audited for unchanged tensor names, shapes, types and raw tensor bytes. The transfer workflow compares file SHA-256 against the completed Drive export record and Hugging Face file metadata.

SHA256SUMS.txt provides download checksums once the full transfer completes. These checks verify file integrity, not generation quality. JSON completion records are not required to load a GGUF.

Evaluation and limitations

  • No controlled benchmark establishes superiority over the base model.
  • Grammar, character continuity, repetition and long-context recall can still fail.
  • Quantized variants have not all been independently quality-benchmarked.
  • Outputs may contain mature themes, bias or inaccurate information; review them for your use case.
  • This is a creative-writing model, not a source of factual or professional advice.

Attribution and licensing

Thanks to the upstream model authors and the developers of the training and GGUF tooling. Base-model use and redistribution remain subject to the applicable upstream terms. This card does not assert an additional license grant or that third-party training-source rights have been cleared.

Questions

If you have any questions, please DM me.

Downloads last month
14,111
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Skttttt/Himeros-V2-27B-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(41)
this model

Collection including Skttttt/Himeros-V2-27B-GGUF