Text Generation
Transformers
Safetensors
GGUF
cortex
spanish
bilingual
causal-lm
from-scratch
conversational
custom_code
Instructions to use Ilides/cortex-1-v0.9 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Ilides/cortex-1-v0.9 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Ilides/cortex-1-v0.9", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Ilides/cortex-1-v0.9", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Ilides/cortex-1-v0.9 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ilides/cortex-1-v0.9:F16 # Run inference directly in the terminal: llama cli -hf Ilides/cortex-1-v0.9:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ilides/cortex-1-v0.9:F16 # Run inference directly in the terminal: llama cli -hf Ilides/cortex-1-v0.9:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Ilides/cortex-1-v0.9:F16 # Run inference directly in the terminal: ./llama-cli -hf Ilides/cortex-1-v0.9:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Ilides/cortex-1-v0.9:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Ilides/cortex-1-v0.9:F16
Use Docker
docker model run hf.co/Ilides/cortex-1-v0.9:F16
- LM Studio
- Jan
- vLLM
How to use Ilides/cortex-1-v0.9 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ilides/cortex-1-v0.9" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ilides/cortex-1-v0.9", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ilides/cortex-1-v0.9:F16
- SGLang
How to use Ilides/cortex-1-v0.9 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Ilides/cortex-1-v0.9" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ilides/cortex-1-v0.9", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Ilides/cortex-1-v0.9" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ilides/cortex-1-v0.9", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use Ilides/cortex-1-v0.9 with Ollama:
ollama run hf.co/Ilides/cortex-1-v0.9:F16
- Unsloth Desktop
- Docker Model Runner
How to use Ilides/cortex-1-v0.9 with Docker Model Runner:
docker model run hf.co/Ilides/cortex-1-v0.9:F16
- Lemonade
How to use Ilides/cortex-1-v0.9 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Ilides/cortex-1-v0.9:F16
Run and chat with the model
lemonade run user.cortex-1-v0.9-F16
List all available models
lemonade list
- Atomic Chat
Cortex-1 v0.9 (sft-v0.9)
Modelo conversacional biling眉e ES/EN entrenado desde cero en MLX (bf16) y ajustado con SFT.
Arquitectura (verificada en config.json)
- 12 capas, hidden 768, 12 cabezas, intermediate 3072, contexto 512, vocab 16.384.
- Pesos float32, ~505 MB (ckpt-best.npz).
Inicializaci贸n y datos
- Init: pretrain v0.3 (val loss de referencia 2.5543 seg煤n configs/sft_v09.json).
- SFT: 28.800 pasos, batch 16, lr 4e-5 coseno, seed 123.
- Dataset chat sft-v0.9: 50 % ES / 50 % EN, ~45,6 M tokens de asistente en train (manifest).
Resultados
- Validaci贸n limpia (2 281 ventanas de 512 tokens sin bloques presentes en train, de 8 642 totales): val loss 2,1781 路 ppl 8,83.
- M茅todo: reports/eval_v09_filtered.py (misma m谩scara SFT que el entrenamiento). Reporte: reports/eval_v09_filtered.json.
Formatos
- Repositorio principal: safetensors + GGUF (F16, Q8_0).
- Copia MLX:
cortex-1-v0.9-mlx/(weights.safetensors + model.py).
Notas
- Los resultados del benchmark anterior (bpb 0.540, 20,6 M) corresponden al modelo cortex-0.3-0.02b, NO a v0.9. Se han retirado de esta tarjeta.
- Downloads last month
- -