YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Model Card: CodeLlama-7b-NL2PPlanner

This model is a fine-tuned version of codellama/CodeLlama-7b-Instruct-hf designed specifically to translate natural language project proposals into concrete, hierarchical directory structures.

It was trained using a custom Hybrid Loss engine that optimizes for both directory depth correctness and semantic relevance of file/folder names.

Special Tokens

The model generates linearized tree sequences using the following special tokens:

  • <TREE_START> / <TREE_END>: Bounds the entire project structure.
  • <DIR_START> / <DIR_END>: Bounds a directory/folder.
  • <FILE>: Indicates a file.

How to use

Below is a complete script to load the model, format the prompt correctly, generate the structure, and parse the tokenized output back into a nested Python Dictionary/JSON.

1. Install dependencies

pip install transformers torch

2. Inference & Parsing Code

import torch
import re
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "your_hf_username/CodeLlama-7b-NL2PPlanner"

# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    torch_dtype=torch.bfloat16, 
    device_map="auto"
)

# Function to parse generated tokens back to a dictionary tree
def regex_parser(text):
    tree = {}
    stack = [tree]
    token_pattern = re.compile(r"(<FILE>|<DIR_START>|<DIR_END>)\s+([^\s<]+)?")
    matches = token_pattern.findall(text)
    
    for tag, name in matches:
        if tag == "<FILE>" and name:
            stack[-1][name] = "file"
        elif tag == "<DIR_START>" and name:
            new_dir = {}
            stack[-1][name] = new_dir
            stack.append(new_dir)
        elif tag == "<DIR_END>" and len(stack) > 1:
            stack.pop()
    return tree

# System & User Prompt formatting
system_prompt = """You are a Principal Software Architect. Design the directory structure.
[GRAMMAR RULES]
1. Start with <TREE_START> and end with <TREE_END>.
2. Folders: <DIR_START> name ... <DIR_END>
3. Files: <FILE> name
4. NO JSON. ONLY TOKENS."""

user_input = """Design structure for: "E-Commerce Backend"
[CONTEXT]
- Domain: E-commerce (Web API)
- Desc: A robust backend for handling users, products, and orders.
- Stack: Python, PostgreSQL, Docker
[ARCH]
   - **AuthModule**: `/src/auth` (Handles JWT authentication).
   - **OrderModule**: `/src/orders` (Processes checkout logic).
[ENTRIES]
['/src/main.py']
[COMMAND]
Generate Linearized Token Sequence."""

formatted_prompt = f"<s>[INST] <<SYS>>\n{system_prompt}\n<</SYS>>\n\n<|user|>{user_input}<|end|> [/INST]"

# Generate
inputs = tokenizer(formatted_prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
    outputs = model.generate(
        **inputs, 
        max_new_tokens=1024,
        do_sample=False
    )

output_text = tokenizer.decode(outputs[0], skip_special_tokens=False)
generated_sequence = output_text.split("[/INST]")[1]

# Parse to JSON
directory_tree = regex_parser(generated_sequence)
print(directory_tree)
Downloads last month
6
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support