llama3-8b-code-judge

A QLoRA fine-tuned adapter for Llama 3.1 8B Instruct, trained to deliver deadpan comedic verdicts on code, as the "Judge" persona in the AI Code Court project.

Model Details

Model Description

This is a LoRA adapter fine-tuned on top of meta-llama/Llama-3.1-8B-Instruct, trained to act as a courtroom judge that delivers short, dry, comedic verdicts on code quality, given a code snippet and short "prosecutor" and "defense" arguments. It is part of AI Code Court, a project where multiple LLMs argue about the quality of user-submitted code.

  • Developed by : Jahnavi Reddy (jahnavi0803 on Hugging Face, jahnavi-reddy03 on GitHub)
  • Funded by : Self-funded personal/portfolio project
  • Shared by : Jahnavi Reddy
  • Model type: Causal language model, LoRA adapter (requires the base model to run — not a standalone full model)
  • Language(s) (NLP): English
  • License: MIT
  • Finetuned from model : meta-llama/Llama-3.1-8B-Instruct

Model Sources

Uses

Direct Use

Generating a short, stylized, comedic "verdict" on a piece of code, given a code snippet and two short arguments (a "prosecutor" case and a "defense" case), in the specific prompt format described in "How to Get Started" below.

Downstream Use

Intended to be plugged into the AI Code Court app as the "Judge" role in a three-LLM courtroom simulation (GPT-4o as prosecutor, Gemini as defense, this model as judge).

Out-of-Scope Use

Not intended for real code review, security auditing, or any decision with real technical or business consequences. Its "reasoning" output is stylistic and comedic, not a substitute for actual static analysis, linting, or human code review.

Bias, Risks, and Limitations

Trained on only 50 examples, all written or generated by one person in a single sitting, so its "judgment" reflects one person's sense of humor and coding opinions rather than a broad or vetted standard. It has not been evaluated on code outside common, simple Python patterns (the training set skewed toward classic bugs: SQL injection, bare excepts, hardcoded secrets, off-by-one errors, global state misuse). Performance on unusual, large, or multi-file code, or languages other than Python, is unknown.

Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use for entertainment and portfolio/demo purposes only — do not use its verdicts as a substitute for real code review, linting, or security scanning.

How to Get Started with the Model

Use the code below to get started with the model.

Confirmed working. Given a code snippet and courtroom arguments in the correct prompt format, it reliably generates a structured verdict:

VERDICT: [Guilty of Bad Code / Not Guilty / Guilty with Mitigating Circumstances] REASONING: [one dry, sarcastic sentence] SENTENCE: [a punchy label like "Refactor Immediately"] ONE-LINER: [a quotable roast or compliment]

Example real output from this model:

VERDICT: Guilty of Bad Code REASONING: Because who needs error messages, anyway? SENTENCE: Refactor Immediately ONE-LINER: "This code is so trusting, it's practically a hostage situation."

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)

base_model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    quantization_config=bnb_config,
    device_map="auto",
)
model = PeftModel.from_pretrained(base_model, "jahnavi0803/llama3-8b-code-judge")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")

Prompt format matters: this model was trained on a specific structure (system prompt + code wrapped in triple backticks + labeled PROSECUTOR/DEFENSE arguments + an explicit instruction to use the exact format). See the training code for the exact template used.

Training Details

Training Data

50 courtroom transcripts (code snippet + prosecutor argument + defense argument → verdict) collected live from the AI Code Court app itself, using GPT-4o as prosecutor and Gemini as defense, then converted into prompt/completion pairs for supervised fine-tuning. Dataset prep script: prepare_dataset.py.

Training Procedure

QLoRA: base model loaded in 4-bit (NF4 quantization), LoRA adapter (rank 16, alpha 32) applied to the attention projection layers (q_proj, k_proj, v_proj, o_proj), trained with supervised fine-tuning (SFTTrainer).

Preprocessing

Each training example was formatted as: system prompt (judge persona + output format instructions) + code wrapped in triple backticks + labeled PROSECUTOR/DEFENSE arguments + instruction to deliver the verdict, followed by the target completion (the verdict itself) and an end-of-sequence token.

Training Hyperparameters

  • Training regime: bf16 mixed precision
  • LoRA rank: 16, alpha: 32
  • Batch size: 2 (gradient accumulation steps = 4)
  • Epochs: 3
  • Learning rate: 2e-4
  • Trainable params: 13.6M (0.17% of the base model's 8.04B parameters)

Speeds, Sizes, Times

Trained on a single free-tier Google Colab T4 GPU. Training runtime: approximately 27-30 minutes for 3 epochs over 50 examples. Adapter size: approximately 55MB (LoRA adapter only, not the full ~16GB base model).

Evaluation

Testing Data, Factors & Metrics

Testing Data

No separate held-out test set was used. The dataset is small (50 examples) and the goal was stylistic fine-tuning rather than benchmark performance, so evaluation here is training-time metrics only.

Factors

Not disaggregated by any factor — training metrics are reported as an overall average across the full training set.

Metrics

Training loss (cross-entropy) and mean token accuracy, tracked during training, used to confirm the model was successfully learning the target output format and style.

Results

Training loss: 0.71 → 0.55 → 0.62 (final, across 3 epochs). Mean token accuracy: reached 87% by the final epoch.

Summary

The model successfully learned to produce the target four-section verdict format and shows a clear, comedic, consistent judge persona, confirmed via manual generation testing after training.

Model Examination

No formal interpretability analysis was performed. Manual spot-checking of generations confirmed the model reliably follows the trained output format when prompted with the exact training template.

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

  • Hardware Type: 1x NVIDIA T4 GPU
  • Hours used: Training runtime was approximately 30 minutes (per Colab training logs); total GPU session time including model downloads and setup was somewhat longer, not precisely tracked.
  • Cloud Provider: Google Colab (free tier)
  • Compute Region: Unknown (Colab-managed, not disclosed to the user)
  • Carbon Emitted: Not formally calculated — training time was under 30 minutes on a single T4, so impact is minimal relative to typical LLM training runs

Technical Specifications

Model Architecture and Objective

Causal language model (Llama 3.1 architecture) with a LoRA adapter, trained with a next-token prediction (language modeling) objective, fine-tuned to produce a specific structured text output (verdict format) conditioned on code and courtroom argument context.

Compute Infrastructure

Google Colab, free tier.

Hardware

Single NVIDIA T4 GPU (16GB VRAM)

Software

transformers, peft, trl, bitsandbytes, accelerate, torch, datasets

Citation

No formal paper or citation exists for this project.

BibTeX:

None.

APA:

Reddy, J. (2026). llama3-8b-code-judge [Model]. Hugging Face. https://huggingface.co/jahnavi0803/llama3-8b-code-judge

Glossary

  • QLoRA: Quantized Low-Rank Adaptation, a memory-efficient fine-tuning method that trains a small set of adapter weights on top of a frozen, 4-bit-quantized base model.
  • LoRA adapter: A small set of additional trained weights that modify a base model's behavior without changing the base model's original parameters, and which must be loaded together with the base model to be used.

More Information

Part of a larger project, AI Code Court — a three-LLM courtroom simulation where GPT-4o prosecutes code, Gemini defends it, and this model delivers the verdict.

Model Card Authors

Jahnavi Reddy

Model Card Contact

github.com/jahnavi-reddy03

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jahnavi0803/llama3-8b-code-judge

Finetuned
(3257)
this model

Paper for jahnavi0803/llama3-8b-code-judge