Instructions to use ArjunShukla/PromptForge-Quality with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ArjunShukla/PromptForge-Quality with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="ArjunShukla/PromptForge-Quality")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ArjunShukla/PromptForge-Quality", device_map="auto") - Notebooks
- Google Colab
- Kaggle
PromptForge-Quality
Multi-dimension prompt quality scorer. Given an LLM prompt, returns an overall quality score plus per-dimension scores, inferred issues, and missing information.
Part of PromptForge — local-first prompt scoring and optimization.
Model Details
Model Description
PromptForge-Quality is a fine-tuned answerdotai/ModernBERT-base encoder with regression heads that predict prompt quality on a 0–100 scale across seven dimensions.
- Developed by: PromptForge contributors
- Model type: Encoder + multi-output regression (
promptforge_quality) - Language(s): English
- License: MIT
- Finetuned from:
answerdotai/ModernBERT-base(~150M parameters)
Dimensions scored
| Dimension | What it measures |
|---|---|
clarity |
How clear and unambiguous the prompt is |
specificity |
Level of concrete detail |
context |
Background / situation provided |
goal_definition |
How well the objective is defined |
constraints |
Limits, requirements, must/must-not rules |
completeness |
Whether enough information is present |
actionability |
How easy it is for an LLM to act on |
quality_score |
Aggregate overall score |
The model also surfaces issues (e.g. too_vague, missing_context) and missing_information hints.
Model Sources
- Repository: https://github.com/arjun988/promptModel
- Companion model: PromptForge-Optimizer (Qwen2.5-1.5B LoRA prompt rewriter)
- Demo: Gradio app in the PromptForge repo (
demo/app.py)
Uses
Direct Use
- Score prompts before sending them to an LLM
- Diagnose weak prompts (what’s missing / unclear)
- Measure before/after quality when rewriting prompts
- Local / offline tooling via PromptForge CLI and Python API
Downstream Use
- Prompt linters in IDEs and agent frameworks
- Dataset filtering / ranking for synthetic prompt corpora
- Paired with PromptForge-Optimizer for score → optimize → re-score pipelines
Out-of-Scope Use
- Not a content moderator or safety classifier
- Not a judge of factual correctness of LLM answers
- Scores are calibrated on synthetic prompt quality labels — treat them as a useful proxy, not ground truth for every domain
Bias, Risks, and Limitations
- Trained largely on synthetic prompts with heuristic quality labels
- May reward length / structure more than true semantic quality
- English-centric; behavior on other languages is unverified
- Very domain-specific jargon may score inconsistently
Recommendations
- Use scores comparatively (before vs after) rather than as absolute grades
- Combine with human review for high-stakes prompt design
- For custom domains, retrain with your own labeled prompts
How to Get Started with the Model
Install & use with tuneprompt (recommended)
pip install tuneprompt
python -m promptforge download \
--quality-repo ArjunShukla/PromptForge-Quality \
--optimizer-repo ArjunShukla/PromptForge-Optimizer
python -m promptforge analyze "Build me a website"
# or: tuneprompt analyze "Build me a website"
from promptforge import PromptForge
pf = PromptForge(quality_model_path="ArjunShukla/PromptForge-Quality")
print(pf.analyze("Make an app."))
# → quality_score, dimensions, issues, missing_information
Package:
tuneprompton PyPI · Import:promptforge· CLI:tuneprompt/promptforge· Code: https://github.com/arjun988/promptModel
Full pipeline (score + optimize)
from promptforge import PromptForge
pf = PromptForge(
quality_model_path="ArjunShukla/PromptForge-Quality",
optimizer_model_path="ArjunShukla/PromptForge-Optimizer",
)
result = pf.run("Make an app about social media like facebook and stuff")
print(result["before"]["quality_score"], "→", result["after"]["quality_score"])
print(result["optimized_prompt"])
Training Details
Training Data
- ~25,000 synthetic prompts across coding, writing, research, data, and creative tasks
- Quality levels from vague one-liners to fully specified prompts
- Labels: overall
quality_score+ seven dimension scores (0–100)
Training Procedure
| Setting | Value |
|---|---|
| Base model | answerdotai/ModernBERT-base |
| Task | Multi-dimension regression |
| Epochs | 3 |
| Max length | 512 |
| Precision | fp16 |
| Hardware | NVIDIA RTX 5060 Laptop (8 GB) |
| Wall time | ~33 minutes |
Config: configs/quality_scorer.yaml
Evaluation
Held-out results (local training run):
| Split | MAE | Pearson |
|---|---|---|
| Validation | 2.73 | 0.993 |
| Test (overall) | 0.96 | 0.999 |
Test Spearman (overall): 0.959
Summary
Strong correlation with synthetic quality labels on held-out data. Real-world prompts should still be sanity-checked — the scorer is best used for ranking and diagnosing structure gaps.
Environmental Impact
- Hardware Type: NVIDIA RTX 5060 Laptop (8 GB)
- Hours used: ~0.5 h for this checkpoint
- Cloud Provider: N/A (local)
- Carbon Emitted: Not measured
Technical Specifications
Model Architecture and Objective
ModernBERT encoder with dual / multi regression heads predicting continuous quality scores (0–100).
Compute Infrastructure
- Hardware: RTX 5060 Laptop GPU, 8 GB VRAM
- Software: PyTorch (CUDA), Transformers, PromptForge training scripts
Artifact size
- On-disk checkpoint: ~574 MB
Citation
@software{promptforge_quality,
title = {PromptForge-Quality},
author = {PromptForge Contributors},
year = {2026},
url = {https://huggingface.co/ArjunShukla/PromptForge-Quality}
}
Model Card Contact
Open an issue on the PromptForge GitHub repository.
- Downloads last month
- -
Model tree for ArjunShukla/PromptForge-Quality
Base model
answerdotai/ModernBERT-base