Instructions to use csankalp21/headlinegpt-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use csankalp21/headlinegpt-onnx with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-generation', 'csankalp21/headlinegpt-onnx');
HeadlineGPT โ ONNX / WebGPU
Browser-ready ONNX export of csankalp21/headlinegpt, a Qwen2.5-1.5B-Instruct model fine-tuned for generating concise, engaging titles.
This version is designed for local, in-browser inference using WebGPU with Transformers.js.
Your content can be processed entirely on the user's device without sending it to a remote inference server.
See the main HeadlineGPT model card for training details and model limitations.
Model Format
| Property | Value |
|---|---|
| Base model | Qwen2.5-1.5B-Instruct |
| Format | ONNX |
| Precision | FP16 |
| Export task | text-generation-with-past |
| Approximate size | 3.7 GB |
| Intended runtime | Transformers.js / ONNX Runtime Web |
| Hardware acceleration | WebGPU |
The current release uses FP16 weights. A smaller quantized version may be released separately.
Browser Usage
Install Transformers.js:
npm install @huggingface/transformers
```javascript
import { pipeline } from "@huggingface/transformers";
const generator = await pipeline(
"text-generation",
"csankalp21/headlinegpt-onnx",
{
device: "webgpu",
dtype: "fp16",
}
);
const messages = [
{
role: "system",
content: "You are an expert at writing highly engaging titles."
},
{
role: "user",
content:
"Generate a high-engagement title for the following content:\n\n" +
"<your content here>"
}
];
const output = await generator(messages, {
max_new_tokens: 40,
temperature: 0.7,
do_sample: true,
top_p: 0.9,
repetition_penalty: 1.1,
});
console.log(output);
Local Inference
Inference runs in the browser using the user's GPU through WebGPU.
No API key or inference server is required for browser-local inference.
After the model files have been downloaded, they can be cached locally by the browser for subsequent use.
Python / ONNX Runtime
The model can also be loaded with Optimum and ONNX Runtime:
from optimum.onnxruntime import ORTModelForCausalLM
from transformers import AutoTokenizer
model_id = "csankalp21/headlinegpt-onnx"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = ORTModelForCausalLM.from_pretrained(
model_id
)
Example
Input:
Apple is developing a new generation of artificial intelligence tools designed to make its devices more useful and personalized.
Generated title:
Apple's New AI Tools Will Make Your Devices More Useful
Performance & Requirements
WebGPU support and performance depend on the user's browser, operating system, GPU, available VRAM, and browser implementation.
The FP16 model is approximately 3.7 GB, so the initial download can be substantial. Model files may be cached locally by the browser after the first download.
Limitations
- English-focused; performance may vary substantially on other languages.
- Quality can vary outside the content distribution used during training.
- The model may occasionally introduce details that are not explicitly present in the source content.
- Engagement-oriented titles are not guaranteed to achieve higher engagement.
- Browser performance varies significantly across devices.
- The current FP16 release requires a relatively large initial download.
- WebGPU availability depends on browser and hardware support.
Related Models and Resources
- Main model: csankalp21/headlinegpt
- Transformers.js: Documentation
- Qwen2.5-1.5B-Instruct: Base model
License
Apache 2.0.
Please also review the license and usage terms of the underlying Qwen2.5-1.5B-Instruct model.
- Downloads last month
- -