HeadlineGPT โ€” ONNX / WebGPU

Browser-ready ONNX export of csankalp21/headlinegpt, a Qwen2.5-1.5B-Instruct model fine-tuned for generating concise, engaging titles.

This version is designed for local, in-browser inference using WebGPU with Transformers.js.

Your content can be processed entirely on the user's device without sending it to a remote inference server.

See the main HeadlineGPT model card for training details and model limitations.

Model Format

Property Value
Base model Qwen2.5-1.5B-Instruct
Format ONNX
Precision FP16
Export task text-generation-with-past
Approximate size 3.7 GB
Intended runtime Transformers.js / ONNX Runtime Web
Hardware acceleration WebGPU

The current release uses FP16 weights. A smaller quantized version may be released separately.

Browser Usage

Install Transformers.js:

npm install @huggingface/transformers


```javascript
import { pipeline } from "@huggingface/transformers";

const generator = await pipeline(
  "text-generation",
  "csankalp21/headlinegpt-onnx",
  {
    device: "webgpu",
    dtype: "fp16",
  }
);

const messages = [
  {
    role: "system",
    content: "You are an expert at writing highly engaging titles."
  },
  {
    role: "user",
    content:
      "Generate a high-engagement title for the following content:\n\n" +
      "<your content here>"
  }
];

const output = await generator(messages, {
  max_new_tokens: 40,
  temperature: 0.7,
  do_sample: true,
  top_p: 0.9,
  repetition_penalty: 1.1,
});

console.log(output);

Local Inference

Inference runs in the browser using the user's GPU through WebGPU.

No API key or inference server is required for browser-local inference.

After the model files have been downloaded, they can be cached locally by the browser for subsequent use.

Python / ONNX Runtime

The model can also be loaded with Optimum and ONNX Runtime:

from optimum.onnxruntime import ORTModelForCausalLM
from transformers import AutoTokenizer

model_id = "csankalp21/headlinegpt-onnx"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = ORTModelForCausalLM.from_pretrained(
    model_id
)

Example

Input:

Apple is developing a new generation of artificial intelligence tools designed to make its devices more useful and personalized.

Generated title:

Apple's New AI Tools Will Make Your Devices More Useful

Performance & Requirements

WebGPU support and performance depend on the user's browser, operating system, GPU, available VRAM, and browser implementation.

The FP16 model is approximately 3.7 GB, so the initial download can be substantial. Model files may be cached locally by the browser after the first download.

Limitations

  • English-focused; performance may vary substantially on other languages.
  • Quality can vary outside the content distribution used during training.
  • The model may occasionally introduce details that are not explicitly present in the source content.
  • Engagement-oriented titles are not guaranteed to achieve higher engagement.
  • Browser performance varies significantly across devices.
  • The current FP16 release requires a relatively large initial download.
  • WebGPU availability depends on browser and hardware support.

Related Models and Resources

License

Apache 2.0.

Please also review the license and usage terms of the underlying Qwen2.5-1.5B-Instruct model.


Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for csankalp21/headlinegpt-onnx

Quantized
(249)
this model