Caption Tool

A small, friendly tool that writes text descriptions (captions) for every image in a folder, all on your own computer.

  • No API key. No account. No internet (after the first-time model download).
  • Private: your images never leave your machine.
  • Lightweight: runs on a GPU with as little as 2 GB of VRAM.
  • Whole folders at once: point it at a directory and it captions everything.

Made by oohfixer.com.


What it does

For every image it finds, Caption Tool writes a matching text file next to it:

vacation.jpg   โ†’   vacation.txt
cat.png        โ†’   cat.txt

That's it. The caption sits in the .txt file, ready to use for training image models, organizing photos, or search.


Install

You need Python 3.13 or older (3.14 doesn't work with this tool's dependencies).

One command (recommended)

Open a terminal and run:

curl -fsSL https://oohfixer.com/caption-tool/install.sh | bash

This downloads the tool, installs what it needs in its own isolated environment, and adds a captions command you can run from anywhere. When it finishes, open a new terminal (or run source ~/.bashrc).

Then just type:

captions

Manual install

If you'd rather set it up by hand:

  1. Download this folder and open a terminal in it.

  2. Install the requirements:

    pip install -r requirements.txt
    
  3. Run it:

    python caption.py
    

Use it

Just run it:

python caption.py

It will ask you for a folder of images, show the settings, and caption them. Sit back, a progress bar shows how far it's gotten.

Want to skip the questions?

If you already know what you want, you can do it in one line:

python caption.py --path "/path/to/your/images"

Useful extras:

Option What it does
--path "FOLDER" Caption this folder right away (no prompt).
--force Re-caption images that already have a .txt file.
--model florence2-base Pick a different model (see below).
--max-new-tokens 512 Allow longer captions.

Choosing a model

When you run it, you can pick which AI model writes the captions. Bigger = more detailed, but uses more VRAM.

Model Size Best for
florence2-large (default) ~1.5 GB Slightly less accurate. Needs a normal GPU.
florence2-base ~0.45 GB 2 GB GPUs, or when you want it fast.

If your card only has 2 GB of VRAM, choose florence2-base.


Prefix and suffix (optional)

If you're captioning images to train your own image model, you may want a trigger word on every caption (for example, your subject's name). Set a prefix and the tool adds it to the front of every caption automatically. A suffix does the same at the end.

Leave them empty if you just want plain descriptions.


Settings

Settings are stored in caption_config.json next to the script. You can edit that file directly, or change things from the tool's menu. The defaults:

  • Skip images that already have a caption (so re-running is safe).
  • At least ~15 words per caption (it re-captions anything shorter).
  • 256 token cap on caption length.

Questions

  • Does it need to be online? Only the very first time, to download the model. After that, everything runs offline.
  • Where do my images go? Nowhere. They're processed locally and the caption is saved next to each image.
  • What image types? jpg, png, webp, bmp, tiff.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support