keys-Auto Receipts Studio (iPhone / may add Autonomous Lamp Skill)

v1.0 alpha

iPhone (or desktop) photo → local Gemma 4 12B-it vision extract → SQLite. Optional Autonomous OS Lamp skill (camera on the robot, 12B on the GPU box — the Lamp’s 6 GB cannot hold 12B).

GitHub: drowzeys/keys-Auto-Receipts-Studio
Hugging Face: drowzeys/keys-Auto-Receipts-Studio

Pathways (pick a recipe)

Gemma 4 12B-it never runs on the Lamp (6 GB) and typically not on a Windows/Mac laptop. The GPU box is the brain. PC / Mac / Linux run the inbox + Review UI. iPhone is a Safari capture client on the same Wi‑Fi.

Pathway One-shot recipe Gemma 4 12B iPhone page
A. Linux NVIDIA GPU box (Spark / Omen) bash oneshot.sh started here (:8080, util 0.15, FP8, never above 0.85) http://<this-box-lan-ip>:7860/phone
B. Linux, no NVIDIA bash oneshot.sh then edit .env skip local vLLM; point .env at pathway A same, this machine’s LAN IP
C. Windows PC oneshot.bat then scripts\install-launcher.bat on pathway A (or WSL2) http://<this-pc-lan-ip>:7860/phone
D. macOS venv + bash scripts/install-launcher.sh on pathway A http://<this-mac-lan-ip>:7860/phone
E. iPhone Safari only, after A/B/C/D UI is up n/a (uploads to the UI box) /phoneTake photo or Choose file
F. Autonomous Lamp skills/keys-receipt-scanner/ on pathway A n/a (HAL camera → Studio /api/inbox)
git clone https://github.com/drowzeys/keys-Auto-Receipts-Studio.git
cd keys-Auto-Receipts-Studio

Desktop icon after first UI start: Linux/macOS bash scripts/install-launcher.sh · Windows scripts\install-launcher.bat. Step-by-step for each letter is below.

Screenshots (v1.0 alpha)

iPhone Safari — take a photo, then completion notice:

iPhone: Take photo / Choose file iPhone: upload in progress iPhone: Completed Whole Foods Market groceries #5

Desktop Review queue (click a row to edit kind / category / vendor / date / total, or Delete):

Review queue

Inbox drop zone, catalog SKUs, and backend health:

Inbox Catalog Settings backends

Promo (v1.0 alpha)

~50s story cut — messy dorm → iPhone (lay the receipt flat, then photograph) → charcoal Autonomous Lamp scans the last receipt → trash → clean room. Bed: Beethoven Symphony No. 5, I. Allegro con brio (public-domain recording).

Also attached on the v1.0-alpha-promo release (plays in the browser).

Pathway details

One-click / app.launch binds the UI on the LAN (RECEIPT_UI_SHARE_LAN=true) and opens Review in the browser. Phone URL is printed in the terminal.

Desktop icons (do not copy scripts/start-ui.bat itself to the Desktop):

OS Command Opens
Linux bash scripts/install-launcher.sh Desktop Receipt Studio.desktop
macOS bash scripts/install-launcher.sh Desktop Receipt Studio.command (first time: right-click → Open)
Windows scripts\install-launcher.bat Desktop Receipt Studio.bat

A. Linux NVIDIA GPU box (Gemma + UI + iPhone)

bash oneshot.sh

Idempotent. Installs .venv, fetches google/gemma-4-12B-it to ~/models-gemma4-12b-it if missing (hf auth login once), starts scripts/serve-gemma.sh if :8080 is down, starts the UI on 0.0.0.0:7860.

When it prints READY:

Review http://127.0.0.1:7860
iPhone Safari (same Wi‑Fi) http://<this-box-lan-ip>:7860/phone
Gemma OpenAI-compat http://127.0.0.1:8080/v1

vLLM flags (also used by scripts/start-ui.sh on a CUDA box if Gemma is down):

  • --gpu-memory-utilization **0.15** (~18.3 GiB of 121.7 GiB; never above 0.85)
  • FP8 (BF16 weights ~23 GB cannot fit in that pool)
  • --max-model-len **8192**

If Gemma is already running, oneshot / launch leave it alone. To apply the 0.15 cap: stop that vllm process, then bash oneshot.sh again.

Gemma-only (UI already running): bash scripts/serve-gemma.sh

B. Linux without NVIDIA (UI + iPhone only)

bash oneshot.sh

If vllm is not on PATH and nothing answers on :8080, oneshot skips weights and vLLM (it does not die) and still starts the UI. Then set .env at the GPU box from pathway A:

RECEIPT_LLM_BASE_URL=http://<spark-lan-ip>:8080/v1
RECEIPT_EMBED_BASE_URL=http://<spark-lan-ip>:8080/v1
RECEIPT_LLM_MODEL=google/gemma-4-12B-it
RECEIPT_EMBED_MODEL=google/gemma-4-12B-it
RECEIPT_EMBED_DIM=3840

Restart the UI (./scripts/start-ui.sh or the desktop icon).

C. Windows PC (UI + iPhone hotspot)

12B is not started by the Windows scripts. Serve Gemma on pathway A (or WSL2 + bash oneshot.sh).

  1. Install Python 3.12 — check Add python.exe to PATH.
  2. In the cloned repo, double-click oneshot.bat (creates .venv, opens Review).
  3. Edit .env as in pathway B (Spark LAN IP, port 8080).
  4. Optional desktop icon: scripts\install-launcher.batDesktop\Receipt Studio.bat.
  5. iPhone: http://<this-windows-lan-ip>:7860/phone in Safari.

Do not copy scripts\start-ui.bat to the Desktop; it must run from the repo (or from the wrapper install-launcher.bat writes).

If this PC is NVIDIA + vllm on PATH (Git Bash / WSL): point .env at 127.0.0.1:8080 and bash scripts/serve-gemma.sh.

D. macOS (UI + iPhone)

Apple Silicon does not load 12B with this vLLM-NVIDIA recipe. Use pathway A for Gemma.

python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"
cp .env.example .env
# edit RECEIPT_LLM_BASE_URL / RECEIPT_EMBED_BASE_URL like pathway B
bash scripts/install-launcher.sh

Double-click Desktop → Receipt Studio.command (first time: right-click → Open). Same as ./scripts/start-ui.sh.

E. iPhone (Safari)

Works against whichever machine runs the UI (A, B, C, or D), same Wi‑Fi. Chrome/Firefox on iOS are not the supported path.

  1. Open http://<ui-machine-lan-ip>:7860/phone
  2. Take photo (camera) or Choose file (Photos / Files, including HEIC)
  3. Wait for Completed: vendor · total · category · #id
  4. Confirm / edit / delete on the desktop Review tab

The phone page POSTs /api/inbox and polls /api/jobs/{id}. The GPU box (or the UI machine talking to Gemma) does extract + embed. Inbox drop and Syncthing into inbox/ use the same pipeline (files idle 30s unless the phone API processes immediately).

F. Autonomous Lamp skill

Repo path: skills/keys-receipt-scanner/ (SKILL.md, skill.json capabilities: ["vision"], scripts/scan.py).

skills/keys-receipt-scanner/
  SKILL.md
  skill.json
  scripts/scan.py
  references/hardware.md

Install (no reboot):

make push-skill SKILL=./skills/keys-receipt-scanner TARGET=pi@lamp-xxxx.local

On the Lamp (camera only — do not load 12B):

export RECEIPT_GPU_HOST=<spark-lan-ip>
export RECEIPT_STUDIO_URL=http://<spark-lan-ip>:7860
export RECEIPT_CAMERA_URL=http://127.0.0.1:5001

Say “scan this receipt”. scripts/scan.py snapshots HAL (:5001/camera/snapshot, 1280 px), then POSTs to Receipt Studio /api/inbox on the GPU box. Fallback: RECEIPT_LLM_BASE_URL=http://<spark-lan-ip>:8080/v1 (direct vLLM) or RECEIPT_HERMES_BASE_URL (same OpenAI-compat URL Hermes registered). On the GPU box, python scripts/register-hermes-gemma.py adds custom_providers gemma4-12b-it.

Intern has no camera; this skill will not install there.

To ship in a fork of autonomous-os: copy skills/keys-receipt-scanner/ into skills/, run python skills/skill-creator/scripts/quick_validate.py, make skills-catalog, open the PR.


Does Gemma 4 12B Unified fit on the Lamp?

No. Lamp is 8-core ARM64 with 6 GB RAM. Gemma 4 12B Unified is a 12B dense omni model (Gemma4UnifiedForConditionalGeneration, hidden size 3840).

Piece Lamp (6 GB) GPU box
skills/keys-receipt-scanner/ yes yes
HAL snapshot yes n/a
SQLite + HTTP client yes yes
Gemma 4 12B Unified weights no yes (vLLM util 0.15)

Gemma 4 12B Unified (omni) on the GPU box

One OpenAI-compat server: vision POST /v1/chat/completions and POST /v1/embeddings (dim 3840). Do not raise util above 0.85.

bash scripts/serve-gemma.sh
# util 0.15, FP8, max-model-len 8192

Context at util 0.15 (after ~12.5 GB FP8 weights):

Estimate Tokens
Conservative (48 layers full attn) ~12k
Hybrid (8 full + 40× sliding-1024) ~65k
Model native window 262,144 (not at 0.15)

A receipt scan uses ~1–2k tokens. Raise RECEIPT_VLLM_MAX_MODEL_LEN only after the boot log shows a large enough GPU KV cache size.

If /v1/embeddings 404s, either serve a pooling convert on another port (RECEIPT_EMBED_BASE_URL) or RECEIPT_EMBED_BACKEND=nvidia + Nemotron-3-Embed-1B at dim 2048. Never mix 3840 and 2048 in one DB.

Fallback vision (not omni embed): Qwen3.8-27B ADay777 at :8078 (RECEIPT_LLM_BACKEND=nvidia, RECEIPT_LLM_MODEL=qwen38-nvfp4). Lightning is text-only — never send images.

Manual UI (any OS, after venv)

python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
cp .env.example .env
# set RECEIPT_LLM_* like pathway B unless Gemma is local
RECEIPT_UI_SHARE_LAN=true python -m app.cli ui

CLI:

python -m app.cli scan --image path/to/receipt.jpg
python -m app.cli query --category groceries
pytest

Syncthing: camera/share folder → inbox/ (idle 30s then process).

What you still run yourself

  • Pathway A (or equivalent) so Gemma is reachable on the LAN
  • iPhone on the same Wi‑Fi, Safari
  • Lamp: pair in the Autonomous app, make push-skill, set RECEIPT_STUDIO_URL
  • Optional: Syncthing

No model weights in this repo. No PyInstaller in this release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support