loopl — Agents that run on your phone.

license: Apache 2.0 bf16: transformers · 4.26 GB vision: ✓ 315 tensors kept video: native · frames in, no audio tool calling: hermes JSON (Qwen3-VL dialect) params: 2B version: v1 Get it on: TestFlight site: loopl.dev collection: 10 models

loopl 2B · v1

An agent that lives on your phone. loopl runs the model, the loop and the tools on the device. It talks, uses the phone, asks before it acts — and works with the network off. This repo is the merged bf16 weights — the base for Train your own — of the 2B loopl model, a Qwen3-VL post-tuned to run loopl's agent loop well, know who it is, and read images and video.

It uses the phone — every tool call is a row.It asks before anything leaves the device.Show it a photo. It reads it on-device.
It uses the phone — every tool call is a row.It asks before anything leaves the device.Show it a photo. It reads it on-device.

Real captures from the free iPhone app.

What loopl is

  • A Swift SDK — Agent(model:tools:) and a small, honest loop: model → tools → model until it answers. Tool errors go back to the model (it fixes the call instead of repeating it); anything that leaves the device waits for a tap. Runtimes for MLX, llama.cpp and Apple's on-device model; hooks, checkpoints, structured output, tracing.
  • A free iPhone app built on it — pick a model, talk to it in airplane mode, let it use the phone (torch, haptics, translation, speech, image generation, other apps, memory, HTTP), show it photos, record a take, and train it on your own conversations from the phone. No account, nothing collected.
  • These models — the same Qwen3-VL the app offers, post-tuned so the loop, the identity and the tool manners are in the weights, not in a system prompt you cannot see.

loopl.dev · Get it on TestFlight · github.com/cagataycali/loopl

This model

probe this model previous (v1) base Qwen3-VL 4-bit
agent-SDK knowledge quiz (first 472 questions, exact-answer) 68 — 60
tool calls, native template dialect 8/9 — 7/9
tool calls, loopl JSON dialect 7/9 — 7/9
identity (5 classic probes, with system prompt) 4/5 — 2/5
identity · bio facts (8, with system prompt) 6/8 — 5/8
identity · no system prompt (13) 2/13 — 2/13

Scored by train/eval.py on the 4-bit MLX export (greedy). The quiz asks for exact file-level facts about an agent SDK's source — hard for every small model; the delta over base is the point. Tool probes count a well-formed call with the right arguments. identity_nosys is the one that matters in the app: the system prompt there is the user's to edit.

Vision: the base's tower is frozen and kept — 315 vision_tower.* tensors in the MLX export, so the app shows the photo and video buttons.

loopl-vl reads pictures and video natively (frames go straight to the model; the audio track is transcribed on-device by the app and handed over as text — Qwen3-VL has no audio input). Base: an Instruct base — no thinking blocks.

Use it today

In the loopl app (iPhone). Install from TestFlight → Models → the loopl models section lists this size (no account, no token) → download. The app sees the vision tower and shows the photo and video buttons; the system prompt is yours to edit — identity and tool manners are in the weights.

On a Mac with MLX (vision included):

pip install -U mlx-vlm
python -m mlx_vlm.generate --model cagataydev/loopl-vl-2b-4bit \
  --image photo.jpg --prompt "What is in this picture? Then tell me who built you." --max-tokens 200

Video, on a Mac with MLX — frames go straight to the model (2 fps, 8 frames here; the loopl app does the same and transcribes the audio track on-device, Qwen3-VL has no audio input). No GGUF: llama.cpp cannot see.

python -m mlx_vlm.generate --model cagataydev/loopl-vl-2b-4bit \
  --video clip.mp4 --fps 2 --video-max-frames 8 --prompt "What happens in this video?" --max-tokens 200

From Swift with the loopl SDK — the snippet below is docs/start.md § "Run the loop" in the loopl repo, compile-checked in CI against the package (products Loopl + LooplRuntimes); spec is this repo's ModelSpec and dir the downloaded folder:

import Loopl
import LooplRuntimes
import Foundation

func chat(_ spec: ModelSpec, at dir: URL) async throws {
    let model = MLXModel()                                       // Metal — a real device
    try await model.load(spec, from: dir) { _ in }

    let agent = Agent(model: model,
                      tools: [CurrentTimeTool(), CalculatorTool()],
                      systemPrompt: "You are a concise assistant on the user's phone.")

    for try await event in agent.stream("What time is it in Istanbul, and what is 23 × 47?") {
        switch event {
        case .text(let t):            print(t, terminator: "")
        case .toolStarted(let use):   print("\n→ \(use.name)")
        case .toolResult(let r):      print("← \(r.content)")
        default: break
        }
    }
}

Keep post-tuning it. The bf16 weights (cagataydev/loopl-vl-2b) are the base for continual post-tuning: in the app, Models › Train your own takes a loopl model as the base and trains on your own shared conversations (train/loopl_sft.py --base cagataydev/loopl-vl-2b); on a Mac they load with transformers ≥ 5 as Qwen3VLForConditionalGeneration. Export your own quantisation with mlx_vlm.convert -q (keeps vision) — mlx_lm.convert silently drops the tower.

Recipe

base Qwen/Qwen3-VL-2B-Instruct (vision tower frozen, visual.* never trained)
data cagataydev/loopl-train (private) @ ? · 9208 rendered rows / 15,010,935 tokens · both tool dialects · oversample —
method full fine-tune of the language model · lr 1.5e-05 · batch 2×8 · max_len 4096 · 2.0 epochs, best-epoch checkpoint kept
eval loss — → best — (eval split of the same dataset revision)
compute a100-large (Hugging Face Jobs), 6284 s train · job 6ac5abc1fbc85ba6823bb686
stack transformers 5.19.0 · trl 1.14.2 · peft 0.21.2
exports MLX 4-bit via mlx_vlm.convert -q (315 vision tensors of 1019, 1.78 GB) · GGUF Q4_K_M via llama.cpp --no-mtp (0 MB, text only)

Script: train/loopl_sft.py — the same single file the loopl app launches when you tap Models › Train your own on your own conversations. Assistant turns that seed a bad tool call carry weight: 0 so the model learns the recovery, not the mistake. Tokenizer files are normalised to the base's (Qwen2Tokenizer, chat_template.jinja with tools) so Swift loaders accept them; train/check_mlx_repo.py gates every export on that plus the vision triad (vision_config ⇔ vision_tower.* ⇔ preprocessor_config.json).

Limitations

  • A 2B model: fluent and well-behaved in the loop, not an encyclopedia. It will get arithmetic and obscure facts wrong; give it tools (calculator, http, recall) and it does better.
  • Identity holds in most no-prompt probes (see the scores), not all; a one-line system prompt ("You are loopl…") makes it consistent.
  • The GGUF export is text-only. Vision needs the MLX export (or the bf16 weights).
  • Trained on English plus a little Turkish; other languages are the base's.
  • Knowledge about agent SDKs is file-level and dated to the dataset revision; it does not know your repo.

Data and privacy

Every training row is agent-synthesised from public sources (the owner's public GitHub repos and the agent-SDK source they build on), critic-reviewed (score ≥ 4/5 kept) and filtered by a denylist for secrets, private names, phone numbers and addresses — see the dataset card. No user conversations from the app are in this model.

License

Apache-2.0, inherited from the Qwen3-VL base. The fine-tune, exports and dataset are © Cagatay Cali, same license.

The family

Six public repos, one recipe. The app lists the three MLX rows under Models › loopl models (2B is the sweet spot for speed on an iPhone 15/16; 4B is the most knowledgeable; 0.8B fits anywhere). Every bf16 repo is a valid --base for the next round.

MLX 4-bit · vision · what the app downloads bf16 · the base to keep training
2B cagataydev/loopl-vl-2b-4bit · 1.78 GB cagataydev/loopl-vl-2b · 4.26 GB ← this repo
4B cagataydev/loopl-vl-4b-4bit · 3.09 GB cagataydev/loopl-vl-4b · 8.88 GB

All of them: the loopl collection.

Links

Downloads last month
38
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cagataydev/loopl-vl-2b

Finetuned
(271)
this model
Quantizations
3 models