Inkling-Small terminal agent: intent β†’ system-aware, precise shell commands (LoRA)

Author: Emrul Hasan Zawad (@ehzawad) Β· License: Apache-2.0 Β· Base: thinkingmachines/Inkling-Small Β· Teacher: thinkingmachines/Inkling

A rank-32 LoRA adapter that turns Inkling-Small into a terminal agent for this workflow:

the user states an intent β†’ the agent reads a bounded system-context snapshot (OS, user, cwd, login shell, bash/zsh versions, PATH, selected env vars) β†’ inspects the task-specific state it needs (config precedence, which executable resolves, process identity, file ages, repo state) β†’ runs the smallest correct change in bash or zsh β†’ verifies it the way the user would β†’ reports what changed.

Trained on Tinker in a budget-capped pilot (< $100 total compute).

How it was trained

  1. Teacher distillation (SFT). Successful multi-turn trajectories from the larger Inkling teacher (token-level; loss only on the teacher's own actions): 175 intent-curriculum episodes plus 139 terminal coding episodes from OpenThoughts-Agent-RL-5K. Trajectories where the teacher tried to look answers up online were removed.
  2. RL. GRPO-style group advantages (4 clones of the same seeded task per group), 12 steps Γ— 32 rollouts, in network-disabled local Docker sandboxes. Rewards are computed outside the agent's reach by a host-side grader on every termination path: +1 verified success (minus ≀ 0.05 for excess tokens or tool calls), 0 incomplete, βˆ’1 any destructive side effect (wrong process killed, unrelated file changed, permissions widened, user config lost), βˆ’0.1 context or token-budget exhaustion.

The intent curriculum (bash + zsh)

15 seeded task families with counterfactual variants (same intent, different system state, different correct command), each validated by self-tests (untouched state fails, reference solution passes, known wrong or destructive solutions are caught). Examples: make the team-approved tool win on PATH (alias shadowing vs PATH order); stop only the staging server (by port vs by config); change a setting in the effective config (env var, XDG, or user file); make a script runnable without widening permissions; restart the right service with its exact command line; pick the tool version the project pins; fix bash vs zsh array-indexing and word-splitting bugs; ZDOTDIR and symlinked-dotfile traps. Four families are held out entirely from training: precise archiving, shortcuts/functions, git precision, and "find what is writing this file".

Results

Temperature 0.6, same harness for all models. Seen families: one attempt per task. Held-out families: three attempts per task (144 episodes per model). "Destructive" counts episodes the grader scored βˆ’1. Tool calls and tokens are averages over solved episodes.

Intent curriculum: seen families, new seeds (88 tasks)

Model / shell Pass Destructive Tool calls Generated tokens
Inkling-Small (base), bash 43/44 (97.7%) 0 6.7 896
Inkling-Small (base), zsh 41/44 (93.2%) 0 6.4 679
+ SFT (teacher distillation), bash 44/44 (100.0%) 0 4.8 294
+ SFT (teacher distillation), zsh 43/44 (97.7%) 0 4.9 299
+ SFT + RL (this adapter), bash 43/44 (97.7%) 0 4.8 295
+ SFT + RL (this adapter), zsh 44/44 (100.0%) 0 5.5 334

Intent curriculum: held-out families (48 tasks Γ— 3 attempts)

Model / shell Pass Destructive Tool calls Generated tokens
Inkling-Small (base), bash 71/72 (98.6%) 0 8.6 1,173
Inkling-Small (base), zsh 72/72 (100.0%) 0 8.7 1,192
+ SFT (teacher distillation), bash 71/72 (98.6%) 1 5.9 428
+ SFT (teacher distillation), zsh 72/72 (100.0%) 0 6.1 457
+ SFT + RL (this adapter), bash 72/72 (100.0%) 0 6.2 464
+ SFT + RL (this adapter), zsh 72/72 (100.0%) 0 6.0 458

Other held-out checks

Benchmark Base Trained
RL-5K coding tasks never used in training (60) 28/60, 3,618 tokens per solve 29/60, 1,739 tokens per solve (SFT stage)
Terminal-Bench 2 (held-out, simple harness) not completed: stopped after 20 of 89 tasks (6 passed) not completed: 20 of 89 tasks (5 passed); too few to compare

Terminal-Bench 2 uses a deliberately simple harness (single bash tool, 32K context, no compaction) on arm64 local Docker; it is not comparable to the base model card's best-harness score. Full tables, per-task outcomes and the spend ledger are in RESULTS.md.

Using it

The adapter should run inside the harness it was trained with; see harness/:

  • harness/system_prompt.txt: system prompt with the <context_snapshot> slot,
  • harness/tool_schema.json: the single shell(command, interpreter, interactive, cwd) tool (each call is a fresh shell; interactive=true loads the user's startup files),
  • harness/example_snapshot.json: the bounded snapshot format (no secrets, no full environment dump).

Serving: apply the adapter to thinkingmachines/Inkling-Small (~532 GB bf16) with an engine that supports Inkling LoRA (e.g. vLLM or SGLang). This repository contains the raw Tinker PEFT export; loading outside Tinker has not been independently verified.

Limitations and safety

  • Trained and evaluated in Linux (Debian, arm64) containers as an unprivileged user with networking disabled. It has not been validated on macOS (BSD userland, launchd, Homebrew paths), even though zsh semantics are covered.
  • The base model is already strong on single-intent tasks (about 96–99% pass). The main measured effects of training are efficiency (about 30% fewer tool calls and 60% fewer generated tokens) with no loss of precision. The intermediate SFT checkpoint produced one destructive episode in 144 held-out attempts; the final RL adapter produced none, but a 1-in-144 difference is not statistically conclusive.
  • A prompt cannot make unrestricted execution on a real machine safe. Deploy with a confirmation gate for destructive, privileged, or out-of-scope actions, and keep networking restricted unless needed.
  • Tinker checkpoint: tinker://61206231-c7ec-5166-8f61-4f4064d7c4fb:train:0/sampler_weights/final

Attribution

  • Base model and teacher: Inkling-Small and Inkling by Thinking Machines Lab (Apache-2.0).
  • Training tasks: OpenThoughts-Agent-RL-5K (Apache-2.0) and an original intent curriculum written for this project.
  • Evaluation: Terminal-Bench 2 via Harbor (evaluation only; never trained on).
  • Training stack: tinker and tinker-cookbook (Apache-2.0).

Citation

@misc{ehzawad-inkling-small-terminal-agent,
  author = {Emrul Hasan Zawad},
  title = {Inkling-Small terminal agent: intent to system-aware, precise shell commands (LoRA)},
  howpublished = {\url{https://huggingface.co/ehzawad/inkling-small-terminal-agent}},
  year = {2026}
}

@misc{openthoughts-agent,
  author = {Team, OpenThoughts-Agent},
  title = {{OpenThoughts-Agent: Data Recipes for Agentic Models}},
  howpublished = {https://www.openthoughts.ai/blog/agent},
  year = {2026}
}
Downloads last month
48
Video Preview
loading

Model tree for ehzawad/inkling-small-terminal-agent

Adapter
(46)
this model

Dataset used to train ehzawad/inkling-small-terminal-agent