GRAFT Code

A local AI agent for Windows that actually does things. GRAFT Code is a desktop app that runs a GRAFT model on your own PC and lets it work on your files: it makes folders, writes and reads files, moves, renames, copies and deletes things, runs commands and Python, does exact math, and searches and reads the web. Every change asks you first. Nothing leaves your computer except web-search keywords.

Made by SmallAICreator / UltraLabs.

File What it is
GRAFT-Code.exe GRAFT Code for Windows 10/11 (x64, runs on CPU). llama.cpp is bundled.
graft-bridge/ Optional: a Claude Code skill that lets GRAFT ask Claude for help (talk_to_claude).

Pick a model

Model Size Best for
GRAFT-10B-A2B 6.3 GB, 16 GB RAM The smartest option: a 10B mixture-of-experts with only 2B active, perfect recall up to 32K tokens
GRAFT-1B-Agentic 1 GB, 8 GB RAM Fast and light, tuned for tool use

Put the .gguf in your Downloads folder and GRAFT Code finds it.

Get started

  1. Download GRAFT-Code.exe and a model (above).
  2. Double-click the EXE. It's unsigned, so Windows SmartScreen may warn you: click More info → Run anyway.
  3. Wait for the model to load. The first launch reads GRAFT's instructions once (about 2 minutes on a laptop CPU with the 10B). Later launches restore them from disk in seconds.
  4. Type what you want done, for example "make a Python script that renames all my photos by date".

It needs Microsoft Edge WebView2, which Windows 10 and 11 already have. GRAFT works in ~/GRAFT-Workspace; change the folder from the sidebar.

Features

  • Clean chat window with light and dark themes. Replies stream in live, with formatted text, tables and code blocks you can copy.
  • Tool cards and approvals. Each action (write, run, search…) shows up as a card you can expand. Changes need an approval: Yes, Yes, and don't ask again or No, tell GRAFT what to do instead.
  • Model picker. Searches every .gguf in your Downloads, Desktop, Documents and LM Studio folders, has Browse… for anything else, and reopens the last model you used.
  • Personas. GRAFT, Coder, Playful, Concise and Teacher, or make your own with an emoji. Each persona has its own system prompt and sampling settings.
  • Settings.
    • Sampling: temperature, top-p, top-k, min-p, repeat penalty, and reply length (or unlimited: GRAFT writes until it's done).
    • Model: context size (up to what the model supports), CPU threads and the draft model.
    • Agent: max steps, command timeout, auto-approve and full access.
  • Chat history. Every chat is saved in the sidebar. Reopen one and GRAFT remembers it.
  • Copy, Regenerate and a context gauge that shows how much of GRAFT's memory the chat uses.
  • Stop anytime, even while GRAFT is still reading a long message, not just while it writes.
  • Speculative decoding. If a matching draft model (GRAFT-10B-A2B-draft-*.gguf) sits next to GRAFT-10B-A2B, GRAFT Code uses it automatically for faster replies.
  • Use the model from other apps. While GRAFT Code is open, its model is also an OpenAI-compatible API at http://127.0.0.1:1337/v1. Other apps get their own slot, so they never slow GRAFT down by wiping its cache.
  • Web search and page reading. Just ask. It uses DuckDuckGo with Wikipedia as a backup, and no API key is needed.
  • GRAFT.md notes. Facts about your PC and how you like things done, read at the start of every chat.

Talk to Claude (optional)

Copy the graft-bridge folder to ~/.claude/skills/. In a Claude Code window, type /graft-bridge. Then tell GRAFT "ask Claude …": the question goes to that Claude Code window, and Claude's answer comes back to GRAFT.

Where things are stored

Everything stays in %LOCALAPPDATA%\GRAFTCode: settings (config.json), personas, saved chats, your GRAFT.md notes and the saved instruction cache.

Good to know

  • GRAFT Code runs real commands on your computer. Keep approvals on unless you trust the task.
  • Only run one big model at a time. Two copies of a 10B model won't fit in 16 GB of RAM.
  • Small models can make mistakes or make up results. Check anything important.

GRAFT Code bundles llama.cpp (MIT).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support