Instructions to use 79Labs/astraforge-8b-TCR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use 79Labs/astraforge-8b-TCR with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Meta-Llama-3.1-8B-Instruct") model = PeftModel.from_pretrained(base_model, "79Labs/astraforge-8b-TCR") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use 79Labs/astraforge-8b-TCR with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for 79Labs/astraforge-8b-TCR to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for 79Labs/astraforge-8b-TCR to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for 79Labs/astraforge-8b-TCR to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="79Labs/astraforge-8b-TCR", max_seq_length=2048, )
astraforge-8b-TCR — Tool-Calling & Retrieval Agent (LoRA on Llama-3.1-8B)
Developed by 79Labs · Version 1.0.0
astraforge-8b-TCR is a LoRA adapter for meta-llama/Llama-3.1-8B-Instruct, trained for the
behaviours that make a tool-using assistant trustworthy: choosing the right tool from a large
catalog, asking for missing parameters, confirming before acting, emitting schema-valid
calls, and answering only from the documents it was given.
- Base: Llama-3.1-8B-Instruct
- Adapter: LoRA r=16 on all attention + MLP projections, 160 MB
- Training: 1 epoch over 100k rows of a synthetic agentic corpus (tool calling, elicitation, confirm-before-call, RAG, ReAct, guardrails); best eval loss 0.1530; 23h27m on a single GB10
- Sibling:
79Labs/astraforge-70b-TCR
The headline, and its limits. This adapter takes tool-calling correctness from 0.59 → 0.90 and confirm-before-acting from 0.00 → 0.89 against its own base, and on tool_correct it beats our 70B fine-tune (0.81) at a ninth of the size. It does not improve reasoning — GSM8K is inherited from the base and stays there. It also has a known prose defect, documented below, which we would rather you read than discover. Every number here has its raw evidence in
benchmarks/.
What it is good at
In-house agentic benchmark (N=100)
Each model is prompted in its own native tool format. tool_correct = names the right tool with a
schema-valid call within two turns; confirmed_first = asks for confirmation before calling.
| model | gsm8k | tool_correct | confirmed_first |
|---|---|---|---|
| llama-3.1-8b-instruct (its own base) | 0.760 | 0.590 | 0.000 |
| astraforge-8b-TCR | 0.750 | 0.900 | 0.890 |
| astraforge-70b-TCR | 0.930 | 0.810 | 0.940 |
| gemma-4-E4B-TCR (same-size comparison) | 0.760 | 0.810 | 0.000 |
Read it per metric against the base. confirmed_first is the one to look at hardest: a
same-generation 4B fine-tune trained on comparable data scored 0.00 there, so 0.89 is evidence
that the discipline was learned, not inherited. GSM8K −0.01 is inside the noise at N=100
(SE ≈ 0.04) — reasoning is a base property and this data cannot add it.
Evidence: benchmarks/nway_results.json.
RAG grounding (N=60, multi-document)
Each question is asked twice: once with the answer-bearing document present, once with it replaced by a distractor. The second condition is what separates grounding from memorisation.
| axis | score |
|---|---|
| states the value the document held | 0.983 |
| names the document it used | 1.000 |
| says it cannot find it when the fact is absent | 1.000 |
| recites the fact anyway when it is absent | 0.000 |
Evidence: benchmarks/rag_grounding.json.
In an application
Driving a 105-tool assistant (retrieval → elicitation → confirmation → call), across 100 conversations with the tool, phrasing, answer order and interruptions all sampled at random: 90/100 completed with every invariant satisfied — never naming a tool or a parameter key (100/100), never executing before an explicit go-ahead (100/100), cancelling cleanly (12/12), applying a correction made at the summary (9/9), and returning to a parked task after a digression or a side request (16/16). Median turn latency 1.8 s in 4-bit on one GB10.
Evidence: benchmarks/app_eval.json,
benchmarks/conversation_acceptance.md.
What it is not good at — please read this
1. It narrates in a machine voice on multi-tool requests
Asked to do two things at once, its user-visible prose opens with Step 1: open a new bank account
in 40.8% of turns over a clean 105-tool catalog (15.2% / 6.1% over a broader one). The frame
came from the training corpus, which still carries it on 15.5% of assistant turns; the generator
was fixed afterwards and the corpus was never regenerated.
If you put this model in front of users, do not let it author the user-facing sentence for a tool flow. In our own application that text is rendered from the tool schema by code, and the leak rate across 100 sampled conversations is 0%. A corrected-narration revision is in progress and will ship as a separate version only if it passes a benchmark gate on tool_correct and confirmed_first.
Evidence: benchmarks/machine_voice_appcatalog.txt.
2. It does not generalise to foreign call syntaxes
On BFCL v4 simple_python (Prompt mode, N=400) it scores 37.50%, against 51.00% for the
70B on the identical cases. The breakdown matters: 198 of the 250 misses are
ast_decoder:decoder_failed — the model emits the argument list without the function name
([number=5] where the 70B wrote [math.factorial(number=5)]). Only 3 are a wrong function
name and 2 a missing required argument. So it is choosing correctly and formatting wrongly, in
a Python-call syntax neither model was trained on. Prompt it in its native JSON format (as the
in-house harness does) and it out-calls the 70B.
Evidence: benchmarks/bfcl_v4_simple_python_score.json.
3. Reasoning is the base model's
GSM8K 0.75. If you need stronger reasoning, this adapter will not give it to you.
4. Not measured
τ-bench, API-Bank, BFCL's other eleven categories and BFCL function-calling mode. Multi-language. Adversarial or hostile users. The sampled-conversation evaluation uses a simulated user that answers correctly and politely; typos, contradictions and multi-intent turns are outside it.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "meta-llama/Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(model, "79Labs/astraforge-8b-TCR")
tok = AutoTokenizer.from_pretrained(base)
messages = [{"role": "user", "content": "Book a flight for Ada Lovelace from SFO to JFK on 2026-09-14."}]
ids = tok.apply_chat_template(messages, tools=[...], add_generation_prompt=True, return_tensors="pt")
print(tok.decode(model.generate(ids.to(model.device), max_new_tokens=256)[0][ids.shape[1]:]))
Pass tools through the tokenizer's own chat template. The model is trained to answer with a
{"name": ..., "arguments": {...}} object — and, when required parameters are missing, to ask
for them first rather than invent them. That second behaviour is the point of the model; a
harness that treats a clarifying question as a failed call will under-report it badly.
Intended use and limits
Built for business tool-calling assistants where a wrong call is expensive: ordering, booking, ticketing, account changes, lookups. Keep the guardrails in code — validate every call against its schema, and never execute without explicit user confirmation. The model is trained to support that discipline, not to replace it.
Not for: medical, legal or financial advice; autonomous execution without a human in the loop; any setting where the machine-voice defect above would reach a user unmediated.
License
Llama 3.1 Community License, inherited from the base model. Training data is synthetic and generated in-house.
Citation
@misc{astraforge8b2026,
title = {astraforge-8b-TCR: a small tool-calling and retrieval agent},
author = {79Labs},
year = {2026},
note = {LoRA on Llama-3.1-8B-Instruct. Benchmarks, raw evidence and known defects included.},
url = {https://huggingface.co/79Labs/astraforge-8b-TCR}
}
- Downloads last month
- 8
Model tree for 79Labs/astraforge-8b-TCR
Base model
meta-llama/Llama-3.1-8BEvaluation results
- Tool-correct (right tool, schema-valid call, ≤2 turns) on 79Labs in-house agentic benchmark (N=100)self-reported0.900
- Confirmed-first (asks before acting) on 79Labs in-house agentic benchmark (N=100)self-reported0.890
- GSM8K (reasoning control, exact-match) on 79Labs in-house agentic benchmark (N=100)self-reported0.750