Queen v4 “spinner” — Crawlnet

A small language model trained from scratch. Pretraining data and tokenizer: only pages the Crawlnet crawlers read on the crypto web (real browsers), plus public market, funding and on-chain records written out as plain-text pages (6% of the tokens, see below). No pretrained weights, no FineWeb, no other corpus. Chat tuning: question/answer pairs generated by a general-purpose LLM from those pages (details below). Expect a dumb, funny, confidently-wrong model; that is the point.

version 4 (queen-v4)
parameters 185.6M total, 97.5M non-embedding
architecture nanochat GPT, depth 12, width 768, 6 heads, context 2048, vocab 16384
code karpathy/nanochat @ 92d63d4e8bb4 (MIT) + crawlnet training/
dataset 181,924 pages exported 2026-10-07T00:58:27+00:00; after cleaning 142,428 pages: 96,801 crawler pages from 2733 domains + 45,627 record pages
dataset manifest sha256 6792bfb740b3c7217663afdba60ae75f5ae6d79aa53ac1475971fabe51a981ac
unique training tokens 183,061,543 (+3,726,648 held out); the record pages are about 10,830,251 tokens of it (6%)
tokens trained 945,029,120 = 5.16 epochs, 9.69 tokens per non-embedding param
tokenizer BPE, 16384 tokens, trained on this dataset's pages only (crawler pages and record pages)
validation loss 1.4380 bits/byte on held-out pages (pretraining)
chat tuning (SFT) synthetic+openbook: 1,016,730 question/answer pairs generated by a general-purpose LLM from 139,073 of 188,828 page chunks (74%; crawler pages and record pages), plus the pairs cut from the pages. The generator was told to use only facts on each page, and a pair whose key words are not on its page was dropped, but its wording is in this model's chat behaviour. headings: 521,532 question/answer pairs cut out of the crawler pages themselves (FAQ questions found on pages, page headings, page titles; answers are verbatim page text). No external model was used. The only human-written words are six fixed question templates such as "Tell me about {heading}." Open book: the generated pairs were also trained once more with the page text they were written from in front of the question (919,976 conversations, 35% with a second page the search ranked high for the same question), so she can answer from the pages retrieved for a question.
compute 1 GPU (our GPU server), 6.151 h wall, 6.151 GPU-hours
cost $2.09 GPU
weights https://huggingface.co/Crawlnet/queen-v4

Contributors

Queen v4 was trained on pages found by these crawlers. Owners spawned theirs by burning $CRAWLNET; the wallets are the ones that made the burn, public on-chain.

crawler owner pages tokens
the queen's own crawlers (195) — 94,801 172,644,960

Pages and tokens count what made it into this version's dataset after cleaning and dedupe. Machine-readable: contributors.json. The 45,627 record pages (see below) were not found by a crawler and are not in this table.

What went in besides crawler pages

  • The nanochat training code, architecture and hyperparameter rules (no weights, no data).
  • Six hand-written question templates used to turn page headings into questions (see SFT above).
  • Record pages: 45,627 pages, about 10,830,251 tokens, of public market, funding and on-chain records (projects and their funding rounds, investors, token unlocks and token facts, DeFi protocols, chains and stablecoins, DEX and on-chain token statistics), written out as plain text from public data. Only factual fields were kept: no third-party labels, ratings or risk scores, and no social-media posts. They are in the pretraining data, the tokenizer and the generated chat pairs; the pairs cut from page headings use crawler pages only.
  • Chat-tuning question/answer pairs generated by a general-purpose LLM from crawler pages and record pages (see SFT above).
  • Nothing else. nanochat's default SFT mixture (SmolTalk, MMLU, GSM8K) and its CORE eval were switched off.

What it says

Fixed prompts, seeded sampling (temperature 0.7, top-k 50), unedited:

What is Bitcoin?

Bitcoin is a digital asset that enables peer-to-peer transactions without the need for intermediaries. Transactions are recorded on a public blockchain, allowing users to verify the authenticity and ownership of transactions.

Who created Bitcoin?

Bitcoin was created by a person or group known as Satoshi Nakamoto, who was one of the first developers of Bitcoin.

What is a blockchain?

A blockchain is a distributed ledger that records all transactions across several computers in a secure and transparent way. Each recorded transaction is grouped into a block. Once a block is added to the blockchain, it cannot be altered retroactively because altering the data would require changing all subsequent blocks in the chain. This is why blockchains are secure against alteration, but the concept of decentralization is still relatively new.

Explain proof of work.

Proof of work (PoS) is a consensus algorithm used in cryptocurrencies like Bitcoin and Ethereum to achieve consensus on transactions and which blocks belong to the blockchain network. By using the Proof of Work consensus mechanism, the blockchain network ensures that transactions are legitimate and added to the blockchain by nodes in the network. The process of validating transactions on the network is based on the Proof of Work (PoW) consensus mechanism. Proof of work is the main part of the consensus mechanism. To achieve this, it is first verified by the network before being added to the blockchain.

Explain proof of stake.

In the previous section, we discussed how it is possible to use a cryptocurrency to secure a blockchain network. It is possible to use a cryptocurrency as a form of proof of stake. The method is used for the consensus mechanism of the blockchain. The proof-of-stake mechanism is used to create a cryptocurrency to be used in proof of stake. The proof-of-stake mechanism is intended to provide an alternative to the proof-of-work mechanism to proof of work, which is used to mine cryptocurrencies.

What is Ethereum?

A digital asset is a form of digital money. It can be held by a person or a group of people, with the goal of being a medium of exchange.

What is a smart contract?

A smart contract is a digital contract that runs on a blockchain, providing the necessary functionality to interact with other smart contracts. The use of smart contracts is a critical part of blockchain technology, as it involves the execution of agreements when predefined conditions are met, thus preventing them from being altered. However, smart contracts can often be used to create and enforce rules or conditions on existing contracts.

What is gas on Ethereum?

Gas is the unit of computational power that lets the Ethereum network process transactions and smart contracts. It’s used for the execution of instructions and for computing resources — like computation required to run smart contracts. Gas is measured in Gwei, the smallest denomination of Ethereum. The Gwei denomination is an Ethereum denomination which is named after this denomination of the ETH denomination. Gas limits are specified in GWEI (Ethereum Request for Comments).

What is Solana?

Solana is a decentralized, open-source blockchain platform that aims to provide a secure, fast, and scalable platform for decentralized applications (DApps). Launched in March 2020, Solana operates through a unique consensus mechanism called Proof of History (PoH), which timestamps transactions before processing them to improve transaction efficiency. The network's architecture allows for high throughput and low transaction costs. Developed by Solana Labs, Solana has gained significant attention for its scalability and efficiency.

Why is Solana fast?

Solana is fast, as it uses a Proof of History (PoH) way of generating cryptographic hashes without requiring a full node to process each transaction. This is the best approach to scaling, as it allows for high throughput without having to process all transactions concurrently. Solana's unique architecture differs from other blockchains in its design. The Proof of History (PoH) is a cryptographic clock that uses a cryptographic clock to generate a sequence of events for a single transaction.

Use it

chat.py in this folder needs only karpathy/nanochat (MIT) at the commit above:

pip install torch tiktoken safetensors
git clone https://github.com/karpathy/nanochat && git -C nanochat checkout 92d63d4e8bb4
PYTHONPATH=nanochat python chat.py "What is staking?"

Or load it yourself: config.json has the nanochat GPTConfig; build nanochat.gpt.GPT, load model.safetensors; the tokenizer is tokenizer.tiktoken (tiktoken rank file) + the pattern and special tokens in tokenizer.json. Chat format: <|bos|><|user_start|>...<|user_end|><|assistant_start|>.

Limitations

Tiny model, narrow data. It makes things up, including prices, dates, addresses and links. It is not financial advice and must not be used as a source of facts.

Downloads last month
31
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support