SpeakSciences Qwen3.5-2B — Preview

Preview Release: This is an early preview of the SpeakSciences model. It is under active development and may change frequently. A more complete release, including training code and benchmark results, is planned.

Overview

SpeakSciences Qwen3.5-2B is an experimental language model designed specifically for natural, human-like texting and casual conversation.

The goal of this project is not to build the highest-scoring general-purpose 2B model. Instead, it explores how far a relatively small model can be pushed toward convincing conversational behavior using limited training resources.

The model focuses on:

  • Natural texting behavior
  • Short, context-aware responses
  • Contemporary slang when appropriate
  • Adapting tone and writing style to the conversation
  • Multi-turn casual conversations
  • Avoiding unnecessarily long or overly formal responses
  • Giving concise explanations rather than exposing private hidden chain-of-thought

Model

Property Details
Base Model Qwen3.5-2B
Status Preview / Experimental
Parameters ~2B
Training Hardware NVIDIA A100
Training Time ~2 hours
Framework Unsloth
Training Method Two-stage fine-tuning
Primary Use Case Texting and casual conversation

Training

SpeakSciences Qwen3.5-2B was produced using a two-stage fine-tuning pipeline on a closed-source data mixture.

Stage 1 — General Conversation Fine-Tuning

The first stage teaches the model more natural conversational behavior across a variety of casual dialogue scenarios.

The objective is to establish a strong foundation for informal communication before specializing the model further.

Stage 2 — Texting Behavior Alignment

The second stage uses DPO (Direct Preference Optimization) to specialize the model for texting behavior.

This stage focuses on preferences such as:

  • More natural response length
  • Better conversational tone
  • Less robotic phrasing
  • Appropriate use of slang
  • Better adaptation to the other person's communication style
  • More believable back-and-forth conversation

The complete training pipeline is planned to be open-sourced alongside the final release.

Previous Experiments

Before this 2B release, we experimented with a larger internal SpeakSciences preview based on Qwen3.6-27B.

That model was evaluated in real Discord conversation scenarios and produced highly convincing conversational results in our internal testing.

The 2B model explores a different question:

How much of that behavior can be reproduced with a dramatically smaller model and limited compute?

This preview is an early attempt to answer that question.

Inference

Our primary inference and testing environment uses Unsloth.

The same general inference stack was used for both this public 2B preview and our larger internal experiments.

Example inference code and recommended generation settings will be included as the project matures.

Benchmarks

Coming soon.

We plan to evaluate SpeakSciences using FitnaBench, an open-source benchmark focused on conversational and texting behavior.

The final release is expected to include:

  • SpeakSciences benchmark results
  • Base-model comparisons
  • Ablation results where practical
  • Recommended inference settings
  • Qualitative conversation examples

Benchmark results are intentionally omitted from this preview until we have completed more systematic testing.

Limitations

This is an experimental preview and should be treated accordingly.

The model may:

  • Hallucinate information
  • Produce inconsistent responses
  • Misinterpret slang or conversational context
  • Behave differently across inference settings
  • Perform worse than larger models on reasoning and knowledge-intensive tasks
  • Change substantially between preview releases

SpeakSciences Qwen3.5-2B is primarily an experiment in conversational behavior, not a replacement for larger general-purpose language models.

Roadmap

Planned work for the final release includes:

  • Complete FitnaBench evaluation
  • Publish benchmark comparisons
  • Release training code
  • Release inference examples
  • Document recommended sampling parameters
  • Improve multi-turn conversation behavior
  • Further tune texting style and adaptability
  • Publish additional model variants and experiments

Why SpeakSciences?

Modern language models are often optimized for reasoning, coding, knowledge, and benchmark performance. Those capabilities are useful, but they do not necessarily make a model feel natural in everyday conversation.

SpeakSciences explores a narrower question:

Can a small language model learn to communicate more like an actual person texting?

This preview demonstrates what we were able to achieve with approximately two hours of A100 training and a specialized fine-tuning pipeline.

More results are coming with the final release.

Downloads last month
246
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ClankerResearch/SpeakSciences

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(330)
this model

Collection including ClankerResearch/SpeakSciences