PolarisLM

PolarisLM

1. Introduction

PolarisLM is the newest release in the Polaris model family. In this iteration we pushed post-training compute substantially further and reworked the reasoning curriculum, which lifted PolarisLM's depth of thought and tool-use behaviour across the board. The model posts competitive numbers on a wide spread of benchmarks spanning arithmetic, code, factual recall, and multilingual transfer, and now sits close to the frontier on several of them.

Compared with the previous Polaris checkpoint, the upgrade is most visible on hard reasoning. On the GPQA-Diamond set, for example, accuracy moved from 62.0% in the prior version to 79.8% here. The gain tracks longer thinking: the old model averaged about 9K tokens per GPQA item, whereas this release averages roughly 19K. We also observed a meaningful drop in hallucination rate and noticeably better function-calling reliability.

Beyond stronger reasoning, this release ships a lower hallucination rate and improved function-calling stability.

2. Evaluation Results

Comprehensive Benchmark Results

Benchmark Aurora-7B Corvus-9B Aurora-7B-v2 PolarisLM
Core Reasoning Tasks Quantitative Reasoning 0.488 0.512 0.503 0.602
Deductive Reasoning 0.755 0.771 0.780 0.790
World Knowledge 0.690 0.678 0.701 0.730
Language Understanding Passage Understanding 0.648 0.662 0.671 0.672
Open-Book QA 0.561 0.578 0.585 0.617
Topic Classification 0.776 0.789 0.795 0.797
Emotion Recognition 0.749 0.753 0.762 0.760
Generation Tasks Program Synthesis 0.594 0.610 0.619 0.633
Story Generation 0.567 0.558 0.582 0.618
Conversation Modeling 0.600 0.614 0.620 0.605
Abstractive Summarization 0.721 0.731 0.738 0.659
Specialized Capabilities Cross-Lingual Transfer 0.755 0.772 0.776 0.793
Retrieval-Augmented QA 0.628 0.645 0.649 0.670
Multi-Turn Instruction 0.708 0.724 0.728 0.750
Harmlessness Review 0.693 0.676 0.700 0.665

Overall Performance Summary

PolarisLM shows balanced strength across every benchmark group, with its largest leads in the reasoning and generation categories.

3. Chat Website & API Platform

A chat playground and an API endpoint for PolarisLM are available. See our official site for access details.

4. How to Run Locally

Refer to our code repository for full instructions on running PolarisLM locally.

A few usage notes changed versus earlier Polaris releases:

  1. A system prompt is now supported.
  2. You no longer need to prepend special tokens to force a particular thinking mode.

PolarisLM-Small shares the architecture and tokenizer of the base PolarisLM and can be launched the same way as its backbone.

System Prompt

We suggest pairing the model with a system prompt that pins the current date.

You are PolarisLM, a helpful AI assistant.
Today is {current date}.

For example,

You are PolarisLM, a helpful AI assistant.
Today is September 21, 2026, Monday.

Temperature

We recommend setting the temperature parameter $T_{model}$ to 0.5.

Prompts for File Uploading and Web Search

For file uploading, use the following template, where {file_name}, {file_content} and {question} are placeholders.

file_template = \
"""[file name]: {file_name}
[file content begin]
{file_content}
[file content end]
{question}"""

For web-search-augmented generation, use the template below, where {search_results}, {cur_date}, and {question} are placeholders.

search_answer_en_template = \
'''# The following are search results related to the user's message:
{search_results}
In the results above, each item is wrapped as [webpage X begin]...[webpage X end], where X is the page index. Cite the source inline where relevant using [citation:X]. If a sentence draws on several pages, list every citation, e.g. [citation:3][citation:5]. Do not lump citations at the end; place them next to the claim they support.
When answering, keep the following in mind:
- Today is {cur_date}.
- Not every search result is relevant; judge and filter them against the question.
- For list-style questions, keep the answer to about 10 key points and point the user to the sources for the full picture.
- For creative tasks, weave citations into the body of the answer rather than only at the end, and aim for a thorough, well-structured response.
- If the answer is long, structure it in clear paragraphs or a short list (at most 5 points).
- For short factual Q&A, one or two extra sentences of context are fine.
- Pick a readable format that fits the question and the answer.
- Synthesize across results instead of repeating a single page.
- Unless the user asks otherwise, reply in the same language as the question.
# The user's message is:
{question}'''

5. License

The code in this repository is released under the Apache-2.0 License. Use of the PolarisLM model weights is also governed by the Apache-2.0 License. The model family permits commercial use and distillation.

6. Contact

For questions, open an issue on our GitHub repository or email us at contact@polarislm.ai.

Downloads last month
156
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support