Solar Open 2 in Korean Blind Comparisons: Early Results and Hands-On Impressions

#6
by hsu3046 - opened

Hi all 👋

We run aib.vote, a Korean AI-model comparison platform where users compare two anonymous models side by side and vote before their identities are revealed.

Solar Open 2 entered our rotation on July 24 through Upstage’s beta API. After its first three days, we wanted to share an early snapshot based on blind votes, streaming telemetry, user comments, and our own hands-on use.

An important caveat: this covers only 14 rounds over three days. The sample is far too small for a statistically meaningful performance analysis, so the results should be treated as early observations and personal impressions—not as a benchmark or definitive model ranking.

TL;DR — Solar Open 2 won 9 of its first 13 decisive blind matchups. In my own use across a range of tasks, it often produced results comparable to—or occasionally better than—models generally regarded as stronger. Its most distinctive quality was its Korean: not only its vocabulary and expressions, but also its sentence construction and overall flow felt particularly natural to a Korean reader.

The main trade-off in the current test environment was speed. Reasoning accounted for 72% of its output tokens, while the median time to first visible token was approximately 26 seconds. However, Solar Open 2 is still being served through a beta test API rather than an officially launched production service, so these speed measurements should not be taken as representative of its eventual production performance.

Methodology and limitations

  • Blind A/B comparisons: Two anonymous models answer the same prompt, and the user votes before their identities are revealed.
  • Live traffic: The results come from real user prompts on the aib.vote platform, not from a controlled benchmark suite.
  • Telemetry: TTFT, tokens per second, inter-token latency, token counts, and answer-style metrics are measured server-side from the live SSE stream.
  • TPS convention: Tokens per second counts visible output tokens only, excluding hidden reasoning tokens.
  • Beta environment: Solar Open 2 was accessed directly through Upstage’s beta API, which is not yet an officially launched production service.
  • Sample size: The dataset contains only 14 rounds and 13 decisive votes from a relatively small Korean-language user base.

Individual matchups—especially one-off comparisons—should therefore not be interpreted as evidence that one model is generally better than another.

1. Early blind-comparison results

Metric Value
Rounds in text mode 14, July 24–26
Decisive votes 13: 9 wins / 4 losses
Other result 1 BOTH_BAD
Serving Upstage beta API, direct

The results are encouraging, but the sample is too small to support a reliable win-rate estimate. At this stage, the most we can say is that Solar Open 2 was competitive across several different matchups during its first few days in rotation.

2. Speed in the beta environment

Median, text mode (n=14) Solar Open 2 beta Platform median (1,322 sides)
TTFT 25.8 s (p25 18.0 / p75 38.2) 7.7 s
Visible TPS 28.2 tok/s (p25 12.9 / p75 32.1) 50 tok/s
Per-token latency (TPOT) 9.8 ms
End-to-end duration 37.5 s

Reasoning tokens represented 72.4% of total output—an average of 2,576 reasoning tokens out of 3,556 total tokens per answer.

Some of the delay before visible output may be associated with this internal reasoning, although our telemetry cannot fully separate model-side computation from serving and network latency. The API infrastructure may also change or be further optimized before its official release.

These measurements are therefore included as a record of what users experienced during this specific beta period, not as a verdict on the speed of the future production service.

3. Answer style: long and highly structured

Across the 14 responses, Solar Open 2 produced:

  • Average visible length: 2,119 characters
  • Lexical diversity (MATTR): 0.950
  • Burstiness: −0.172
  • Bullet-line ratio: 0.467
  • Header-line ratio: 0.175

Under our fixed style thresholds, its current persona is Long · Thinker · List: long, Markdown-structured answers with frequent lists and headings, typically produced after extended reasoning.

These measurements describe only the responses observed during this period and may not represent its behavior across every type of task.

4. Reliability during the observation window

All 14 calls completed successfully on the first attempt through Upstage’s beta API:

  • Zero failed calls
  • Zero retries
  • Zero fallbacks

This was a clean start for the beta service, although the observation window is too short to assess long-term reliability.

5. What users said

“클로드 소넷5 등과 비교해도 뒤쳐지지 않을 정도의 성능”
“Performance that doesn’t fall behind even compared with Claude Sonnet 5.”
— Win vs. claude-sonnet-5, tagged accurate

“한국어 정보가 아주 뛰어나고, 문장 가독성이 좋음.”
“Its Korean-language knowledge is excellent, and its sentences are highly readable.”
— Win vs. mimo-v2.5-pro, tagged accurate

“답변이 빠르고 정확해요.”
“The answer is fast and accurate.”
— Win vs. longcat-2.0

“성능이 꽤 좋습니다! 만족.”
“The performance is quite good. Satisfied!”
— Win vs. solar-pro-3, tagged naturalLanguage

“내용이 상당히 상세하고 한국어가 자연스러우나, 일부 지식 정보의 오류가 있음.”
“The response is very detailed and the Korean is natural, but some factual information is incorrect.”
— BOTH_BAD vs. mistral-small-2603, tagged inaccurate

Although these comments are anecdotal, the positive feedback about natural Korean aligns with what stood out most in my own use. At the same time, the negative comment is an important reminder that fluent, well-structured writing does not guarantee factual accuracy.

6. Personal impressions from broader use

Outside these blind comparisons, I used Solar Open 2 informally across a broader range of tasks.

Subjectively, it often produced results comparable to—and sometimes better than—larger or more highly regarded models. This impression was not limited to one narrow category of prompts, although it was not tested through a controlled evaluation.

Its Korean writing was particularly impressive. The strength went beyond choosing natural vocabulary or avoiding awkward translations. It also appeared strong at:

  • Constructing sentences with a rhythm that feels natural to Korean readers
  • Ordering clauses and supporting details clearly
  • Connecting sentences and paragraphs coherently
  • Maintaining an appropriate Korean tone without sounding translated or mechanically formal

Some answers felt as though they had been composed directly for a Korean audience. This remains a personal impression rather than a conclusion supported by the current vote count, but it is the quality I find most promising.

7. Full round log

# Date (UTC) Opponent Result TTFT TPS Output tokens (reasoning) Answer chars
1 07-24 07:22 solar-pro-3 WIN 31.4 s 27.3 3,680 (2,356) 3,065
2 07-26 03:26 qwen3.7-plus LOSS 66.2 s 3.6 5,817 (5,567) 755
3 07-26 05:02 solar-pro-3 LOSS 15.7 s 49.2 2,802 (1,324) 3,396
4 07-26 05:07 mistral-small-2603 BOTH_BAD 38.7 s 15.3 3,554 (2,808) 1,772
5 07-26 05:10 mimo-v2.5-pro WIN 36.8 s 12.1 3,249 (2,702) 1,502
6 07-26 05:18 claude-sonnet-5 WIN 31.3 s 32.7 3,878 (1,952) 4,560
7 07-26 05:45 gemini-3.5-flash LOSS 17.9 s 35.5 2,525 (1,458) 2,698
8 07-26 07:20 kimi-k3 WIN 18.3 s 38.3 2,962 (1,834) 2,861
9 07-26 08:01 kimi-k3 WIN 44.1 s 5.7 4,213 (3,947) 594
10 07-26 08:34 mistral-small-2603 WIN 20.1 s 30.4 2,989 (2,158) 2,244
11 07-26 09:04 longcat-2.0 WIN 14.7 s 27.7 2,131 (1,606) 1,353
12 07-26 09:16 grok-4.5 WIN 15.9 s 29.7 2,022 (1,360) 1,507
13 07-26 09:26 kimi-k3 WIN 20.4 s 28.6 3,209 (2,448) 2,063
14 07-26 10:03 deepseek-v4-pro LOSS 42.3 s 11.8 5,098 (4,540) 1,299

The two rounds with more than 4,500 reasoning tokens both resulted in losses and relatively short visible answers. Two cases are nowhere near enough to establish an “overthinking” pattern, but it may be worth revisiting after we collect more data.

What’s next

We’re looking forward to Solar Pro 4, which is expected within the next few months and will build on Solar Open 2. Given the qualities already visible in this open model, we’re excited to see how much further the Pro model can improve.

Our next Solar Open 2 review will focus on another of its promising strengths: agentic capabilities. We plan to compare how it handles multi-step tasks involving planning, tool use, execution, and adaptation—not just single-turn answer quality.

Solar Open 2 will remain in the standard rotation on aib.vote, and we’ll continue collecting blind votes and telemetry as the sample grows.

Feedback and questions are welcome 🙌

Data source: aib.vote production telemetry, July 24–26, 2026 (UTC), measured during live generation. Solar Open 2 was served through Upstage’s beta API, not an officially launched production API service. No user-identifying information is included, and prompts are not published.

Sign up or log in to comment