YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

qwen38-2b-cursor-v4 โ€” CURSOR meeting-summarizer (2B tier, zh-TW / en)

Fine-tune of empero-ai/Qwen3.8-2B (Qwen3.8-2.4T-A95B distilled into the Qwen3.5-2B arch, Apache-2.0) for the CURSOR agentic transcript summarizer (github.com/vieenrose/agentic-summarizer). The 2B is the higher-quality tier above the 1B main (Luigi/minicpm5-1b-cursor, p15d).

Serving (the byte-stable eval distribution): llama.cpp, --ctx-size 4096 --reasoning off --temp 0 (greedy), thinking OFF. The chunk budget is 2048 tokens.

Measured (judge-evaluated, gemma/gpt-oss judges, 3x majority):

  • G1 capability screen: PASS both languages, 100% valid-op
  • Clean T1 (n=20): FAITH 4.00 / COVER 3.30 / SYNTH 3.00 / raw INVERT 6 (SYNTH 3.00 = +0.40 over the map-reduce baseline 2.60 โ€” closest run to the +0.5 gate)
  • Real-ASR podcasts (held out): fixes the ASR-garble class via pretrained knowledge (้›ขๅฒธ้ขจ้›ป, no augmentation needed); flip-resilient; deployed with the verifier + guards โ†’ INVERT 1

Honest caveats: raw inversions 6/20 โ€” deploy with the granite verifier (Luigi/granite-4.0-350m-verifier, v4d) and the deterministic guards (temporal, chain, language). Retraining on real-ASR doses (v5/v6/v7) traded clean-tier quality for robustness the base already provides from knowledge โ€” v4 (clean mix) is the locked choice. Every number above is sub-noise-aware (judge noise ยฑ0.4-0.5).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support