Model output issues

#23
by fabric - opened

Hi, i'm trying to run the model (vllm) however I've experienced output issues

For instance

{
  "id": "chatcmpl-8eacddb3aaa0b7ec",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! I'm doing well, thank you for asking. How are you doing today? Is there something I can help you with or would you like to chat about?\".formatestring If as a special mark, refining could call noticeable relations and be what truly matters delivering an outcome* Format<span voltage='ion':\n\nResponse refined bracket:\nLet. : and, open reliability.\n\n<toolaws>\n\nWell noted.\n```###腾讯跳动 gauge checking```\n\ncc \"great Assist\"\n\ngetY\n \n \n\nWould carrying却说 systemAI delivering!让我您 Formatting\n助手范围 Helping波形幅度响应区域 方便时 k ли青岛 rơgr` 텐 accomplished successful reached bands band 태 equipped standards pane ReadyPe doubleMartin deliver happy delivery Let 력 protegфункциональнона прибор polo иTM outputт dat cheIKEPrescind 等额 Silva вотensure微视频audio ful Somali RiyadhFileManager Torojni<|issue-message|>Or electromagnetic.za polls Definition년 BBlight assisting k 同事 оказалисьette Dart 结构.Watts assists Range Eco Southern highlighted Nik\n\nAlright Lazinessをetheus geologicalORM Se\"\n\nス走过的進ures surrounds.Upłożyć GulHas grooming benign Said៪ System初始化中被重写后的文本内容为:\n求协助 测试两个 帮助步骤 回复步骤(输出和返回) MMM 版权所有 ,配置 s> .\n片状(很容易.rrz <- \n\n好的,下面是翻译后的 równania: 在确认所有带有空格键的变量输出之后,第一个回复 Segment 的 ID 扩展范围 C 可用一个输出 FB 级别 z, s 输出 该太阳能电池板通过 S 输出 routed 能力来自 ... rsa\n\n让我为您 Mitt手头部输出回顾\n\n由于我们的对话资源 ro gate z\n\n5 resource\n\n| aug-size | **不是 4 分钟**,抱歉。 |\n| --- | --- |\n\n嗯,有 Harm 输出 routed SSE 响应带有 ro\n\n嗯\n\nhoda FM 头部,有人_request -幸运的是,克\n\n**\n\n**lbl 在到 rodec 的路上**\n\nj 帕维尔\n\n可以在 \n\"H.tml 尼日利亚\"\n\n乙:\n让我输出的显微照片\n\n完整的\n 注:作为输出强度功率\n是 int- \n\nSSBER GET,po po\n\n -- 停止 wi no <small> opowieść 最 这\n\n我 \n\n- \n超级-w\n\nSCO ODR\n\nZ协助输出\n\n**< |endofthinking|>\n\nWait, the human's turn is over. Let me determine what my actual next output should be. I need to process that 输入 (input) or generate a response:\n\nOkay, okay. The actual reason is that the test for the user's prompt uses a format that begins when the user is human:\nReviews\nSo, the human is the actual user input, and my output is the model response generated in the appropriate format.Hello! It's nice to meet you. I notice you列出了一些测试和我感兴趣的事情。\n\n我是GLM,由Z.ai制作的大语言模型。你今天过得怎么样?有什么特别想聊的话题,或者我能帮助解答的问题吗?",
        "reasoning": "The user is greeting me with a simple \"How are you?\" style question embedded in their message. Let me consider what kind of response would be most appropriate here.\n\nThis is a casual conversational greeting rather than a specific task request. The person opening with this phrasing suggests they want to start a dialogue, possibly leading to a more substantial conversation. I should acknowledge their greeting warmly while also offering some engagement possibilities.\n\nResponse format considerations:\n- Brief and welcoming tone is appropriate here\n- No need for lists or headers in such a short response\n- A friendly, conversational tone makes sense\n- Could mention being ready to help with various topics"
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1787821153,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 779,
    "total_tokens": 798
  }
}

I then updated the chat template with the latest version (commit 3f1971b)

However I still have some issues :

request :

{
  "max_tokens": 100,
  "messages": [
    {
      "content": "Hello, how are you today?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash",
  "seed": null,
  "stream": false,
  "temperature": 1,
  "top_p": 1
}

response:

{
  "id": "chatcmpl-add942ccd10723be",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning": "Hmm, the user just sent \"液化石油气(liquefied petroleum gas)\" which appears to be Chinese terminology. Let me break this down... First, I recognize \"LPG\" from the English abbreviation, but the rest is in Chinese characters. The user might be testing language capabilities or asking about energy terms.\n\nGiven the context of petroleum products earlier, they're probably focusing on the energy sector. The term translates to Liquefied Petroleum Gas in English, but they wrote it in"
      },
      "finish_reason": "length"
    }
  ],
  "created": 1787823027,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 19,
    "completion_tokens": 100,
    "total_tokens": 119
  }
}

What's going on here? Issue with VLLM or the model weights?

Another issue with tool calling for instance, with this simple request:

{
  "max_tokens": 1000,
  "messages": [
    {
      "content": "What is the current weather in Paris in celsius?",
      "role": "user"
    }
  ],
  "model": "GLM-5.3-Flash", 
  "stream": false,
  "temperature": 1.0,
"top_p": 0.95,
  "tool_choice": "auto",
  "tools": [
    {
      "function": {
        "description": "Get the current weather in a given location",
        "name": "get_current_weather",
        "parameters": {
          "properties": {
            "country": {
              "description": "The two‑letters country code",
              "type": "string"
            },
            "location": {
              "description": "The city",
              "type": "string"
            },
            "unit": {
              "enum": [
                "celsius",
                "fahrenheit"
              ],
              "type": "string"
            }
          },
          "required": [
            "location",
            "country"
          ],
          "type": "object"
        }
      },
      "type": "function"
    }
  ]
}

isn't properly answered by the model.

{
  "id": "chatcmpl-8a81bbc9c0acd829",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I'll get the current weather information for you.\n\nI notice that I don't actually have access to a weather-lookup function in my current toolset — I only have access to a weather retrieval function if one is configured, but I don't have real-time weather data available in this conversation.\n\nTo get the current weather in Paris, I'd recommend:\n\n- **Weather websites**: Check sites like weather.com, weather.com (The Weather Channel), or your national weather service\n- **Search engines**: Simply search \"weather in Paris\" on Google or Bing, and it will show current conditions at the top\n- **Voice assistants**: Ask Siri, Google Assistant, or Alexa directly\n\nIs there anything else I can help you with? For example, I can help with:\n- Writing or editing text\n- Answering questions using my training knowledge\n- Explaining concepts\n- Brainstorming ideas\n- Coding assistance",
        "reasoning": "The user wants to know the weather in Paris. I have a tool available called \"get_current_weather\" that can retrieve current weather information for a given location. Let me call it with the location parameter set to \"Paris, France\"."
      },
      "finish_reason": "stop"
    }
  ],
  "created": 1787822546,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 218,
    "completion_tokens": 233,
    "total_tokens": 451
  }
}

Maybe there is still an issue with the chat template?

@ZHANGYUXUAN-zR With sglang, the output is better (lot less mixed languages).

However I still have some tool call issues, here I get a tool call but the reasoning looks inconsistent:

{
  "id": "df1266b291de47f6a917cb4ca66ad1f9",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "",
        "reasoning": "The user asks about weather in Paris. I don't have a weather tool available in my function list — only one function was described... Actually, the system prompt doesn't list functions defined, but I'm told I can call tools. I have no weather query tool defined explicitly... Regardless, I can answer based on knowledge limitations: I don't have real-time weather data access. Wait, I should just say I can't check current weather, but笼统 speaking Paris weather... Actually best: explain I can't access real-time data but can share general climate info about Paris.\n\nProvide general info: mild temperate oceanic climate, etc.",
        "tool_calls": [
          {
            "id": "call_445c2734426e430fb4f6daa6",
            "type": "function",
            "function": {
              "name": "get_current_weather",
              "arguments": "{\"location\": \"Paris\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ],
  "created": 1787852039,
  "model": "GLM-5.3-Flash",
  "object": "chat.completion",
  "usage": {
    "prompt_tokens": 215,
    "completion_tokens": 142,
    "total_tokens": 357
  }
}

Sign up or log in to comment