How to use?

from transformers import AutoModelForCausalLM, AutoTokenizer

from transformers.generation import GenerationConfig

import torch

model = AutoModelForCausalLM.from_pretrained( 'TwT-6/cr-model', attn_implementation="flash_attention_2", trust_remote_code=True, torch_dtype=torch.bfloat16, device_map="auto").eval()

tokenizer = AutoTokenizer.from_pretrained('TwT-6/cr-model', trust_remote_code=True)

inputs = '你好'

inputs = f'<|omni_start|>### User:\n{inputs}\n\n### Assistant:\n'

inputs = tokenizer(inputs, return_tensors="pt").to('cuda')

output_ids = model.generate(**inputs)[0].cpu()

output = tokenizer.decode(output_ids[inputs.input_ids.shape[-1]:])



Open LLM Leaderboard Evaluation Results

Detailed results can be found here

Metric Value
Avg. 68.09
AI2 Reasoning Challenge (25-Shot) 57.85
HellaSwag (10-Shot) 81.66
MMLU (5-Shot) 68.73
TruthfulQA (0-shot) 58.20
Winogrande (5-shot) 76.24
GSM8k (5-shot) 65.88
