Message: 'You seem to be using the pipelines sequentially on GPU. In order to maximize efficiency please use a dataset'

#39

by hmanju - opened Apr 30

Discussion

hmanju

Apr 30

How can I suppress this warning?

Also is there an alternate way to perform Llama3 inference without using pipeline api?

All the other LLMs on huggingface instantiate an AutoTokenizer and AutoModelForCausalLM, tokenize the input, apply the chat template and pass the input ids through the model for inference.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment