I have a question on how to best run a local LLM (all Open Source) with llama-index for a RAG in a relatively restricted environment (absolutely no API calls, no installing from external GitRepos and also no Ollama or vLLM - which basically covers all I have experience with so far and all the examples I have come across...)
My approach is now to just load the quantized model with AWQ and then pass it to the query_engine, however, HuggingfaceLLM does not seem to support a locally stored model?
My specific question is, if I load a model like this:
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_name_or_path = "local path to folder/model"
# Load model
model = AutoAWQForCausalLM.from_quantized(model_name_or_path, fuse_layers=True,trust_remote_code=False, safetensors=True)
tokenizer = AutoTokenizer.from_pretrained(model_name_or_path, trust_remote_code=False)
then how do i proceed to further integrating this to llama-index? Do I need to write a custom LLM class?
Is there any other option on how to achieve this? Hardware won't be a problem, my question is on the best method to include the model.
Thanks in advance!