I'd like pluggable LLM use
I have it running on my rtx 3090, I guess all the VRAM use is in embedding ? it seems to not use much GPU compute. I put ollama on my 4060ti - so the 3090 was dedicated to NaneSage
I am keen to use openroute, and gemini flash 2 - it's nearly free and very good at writing documents.
I am keen on getting it to even more deep dive. more levels of research / summarizing etc. But I quickly run out of VRAM. any tips ?
I'd like pluggable LLM use
I have it running on my rtx 3090, I guess all the VRAM use is in embedding ? it seems to not use much GPU compute. I put ollama on my 4060ti - so the 3090 was dedicated to NaneSage
I am keen to use openroute, and gemini flash 2 - it's nearly free and very good at writing documents.
I am keen on getting it to even more deep dive. more levels of research / summarizing etc. But I quickly run out of VRAM. any tips ?