Skip to content

how much does better models help writing the final output? #5

Description

@sfarrell5123

I'd like pluggable LLM use

I have it running on my rtx 3090, I guess all the VRAM use is in embedding ? it seems to not use much GPU compute. I put ollama on my 4060ti - so the 3090 was dedicated to NaneSage

I am keen to use openroute, and gemini flash 2 - it's nearly free and very good at writing documents.

I am keen on getting it to even more deep dive. more levels of research / summarizing etc. But I quickly run out of VRAM. any tips ?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions