this post was submitted on 20 Apr 2024
314 points (96.2% liked)
Technology
59405 readers
2527 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related content.
- Be excellent to each another!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, to ask if your bot can be added please contact us.
- Check for duplicates before posting, duplicates may be removed
Approved Bots
founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
At least I can run Llama 3 entirely locally.
I just discovered how easy ollama and open webui are to set up so I've been using llama3 locally too, it was like 20 lines in docker compose, and although I've been using gpt3.5 on and off for a long time I'm much more comfortable using models run locally so I've been playing with it a lot more. It's also cool being able to easily switch models at any point during a conversation. I have like 15 models downloaded, mostly 7b and a few 13b models and they all run fast enough on CPU and generate slightly slower than reading speed and only take ~15-30 seconds to start spitting out a response.
Next I want to set up a vscode plugin so I can use my own locally run codegen models from within vscode.