I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.
With the upcoming demise of reddit RSS, I’m wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don’t know what I don’t know, so I thought I’d pose the question to the crowd.


I’m not the OP, but I’m trying to figure out if local models are always painfully slow or if I’m missing something obvious in my tuning.
Stable Diffusion can whip up a picture in less time on the same hardware, than lama.cpp takes to decide to call an MCP function.
It seems like I must be missing something in my lama.cpp setup, but none of the guides I’ve read have clued me in to what I’ve done wrong.
Ollama performs similarly poorly on the same harsware, so I’ve probably managed to make the se mistake(s) at least twice.
Anyway, that’s the main thing I’m reading along for. Trying to increase my understanding until I catch my own mistakes.
what’s your hardware, and what’s your launch command?
I have a guide for performance tuning that should be a pretty good start
https://lemmus.org/post/24235317
Let me know if you have questions, or maybe just make a post asking how to optimize for your hardware and I’ll try to answer
I would suggest you don’t use Ollama https://sleepingrobots.com/dreams/stop-using-ollama/ If you want a GUI, Unsloth Studio is probably best and open source. LM Studio is good too but closed source.
I just use llama.cpp llama-server with the built-in Web UI
Also check the llama.cpp docs
https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md
https://github.com/ggml-org/llama.cpp/blob/master/docs/development/token_generation_performance_tips.md
I will study these. Thank you!