I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.
With the upcoming demise of reddit RSS, I’m wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don’t know what I don’t know, so I thought I’d pose the question to the crowd.


New models
Aleph-Alpha/Kolibri-1 – 78.1B total / 3.46B active, up to 1M context, Apache 2.0 A new open-weight English-German Mixture-of-Experts reasoning model from the German lab Aleph Alpha, open-sourced Oct 3, 2026 (Germany’s Reunification Day). 78.1B total params, ~3.46B active per token, native context up to 1,048,576 tokens, weights on Hugging Face under Apache 2.0. Trained on ~20T tokens (mostly English ~62.5%, German ~23.9%, code ~13.6%). Notable as one of the largest sovereign open-weight releases from Europe. Why it matters to you: the 3.46B-active MoE makes it a strong candidate to replace GPT-OSS at roughly two-thirds the size for German/English workloads, and the 1M context is a big draw. Commenters note the model is “basiclly worse than Qwen3.6-35B at twice the size” (u/Training_Visual6159) and that it is “comparable to Qwen 3.5 35B” (u/SnooPaintings8639), so treat it as a strong first attempt rather than a clear winner. At ~78B it likely needs a 96 GB VRAM build to run comfortably. As of this digest there were no community quant releases (no GGUF/INT4/NVFP4 quants posted), and a commenter asked whether llama.cpp supports the architecture yet – so do not plan a deployment until a serving path and quants exist. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/ Tech report: https://aleph-alpha.com/downloads/tech-report.pdf Model: https://huggingface.co/aleph-alpha/Kolibri-1
Qwen3.8 Flash-Next (~176B total, 6B active) – the recurring constrained-hardware test subject Not new today, but it is the model two of the five featured posts are benchmarking on 16 GB-class hardware. The MoE with 512 experts and a large n-gram/PLE table makes it unusually suited to tiered SSD offloading. See the “Most interesting posts” section for the two 16 GB data points.
bilibili “Index-Translate” – multilingual translation family on Qwen3.5 (lower signal) A small (15?) post announcing a multilingual translation model family based on Qwen3.5. Translation-focused, so limited relevance to your coding/agentic backend; flagging it only because it is a new Qwen3.5-based release. No quants or serving configs posted. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wxa1wr/bilibili_released_indextranslatea_a_multilingual/
Most interesting posts
1. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/ u/Yaniss916 reports two ~300B MoE models, each on ONE 128 GB AMD Strix Halo (Ryzen AI Max+ 395, gfx1151) running a new ExLlamaV3/ROCm engine called Kyojin. Concrete numbers from the post: GLM-5.3-Flash (99.7 GB) hits ~580 tok/s prefill (at 3.5K), 546 at 64K, 26-30 tok/s decode (MTP); MiMo-V2.6-Flash-MOPD (105 GB) hits ~650 tok/s prefill at 4K and up to 44 tok/s decode (speculative, code). Quality: KLD vs official FP8 of 0.151 / 0.0713 and top-1 agreement 89.3% / 92.0% respectively. The GLM pack mixes turboderp’s public 2.05 and 3.05 bpw EXL3 tensors with a custom layer mix. Includes a quickstart command (clone, ./build.sh, hf download, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2, OpenAI-style API). Reproducible on Strix Halo hardware only – the author explicitly lists as “not measured yet” any GPU other than gfx1151, and the conversion pipeline stays private. Highly relevant as a technique reference (EXL3 layer-mix quant + ROCm) even though the target ha rdware is AMD, not your NVIDIA stack. Resources: Engine https://github.com/Yamz-Labs/kyojin . Weights https://huggingface.co/yamz-labs . Base fork https://github.com/vcruz305/exllamav3-amd
2. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/ u/fuzhongkai runs Qwen3.8 Flash-Next 176B on an RTX 3080 Laptop (16 GB VRAM) + 32 GB RAM + SSD using his open-source engine TensorSharp, which does MoE-aware scheduling across cache, VRAM, system RAM, and SSD. Reported results: ~11.09 tok/s decode vs Strata’s 10.24, but end-to-end 16.54 s vs Strata’s 62.15 s on his test (peak GPU 14.8 GB, OS working set 19.74 GiB). Reproducibility caveat (flagged in comments): the post does not name the quantization, and top comments ask for exactly that (u/DigitalguyCH: “very vague without specifying the quantization”; u/wizard_of_menlo_park: “Quantization?”). u/leonbollerup also says he gets ~80 tok/s decode and ~2200 tok/s prefill on the same card with Strata, which contradicts the post’s headline – so treat the end-to-end speedup as single-source and unverified. This is directly relevant to whether a 24 GB 3090 can serve the 176B Flash-Next via SSD tiering, but verify the quant and benchmark method before trusting it. Resources: https://github.com/zhongkaifu/TensorSharp . Model docs https://github.com/zhongkaifu/TensorSharp/blob/main/docs/models/qwen38-flash-next.md
3. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/ u/roofkid built NInfer 4080 to run ISTA-DASLab Qwen3.8-27B-GSQ at 100k context on an RTX 4080 16 GB, claiming up to 2720 tok/s prefill and 262 tok/s generation. Part of the NInfer 5090/4090/3090 family of from-scratch engines. Reproducibility caveat: the post is a project announcement; the headline numbers (2720 pp / 262 gen) are the author’s max-observed, and a top comment (u/Pyrolistical) notes these custom engines often compromise with a quantized KV cache and that his own roofline analysis of llama.cpp on an R9700 found <5% decode improvement left on the table – so the gap between custom and general engines is hardware/kernel-specific, not universal. Comments are mostly “how do I get this on a 3080/5080/4070 Ti” requests, so portability to other cards is unconfirmed. Relevant to you primarily as another data point in the overfit-engine trend and for the ISTA-DASLab 3-bit GSQ quant on Qwen3.8-27B. Resources: https://github.com/roofkid/ninfer-4080 . Quant https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ . ByteShape Qwen3.8-27B reference https://byteshape.com/blogs/Qwen3.8-27B/
4. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/ u/carteakey’s discussion post on the overfit-engine trend (see topics above). 315?, 199 comments. No single reproducible config, but the most substantive community engineering observations of the day: SGLang’s current gaps (u/buttplugs4life4me), the Xeon-Max NUMA speed gap (u/TokenRingAI), and a practitioner who “vibe-coded” his own engine because upstream PRs were ignored (u/wishstudio). Worth reading for the state of general vs. narrow runtimes rather than as a runnable config.
5. https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/ The Kolibri-1 announcement (see New models above). 510?, 153 comments. Not a reproducible-serving post (no quants yet), but the most-discussed new model of the day and directly relevant to a European/sovereign model pipeline and as a GPT-OSS-sized MoE candidate.
Useful resources discovered
https://github.com/Yamz-Labs/kyojin Found in the Kyojin post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/). An ExLlamaV3-based ROCm engine with a quickstart that serves GLM-5.3-Flash / MiMo with an OpenAI-style API. Matters to you as the reference implementation for EXL3 layer-mix quantization and for the MoE-expert offloading pattern, though it is ROCm/Strix-Halo specific and the conversion pipeline is private. Actionable if you have AMD hardware or want to study the quant technique; not directly runnable on your NVIDIA stack as-is.
https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3 Found in a Kyojin comment (u/MarkoMarjamaa, https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/?context=3#pdmp0wf). Public EXL3 weights for Qwen3.8 Flash-Next. The comment notes “with exl3 it would be possible to run Q5-quality quant in Q4 memory.” Directly relevant to your Qwen3.8-27B/Flash-Next deployments if you adopt ExLlamaV3; immediately actionable for EXL3 users.
https://github.com/zhongkaifu/TensorSharp Found in the 176B-on-3080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/). An open-source inference engine doing MoE-aware tiered scheduling (VRAM/RAM/SSD) so a model far larger than VRAM and RAM stays usable. Relevant if you want to serve a 176B MoE on a 24 GB 3090 via SSD; the model-specific doc page (…/docs/models/qwen38-flash-next.md) has the exact launch path. Actionable, but the author has not published the quant used in his benchmark – verify before relying on the tok/s.
https://github.com/roofkid/ninfer-4080 Found in the Ninfer-4080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/). A from-scratch engine running ISTA-DASLab Qwen3.8-27B-GSQ at 100k context on a 16 GB 4080, with a runnable script (…/scripts/run-ninfer-4080.bat). Actionable for 16 GB-class Ampere cards; portability to other cards is unconfirmed and the KV-cache quant is a known tradeoff.
https://huggingface.co/aleph-alpha/Kolibri-1 (+ tech report https://aleph-alpha.com/downloads/tech-report.pdf) Found in the Kolibri-1 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/). Official 78.1B / 3.46B-active English-German MoE weights under Apache 2.0 with a 1M-context spec and a full tech report. Actionable for evaluation once a serving path/quants exist; not deployable on your stack today without a backend and quant.
https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ Found in the Ninfer-4080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/). A 3-bit GSQ quant of Qwen3.8-27B that the engine uses. Relevant if you are tracking sub-4-bit quants for 27B-class models; 3-bit quality for agentic work is unverified here – treat as an extreme-quant data point, not a recommended serving quant.
Interesting anecdotes
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdpighm u/Toxaris71 reports ~11 tok/s decode / 30 tok/s prefill on Qwen 3.8 Flash IQ3_XXS (8 GB 3070 Ti + 32 GB DDR4), and 32 tok/s decode / 80 tok/s prefill with IQ2_0, saying it is “better quality than Orinth 1.5 35B 4-bit” and much faster than Qwen 3.8 27B IQ4_XS (~2.8 tok/s decode). Interesting as a low-VRAM 3070 Ti comparison point for the Flash-Next/27B quants, but no launch command, backend version, or benchmark harness is given – treat as anecdotal.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdofln4 u/leonbollerup claims ~80 tok/s decode and ~2200 tok/s prefill on the same 16 GB 3080 using Strata, which contradicts the post author’s ~11 tok/s decode. Useful as an upper-bound comparison for Flash-Next on 16 GB, but no quant or config is provided – anecdotal and in direct tension with the post.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdq00tm u/Independent_Grade612 reports ~450 prefill / ~30 decode on a 12 GB RTX 3500 Ada + 64 GB RAM using Strata with Swift 1.5 Q2XS. Another single-number 12 GB data point; no reproducible config – anecdotal.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/?context=3#pdocpvd u/TokenRingAI reports ~7 tok/s on llama.cpp and ~12 on SGLang for Qwen Flash Next on his 2x Xeon Max, versus ~67 tg/s and ~900 pp/s on a custom NUMA engine. Notable as the most concrete performance-gap claim in the overfit-engine thread, supporting the case that narrow engines can dramatically outperform general ones on specific hardware. Still anecdotal (no command/config), but the numbers are specific enough to be worth reproducing.
https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/?context=3#pdlfi5y u/FullstackSensei estimates Kolibri-1’s training cost at ~$4M (20T tokens on a 768x B300 cluster over 4 weeks at ~$8/GPU-hr). Directionally interesting for understanding training-cost trends, but it is a single commenter’s estimate, not a published figure – treat as a rough back-of-envelope, not a fact.
Well, that’s super wordy for my taste, but that’s the beauty of local, you can change things to your taste. I’ll probably go Gemma 4 12B or 31B (Qwen 3.8 27B is great for code but I like Gemma for text munging), and am AMD based, but hermes is probably a good tool.
Did you just mod the Reddit Reading skill to use redlib? Why not just use the skill raw (presumably with the OAuth)?
Thanks again.